Skip to content

Latest commit

 

History

History
387 lines (324 loc) · 11 KB

File metadata and controls

387 lines (324 loc) · 11 KB

ML Services Documentation

Overview

The Photobook system uses three specialized microservices for ML-powered image analysis. Each service is stateless and exposes a REST API. The backend orchestrates these services through background jobs with circuit breaker protection.

┌─────────────────────────────────────────────────────────────────────┐
│                           Backend                                    │
│                                                                      │
│  ┌──────────────────┐  ┌───────────────────┐  ┌──────────────────┐  │
│  │ VLM Tasks        │  │ Location Tasks    │  │ Face Tasks       │  │
│  │ (jobs/tasks/vlm) │  │ (jobs/tasks/loc)  │  │ (jobs/tasks/face)│  │
│  └────────┬─────────┘  └─────────┬─────────┘  └────────┬─────────┘  │
│           │                      │                      │            │
│           │ HTTP (Circuit Breaker)                      │            │
└───────────┼──────────────────────┼──────────────────────┼────────────┘
            │                      │                      │
            ▼                      ▼                      ▼
     ┌──────────────┐       ┌──────────────┐       ┌──────────────┐
     │ VLM Service  │       │ Geo Service  │       │ Face Service │
     │ Port: 8031   │       │ Port: 8030   │       │ Port: 8033   │
     │              │       │              │       │              │
     │ Qwen3-VL-32B │       │ Google Places│       │ InsightFace  │
     │ llama.cpp    │       │ API          │       │ buffalo_l    │
     └──────────────┘       └──────────────┘       └──────────────┘

VLM Service (Port 8031)

Purpose

Vision-Language Model service using Qwen3-VL-32B for comprehensive image understanding:

  • Object detection with bounding boxes
  • People detection
  • Face localization
  • Image description and tagging
  • Scene context extraction
  • Semantic relationship extraction
  • Caption generation
  • Landmark identification

Directory Structure

services/vlm-service/
├── api.py              # FastAPI application
├── config.py           # Service configuration
├── llama_server.py     # llama.cpp server manager
├── clients/
│   └── llama_cpp.py    # llama.cpp client
└── requirements.txt

API Endpoints

Endpoint Method Description
/health GET Service health check
/analyze POST Custom prompt analysis
/describe POST Comprehensive image description
/people POST Identify and describe people
/detect/objects POST Object detection with bboxes
/detect/people POST People detection with bboxes
/detect/faces POST Face detection with bboxes
/scene-graph POST Extract scene graph with relationships
/extract POST Unified extraction (PREFERRED)
/caption POST Generate caption from facts
/identify_landmark POST Match image to landmark candidates

Key Endpoint: /extract

The unified extraction endpoint makes a single VLM call for comprehensive analysis:

Request:

curl -X POST http://localhost:8031/extract \
  -F "image=@photo.jpg" \
  -F "include_relationships=true"

Response:

{
  "objects": [
    {"bbox_2d": [100, 200, 300, 400], "label": "car"}
  ],
  "people": [
    {"bbox_2d": [500, 100, 700, 600], "label": "woman", "face_bbox": [520, 110, 580, 180]}
  ],
  "description": "A woman standing next to a red car in a parking lot.",
  "tags": ["outdoor", "parking", "vehicle", "person"],
  "scene_context": ["parking lot", "daytime", "urban"],
  "relationships": [
    {"subject": "person_1", "predicate": "standing_next_to", "object": "object_0"}
  ],
  "actions": [
    {"actor": "person_1", "action": "standing"}
  ],
  "background_objects": ["building", "trees", "sky"]
}

Bounding Box Format

All bounding boxes use normalized 0-1000 scale:

  • bbox_2d: [x1, y1, x2, y2] where (x1,y1) is top-left, (x2,y2) is bottom-right
  • Convert to pixel coordinates: pixel_x = bbox_x * image_width / 1000

Configuration

# config.py
llama_server_url = "http://localhost:8089"  # llama.cpp server
service_port = 8031
model_name = "qwen3-vl-32b"

Geo Service (Port 8030)

Purpose

Stateless geo-resolution service that converts GPS coordinates to rich location data:

  • Reverse geocoding (GPS → address)
  • Nearby place discovery
  • Landmark identification
  • Scene-based place type filtering

Directory Structure

services/geo-service/
├── api.py              # FastAPI application
├── config.py           # Service configuration
├── clients/
│   └── google_places_client.py
└── requirements.txt

API Endpoints

Endpoint Method Description
/health GET Service health check
/resolve POST Resolve GPS to location data

Key Endpoint: /resolve

Request:

{
  "latitude": 28.272461,
  "longitude": -16.642578,
  "include_landmarks": true,
  "landmark_radius": 500,
  "max_landmarks": 3,
  "scene_hints": ["volcano", "mountain", "hiking"]
}

Response:

{
  "success": true,
  "location": {
    "place_id": "ChIJ...",
    "name": "Teide National Park",
    "formatted_address": "38300 La Orotava, Santa Cruz de Tenerife, Spain",
    "types": ["park", "natural_feature", "tourist_attraction"],
    "geometry": {"lat": 28.272, "lng": -16.642},
    "city": "La Orotava",
    "state": "Santa Cruz de Tenerife",
    "country": "Spain",
    "country_code": "ES",
    "nearby_landmarks": [
      {
        "name": "Mount Teide",
        "types": ["natural_feature", "tourist_attraction"],
        "distance_meters": 1200,
        "rating": 4.8
      }
    ]
  }
}

Scene-to-Place Type Mapping

VLM scene hints are mapped to Google Places API types:

Scene Hint Google Types
zoo zoo, tourist_attraction, amusement_park
restaurant restaurant, cafe, bar
hotel lodging
beach natural_feature, park, tourist_attraction
museum museum, art_gallery, tourist_attraction
church church, place_of_worship

Configuration

# config.py
GOOGLE_PLACES_API_KEY = os.getenv("GOOGLE_PLACES_API_KEY")
service_port = 8030

Face Service (Port 8033)

Purpose

Face detection and embedding service using InsightFace buffalo_l model:

  • Face detection with quality scoring
  • 512-dimensional embedding extraction
  • Face matching/comparison
  • Batch embedding for multiple faces

Directory Structure

services/face-service/
├── api.py              # FastAPI application
├── config.py           # Service configuration
├── clients/
│   └── face_model.py   # InsightFace wrapper
└── requirements.txt

API Endpoints

Endpoint Method Description
/health GET Service health check
/detect POST Detect faces with bboxes
/embed POST Extract embedding for one face
/embed-batch POST Extract embeddings for all faces
/match POST Match embedding against known set
/compare POST Compare two embeddings

Key Endpoint: /embed-batch

Request:

curl -X POST http://localhost:8033/embed-batch \
  -F "image=@group_photo.jpg"

Response:

{
  "embeddings": [
    {
      "embedding": [0.023, -0.145, ...],  // 512 floats
      "quality_score": 0.92,
      "bbox_2d": [100, 150, 200, 280],
      "det_score": 0.98
    },
    {
      "embedding": [...],
      "quality_score": 0.78,
      "bbox_2d": [400, 160, 510, 300],
      "det_score": 0.95
    }
  ],
  "count": 2
}

Key Endpoint: /match

Request:

{
  "query_embedding": [0.023, -0.145, ...],
  "known_embeddings": [
    {"id": "person_1", "embedding": [...]},
    {"id": "person_2", "embedding": [...]}
  ],
  "threshold": 0.4
}

Response:

{
  "matches": [
    {"id": "person_1", "similarity": 0.87}
  ],
  "count": 1
}

Configuration

# config.py
model_name = "buffalo_l"
match_threshold = 0.4
quality_threshold = 0.6
service_port = 8033

Shared Module

Common utilities used across services:

services/shared/
├── logging.py      # Structured logging with structlog
├── validation.py   # Image validation helpers
└── errors.py       # Error response utilities

Structured Logging

from shared.logging import configure_logging, get_logger

configure_logging(service_name="vlm-service", log_level="INFO")
logger = get_logger(__name__)

logger.info("Processing image", image_id="abc123", size=1024000)

Image Validation

from shared.validation import read_and_validate_image

image_data = await read_and_validate_image(upload_file)
# Raises HTTPException if invalid

Running Services

Development

Each service can be run independently:

# VLM Service
cd services/vlm-service
source ../../backend/.venv/bin/activate
uvicorn api:app --host 0.0.0.0 --port 8031 --reload

# Geo Service
cd services/geo-service
source .venv/bin/activate
uvicorn api:app --host 0.0.0.0 --port 8030 --reload

# Face Service
cd services/face-service
source ../../backend/.venv/bin/activate
uvicorn api:app --host 0.0.0.0 --port 8033 --reload

Health Checks

curl http://localhost:8031/health  # VLM
curl http://localhost:8030/health  # Geo
curl http://localhost:8033/health  # Face

Backend Service Status

The backend provides an aggregated status view:

curl http://localhost:8000/services/status

Response includes:

  • Individual service health
  • Circuit breaker state
  • Recent failure counts

Circuit Breaker Integration

The backend protects service calls with circuit breakers:

Service Circuit Name Failure Threshold Recovery Timeout
VLM vlm 3 60s
Geo geo 3 30s
Face face 3 60s

States:

  • CLOSED: Normal operation
  • OPEN: Fast-fail (service down)
  • HALF_OPEN: Testing recovery

Monitor via: GET http://localhost:8000/circuits