The Photobook system uses three specialized microservices for ML-powered image analysis. Each service is stateless and exposes a REST API. The backend orchestrates these services through background jobs with circuit breaker protection.
┌─────────────────────────────────────────────────────────────────────┐
│ Backend │
│ │
│ ┌──────────────────┐ ┌───────────────────┐ ┌──────────────────┐ │
│ │ VLM Tasks │ │ Location Tasks │ │ Face Tasks │ │
│ │ (jobs/tasks/vlm) │ │ (jobs/tasks/loc) │ │ (jobs/tasks/face)│ │
│ └────────┬─────────┘ └─────────┬─────────┘ └────────┬─────────┘ │
│ │ │ │ │
│ │ HTTP (Circuit Breaker) │ │
└───────────┼──────────────────────┼──────────────────────┼────────────┘
│ │ │
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ VLM Service │ │ Geo Service │ │ Face Service │
│ Port: 8031 │ │ Port: 8030 │ │ Port: 8033 │
│ │ │ │ │ │
│ Qwen3-VL-32B │ │ Google Places│ │ InsightFace │
│ llama.cpp │ │ API │ │ buffalo_l │
└──────────────┘ └──────────────┘ └──────────────┘
Vision-Language Model service using Qwen3-VL-32B for comprehensive image understanding:
- Object detection with bounding boxes
- People detection
- Face localization
- Image description and tagging
- Scene context extraction
- Semantic relationship extraction
- Caption generation
- Landmark identification
services/vlm-service/
├── api.py # FastAPI application
├── config.py # Service configuration
├── llama_server.py # llama.cpp server manager
├── clients/
│ └── llama_cpp.py # llama.cpp client
└── requirements.txt
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Service health check |
/analyze |
POST | Custom prompt analysis |
/describe |
POST | Comprehensive image description |
/people |
POST | Identify and describe people |
/detect/objects |
POST | Object detection with bboxes |
/detect/people |
POST | People detection with bboxes |
/detect/faces |
POST | Face detection with bboxes |
/scene-graph |
POST | Extract scene graph with relationships |
/extract |
POST | Unified extraction (PREFERRED) |
/caption |
POST | Generate caption from facts |
/identify_landmark |
POST | Match image to landmark candidates |
The unified extraction endpoint makes a single VLM call for comprehensive analysis:
Request:
curl -X POST http://localhost:8031/extract \
-F "image=@photo.jpg" \
-F "include_relationships=true"Response:
{
"objects": [
{"bbox_2d": [100, 200, 300, 400], "label": "car"}
],
"people": [
{"bbox_2d": [500, 100, 700, 600], "label": "woman", "face_bbox": [520, 110, 580, 180]}
],
"description": "A woman standing next to a red car in a parking lot.",
"tags": ["outdoor", "parking", "vehicle", "person"],
"scene_context": ["parking lot", "daytime", "urban"],
"relationships": [
{"subject": "person_1", "predicate": "standing_next_to", "object": "object_0"}
],
"actions": [
{"actor": "person_1", "action": "standing"}
],
"background_objects": ["building", "trees", "sky"]
}All bounding boxes use normalized 0-1000 scale:
bbox_2d: [x1, y1, x2, y2]where (x1,y1) is top-left, (x2,y2) is bottom-right- Convert to pixel coordinates:
pixel_x = bbox_x * image_width / 1000
# config.py
llama_server_url = "http://localhost:8089" # llama.cpp server
service_port = 8031
model_name = "qwen3-vl-32b"Stateless geo-resolution service that converts GPS coordinates to rich location data:
- Reverse geocoding (GPS → address)
- Nearby place discovery
- Landmark identification
- Scene-based place type filtering
services/geo-service/
├── api.py # FastAPI application
├── config.py # Service configuration
├── clients/
│ └── google_places_client.py
└── requirements.txt
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Service health check |
/resolve |
POST | Resolve GPS to location data |
Request:
{
"latitude": 28.272461,
"longitude": -16.642578,
"include_landmarks": true,
"landmark_radius": 500,
"max_landmarks": 3,
"scene_hints": ["volcano", "mountain", "hiking"]
}Response:
{
"success": true,
"location": {
"place_id": "ChIJ...",
"name": "Teide National Park",
"formatted_address": "38300 La Orotava, Santa Cruz de Tenerife, Spain",
"types": ["park", "natural_feature", "tourist_attraction"],
"geometry": {"lat": 28.272, "lng": -16.642},
"city": "La Orotava",
"state": "Santa Cruz de Tenerife",
"country": "Spain",
"country_code": "ES",
"nearby_landmarks": [
{
"name": "Mount Teide",
"types": ["natural_feature", "tourist_attraction"],
"distance_meters": 1200,
"rating": 4.8
}
]
}
}VLM scene hints are mapped to Google Places API types:
| Scene Hint | Google Types |
|---|---|
| zoo | zoo, tourist_attraction, amusement_park |
| restaurant | restaurant, cafe, bar |
| hotel | lodging |
| beach | natural_feature, park, tourist_attraction |
| museum | museum, art_gallery, tourist_attraction |
| church | church, place_of_worship |
# config.py
GOOGLE_PLACES_API_KEY = os.getenv("GOOGLE_PLACES_API_KEY")
service_port = 8030Face detection and embedding service using InsightFace buffalo_l model:
- Face detection with quality scoring
- 512-dimensional embedding extraction
- Face matching/comparison
- Batch embedding for multiple faces
services/face-service/
├── api.py # FastAPI application
├── config.py # Service configuration
├── clients/
│ └── face_model.py # InsightFace wrapper
└── requirements.txt
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Service health check |
/detect |
POST | Detect faces with bboxes |
/embed |
POST | Extract embedding for one face |
/embed-batch |
POST | Extract embeddings for all faces |
/match |
POST | Match embedding against known set |
/compare |
POST | Compare two embeddings |
Request:
curl -X POST http://localhost:8033/embed-batch \
-F "image=@group_photo.jpg"Response:
{
"embeddings": [
{
"embedding": [0.023, -0.145, ...], // 512 floats
"quality_score": 0.92,
"bbox_2d": [100, 150, 200, 280],
"det_score": 0.98
},
{
"embedding": [...],
"quality_score": 0.78,
"bbox_2d": [400, 160, 510, 300],
"det_score": 0.95
}
],
"count": 2
}Request:
{
"query_embedding": [0.023, -0.145, ...],
"known_embeddings": [
{"id": "person_1", "embedding": [...]},
{"id": "person_2", "embedding": [...]}
],
"threshold": 0.4
}Response:
{
"matches": [
{"id": "person_1", "similarity": 0.87}
],
"count": 1
}# config.py
model_name = "buffalo_l"
match_threshold = 0.4
quality_threshold = 0.6
service_port = 8033Common utilities used across services:
services/shared/
├── logging.py # Structured logging with structlog
├── validation.py # Image validation helpers
└── errors.py # Error response utilities
from shared.logging import configure_logging, get_logger
configure_logging(service_name="vlm-service", log_level="INFO")
logger = get_logger(__name__)
logger.info("Processing image", image_id="abc123", size=1024000)from shared.validation import read_and_validate_image
image_data = await read_and_validate_image(upload_file)
# Raises HTTPException if invalidEach service can be run independently:
# VLM Service
cd services/vlm-service
source ../../backend/.venv/bin/activate
uvicorn api:app --host 0.0.0.0 --port 8031 --reload
# Geo Service
cd services/geo-service
source .venv/bin/activate
uvicorn api:app --host 0.0.0.0 --port 8030 --reload
# Face Service
cd services/face-service
source ../../backend/.venv/bin/activate
uvicorn api:app --host 0.0.0.0 --port 8033 --reloadcurl http://localhost:8031/health # VLM
curl http://localhost:8030/health # Geo
curl http://localhost:8033/health # FaceThe backend provides an aggregated status view:
curl http://localhost:8000/services/statusResponse includes:
- Individual service health
- Circuit breaker state
- Recent failure counts
The backend protects service calls with circuit breakers:
| Service | Circuit Name | Failure Threshold | Recovery Timeout |
|---|---|---|---|
| VLM | vlm | 3 | 60s |
| Geo | geo | 3 | 30s |
| Face | face | 3 | 60s |
States:
- CLOSED: Normal operation
- OPEN: Fast-fail (service down)
- HALF_OPEN: Testing recovery
Monitor via: GET http://localhost:8000/circuits