Skip to content

Fail over ColPali query embeddings to the next endpoint - #445

Merged
Adityav369 merged 2 commits into
mainfrom
fix/colpali-query-failover
Oct 5, 2026
Merged

Adityav369 merged 2 commits into
mainfrom
fix/colpali-query-failover

Conversation

@Adityav369

@Adityav369 Adityav369 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

With multiple ColPali API endpoints configured (morphik_embedding_api_domain), the query paths — embed_for_query (text) and generate_embeddings (image queries):

  • picked next(iter(self.healthy_endpoints)) — an arbitrary element of a set, not config order; and
  • never marked an endpoint unhealthy on failure, and never tried another endpoint.

So if one embedding server became unreachable, every use_colpali=true query could fail with EmbeddingUnavailableError, even though a healthy endpoint was configured. Ingestion was unaffected because the distributed ingestion path already fails over.

Fix

  • Text and image queries share one _embed_query_input path: try endpoints in config order (healthy first, unhealthy as a last resort), mark an endpoint unhealthy when it is unavailable, and fall through to the next. An endpoint that succeeds is marked healthy again. The existing 60s cooldown/re-probe logic is reused.
  • Queries use a 5s connect timeout so an unreachable host is detected quickly instead of waiting out the 30s connect timeout (×2 attempts). Read/write/pool timeouts are unchanged from the client default: query embeddings can legitimately take over a minute to return under heavy ingestion load, so those must not be shortened.
  • Non-retryable client errors (e.g. 401) still raise immediately without failover or health changes.
  • _call_api_endpoint gains an optional timeout override; the ingestion path is unchanged.

Tests

core/tests/unit/test_week1_ingest_fixes.py (20 passed):

  • updated test_query_path_retries_once_only: retry once per endpoint
  • new: text query fails over on a dead endpoint + skips it during cooldown
  • new: image query fails over on a dead endpoint
  • new: config order is respected among healthy endpoints
  • new: no failover on a 401
  • new: query timeout only shortens connect

🤖 Generated with Claude Code

Adityav369 and others added 2 commits October 5, 2026 08:36
embed_for_query picked an arbitrary element of the healthy-endpoint set and
never marked an endpoint unhealthy on failure, so with multiple endpoints
configured a single unreachable embedding server failed every ColPali query
even though another endpoint was healthy.

Queries now try endpoints in config order (healthy first, unhealthy as a last
resort), mark an endpoint unhealthy when it is unavailable, and fall through
to the next one. Queries use a 5s connect timeout so a dead host is detected
quickly instead of waiting out the 30s ingestion connect timeout. Non-retryable
client errors (e.g. 401) still raise immediately without failover.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Only shorten the connect timeout for queries: production logs show query
embeddings occasionally take over 60s to return under ingestion load, so a
60s read timeout would have failed real queries. Image queries
(generate_embeddings) had the same no-failover bug as text queries; both now
share one failover path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Adityav369
Adityav369 merged commit fc6dd93 into main Oct 5, 2026
9 checks passed
@Adityav369
Adityav369 deleted the fix/colpali-query-failover branch October 5, 2026 16:13
Adityav369 added a commit that referenced this pull request Oct 5, 2026
* Fail over ColPali query embeddings to the next endpoint

embed_for_query picked an arbitrary element of the healthy-endpoint set and
never marked an endpoint unhealthy on failure, so with multiple endpoints
configured a single unreachable embedding server failed every ColPali query
even though another endpoint was healthy.

Queries now try endpoints in config order (healthy first, unhealthy as a last
resort), mark an endpoint unhealthy when it is unavailable, and fall through
to the next one. Queries use a 5s connect timeout so a dead host is detected
quickly instead of waiting out the 30s ingestion connect timeout. Non-retryable
client errors (e.g. 401) still raise immediately without failover.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep query read timeout; fail over image queries too

Only shorten the connect timeout for queries: production logs show query
embeddings occasionally take over 60s to return under ingestion load, so a
60s read timeout would have failed real queries. Image queries
(generate_embeddings) had the same no-failover bug as text queries; both now
share one failover path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit fc6dd93)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant