Repository navigation
Fail over ColPali query embeddings to the next endpoint - #445
Merged
Merged
Conversation
embed_for_query picked an arbitrary element of the healthy-endpoint set and never marked an endpoint unhealthy on failure, so with multiple endpoints configured a single unreachable embedding server failed every ColPali query even though another endpoint was healthy. Queries now try endpoints in config order (healthy first, unhealthy as a last resort), mark an endpoint unhealthy when it is unavailable, and fall through to the next one. Queries use a 5s connect timeout so a dead host is detected quickly instead of waiting out the 30s ingestion connect timeout. Non-retryable client errors (e.g. 401) still raise immediately without failover. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Only shorten the connect timeout for queries: production logs show query embeddings occasionally take over 60s to return under ingestion load, so a 60s read timeout would have failed real queries. Image queries (generate_embeddings) had the same no-failover bug as text queries; both now share one failover path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adityav369
added a commit
that referenced
this pull request
Oct 5, 2026
* Fail over ColPali query embeddings to the next endpoint embed_for_query picked an arbitrary element of the healthy-endpoint set and never marked an endpoint unhealthy on failure, so with multiple endpoints configured a single unreachable embedding server failed every ColPali query even though another endpoint was healthy. Queries now try endpoints in config order (healthy first, unhealthy as a last resort), mark an endpoint unhealthy when it is unavailable, and fall through to the next one. Queries use a 5s connect timeout so a dead host is detected quickly instead of waiting out the 30s ingestion connect timeout. Non-retryable client errors (e.g. 401) still raise immediately without failover. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep query read timeout; fail over image queries too Only shorten the connect timeout for queries: production logs show query embeddings occasionally take over 60s to return under ingestion load, so a 60s read timeout would have failed real queries. Image queries (generate_embeddings) had the same no-failover bug as text queries; both now share one failover path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit fc6dd93)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
With multiple ColPali API endpoints configured (
morphik_embedding_api_domain), the query paths —embed_for_query(text) andgenerate_embeddings(image queries):next(iter(self.healthy_endpoints))— an arbitrary element of a set, not config order; andSo if one embedding server became unreachable, every
use_colpali=truequery could fail withEmbeddingUnavailableError, even though a healthy endpoint was configured. Ingestion was unaffected because the distributed ingestion path already fails over.Fix
_embed_query_inputpath: try endpoints in config order (healthy first, unhealthy as a last resort), mark an endpoint unhealthy when it is unavailable, and fall through to the next. An endpoint that succeeds is marked healthy again. The existing 60s cooldown/re-probe logic is reused._call_api_endpointgains an optionaltimeoutoverride; the ingestion path is unchanged.Tests
core/tests/unit/test_week1_ingest_fixes.py(20 passed):test_query_path_retries_once_only: retry once per endpointconnect🤖 Generated with Claude Code