Skip to content

fix(rag): bound cloud embedding request batches - #382

Draft
Milbaxter wants to merge 1 commit into
singnet:mainfrom
Milbaxter:fix/bound-cloud-embedding-batches
Draft

Milbaxter wants to merge 1 commit into
singnet:mainfrom
Milbaxter:fix/bound-cloud-embedding-batches

Conversation

@Milbaxter

Copy link
Copy Markdown

Description

Fixes #366.

Runtime knowledge loading sends all chunks of one source file to the embedding API in one request. A large file can exceed the request limit and stop the knowledge load before later files are indexed.

cloud_embed_batch now sends at most 32 texts per request and joins the returned vectors in input order. The new embeddingBatchSize setting lets operators tune the request size for their provider. It accepts positive integers from YAML or command-line/environment strings. Empty input makes no request; a failed batch raises the existing error without returning a partial result.

The OpenAI embedding API reference sets an 8192-token limit per input and a 300,000-token limit per request. With individually valid inputs, the default caps a request at 262,144 tokens and 32 inputs. Custom sizes must fit the provider's limits. Individual overlong texts are still the chunker's responsibility; this change does not truncate them.

How Has This Been Tested?

  • Reproduced the original failure with a mock endpoint that enforces the 300,000-token request limit for simulated 8192-token inputs. The 70-input call failed before the change; it now sends batches of 32, 32, and 6 and returns all vectors in order.
  • python -m pytest -q tests: 128 passed, 5 skipped on macOS/Python 3.12.12. The skips are four Linux-only cases and the optional import-kb model comparison (dependency not installed).
  • python -m pytest -q --noconftest Autotests/test_openai_runtime_embeddings.py: 6 passed. These retain provider/model routing and single-memory embedding behavior. --noconftest omits the live-agent cleanup fixture.
  • The new tests also cover ASICloud routing with custom batch sizes, CLI string values, invalid settings, empty input, and failure in a later batch.
  • git diff --check: passed.

No live embedding API requests were made. The Docker daemon is unavailable on this host, so the mandatory container scenario suite and MeTTa tests were not run locally. This PR is a draft pending that CI validation.

Checklist

  • PR contains autogenerated code
  • Self-review completed
  • Test scenarios above are passed with the version of the code from PR

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

cloud_embed_batch sends one request per source file, which exceeds the embeddings API's 300,000-token per-request limit

2 participants