Repository navigation
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes #366.
Runtime knowledge loading sends all chunks of one source file to the embedding API in one request. A large file can exceed the request limit and stop the knowledge load before later files are indexed.
cloud_embed_batchnow sends at most 32 texts per request and joins the returned vectors in input order. The newembeddingBatchSizesetting lets operators tune the request size for their provider. It accepts positive integers from YAML or command-line/environment strings. Empty input makes no request; a failed batch raises the existing error without returning a partial result.The OpenAI embedding API reference sets an 8192-token limit per input and a 300,000-token limit per request. With individually valid inputs, the default caps a request at 262,144 tokens and 32 inputs. Custom sizes must fit the provider's limits. Individual overlong texts are still the chunker's responsibility; this change does not truncate them.
How Has This Been Tested?
python -m pytest -q tests: 128 passed, 5 skipped on macOS/Python 3.12.12. The skips are four Linux-only cases and the optionalimport-kbmodel comparison (dependency not installed).python -m pytest -q --noconftest Autotests/test_openai_runtime_embeddings.py: 6 passed. These retain provider/model routing and single-memory embedding behavior.--noconftestomits the live-agent cleanup fixture.git diff --check: passed.No live embedding API requests were made. The Docker daemon is unavailable on this host, so the mandatory container scenario suite and MeTTa tests were not run locally. This PR is a draft pending that CI validation.
Checklist