Skip to content

feat: add support for LoRA adapters - #25

Merged
daavoo merged 3 commits into
mozilla-ai:mainfrom
oglego:feature/support-lora-adapters
Oct 1, 2026
Merged

daavoo merged 3 commits into
mozilla-ai:mainfrom
oglego:feature/support-lora-adapters

Conversation

@oglego

@oglego oglego commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Summary

This PR adds support for LoRA adapters and closes #18. If there is anything that I should modify or change on this just let me know!

Changes

The main items that I have added on this PR are:

  • LoraAdapterConfig (path, scale) and ModelConfig::loras (std::vector<LoraAdapterConfig>) for configuring one or more adapters.
  • Model::initialize_context() loads and applies LoRA adapters to the model context with individual scaling via llama_set_adapters_lora().
  • Model::get_loras() getter for inspecting loaded adapter pointers.
  • examples/lora/ — Example application with documentation for converting adapters and using task prompt templates (including a worked example using IBM Granite RAG query rewriting).
  • tests/test_lora.cpp — unit tests covering config defaults, default scale, and multiple adapter stacking.
  • Root and examples README.md updates documenting LoRA usage and multi-agent context sharing.

For testing this locally and interacting with it I used the recommended model (granite-4.0-micro-Q8_0.gguf) along with a converted Granite query-rewriter LoRA adapter (query_rewrite_lora.gguf).

Testing

  • ctest — all 5 suites pass (ToolTests, CallbacksTests, GrammarTests, LoraTests, ChatParserTests)
  • clang-format passes
  • Manual verification against a real adapter: granite-4.0-micro-Q8_0.gguf + IBM Granite RAG's query-rewrite LoRA, confirmed the adapter correctly rewrites contextual follow-ups (e.g., "What about Germany?" → "What is the capital of Germany?") that the base model alone does not

Thanks in advance for the review!!

@daavoo daavoo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks! I pushed a small README note (7a68025) explaining that prompt cache files don't record which LoRA adapters were active. Storing an adapter fingerprint with the cache can be a follow-up.

@daavoo
daavoo force-pushed the feature/support-lora-adapters branch from 7a68025 to f3b1679 Compare October 1, 2026 03:02
@daavoo
daavoo merged commit fd62492 into mozilla-ai:main Oct 1, 2026
5 checks passed
@oglego

oglego commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Thank you!!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature Request: Support for GBNF grammar and LoRA adapters in multi‑agent setups

2 participants