A deterministic Rust memory and reliability kernel for AI agents.
Hikmah Stack is an open-source reference implementation for keeping selected agent responsibilities outside a generative model's hidden state. Its Rust kernel provides local, inspectable mechanisms for provenance-bearing memory, contradiction detection, bounded recall, commitments, symbolic planning, decision controls, and narrow completion checks.
Portfolio summary: This repository demonstrates systems design and hands-on Rust implementation for deterministic Agentic AI memory and reliability controls. It is a working local proof of concept, not an end-to-end enterprise GenAI platform.
Maintained by Juber Shaikh · MIT licensed
The table below separates executable evidence from architectural intent.
| Demonstrated capability | Inspectable proof | Evidence level |
|---|---|---|
| Typed agent memory with provenance, confidence, privacy, deadlines, claims, and correction links | Trace, Provenance, and validation |
Implemented |
| Append-only, sequence-numbered, hash-chained local ledger | MemoryStore::open, integrity replay, append, and verify |
Implemented; persistence and replay are tested |
| Contradiction-aware structured claims | Conflict detection and integration test | Implemented and tested |
| Contextual bounded recall using lexical, tag, recency, salience, confidence, provenance, and deadline signals | Recall scoring and redundancy suppression and persistence/recall test | Implemented and tested |
| Evidence-preserving consolidation proposals | consolidation_proposals |
Implemented; no automatic promotion |
| Prospective commitment recall | commitments_due |
Implemented |
| Bounded symbolic planning | Planner and integration test | Implemented and tested |
| Evidence-adjusted decision ranking with hard blocks and reversibility | Decision evaluator and example frame | Implemented with a runnable example |
| Independent evidence, memory, risk, human-impact, and delivery checks | deliberate |
Implemented as deterministic concurrent checks; not LLM agents |
| Narrow completion-claim hygiene | Rust Truth Gate, shell launcher, and Python compatibility fallback | Implemented; deliberately not a fact-checker |
| Reusable host packaging | Codex manifest, Claude manifest, Kimi manifest, and portable skills | Configuration and instruction layer |
| Automated validation | GitHub Actions workflow runs formatting, Clippy with warnings denied, Rust tests, package validation, and Python syntax compilation | CI-backed repository validation |
| Area | Current state |
|---|---|
| Core runtime | Working Rust CLI and library reference implementation |
| Persistence | Local append-only JSONL file with tamper-evident hash chaining |
| Retrieval | Deterministic token/tag/time/provenance scoring; no embeddings |
| Model integration | A ProposalEngine extension boundary plus NoModel; no provider adapter ships today |
| Agent packaging | Portable instruction skills and thin Codex, Claude Code, and Kimi manifests |
| Tests | Focused unit/integration coverage for recall, ledger replay, contradictions, and planning |
| Deployment | Local source/CLI use; no hosted service or public production deployment is claimed |
Hikmah Stack does not currently implement or claim:
- an LLM inference application or model-training pipeline;
- RAG, document ingestion, embeddings, reranking, or a vector database;
- an LLM multi-agent runtime or orchestration through LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or Copilot Studio;
- a custom LLM tool/function-calling runtime or autonomous execution against external systems;
- Azure OpenAI, Azure AI Foundry, AWS Bedrock, or another cloud-model integration;
- a Python AI/GenAI application—the Python file is only a small compatibility fallback for the Truth Gate;
- an HTTP API, MCP server, enterprise application/database/RPA connector, or multi-tenant service;
- production-scale security, encryption, access control, observability, load testing, or deployment automation.
These are integration opportunities, not hidden capabilities. The current value is the deterministic kernel and the explicit control boundary it gives future model- and tool-driven systems.
Hikmah separates generative proposals from durable state and deterministic controls.
flowchart TD
H["Agent host and portable skills"] --> K["Deterministic Rust kernel"]
P["Optional proposal engine"] --> K
K --> M["Local hash-chained memory"]
K --> C["Recall, planning, decisions, gates"]
The current release ships the kernel and a narrow ProposalEngine interface. Its only concrete engine is NoModel, so model integration remains outside the present implementation. A future LLM, rules engine, search system, or local model can propose outputs through that boundary without automatically gaining authority over durable memory or policy.
See Architecture, Cognitive Kernel, and Co-Model Architecture.
| Capability | Responsibility |
|---|---|
| Operator Core | Human judgment, leadership, ethics, communication, recovery, and stewardship |
| Agent Radar | Diagnose hallucination, loops, context loss, sycophancy, opacity, cost, and memory failures |
| Decision Forge | Structure options, evidence coverage, hard constraints, reversibility, and action |
| Ship Guard | Define acceptance criteria, verification, rollback, handoffs, and completion discipline |
| Hikmah Orchestrator | Route across the portable skills and synthesize one response |
| Cognitive Kernel | Maintain local typed memory, contradictions, commitments, and deterministic controls |
| Truth Gate | Catch a narrow class of contradictory completion claims at host stop time |
The first five capabilities are primarily portable instruction skills. The Cognitive Kernel and Rust Truth Gate contain the executable runtime behavior.
A memory is an immutable trace with a kind, content, provenance, confidence, salience, privacy class, creation time, optional deadline, optional structured claim, and optional correction link.
The local ledger is append-only and hash-chained. Corrections can supersede earlier traces without rewriting history, and conflicting structured claims remain visible rather than silently replacing one another.
At recall time, the kernel scores active traces using:
- lexical overlap;
- explicit tags;
- recency;
- salience;
- confidence;
- provenance authority and verification state;
- prospective urgency for commitments.
It then suppresses redundant results to produce a bounded working set. This is a transparent deterministic baseline, not semantic embedding retrieval. Read Memory and the Remember → Recall → Consolidate playbook.
- Rust stable toolchain with Cargo
- Python 3 only if the zero-install compatibility hook is needed
git clone https://github.com/CodeWithJuber/hikmah-stack.git
cd hikmah-stack
cargo fmt --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p hikmah-kernel -- validate --root .cargo run -p hikmah-kernel -- init
cargo run -p hikmah-kernel -- remember \
--kind observation \
--content "Migration failed because the lock timed out" \
--source incident-review \
--tag database \
--tag deployment \
--salience 0.9 \
--confidence 0.9 \
--verified
cargo run -p hikmah-kernel -- recall \
--query "why did the deployment fail"
cargo run -p hikmah-kernel -- consolidate
cargo run -p hikmah-kernel -- commitments --within-hours 168
cargo run -p hikmah-kernel -- verify-ledgerBy default, local memory is written to .hikmah/memory.jsonl.
cargo run -p hikmah-kernel -- plan \
--problem examples/plan-problem.json
cargo run -p hikmah-kernel -- decide \
--frame examples/decision-frame.json
cargo run -p hikmah-kernel -- deliberate \
--unverified-claims 1 \
--memory-conflicts 0 \
--irreversible-actions 0 \
--human-impact-questions 0 \
--missing-acceptance-criteria 1The files under skills/ are the portable layer. Host adapters package or preload those instructions; they do not turn Hikmah into a model provider or multi-agent runtime.
The read-only Claude adapter declares a small set of host-provided repository tools. That configuration should not be confused with an implemented custom function-calling pipeline or enterprise action layer.
Register the repository as a marketplace source:
codex plugin marketplace add CodeWithJuber/hikmah-stackThis command adds the catalog source; install and test the plugin from the ChatGPT desktop app's Plugins Directory. The included Codex manifest references the portable skills and the command-based completion hook.
/plugin marketplace add CodeWithJuber/hikmah-stack
/plugin install hikmah-stack@hikmah-stack
If the install summary asks you to activate the plugin, run /reload-plugins. The repository also includes a read-only Claude orchestrator adapter and a deterministic-plus-prompt completion check.
The root kimi.plugin.json points to ./skills/ and supplies routing guidance through skillInstructions. Add the repository as a Kimi plugin, or package the repository for the relevant catalog flow.
Reuse the required directories under skills/. Keep host-specific adapters thin and review executable hooks before enabling them.
See Compatibility.
Hikmah borrows software design ideas from research on selective consolidation, replay, temporal structure, correction, bounded attention, and prospective memory. The project does not claim to reproduce a brain or prove a software mechanism from a biological analogy.
Research sources and limitations are maintained in Research Notes. Proposed measurements are listed in the Evaluation Contract; that document is a measurement specification, not evidence that all benchmarks have already been run.
- Evidence before confidence.
- Memory is typed, provenance-bearing state, not hidden chain-of-thought.
- Corrections supersede; contradictions remain visible.
- Models propose; explicit controls govern durable state transitions.
- Human-impact, privacy, consent, and safety blocks are not averaged away by a high score.
- Independent challenge surfaces failures more clearly than self-grading alone.
- Observed outcomes are more useful than unverified plans.
- Unknown is a valid serialized state.
- New architecture should beat a measurable baseline before it is preferred for novelty.
- Discovered failures should become tests, controls, or documented limitations.
runtime/hikmah-kernel/ deterministic Rust library and CLI
runtime/hikmah-kernel/tests focused integration tests
skills/ portable judgment and cognition instructions
agents/ thin host-specific agent adapter
playbooks/ operational cognitive loops
lenses/ reusable diagnostic perspectives
docs/ architecture, memory, research, ethics, and evaluation
examples/ symbolic planning and decision-frame inputs
hooks/ Rust-first completion hook plus Python fallback
.codex-plugin/ Codex plugin manifest
.claude-plugin/ Claude Code plugin metadata
kimi.plugin.json Kimi plugin manifest
.github/workflows/ repository validation CI
Hikmah ships no credentials, privileged remote service, or external database connection.
- The reference store is local JSONL.
- Hash chaining provides tamper evidence; it does not encrypt content or provide access control.
sensitivepersistence is refused by default.- The append-only reference ledger is not a complete right-to-delete implementation.
- A production system handling sensitive data needs an encrypted, access-controlled, deletion-capable storage adapter and an explicit retention policy.
- The narrow Truth Gate does not fact-check arbitrary model output.
Review Security before enabling executable hooks or adapting the memory layer for sensitive environments.
Hikmah is decision-support infrastructure. It is not a substitute for current qualified medical, legal, financial, security, religious, or other professional judgment.
- Architecture
- Cognitive Kernel
- Co-Model Architecture
- Memory
- Evaluation Contract
- Research Notes
- Ethics
- Compatibility
- Changelog
- Contributing
- Security
MIT. See LICENSE.