How Mneme works, precisely
Technical deep dives into the pipeline — retrieval mechanics, scoring, decision memory, and the deliberate architectural choices behind deterministic governance. No embeddings, no ML, no approximations in the enforcement path.
An autonomous agent system is not one layer. Models produce candidate output. Harnesses coordinate execution, retries, and tool use. Execution systems maintain long-running loops, sessions, and memory. Governance infrastructure defines and enforces the architectural constraints the output must satisfy. Verification confirms the resulting system still passes its objective checks. Mneme operates in the governance infrastructure layer as the decision and control layer — logically separate from harnesses and execution systems, but integrated directly into their execution boundaries: before tool calls, file writes, model calls, and commits.
Harnesses coordinate execution; governance defines constraints; verification enforces invariants. None of those layers can do the others' jobs. The argument in full: Harness Engineering Still Needs Governance. The concept page that anchors this stack: Governance Infrastructure.
Layer 4 above answers what has been decided and whether it applies. A separate class of tooling — code review, PR prioritization, change-management systems — answers what changed and whether it is safe to ship. Both are necessary. Neither can do the other's job.
Decision plane · Mneme
- What architecture has been decided?
- Which decision applies here?
- Is it authoritative?
- What constraint follows?
- Can this action proceed?
Change plane · review & change management
- What changed?
- Is the implementation correct?
- What risk does it introduce?
- Which PR deserves attention?
- Is it ready to merge?
The two planes compose over MCP rather than merge: a change-management system could ask Mneme's Decision Index which decisions govern a proposed change, while keeping full ownership of review, risk scoring, and merge readiness. MCP is the interoperability boundary that would make that composition possible; the 0.9.0 server is a local stdio process. See the MCP integration overview for how that boundary is defined.
Loads project_memory.json → scores decisions by field weights → injects top-K=3 into prompt → checks model output → emits PASS / FAIL / WEAK_RETRIEVAL. Same query, same corpus, same result every time.
The full DecisionRetriever walkthrough — tokenization, field weights (title×3.0, tags×2.5, constraint×1.5, content×1.0), tag boosting, top-K=3 selection, tie-break determinism, and why there are no embeddings. Includes Layer 1 vs. Layer 2 metric distinctions and WEAK_RETRIEVAL semantics.
→The three-tier model — documentation (prose, wikis, ADR bodies), prompt memory (CLAUDE.md, rules files, RAG injection), and decision memory (typed schema with scope, status, precedence, constraint fields). Why documentation retrieval ends in suggestion. Why decision memory enables enforcement.
→Architectural Governance, Deterministic Enforcement, Architectural Compiler, Decision Continuity, and seven more. Systems-level explanations of why each concept exists in AI-native software delivery.
Read the source
DecisionRetriever, MemoryStore, ContextBuilder, and the benchmark harness are all open source under MIT.