DSH Plugins Marketplace

DSH Plugins

Plugins

/

mnemon-memory-agent

G

mnemon-memory-agent

Manifest valid★ 4

Long-term memory for AI agents on Jev. Keep raw records, judge them with a fast System 1 model and answer from under 4k tokens of context.

UI (client)hasBundlePatch

Mnemon

Raw Records, Fast Judgments, Slow Thoughts

arXiv 2609.36059 Hugging Face Paper

Paper · Results · Reproduce · Run · 中文

Mnemon is a long-term memory agent for LLM assistants. It keeps conversations as raw, dated records and does its work when a question arrives, dividing that work the way dual-process accounts divide thinking. A fast decision model (System 1) makes many small yes/no judgments about the records the searches return. An LLM (System 2) writes a few search queries, names what the reply needs and composes the answer. A background pass consolidates each record once into an index that points back to the records, so that questions about a whole conversation reach evidence their own searches miss. The results of this research will be brought step by step into the official projects mnemon and dsh-mnemon (see From research to product).

Accuracy against context per question on LoCoMo and LongMemEval-S: Mnemon and the 14 systems re-evaluated by OmniMemEval

Accuracy against the context sent to the answering model per question. Mnemon (star) and the 14 systems re-evaluated by OmniMemEval all use gpt-4.1-mini to answer. Dashed lines join points of equal effective cost index; up and to the left is better.

Highlights

  • Most accurate on LoCoMo, from under 4k tokens of context. Compared with the 14 systems OmniMemEval re-evaluated, all with gpt-4.1-mini answering, Mnemon scores 91.7% on LoCoMo (first of 15 systems) and 83.8% on LongMemEval-S (second of 13). It sends the answering model about 3.8k tokens per question. It is the only system above 80% on both benchmarks below 4k tokens.
  • Lowest effective cost index on LoCoMo: 0.259, against 0.337 for the next system.
  • On par with the best published results on LongMemEval-S. With DeepSeek-V4.1-Flash as System 2, Mnemon reaches 94.4% on LongMemEval-S and 92.2% on LoCoMo (95.3% on revised labels).
  • Bounded cost at 10M tokens. From BEAM-100K to BEAM-10M, with 80 times as many records, the cost per question grows by a factor of 1.11, and the work on a question's critical path stays about the same.
  • System 1 judges better. On the same 14,359 records, Jev separates gold evidence with an AUC of 0.942. DeepSeek reaches 0.900 and gpt-4.1-mini 0.853, at 3–11 times Jev's latency.
  • Raw records beat a write-time graph built by the same model. Under one protocol, Jev-Mem, which uses Jev to organize memory into a graph as turns are written, scores 84.4% on LoCoMo against Mnemon's 91.7% (7.3 points, 95% CI 5.5–9.2).

Approach

Fast judgments, slow thoughts

Most of the read-time work of memory is System 1 work: small, independent yes/no judgments with explicit criteria.

  • Should the reply use this record?
  • Is it no longer current?
  • Does it give the second of the two dates the question needs?

A decision model such as Jev makes dozens of these judgments in a third of a second.

Only a little is System 2 work: writing a few search queries, naming what the reply needs and composing the answer. An LLM does this well but slowly.

Because the judging is fast, Mnemon can afford to read raw records when a question arrives instead of rewriting them in advance. The rules between the two systems read only what System 1 reliably gives: its ranking and its yes/no decision. They call on System 2 only when no judged record satisfies a need.

Mnemon at one user turn: plan (System 2), retrieve, screen and judge (System 1), loop, compose the View; consolidation in the background

Mnemon at one user turn. It runs as a second instance of DeepSeek Harness beside the main agent, which it leaves unchanged, and publishes one View per turn.

Memory without a write-time schema

Systems that extract at write time must decide in advance what counts as a fact, an entity or a preference. Each new kind of data then needs a new extraction schema. Mnemon decides nothing about a record when it is written, so it needs from a store only a search route that returns dated records.

The consolidated index (topic timelines, value histories, standing instructions) sits on top of the records and points back to them. It is an overlay on the records, not their schema. Mnemon reads the records with or without it.

Keeping records raw also lets the memory grow without rewriting what is stored. A record can be searched as soon as it is embedded; organizing runs in the background and can be redone from the records. Longer histories, stronger answering models and new rules apply to everything already stored.

Memory of this kind can be added wherever records can be searched. The paper evaluates conversational memory, the setting with public benchmarks; other stores are untested.

Results

All numbers come from docs/paper/data/results.json, which is computed from the run records in runs/. Our runs are graded by gpt-4.1-mini and DeepSeek-V4.1-Flash; OmniMemEval grades with gpt-4o-mini, a grader difference of 1–2 points.

Under one protocol (gpt-4.1-mini answering)

Accuracy (%), context sent to the answering model per question, and the effective cost index. The index is ECI = (1 − accuracy) + context / full-context tokens: the expected cost of a question when each error is repaired by one full-context answer. Lower is better. Other systems' numbers are OmniMemEval's.

SystemLoCoMocontextECILongMemEval-ScontextECI
Mnemon91.73.8k0.25983.83.8k0.198
MemOS88.835.4k0.36289.24.2k0.147
Cognee83.4832.5k1.67051.810.3k0.580
EverMemOS82.758.6k0.56980.412.4k0.314
Hindsight81.9924.7k1.32272.229.8k0.561
Mem077.6817.4k1.02856.00.9k0.448
Letta77.1214.2k0.88577.6749.4k0.693
MemMachine73.92.6k0.38063.62.8k0.391
mem973.641.6k0.33778.03.8k0.256
Supermemory73.5315.2k0.97066.076.6k0.402
MemoryLake72.495.2k0.516–––
Viking Memory69.336.0k0.58361.072.3k0.411
Zep63.831.9k0.44879.8117.1k1.316
Memori41.348.1k0.96320.82.8k0.818
Backboard.io22.41.2k0.831–––

Only the context sent to the answering model is compared, because it is the one cost every system reports. Mnemon's other costs (planner, Jev, consolidation) are listed below and are not folded into the index.

Against each project's best published result

Each project's best claim, with whatever answering model, grader and protocol it used. The settings differ widely, so this ranks claims, not systems. The paper's Table 3 has all 20 entries and their sources.

ProjectLoCoMoLongMemEval-SAnswering model / grader
Mnemon95.3†94.4DeepSeek-V4.1-Flash (thinking) / DeepSeek
Zep / Graphiti94.790.2gpt-5.4 (medium reasoning) / gpt-5.4
EverMemOS93.0583.0gpt-4.1-mini / three graders averaged
Mem092.594.4gpt-5 / gpt-5
memU92.09–not stated (early version)
Hindsight92.094.6not stated (paper: gemini-3-pro 89.6 / 91.4)
MemMachine91.6993.0gpt-4.1-mini; LongMemEval-S gpt-5-mini / gpt-4o-mini
MemOS88.8389.2gpt-4.1-mini / gpt-4o-mini (OmniMemEval)

† On the revised LoCoMo labels, which drop 44 unusable questions and correct 25 answers; 92.2 on the original labels, which every other entry uses.

Every benchmark and tier, with the full cost

gpt-4.1-mini answering. Score under the gpt-4.1-mini / DeepSeek graders (accuracy; HaluMem: share correct; BEAM: rubric score). The rank is among the systems OmniMemEval re-evaluated. Cost per 1,000 questions covers the answer, the planner and Jev at list prices; consolidation is a one-time cost per history. Jev calls (sequential waves) and searches are the median question's; the last column is the warm latency of one search on the largest history.

BenchmarkScoreRankContextCost / 1k questionsConsolidation / historyJev calls (waves)SearchesSearch
LoCoMo91.7 / 91.41/153.8k$3.27$0.0135 (4)714 ms
LongMemEval‑S83.8 / 85.42/133.8k$3.33$0.0175 (5)714 ms
HaluMem73.3 / 65.88/133.4k$3.60$0.0897 (6)1018 ms
BEAM‑100K64.5 / 60.510/123.8k$4.80$0.0129 (7)1317 ms
BEAM‑10M51.2 / 48.810/123.8k$5.32$1.6210 (7)14248 ms

Nothing on the read path grows with the history except the search index. The planner reads the recent dialogue, Jev screens at most 48 records a round, and the View has fixed budgets. System 1 does the broad reading: per question, Jev reads 35–73k tokens of records, 9–19 times what the answering model reads, at about a tenth of its price per token.

Work per question. Our runs shared one laptop and public model APIs, so we state latency as critical-path work (medians):

  • System 2 plans once, with two calls in parallel, and answers once; a question whose loop asks for a new search (6–31% of them) makes one more call.
  • System 1 makes 5–10 Jev calls in 4–7 sequential waves of about 0.34 s each: 1.4–2.4 s in all.
  • A warm search takes 14–18 ms, mostly to embed the query, and all of a question's reads of the journal 0.1–0.2 s.

These counts stay about the same from BEAM-100K to BEAM-10M. Only the search grows with the history, to 248 ms a search and 3.5 s of reads a question on BEAM-10M's largest history (108,810 records), because it scores every record; inverted and approximate nearest-neighbor indexes would avoid this.

System 1 against LLMs

ROC curves and per-call latency of Jev, DeepSeek and gpt-4.1-mini judging the same records

Jev, DeepSeek and gpt-4.1-mini judged the same proposition about each of the same 14,359 records. Jev separates the gold evidence best and answers two propositions per record in the time an LLM takes for one.

Raw records or a write-time graph, with the same Jev

Jev-Mem, concurrent work, uses Jev at write time: it types each turn and links it into a relation graph, then steers retrieval over that graph. We ran its released code under our protocol: the same 1,540 LoCoMo questions, gpt-4.1-mini answering once per question from the question alone, and both graders. (Its released runner, by default, picks the best of three answers against the gold answer; we did not.)

LoCoMo, gpt-4.1-mini answeringMnemonJev-Mem
accuracy, gpt-4.1-mini grader91.784.4
accuracy, DeepSeek grader91.482.1
multi-hop / temporal (gpt-4.1-mini grader)91.8 / 91.377.7 / 82.9
context per question3.8k2.6k
write-time cost per history$0.013$0.12

The paired difference is 7.3 points (95% CI 5.5–9.2). Given each question's category, Jev-Mem scores 84.0%. The two systems differ in more than where Jev works, so this is not an ablation, but with the decision model held fixed, judging raw records once the question is known was the more accurate. The adapter is docs/paper/scripts/jevmem_locomo.py, and the run records are in runs/jevmem-locomo-20260928.

Components and versions

Mnemon is the research branch codex/jev-replica-practice of dsh-mnemon. It was forked from the dsh-mnemon release v0.5.13 (commit 84d469ff, 2026-09-22) and developed over 225 commits up to this snapshot, 50a6831d. It runs on DeepSeek Harness 0.1.5-rc.1, which it does not modify. Apart from 115 changed lines in dsh-mnemon's existing code, the memory agent consists of new plugins.

ComponentVersionRole in Mnemon
dsh-mnemonv0.5.13 + 225 research commits (50a6831d)memory plugins, the replica and the benchmark harness
DeepSeek Harness (DSH)0.1.5-rc.1agent harness; Mnemon runs as a second instance beside the main agent
Jev (TypeSafe System One)jev-1.13.0, through @typesafe-ai/sdk 0.6.0System 1: screens and judges records and index items
gpt-4.1-mini2025-04-14 snapshot, temperature 0System 2 in the standard setting (planner and answering model); grader, primary in the standard setting
DeepSeek-V4.1-FlashAPI model deepseek-flashSystem 2 in the reasoning setting (answers with thinking, plans without); consolidation, without thinking; grader, primary in the reasoning setting
nomic-embed-textserved locallyembeddings for hybrid search and index items
Node.js / pnpmv25.1.0 / 11runtime and package manager

Each run directory records the commit and the models it ran with (docs/run-commits.json); VERSIONS.md lists every version, including the build tools and the datasets.

Reproduce the numbers

Every number in the paper comes from docs/paper/scripts/collect.py, which reads the run records. No API keys are needed.

python3 tools/restore_runs.py                     # expands runs/ into runs-expanded/ and checks every file
# Put the two public datasets beside them:
#   runs-expanded/benchmarks/locomo10.json              LoCoMo, from github.com/snap-research/locomo
#   runs-expanded/benchmarks/longmemeval_s_cleaned.json LongMemEval-S (cleaned), from the LongMemEval release
MNEMON_RUNS=$PWD/runs-expanded python3 docs/paper/scripts/collect.py   # rewrites docs/paper/data/results.json
git diff --stat docs/paper/data/results.json      # no change: the records reproduce the committed numbers
TECTONIC=tectonic PYTHON=python3 bash docs/paper/build.sh   # tables, figures (matplotlib) and main.pdf
python3 docs/paper/scripts/arxiv.py --out out/arxiv         # arXiv source package (pdfLaTeX, TeX Live 2025)

HaluMem's run records are not included (see DATA-LICENSES.md). Without them, collect.py keeps the committed HaluMem entries, so those numbers cannot be recomputed here; every other number can.

beam_evidence.py additionally needs pyarrow and BEAM's 100K.parquet. work.py and retrieval.ts, which measure the work per question, read the replica traces and journals, which this snapshot does not include; their results are in docs/paper/data/.

Run the system

Requirements: Node.js 25 (the runs used v25.1.0) and pnpm 11.

pnpm install --frozen-lockfile        # installs DSH 0.1.5-rc.1 exactly as the runs did
pnpm run build && pnpm run build:plugins
pnpm -r --filter 'dsh-mnemon-*' test --passWithNoTests

Live runs call external models. Provide the keys in the environment or in a local .env (git-ignored), never in the repository:

VariableFor
TYPESAFE_API_KEYJev (System 1)
DEEPSEEK_API_KEYDeepSeek answering, planning, notes and consolidation
OPENAI_API_KEYgpt-4.1-mini runs and judging
OPENAI_BASE_URLoptional; defaults to the OpenAI API

A two-question smoke run of the final configuration's core:

node --env-file=.env --experimental-transform-types scripts/bench/run.ts --dataset locomo \
  --file runs-expanded/benchmarks/locomo10.json --cases conv-26 --questions 2 \
  --arms replica-cuefill:raw-records --no-thinking --concurrency 1 --out out/smoke

The full system adds --simple, hybrid search with --embed-url <local nomic-embed-text server>, and --consolidate deepseek-flash. The header of scripts/bench/run.ts documents every option.

About this snapshot

This repository is a frozen, history-free snapshot of the research branch at 50a6831d (2026-09-29). Every file can be traced to its source through PROVENANCE.json.

Repository layout
PathWhat
docs/paper/The paper: LaTeX sources, bibliography, data, generated tables and figures, main.pdf, and its scripts
docs/reports/The pre-registrations and study reports behind the paper
docs/pr-assets/Only the report assets those reports link to
docs/plans/The two design notes the reports link to
runs/The run records collect.py reads, gzip-compressed, with MANIFEST.json
docs/run-commits.jsonFor each run directory, the code commit and models its runs recorded
src/The dsh-mnemon kernel at the snapshot commit
plugins/The 18 plugins the scripts and the kernel build need
scripts/The evaluation harness (scripts/bench), the replica launcher and their libraries
assets/The figures in this README, rendered from the paper; announcement/ holds the release post's images, drawn from the paper and its data
tools/How this snapshot is made and checked
VERSIONS.mdDSH, dsh-mnemon, model and dataset versions
DATA-LICENSES.mdThe license of each benchmark's text in the run records, and what is left out
PROVENANCE.jsonSource commit, and each file's git blob id and SHA-256
How the snapshot is made and checked
ToolDoes
tools/snapshot.pyCopies the paper's part of the research branch byte for byte: the kernel, the plugins the scripts import (found by following imports), the scripts, the paper, the reports and what they link to; writes PROVENANCE.json
tools/scrub.pyReplaces machine-specific absolute paths with <repo>, <tmp> or ~ and records every rewritten file under modified in PROVENANCE.json
tools/collect_runs.pyRuns collect.py on the full records, keeps exactly the files it opens except HaluMem's, and checks the recomputed results.json against the committed one
tools/restore_runs.pyExpands runs/ and verifies each file against runs/MANIFEST.json
tools/run_commits.pyWrites docs/run-commits.json
tools/audit.pyFails on credentials, local paths, non-loopback endpoints, files over 50 MB, HaluMem run records, and the words of a local, git-ignored .audit-deny; --history checks every commit

What has been checked:

  • results.json recomputed from runs/ and the two datasets alone equals the committed one, byte for byte. Its HaluMem entries are kept as committed, since HaluMem's run records are not included.
  • The frozen-lockfile install, the kernel and plugin builds, and the plugins' 223 tests pass.
  • The smoke run above answered both questions.
  • tools/audit.py is clean, over the working tree and over every commit (--history).

The install, builds, tests and smoke run were checked on the snapshot of e5c7954a. The later refreshes, to f97c5679, 3bbf7835, c461581c, 7649d8ff, 65145f69, aea19abc, 44fb4e71, 8843a5ea and 50a6831d, changed only the paper and its run records: its text, bibliography, table and figure scripts, generated tables, figures and PDF, the two scripts that measure the work per question with their data, collect.py, which keeps the HaluMem entries when their records are absent and compares Jev-Mem with the final version, arxiv.py, which packages the sources for arXiv, and the Jev-Mem adapter with its run records. The system's code is unchanged. The recomputation and the audit were repeated at 50a6831d.

What the snapshot does and does not claim:

  • Files listed in PROVENANCE.json equal the source commit, except the ones under modified, which differ only in path placeholders.
  • The paper's run directories ran at 25 different commits (docs/run-commits.json). The snapshot is the branch's latest commit, not each of those commits.
  • Three plugins (Memory Spaces, Runtime, three-tier) are here only because the kernel's client bundles their pages. They ship without their tests, which exercise product components outside this snapshot.
  • Git history is not included.

From research to product

The results of this research will be brought step by step into the two official projects, mnemon and dsh-mnemon, to give their users the best product experience. Next, DSH will serve as a micro agent kernel, with Mnemon running on it as a memory agent, the form the replica in this study already takes.

This repository itself stays a frozen research snapshot:

  • It is not the mnemon CLI and not Mnemon Agency, and the evaluated system does not use the mnemon binary.
  • It is not a release of the dsh-mnemon package.

Citation

@misc{wang2026mnemon,
  title         = {Mnemon: Raw Records, Fast Judgments, Slow Thoughts},
  author        = {Wang, Guangren},
  year          = {2026},
  eprint        = {2609.36059},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.36059}
}

Licenses and data

The code is dsh-mnemon's, under the MIT license in LICENSE. That license does not cover the benchmark text in the run records and report assets, which stays under each benchmark's license: LoCoMo CC BY-NC 4.0, LongMemEval MIT, BEAM CC BY-SA 4.0. HaluMem's run records are not included: its license (CC BY-NC-ND 4.0) does not allow sharing adapted material. DATA-LICENSES.md has the details and the attributions.

The paper itself (the manuscript in docs/paper: its text, figures, tables and PDF) is © 2026 Guangren Wang, all rights reserved, and not covered by the MIT license; the scripts in docs/paper/scripts are.

Versions

Latest versionPublishedSize
0.5.22——
0.5.23——
0.5.24——
0.4.1——
0.4.2——
0.4.3——
0.4.4——
0.4.5——
0.4.6——
0.4.7——
0.5.0-rc.1——
0.5.0——
0.5.1——
0.5.2——
0.5.3——
0.5.4——
0.5.5——
0.5.6——
0.5.7——
0.5.8——

Comments

Loading…

Similar plugins

mnemon

by mnemon-dev

LLM-supervised persistent memory for AI agents — graph-based recall, cross-session knowledge, single binary. Works with DeepSeek Harness, Claude Code, OpenClaw, and any agent runtime.

Memory & ContextManifest valid

★ 611

Apache-2.0

Go

Oct 3, 2026

dsh plugin --profile agent add @mnemon-dev/dsh-mnemon

by djasdh

Low-footprint memory backend for AI agents — single binary, ~50MB RAM, verify-augmented accuracy

Manifest valid

★ 3

MIT

Go

Aug 21, 2026

dsh plugin --profile web add @djasdh/interest-memory

by MemTensor

Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings and DeepSeek Harness support.

Memory & ContextManifest valid

★ 11.7k

↓ 2.4k/wk

Apache-2.0

TypeScript

Sep 29, 2026

dsh plugin --profile web add @memtensor/memos-local-plugin

by zilliztech

A persistent, unified memory layer for all your AI agents (e.g. Claude Code, Codex, DSH), backed by Markdown and Milvus.

Memory & ContextManifest valid

★ 2.7k

↓ 480/wk

MIT

Python

Sep 24, 2026

dsh plugin --profile agent add @zilliz/memsearch-dsh

by jonah791

Agent-driven long-term memory for DeepSeek Harness: scoped memory (global + per-workspace), layered entries (fact/knowle

Manifest valid

★ 3

↓ 97/wk

MIT

TypeScript

Sep 7, 2026

dsh plugin --profile web add dsh-agent-memory

by menotbobbybrown

Persistent long-term memory for DeepSeek Harness: a JSON store of typed entries (fact, preference, entity, rule, episodic) exposed as the memory_remember and memory_recall agent tools, with recall ran

Memory & ContextManifest valid

★ 0

MIT

TypeScript

Sep 8, 2026

dsh plugin --profile web add @modelnorth/dsh-plugin-memory