DSH Plugins Marketplace

DSH Plugins

Plugins

/

open-dream-rsi

p

open-dream-rsi

Discovered★ 10

Recursive Self-Improvement through Evolving Worlds (Dream-RSI) for LLM agents

Open Dream-RSI

📄 Preprint: Open Dream-RSI: An Open-Source Library for Recursive Self-Improvement Around a Frozen LLM, with Replay-Gated Learned Policies and a Curated Knowledge Base — P. Orwat, 2026. PDF · LaTeX source

License: MIT Python Version Tests

What is this?

Every AI agent pays the same tax over and over: it fails a task, works out what went wrong, and then throws that understanding away when the session ends. Open Dream-RSI keeps what was learned — without touching the model.

The LLM stays frozen. What improves is everything around it:

  • the exploration policy (even re-written as Python code by the LLM itself, gated by replay before it is trusted),
  • a library of verified solutions (recipes) reused as warm starts,
  • curated lessons distilled from failures — every one of them promoted only after it measurably helps on replay,
  • and the world model the policy is scored against: a Discovery Tree of every attempt, enriched with replay-safe error-class facts by the Sentinel engine.

Between runs the system "dreams": it replays thousands of strategies over the recorded history in-process, at zero API cost, and keeps only what beats the incumbent. An open implementation of Recursive Self-Improvement through Evolving Worlds (Dream-RSI, Zheng et al., arXiv:2609.14858).

Zero runtime dependencies — Python 3.10+ stdlib only, ~10.2k lines of library code, ~5.1k lines of tests, no build step.

Not another memory plugin. Experience-memory systems record what the agent learned and reuse it as-is. ODR is an evolving worlds loop: every attempt is recorded into a Discovery Tree, and between runs the system dreams — replaying thousands of strategy variants against that recorded world. Nothing is promoted on the model's word alone: a candidate policy, recipe or lesson replaces the incumbent only after surviving counterfactual replay.

Evaluation policy: this project makes performance claims on exactly one benchmark — goose × TravelPlanner, public and scored by the benchmark's own evaluators (below). The scripted suites (bench, bench-policy) are deterministic self-checks — behavioural regression contracts for the machinery — and carry no model-quality claim.


How it works

                        ┌───────────────┐
                        │  Frozen LLM   │
                        └──────┬────────┘
                               │ proposals
              online loop      ▼      offline dream
        ┌──────────── DiscoveryTree ──────────── ReplaySimulator ──────────┐
        │                     │                          │                 │
        │              node.sentinel              counterfactual          │
        │              (Sentinel facts,           rollout scoring         │
        │               replay-safe)                    │                 │
        │                     │                         ▼                 │
        │                     └────→ Policy improvement ────→ next cycle  │
        │                                                                   │
   SENTINEL MODE (tool-execution layer)                                     │
        │          │             │                                          │
   error-class   nudge       permute-gate ──→ block tool calls             │
   ledger        (once/session)              past 80% of budget ───────────┘
  1. Online execution — the agent attempts each task with the real LLM and a sandboxed verifier, appending every attempt to a per-task Discovery Tree (code, test feedback, score, one-line plan, and — when a Sentinel engine is wired — structured error-class facts and an explicit termination reason).
  2. Offline dreaming — instead of paying for real-world rollouts, the agent "dreams" over the recorded history: thousands of strategy variants are scored in an in-process replay simulator at zero external-call cost. The best parameters, programs and recipes are persisted and steer the next cycle.

Beyond parameter-level dreaming, the loop closes section 3 of the paper ("dreaming with code"): each cycle the LLM may rewrite the exploration policy itself as a small Python program, choose_action(frontier, step). Candidates are statically validated (AST gate), executed only in a hardened subprocess sandbox, and scored by counterfactual replay rollout on the recorded tree. A candidate replaces the incumbent only on evidence — and the incumbent is re-scored on the current tree, not trusted at its stored score. A crashing or cheating policy can never break the loop: expansion falls back to the greedy baseline.


Quick start

git clone https://github.com/patrykorwat/open-dream-rsi.git
cd open-dream-rsi
pip install -e .

Point it at a model — or don't. Any OpenAI-compatible endpoint works; keys are read only from the environment. The model name is auto-detected from the endpoint (GET /v1/models), so pointing at a local vLLM/Ollama needs only the base URL:

export OPENAI_BASE_URL=http://127.0.0.1:8000/v1   # vLLM / Ollama / any gateway
export OPENAI_API_KEY=***                         # any token for local servers
# no ODR_LLM_MODEL needed: the first model the endpoint serves is used.
# Pin ODR_LLM_MODEL only to select among several served models.

Queue a task in tasks.json:

[{
  "task_id": "add1",
  "category": "math",
  "prompt": "Implement add(a, b) returning the sum.",
  "tests": [{"call": "add(2, 3)", "expected": 5}],
  "max_attempts": 3
}]

Run the supervisor — once, as a daemon, or as the recommended automated dream cycle (one command that imports session evidence, runs the promotion gate when the evidence justifies it, and publishes earned lessons as skills):

python -m open_dream_rsi --memory .dream_rsi dream \
    --sessions ~/.hermes/state.db --skills-out ~/.hermes/skills/odr-curated
python -m open_dream_rsi loop --tasks tasks.json --interval 300   # daemon mode
python -m open_dream_rsi loop --tasks tasks.json --once           # single cycle
python -m open_dream_rsi status                                   # what it learned

No key? Everything deterministic still runs: bench, bench-policy, gate-replay and the dashboard all use a scripted mock client. A full in-process walkthrough with a mock OpenAI server: examples/live_loop_demo.py.

Don't want the CLI at all? The same loop runs as an MCP server inside your existing coding agent — that is the fastest way in:

python3 -m open_dream_rsi mcp --tasks ./tasks.json --memory ./.dream_rsi

…see the next section for the exact block for your agent.


Plug in your favourite agent

The loop speaks MCP (Model Context Protocol). One config block and your everyday agent can queue tasks for the dreamer, pull back verified recipes and consult curated lessons — zero extra LLM setup: the dreamer resolves its own brain (env → local goose config → localhost vLLM), and asks the endpoint itself which model to use.

ToolWhat it does
odr_statuswhat the loop has learned (policies, recipes, events)
odr_recipesbest verified solution for a category — use as warm start
odr_lessonscurated failure lessons for a category / search query
odr_add_taskqueue a task (prompt + tests, or prompt + criteria) — test-less tasks are completion-judged
odr_run_oncerun one improvement cycle now (bounded API budget)
odr_dreamthe automated cycle: import host session evidence, distill curated lessons, run the promotion gate, publish skills

Tools are registered prefixed per client (e.g. Hermes mcp_open_dream_rsi_odr_run_once). No model calls tools on its own initiative — trigger them via a recipe's instructions:, a hook, or an explicit ask ("check odr_recipes before you try again"). A tip that makes the loop proactive: add to your project AGENTS.md:

## Self-improvement loop (MCP: open-dream-rsi)
- Before implementing a self-contained Python utility, call `odr_recipes`
  and `odr_lessons` for its category; reuse a verified recipe verbatim.
- When you finish a hard, testable function, queue it with `odr_add_task`
  (task_id = function name, tests = your test cases) so the loop can
  dream over it offline.

OpenCode

Create (or merge into) opencode.json in the project root where you work:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "open-dream-rsi": {
      "type": "local",
      "command": ["python3", "-m", "open_dream_rsi", "mcp",
                  "--tasks", "./tasks.json", "--memory", "./.dream_rsi"],
      "enabled": true
    }
  }
}

Start OpenCode and type /mcps — you should see open-dream-rsi: connected.

Goose

Recommended — the script wires everything (diagnoses the install, resolves your goose provider, optionally starts the model-borrowing proxy, and writes the extension block into ~/.config/goose/config.yaml itself):

./scripts/odr_goose_setup.sh            # diagnose + configure goose + start proxy
./scripts/odr_goose_setup.sh --check    # diagnose only, change nothing
./scripts/odr_goose_setup.sh --no-proxy # skip the proxy (use provider='goose')

Or add the extension manually to ~/.config/goose/config.yaml:

extensions:
  open-dream-rsi:
    enabled: true
    type: stdio
    name: open-dream-rsi
    description: "Dream-RSI self-improvement loop. Use odr_add_task to queue a
      Python task with tests, odr_run_once to run an improvement cycle,
      odr_recipes/odr_lessons to reuse verified solutions and failure lessons,
      odr_status to inspect what the loop has learned."
    cmd: python3
    args: ["-m", "open_dream_rsi", "mcp",
           "--tasks", "/ABSOLUTE/PATH/projects/myproject/tasks.json",
           "--memory", "/ABSOLUTE/PATH/projects/myproject/.dream_rsi"]
    timeout: 300

Goose spawns extensions without a shell and with a scrubbed environment: use absolute paths (no ~, no relative ./) and pass any env the server needs via envs: — it will not inherit your export'ed OPENAI_*. After any config edit: fully quit goose (the desktop app caches config at startup), relaunch, and activate open-dream-rsi in the session's extensions picker. One-shot alternative (no config edit): goose session --with-extension "python3 -m open_dream_rsi mcp --tasks ./tasks.json --memory ./.dream_rsi".

Two first-class ways to run the dreamer on exactly the model goose uses (no second API key): provider: "goose" reads ~/.config/goose/config.yaml itself (CLI and desktop dialects, custom_providers/*.json, secrets.yaml, macOS keychain), and python3 -m open_dream_rsi proxy --port 8799 exposes goose's brain as a plain OpenAI-compatible endpoint — details in docs/integrations.md.

Hermes Agent

Add under mcp_servers in ~/.hermes/config.yaml (or via the dashboard's MCP catalog):

mcp_servers:
  open-dream-rsi:
    command: "python3"
    args: ["-m", "open_dream_rsi", "mcp",
           "--tasks", "/ABS/PATH/open-dream-rsi/tasks.json",
           "--memory", "/ABS/PATH/open-dream-rsi/.dream_rsi"]
    timeout: 300
    env:
      OPENAI_BASE_URL: "http://YOUR-VLLM-HOST:8000/v1"
      OPENAI_API_KEY: "***"        # any token for local servers

Restart Hermes; tools appear as mcp_open_dream_rsi_* in every platform toolset. Hermes is also the one host with MCP sampling support, if you ever want the dreamer to borrow the host's model through the protocol (today ODR resolves the endpoint itself instead).

Hermes in Docker: mount the repo read-only and export PYTHONPATH (no pip layer in the image), then add the server block with an explicit env: (the container has no goose config):

ENV PYTHONPATH=/opt/open-dream-rsi
# compose.yaml -> services.hermes
volumes:
  - /home/USER/git/open-dream-rsi:/opt/open-dream-rsi:ro
  - ./data:/opt/data

Point --tasks/--memory at paths under the mounted data volume (e.g. /opt/data/dream-rsi/...) so memory survives rebuilds. Updating = git pull on the host + restart the container — no image rebuild.

Codex CLI

codex mcp add open-dream-rsi -- python3 -m open_dream_rsi mcp \
  --tasks /ABS/PATH/tasks.json --memory /ABS/PATH/.dream_rsi

Verify with codex mcp list (or [mcp_servers.open-dream-rsi] with command/args in ~/.codex/config.toml).

Claude Code

claude mcp add open-dream-rsi -- python3 -m open_dream_rsi mcp \
  --tasks /ABS/PATH/tasks.json --memory /ABS/PATH/.dream_rsi

Check with /mcp inside a session.

Claude Cowork / claude.ai (remote only)

Cowork connects to remote MCP URLs brokered from Anthropic's cloud. Serve HTTP, tunnel, add the URL as a custom connector:

python3 -m open_dream_rsi mcp --http --host 0.0.0.0 --port 8800 \
  --tasks /ABS/PATH/tasks.json --memory /ABS/PATH/.dream_rsi
# tunnel 8800 (tailscale funnel / cloudflared) -> Settings -> Connectors -> Add custom

The endpoint implements the stateless Streamable-HTTP profile (POST /mcp, one JSON-RPC message per request; GET /health). Treat the URL as a credential: whoever reaches it can queue tasks and spend your LLM budget.

DeepSeek Harness

DeepSeek Harness (dsh) runs external MCP servers through its first-party bridge (@deepseek-ai/dsh-mcp-client). The bundle is published on npm:

dsh plugin --profile web add dsh-open-dream-rsi

The tools appear as mcp__open_dream_rsi__<tool> (odr_status, odr_recipes, odr_lessons, odr_add_task, odr_run_once, odr_dream). Bundle defaults point --tasks/--memory at tasks.json/.dream_rsi beside the directory dsh launches from; override the mcp-open-dream-rsi row from your profile's cordis.patch.yml (same id). DSH scrubs the spawned child's environment (names matching KEY|PASSWORD|SECRET|TOKEN and all DSH_*), so the bundle re-passes OPENAI_BASE_URL / OPENAI_API_KEY / ODR_LLM_* / PYTHONPATH from the harness process's env — without them the dreamer falls back to http://127.0.0.1:8000/v1 with model auto-detection (fine for a local vLLM/Ollama, needs no key). toolCallTimeoutMs: 1800000 replaces the 60 s default because odr_run_once/odr_dream cycles run for many minutes.

Prefer a hand-rolled overlay? dsh --patch /ABS/PATH/TO/open-dream-rsi/plugins/deepseek_harness/odr.cordis.yml web inserts the same row with explicit absolute paths (a remote Streamable-HTTP variant, odr-http.cordis.yml, points at the mcp --http endpoint — the URL is a credential, same caveat as Cowork below). Full walkthrough: plugins/deepseek_harness/README.md.

Zero configuration (all of them)

odr_run_once resolves the brain automatically, in this order:

  1. OPENAI_BASE_URL / OPENAI_API_KEY in the server's environment,
  2. your local goose config (~/.config/goose: active_provider + providers: block + custom_providers/*.json + secrets.yaml + macOS keychain) — read from disk, so it survives the env scrubbing that Hermes/Codex/Claude Code apply to spawned servers,
  3. http://127.0.0.1:8000/v1 (vLLM/Ollama default).

The served model is auto-detected (GET /v1/models) unless pinned. Self-hosted endpoints run with hidden reasoning disabled by default (chat_template_kwargs.enable_thinking=false — a thinking-on Qwen otherwise burns the whole completion budget); endpoints that reject the flag get a clean automatic retry without it. An empty task queue is a benign no-op with a hint, never an error.

Full per-host instructions (including the launchd keep-alive for the goose proxy and troubleshooting): docs/integrations.md.


Providers

The agent talks to any OpenAI-compatible POST {base_url}/chat/completions server. API keys are read only from environment variables.

ProviderPresetBase URLKey env var
OpenAIopenaihttps://api.openai.com/v1OPENAI_API_KEY
Cursor Models APIcursorhttps://api2.cursor.shCURSOR_API_KEY
Local (vLLM / Ollama / LM Studio)localhttp://127.0.0.1:8000/v1OPENAI_API_KEY
Borrowed from goose configgooseresolved from ~/.config/goose—

Precedence is always: explicit overrides → environment → endpoint model discovery (first id from GET /v1/models) → preset defaults. A redeployed local server is picked up automatically — and a pinned model the endpoint no longer serves self-heals: the client re-discovers once and retries instead of failing the whole run. Self-hosted reasoning endpoints get enable_thinking: false automatically; an HTTP 400 rejection of that flag triggers a self-heal retry.

Library-level use (no CLI):

from open_dream_rsi import (DiscoveryTree, DreamEngine, ReplaySimulator,
                            DreamAgent, LLMConfig, OpenAICompatibleClient)

client = OpenAICompatibleClient(LLMConfig.from_preset("local"))
tree = DiscoveryTree()
dreamer = DreamEngine(simulator=ReplaySimulator(tree))
dreamer.run_offline_optimization(iterations=100)
agent = DreamAgent(dreamer=dreamer, client=client)

What persists between runs (--memory dir)

ArtifactFileEffect
Dreamed policy parameterspolicies.jsonnext cycle starts from the last one's optimum
Promoted policy programspolicy_codes.jsonLLM-written choose_action steers expansion
Recipe libraryrecipes.jsonverified solutions replayed as warm starts
Knowledge baselessons.jsoncurated, replay-gated lessons (declarative text)
Discovery treestrees/offline dreaming always has history; nodes optionally carry replay-safe Sentinel facts
Audit logevents.jsonlappend-only record of every decision

Switches worth knowing: --no-policy-code, --no-knowledge, --no-thoughts, --no-judge ablate the four optional machinery layers; --budget N caps API calls per cycle (dreaming stays free).


Automated dreaming: when does the loop actually run?

The seam is dream_once() (CLI odr dream, MCP odr_dream, or a host hook). You do not set a schedule — the module decides the tier per call, so firing it after every task is safe and costs nothing extra when there is nothing new to learn:

TierTriggered whenWhat runs
maintenancequiet call (no new evidence, recent last dream)ingest scan (cheap), ALM stale/supersede/archive/GC, skills republish — zero LLM
fullmin_new_evidence (default 3) new failure-shaped sessions, OR max_age_seconds (default 6 h) since the last full dream+ online attempts, dreaming, paired-replay promotion gate (budget, default 20 API calls)

A busy day dreams a few times, a quiet day once, an idle day never; concurrent calls collapse on a single-instance lock. Probes are queued automatically: every failure-shaped session appends its own request to tasks.json as a probe:<session> task (max_attempts=1, appended after manual tasks), so the promotion gate never depends on hand-written task files — the evidence you already produced is the counterparty.

For a mass install the whole thing is two artifacts, no host glue:

python3 -m pip install git+https://github.com/patrykorwat/open-dream-rsi  # library + CLI + MCP server
cp -r plugins/hermes_odr_trigger $HERMES_HOME/plugins/odr-trigger
hermes plugins enable odr-trigger                # restart to load the hook
# optional overrides:
hermes config set plugins.entries.odr-trigger.settings.provider local
hermes config set plugins.entries.odr-trigger.settings.base_url http://.../v1

Benchmark: goose × TravelPlanner (public, externally scored)

This is the only benchmark in this project — everything else below the heading line is machinery self-check. A real agent (goose v1.53) plans real itineraries against the official TravelPlanner offline database (osunlp/TravelPlanner) over a stdlib MCP sandbox, scored by the benchmark's own commonsense and hard-constraint evaluators (nothing in the scoring path is ours), on the model Qwen3.8-Flash-Next served by vLLM. Same endpoint and model across every study in this repo and the paper.

Frozen protocol: iterations only on train.csv (45 tasks); one preregistered eval pass on official validation rows 102–181 (79 tasks, 60 hard), configuration unchanged from the dev split; official test split never touched. Arms interleaved task-by-task per pass (paired, McNemar-ready). The gate arm is choice removal, not persuasion: past 36 of 45 tool calls the sandbox refuses to serve (all text channels off).

splitarmdelivered planscommonsense passhard passfinal pass
train (45)cold15/45 (33%)436.7%
train (45)gate41/45 (91%)18920.0%
eval (79)cold18/79 (23%)1267.6%
eval (79)gate74/79 (94%)401721.5%

McNemar on paired delivery: eval 56↔0 discordant, p=1.4e-17 (train 28↔2, p=4.3e-7).

goose × TravelPlanner: delivery + failure anatomy

Why it works: every single undelivered cold episode (91/91 across both splits) died pinned at the harness call cap with the sandbox data already in hand — the binding failure under budget pressure is termination, not error recovery, and three text-in-context channels measured earlier (pull tool, Stop-hook note, budget nudge) converted zero of those episodes. Removing the option to keep exploring is the only lever that moved delivery. Honest caveats: the choice of gate over persuasion was made knowing results on the earlier 101-task validation pool (selection era, archived — the eval pass is a preregistered confirmation, not a first exposure); the gate's 5 residual eval failures are episodes that called through the refusals to the cap anyway. Everything behind these numbers is in this directory: benchmarks/travelplanner/ — the sandbox MCP server, the episode runner, the official-evaluator wrapper, the exact split CSVs and the raw per-episode records. python3 score.py results/submissions/… reproduces the table above to the task from the committed records; the directory README documents prerequisites and the one-command rerun. The figure re-renders from fixtures/tp_v3_summary.json via paper/make_figures.py. Paper §Evaluation.


Machinery self-checks (deterministic, key-free)

Not benchmarks — behavioural regression contracts: a scripted solver makes the model arm a constant, so every delta is attributable to the library's gates, sandboxes, promotion and retrieval logic. They exist so a broken mechanism fails CI instead of silently mutating a published number, and they re-run from a fresh clone in seconds.

bench — the dreaming loop against a cold baseline

armsolvesAPI callscalls / task·cycledream its
cold_baseline (fresh memory each cycle)501503.00
dream_rsi_loop (policies + recipes + dreaming)50691.383000

54% fewer API calls at equal solve quality (re-measured on this commit: identical to the published table).

python -m open_dream_rsi bench --cycles 10 --markdown

bench-policy — decoy traps (exploration contracts)

Solve-rate self-checks saturate on easy suites and cannot tell good exploration from luck, so bench-policy hides each fix behind a low-scoring branch (passes 1/3 tests) past a plausible decoy (passes 2/3 forever). A score-greedy loop locks onto the decoy and starves; escaping requires structural exploration. Five arms, one scripted solver, one budget:

python -m open_dream_rsi bench-policy --cycles 8 --format md
armsolvessolve rateAPI callscalls / solvepolicy callscurator calls
greedy0/1200%480∞00
epsilon_greedy29/12024%40513.9700
evolved_policy92/12077%2522.74270
knowledge_curator88/12073%3323.77034
thought_guided120/120100%1641.3700
  • evolved_policy — the loop asks the LLM for a policy program per category, validates and replay-scores it, and promotes only on evidence. 100% solve rate by cycle 5 while spending fewer total calls than ε wasted on luck.
  • knowledge_curator — ε-greedy with zero policy calls: every escape above the ε baseline came from remembered knowledge (73%, at parity with the gated policies at ~80% of ε's calls).
  • thought_guided (library default) — each attempt records its one-line PLAN:; expansion leaves a branch as soon as its idea repeats itself. With ε forced to 0 and no policy/curator calls it escapes every trap from cycle 1; the ledger-only ablation (same text in prompts, pick disabled) collapses back to the baselines — the win is the expansion rule, not prompt length.

Decoy-trap suite: per-cycle solve rate, five arms

gate-replay — would a lesson set have been promoted?

Applies the shipped promotion rule (lesson_gate_verdict) to per-task outcome pairs from any recorded replay harness — no model, no GPU:

python -m open_dream_rsi gate-replay \
  --baseline cold=/tmp/arms.json?label=cold \
  --compare lessons=/tmp/arms.json?label=warm \
  --key task_id --format md

What the loop learned about memory (honest results)

The knowledge curator distils verifier failures into short, validated, deduplicated lessons — pure text, never executed. Live-model replay, however, proved that fresh lessons are hazardous by default:

  • Structurally perfect lessons (passed every schema/caps gate) dropped solve-rate 18/20 → 6/20 on a greedy decoder: "keep checking"-style advice turns greedy decoding into an over-exploration loop.
  • The same facts rewritten as declarative background (no imperative verbs, explicit stop clause): 18/20 — indistinguishable from cold (McNemar p=1.0). The hazard is imperative mood, not content.
  • Seven further published injection strategies (abstract workflows, exemplars, evidence certificates, reactive-on-error notes, end-of-prompt position) scored zero paired gains against a contemporaneous cold baseline; moving safe text to the end of the prompt was actively harmful (11/20, p=0.008 — recency makes the model act on facts that were inert mid-prompt).

The shipped design encodes all of that: lessons are staging candidates that must pass a paired-replay gate (net ≥ 1, zero solve→fail regressions, explicit stop clause) before any prompt sees them; proposal prompts frame the KB as declarative background; and a per-category headroom verdict (skip curation when recent solve-rate ≥ 0.7 and calls/solve ≤ 5) spends zero curator/gate calls where the cold baseline already wins. Guidance pays a perturbation tax — the loop only pays it where there is measurable headroom to spend it against.


Learning vs lifecycle vs history

The repository keeps four planes strictly apart:

PlaneOwnerQuestion it answers
LearningDream-RSI pipeline (dreamer, replay gate, lesson gate)Does this artifact deserve to be active?
Current artifact stateArtifactLifecycleManager (lifecycle.py)What is true about this artifact now?
Transition historyappend-only artifacts/events.jsonlHow did it get here?
Immutable evidenceDiscovery Trees + the event logsWhy do these artifacts exist, and what world produced them?

The ALM manages the four artifact types (policy_parameters, policy_program, recipe, lesson) through an explicit state machine (CANDIDATE → VALIDATED → ACTIVE → STALE/SUPERSEDED/QUARANTINED → ARCHIVED, terminal REJECTED), with hard rules: activation is only an explicit promotion call carrying the evidence; every state change appends exactly one immutable event before exposing the new state; the materialized view must always equal rebuild_artifact_state(events) (a hand-edited view is rebuilt from the log, never trusted); merges create NEW artifacts with lineage and never mutate sources; deletion never prunes the event log. Discovery Trees and audit events are evidence, not artifacts — the ALM refuses to register them. Inspect the store: odr artifacts [--type lesson --state ACTIVE --history].


Importing real sessions: Hermes and Cursor (curator's second door)

The loop distils lessons from failures it produces itself. sessions.py opens a second legitimate door: transcripts from host agents are evidence you already paid for. Imports feed the same pipeline as loop-collected failures — structural gate, staging-by-default, paired replay promotion. Nothing imported is trusted; activation is still earned. Transcripts are treated as hostile input: opened read-only, credential-scrubbed, parsed defensively — and lessons that never passed the promotion gate never reach a prompt.

From Hermes sessions

Hermes persists every session to <HERMES_HOME>/state.db (SQLite: sessions + messages). Export, then distil:

python -m open_dream_rsi --memory .dream_rsi sessions export \
    --source hermes --db ~/.hermes/state.db --limit 50 \
    --out episodes.jsonl

# inspect what WOULD be distilled — no LLM calls, no writes:
python -m open_dream_rsi --memory .dream_rsi sessions distill \
    episodes.jsonl --dry-run --limit 5

# real run: staging lessons into the KB (per-category = episode cwd name)
python -m open_dream_rsi --memory .dream_rsi sessions distill episodes.jsonl

The next odr loop cycle gates the staging entries exactly like loop-collected ones; odr artifacts --type lesson --history shows the lifecycle events behind every activation.

From Cursor chats

Cursor keeps conversations in globalStorage/state.vscdb (Linux: ~/.config/Cursor/User, macOS: ~/Library/Application Support/Cursor/User, Windows: %APPDATA%\Cursor\User), with workspaceStorage mapping chats to project folders for the category. The reader handles the undocumented layout and the bubble-ordering rules; re-verify after major Cursor upgrades.

python -m open_dream_rsi --memory .dream_rsi sessions export \
    --source cursor --out episodes.jsonl --limit 100
python -m open_dream_rsi --memory .dream_rsi sessions distill episodes.jsonl

Notes: the live DB is write-ahead-logged — immutable=1 may miss the newest chats, so quit Cursor or copy state.vscdb+-wal+-shm together and point at the copy for a complete snapshot. Chat history contains everything ever pasted into it: export redacts credential-shaped substrings by default (--no-redact exists for local processing only) — review episodes.jsonl before sharing it.

Curator lessons as Hermes skills

python -m open_dream_rsi --memory .dream_rsi skills export \
    --out ~/.hermes/skills/odr-curated

Renders only active (promotion-gated) lessons into Hermes-style skills — one directory per category, declarative bullets, never imperative commands (the measured hazard). The export is one-way and regenerated wholesale: edit lessons.json (or distil again), not the skills.

Sharing lessons: git-backed stores

share.py turns curated knowledge into a portable, verifiable store that other instances (or other people) can consume — a directory of JSON files, optionally a git checkout, so GitHub is just storage:

# one-time: point the store at a repo (clone/pull handled for you)
python -m open_dream_rsi --memory .dream_rsi lessons export \
    --store git@github.com:you/dream-rsi-lessons.git

# on the consumer side — someone else's store becomes your staging:
python -m open_dream_rsi --memory .dream_rsi lessons import \
    --store https://github.com/you/dream-rsi-lessons.git

Trust model: export ships only what earned activation (staging and rejected lessons excluded by construction, secrets scrubbed, sha256 manifest); import lands as staging, never active (checksum mismatch aborts the whole import — tampering detected, not trusted; each lesson still needs your paired-replay gate to reach a prompt). A local directory works identically (--store ./store) — git is just the transport.

Full Hermes wiring (what exists today)

PathDirectionWhat it does
odr mcp (stdio) / --httpHermes → ODRHermes calls odr_run_once / odr_dream / odr_status as tools
plugins/hermes_sentinel/Hermes → ODR engineSentinel observes tool results, annotates, gates
plugins/hermes_odr_trigger/Hermes → ODR enginesession-end hook spawns the detached odr dream cycle (no shell/cron)
plugins/deepseek_harness/DeepSeek Harness (dsh) → ODRCordis overlay: the ODR MCP server behind @deepseek-ai/dsh-mcp-client
odr sessions export/distillHermes → curatorsession history becomes staging lessons
odr skills exportcurator → Hermesactive lessons become discovered SKILL.md files
odr dreamany triggerone automated cycle: evidence → gate → ALM → skills

Security model

Candidate code — both task solutions and policy programs — runs in a hardened subprocess (open_dream_rsi/sandbox.py): python -I, scrubbed environment (only PATH — API keys never cross the wall), a timeout that kills the entire process group, and POSIX address-space / CPU / thread-count limits. Policy code additionally passes an AST gate before any spawn: required entry point, no imports, no dunder access, no __builtins__ reference (name plus constructed-key indexing is itself an escape route).

This is defence in depth, not a container. The subprocess does not restrict filesystem or network access, and an AST gate cannot be provably complete against code that is not statically gated. For hostile task sources set ODR_SANDBOX_CMD to a bubblewrap/nsjail wrapper (e.g. bwrap --unshare-all --die-with-parent --ro-bind / /) or run the loop inside a container.


Live dashboard

One command, no dependencies, stdlib http server:

python -m open_dream_rsi dashboard                   # mock LLM, port 8765
python -m open_dream_rsi dashboard --provider cursor # real model
python -m open_dream_rsi dashboard --host 0.0.0.0    # LAN access

Dark single-page UI: KPIs (cycles / solved / API budget / dream iterations), live event feed, per-task attempt boards with score bars and code diff-downs, dreamed-policy gauges per category and the learned recipe library.


Positioning

  • The paper this implements: Dream-RSI (arXiv:2609.14858); official repo zhengkid/Dream-RSI, explainer dream-rsi.com.
  • Peer implementations (this repo is an independent open one, not the first): TheAstrayDev/dream-rsi-sdk (adapter SDK, strict replay; LLM-written policies on their roadmap), robinber/dream-rsi-spark (local Qwen + CUDA kernel exploration), plus skill variants (hermes-dream-rsi, dream-rsi-skill, Pi-RSI). What differentiates this repo: an always-on autonomous supervisor with budget guards, LLM-written exploration policies gated by counterfactual replay (section 3 closed), a public-benchmark measurement (goose × TravelPlanner) rather than self-scored claims, a deterministic decoy-trap self-check suite, persistent cross-run memory, and a negative-results section on memory that the loop itself enforces.
  • OpenRSI / Frontis-MA1 (FrontisAI/OpenRSI) is a different layer: they post-train weights; this library optimises the exploration policy around a frozen model — no training, no GPU, no weight access. Complementary: their trained improver could be the frozen LLM behind our client.

Module architecture

ModuleRole
core.treeDiscoveryTree — hypotheses, results, actions, thoughts, Sentinel facts, termination reason
core.simulatorreplay simulation without touching the environment
core.dreameroffline optimisation over recorded history
core.agentLLM agent driven by the rewarded policy
core.policygenpolicy generation, AST gate, rollout scoring, promotion
core.curatorlesson distillation, headroom verdict, gate verdict
core.judgecompletion judge for test-less (criteria) tasks
loopAutoRSIRuntime — the autonomous supervisor (opt-in sentinel_engine)
dreamdream_once — automated cadence seam: evidence ingest, tier decision, gate + ALM, skills publish; lock-guarded
lifecycleArtifactLifecycleManager — state machine, event-sourced history, lineage, GC
sessionsHermes state.db / Cursor state.vscdb readers, redaction, lessons→skills bridge
memoryDreamMemory — policies, recipes, lessons, trees, events
sandboxrun_isolated — the single untrusted-code execution boundary
toolsCodeVerifier — sandboxed verification of candidate solutions
llmOpenAI-compatible client (stdlib urllib transport; model discovery + 404 self-heal)
cliloop / status / dashboard / bench / bench-policy / gate-replay / mcp / proxy / sentinel
mcp, mcp_serverMCP stdio server for loop tools and sentinel checks
proxyloopback OpenAI-compatible proxy borrowing goose's upstream
dashboardzero-dependency live web UI
gate_replayoffline validation of the lesson promotion rule
sentinelhost-agnostic error-class engine + SentinelObservation world contract
utils.goosegoose config resolver (CLI + desktop dialects, keychain)
plugins/*Hermes / Claude Code / goose / DeepSeek Harness adapters
benchmarks/travelplannerthe benchmark: sandbox MCP server, episode runner, official-evaluator wrapper, split CSVs, raw published episode records

Tests

python -m unittest discover -s tests          # or: pytest tests/ -q

321 tests (+10 subtests), stdlib-only: benchmark behavioural contracts (greedy must solve nothing; the gated arm must dominate ε; the rollout must be prefix-only), sandbox escape regressions, gate semantics, MCP protocol, goose config dialects, sentinel channel semantics (nudge fires once per budget, merges with recurrence notes, clean calls count toward the budget; the permute-gate serves until ceil(0.8·budget) then refuses forever, per session; the 30-day production replay must reproduce its recorded note count exactly), model discovery and stale-pin self-heal contracts, and the structured seam (cross-session counts never reach to_world_dict(); the frontier-feature attribution pattern — beats greedy with metadata, ties without; the vendored Hermes engine pinned byte-for-byte to upstream by test_vendored_engine_matches_upstream).

License

MIT — see LICENSE.

Comments

Loading…

Similar plugins

deepseek-desk-rsi

by huchunlinnk

Recursive Self-Improvement engine for DeepSeek Harness — DSH maintains DSH: bounded perceive→integrate→verify→parity→repair→propose loop enforcing 1:1 upstream feature parity

Manifest valid

★ 4

MIT

JavaScript

Aug 18, 2026

dsh plugin --profile web add dsh-desk-rsi

by pk14742952-AD

DSH host RSI for local LLMs (currently Qwen3.8 27B): per-use continual self-improvement, counter-reason self-check, critic escalation, and memory retrieval.

Memory & ContextManifest valid

★ 17

MIT

JavaScript

Oct 7, 2026

dsh plugin --profile web add dsh-host-rsi

by STARDUSTLC666

Dreaming for agents: replay recent conversations between sessions (dream_digest), reflect and write a dream journal (dream_save/journal/recall), and idempotently bridge recurring lessons into AGENTS.m

Sessions & MessagesDevelopment & InfrastructureManifest valid

★ 2

MIT

JavaScript

Oct 8, 2026

dsh plugin --profile web add @stardustlc/dsh-dream

by jasen215

DeepSeek Harness (DSH) plugin for self-improving AI agents: continual learning, persistent memory, cross-session knowledge, review-and-refine workflows, and automatic rollback.

Memory & ContextManifest valid

★ 9

↓ 444/wk

MIT

TypeScript

Oct 8, 2026

dsh plugin --profile agent add dsh-continual-harness

DeepSeek Harness plugin for Reactive Resume: bridges your resumes and job applications into a Harness session over MCP.

Tools & CapabilitiesManifest valid

★ 41.7k

↓ 156/wk

MIT

Aug 24, 2026

dsh plugin --profile web add dsh-plugin-reactive-resume

by TaurenMountain

LLM-as-a-Verifier for DeepSeek Harness: fine-grained reward, Probabilistic Pivot Tournament best-of-N selection, and per-step progress tracking as agent tools.

Tools & CapabilitiesManifest valid

★ 5

↓ 106/wk

TypeScript

Aug 24, 2026

dsh plugin --profile web add dsh-llm-as-a-verifier