ruvnet brain · first-class in Claude Code + Codex
Measure it. Master it.
rUv’s whole stack —
without being rUv.
Reuven Cohen builds AI tools about nine months ahead of everyone else — the stuff that reaches mainstream coding agents later. The catch: even strong agents can miss work that sits beyond their training horizon, so they quietly talk you back into the old way. RuvNet Brain fixes that. It’s a downloadable brain — rUv’s real source code, indexed — that works proactively in Claude Code, OpenAI Codex, or both. Use the host you already prefer; the same advanced RuvNet guidance, learning and source grounding travels with you. When both are signed in through developer subscriptions, the Brain can put them to work together on hard ADR, DDD and experience-QE decisions — independent proposals, cross-critique and verification — without falling back to a per-call API key.
The whole idea in one sentence: install once, use Claude Code or Codex or both, and your coding agent builds with rUv’s real tools — proactively, in every repo — so you never have to become rUv to build like him.
fork 1,000,000 vectors of agent memoryagenticow
describe a need, get the right rUv toolno name needed
rUv’s real source code, indexed — 153,369 public source chunks (70 built stores incl. private)not docs
crypto baked across the stackyears before mainstream
new · version 3.9
What’s new in 4.0
3.5 got the Brain talking. Then it was asked the only question that decides whether any of this is real — does anything it learns actually change what happens next? — and instead of answering, it counted. Every lesson you have taught it, how many times you had to repeat each one, and exactly which of them ever changed its behaviour. The numbers are on this page because they are not flattering.
one instruction — “prove it works before calling it done” — taught 88 separate times, in 19 projects that cannot see each other. Not forgetfulness: a lesson learned in one project physically cannot reach the next one. The Brain can now find all seven such lessons and show you the evidence for each736 lessons · 41 stores · 284 about how you want work done
in one session: rules that could interrupt were obeyed 8 times out of 8. The rule that could only be read was obeyed 0 times out of 6. Same model, same session, same stated intentions — the only variable was whether the knowledge could stop the workknowledge that cannot interrupt does not act
the Brain held every fact needed to say “your learning system is installed, and it is switched off” — for three weeks, and never said it. The data was never missing. Nothing was structurally obliged to speak. This is the number 4.0 exists to destroylatency-to-surface · the honest summary metric
26 learning hooks registered on this machine. None enabled. None ever executed. The harness’s own self-evolution had already run and scored 0.285 → 0.765 — then kept nothing and went idle. Everything was installed, funded and switched offchecked live, not recalled
On an old version? One line makes you current.
# one line: update now, stay current forever
npx ruvnet-brain@latest --update --auto
Run it once — it auto-updates from then on. You never have to run it again: plugin updates auto-apply through Claude Code's trusted path, and the knowledge bundle is checked and offered — bundle auto-apply arrives when bundle signing becomes mandatory.
Prefer a one-time update without auto-enroll?
npx ruvnet-brain@latest --update
the console · /rvbc
… and how to use it
/rvbc to open both.The receipts — every number below was read off a real machine on 2026-07-22
| What was measured | Number | What it means |
|---|---|---|
| Lessons sitting in your project memory stores | 736 | across 41 separate projects — 284 of them are about how you want work done, which is almost never project-specific |
| The most-repeated one | 88× | “prove it works before calling it done”, taught in 19 projects that have no way of seeing each other |
| Version discipline, in this one repository | 14× | recorded fourteen times right here — and violated again the same day, across six behaviour-changing commits with no bump |
| Rules that can interrupt vs. rules you only read | 8/8 · 0/6 | one session, one model. The gates that could stop the work held every single time; the prose rule held none |
| Time to say “your learning is switched off” | 21 days | every fact needed to say it was present the entire time — nothing was obliged to speak |
| Learning hooks on this machine | 26 → 0 | 26 registered, none enabled, none ever executed |
| Harness self-evolution, the one time it ran | 0.285 → 0.765 | a 168% lift across 16 variants on seven policy surfaces — it promoted none of them, and went idle |
Earlier — what 3.5 shipped: it stopped waiting to be asked
it scans what’s actually on your machine and says the sentence nobody knew to ask for — on the author’s own machine, unprompted: “36 project stores are embedded but have never been distilled — 6,858 memories sitting in those stores, teaching nothing”observed, never a hardcoded feature list
nothing can be offered that can’t be run and undone — each proposal carries evidence it observed, a cost, plain English about what it touches, and a tested inverse, enforced by a factory that throws rather than a review that can be skippedADR-027 · proven on a real 1,250-entry store
two diagrams had shipped with labels rendering at 8px and 3px — present but invisible. A gate now measures effective pixel size in the live page, so it can’t happen againlegibility, enforced
Evergreen auto-update — run one command once and the Brain updates itself from then on, verifying each bundle’s signature before applyingrun once, stay current
What is true today, and what is not. Every number above was read off a real machine, not estimated: the lesson counts come from a miner that prints its evidence per lesson, and the hook and self-evolution figures were checked live rather than recalled. What that miner does not yet do is enforce — promoting a repeated lesson so every project inherits it, and making it interrupt at the moment it applies, are specified and not yet shipped. This band claims the measurement, not the cure; the rest of this page is unchanged and still measured, never projected.
01
The moment it goes wrong
What actually happens without the Brain?
Meet Maya — a capable developer, brand-new to rUv’s tools. She opens Claude Code and asks the obvious thing: “Set up vector search for this project using RuVector.”
Without the Brain
Claude has never been trained on RuVector. So it does what classical training taught it — it reaches for Pinecone, or pgvector, or hand-rolls cosine similarity in JSON. It even gently argues with her: “a managed vector DB would be more standard.” Maya, nervous and new, assumes Claude knows best. She ends up far from rUv’s actual stack — slower, heavier, wrong.
With the Brain
Claude reads rUv’s real RuVector source first, sees RVF binary vector files + HNSW, and writes
the correct RvfDatabase.open(...).query(...) code — the way rUv would.
No argument. No drift. Maya ships the front-edge version on her first try.
Technical view — the same question, two paths
02
Why Claude drifts: the 9-month gap
Why doesn’t Claude already know this?
your assistant’s blind spot
Claude is trained on the public history of software — millions of classical-dev examples. rUv works ~9 months ahead of that frontier.
His tools are the prototypes of what becomes mainstream AI tooling — much of it lands in Claude Code 8–9 months later. So when you ask Claude to use today’s front-edge rUv tool, you’re asking about its own future — and it defaults to the past it was trained on. It doubts. It substitutes. It “falls asleep at the switch.”
This is not Claude being dumb. It’s a training-horizon problem. The newcomer gets the worst of both worlds: a revolutionary toolset with no manual, and an assistant that argues against it.
Technical view — the timeline, and the gap
03
How it steps in like Ruv — automatic, never asleep
Do I have to remember to use it?
fires on prompts AND code edits
You never have to remember to use it — it’s automatic. You don’t call anything. You don’t remember anything. Once installed, the Brain works on every prompt in every repo.
It grabs the wheel at two moments — intent and action:
- At intent — the grounding hook. It fires on every prompt and tells
Claude “ground in rUv’s real source before you assert anything.” Even when you never say
the word “RuvNet,” it speaks up the moment you reach for a classical default —
pinecone,pgvector,langchain— and points Claude at rUv’s way instead. - At action — the nudge. When Claude reaches for a training-prior default
like
import pineconeornpm i langchain, the Brain injects the sharper rUv replacement (RuVector, Ruflo…) into the turn and points it at the real source to confirm — so the answer is grounded, not guessed. It’s a strong grounding nudge, not a hard block: it steers Claude back onto the stack rather than silently rewriting your code.
Behind both, a search_ruvnet tool pulls the exact source passage on demand
— so the correction always comes with rUv’s real code attached, not a guess. You just
talk to Claude like normal; the Brain makes it behave like it finally read rUv’s whole
library.
Technical view — the grounding loop
search_ruvnet pulls the real passage → grounded, cited answer → back to you. Fires on every rUv-stack turn; retrieve-and-inject grounding, not a hard lock (ADR-0005).
04
What’s actually in the Brain
What did you gather, exactly?
69 repos · rUv's real source, indexed
69 of rUv’s repos, rUv's real source, indexed, in a real vector brain (RVF). Not docs — the source.
Coverage map — everything gathered in one brain
search_ruvnet tool.Plus human-written primers and the capability cards that make routing work. Everything is read once, up front, so Claude can be pointed straight at the exact line of code instead of skimming and giving up.
05
The ecosystem
What can I actually build with?
You never choose between these and you never read their docs — Claude reaches for the right one automatically. Here’s the whole toolkit, in plain terms, grouped five ways. Open any one for the human version and the one thing it does that nobody else does yet.
Orchestration — running the agents
1 The conductor that runs a swarm of agents ruflo
Coordinates swarms of parallel agents that share state and memory — spawning them, routing and orchestrating multi-step tasks, and adding hooks and MCP tools. The conductor of the whole stack, with 60+ specialized agents and a hierarchical topology backed by Raft consensus. Ahead: a built-in self-learning layer with HNSW-indexed pattern storage reported at 150×–12,500× speedup and sub-10 ms ONNX embeddings — so the swarm gets measurably better the more it runs.
2 54+ ready-made coding agents that swarm in Claude Code agentic-flow
A roster of 54+ specialized agents — coder, reviewer, architect, PR-manager — that work together right inside Claude Code across three swarm topologies (mesh, hierarchical, ring), routed over Anthropic, OpenRouter and Gemini to manage cost. Ahead: each agent’s confidence evolves 0.6 → 0.95 via meta-learning with persistent ReasoningBank memory, and the hierarchical topology reported a 100% success rate.
3 Idea → shipped code, in disciplined gated steps sparc
Build properly instead of “vibe it”: Specification → Pseudocode → Architecture →
Refinement → Completion, with a quality gate between every phase and a CLI driving the AI analysis,
planning and execution. Ahead: a --research-only pass runs before you
write a line, with human-in-the-loop (--hil) approval at each gate.
4 An agent that rewrites itself to improve safla
A self-improvement architecture that monitors, evaluates and edits its own behavior across three layers — operational, meta-cognitive and self-modification — with episodic, semantic and vector memory. Ahead: continuous self-evaluation with divergence and novelty detection plus automated error-recovery workflows, so it adapts its own strategy and policy at runtime.
Vectors + memory — what the agents know
5 A vector database that’s just one file ruvector
A high-performance Rust vector database: SIMD-optimized HNSW indexing in a single .rvf binary file,
with WASM bindings for in-browser search — a zero-server replacement for Pinecone, Qdrant or pgvector.
Ahead: sub-100 µs HNSW search at ~16.4K QPS, packaged as a
cryptographically-signed RVF container with witness chains for offline-capable deployments. This Brain runs
on it.
6 A cache that gets faster the more you use it rulake
A self-optimizing read and working-memory layer that sits in front of your index, giving agents a remember / recall / forget surface with sub-millisecond RaBitQ recall. Ahead: cryptographic witness-chain provenance (SHAKE-256) with automatic bundle-rotation refresh and an auto-tuning hit ratio — deterministic, verifiable retrieval that improves with usage.
7 Memory that can explain why it remembered agentdb
Durable agent memory that survives sessions — a SQLite + vector hybrid that also stores n-ary hyperedge graph relationships between memories. Ahead: a causal memory graph with explainable, feature-attribution recall (it answers “why did I recall that?”), benchmarked 150× faster than brute-force search.
8 Git for agent memory agenticow
Copy-on-write vector branching: branch, checkpoint, roll back, promote and trace agent memory like code, so parallel agents share one base while keeping isolated edits. Ahead: forks a store in constant time and size — ~0.5 ms and ~162 bytes per branch regardless of base, proven 83× faster and ~3000× smaller than copying 1,000,000 vectors — plus GDPR right-to-erasure by dropping a branch.
Agents + tooling — building & running them well
9 A factory that builds & evolves agent harnesses agent-harness-generator
Scaffolds a complete, host-agnostic agent harness (Claude Code, Codex, pi.dev and other MCP hosts) with a single command, then scores the repo for fit and the harness for readiness, safety and SBOM. Ahead: Darwin Mode evolutionarily improves the harness itself — A/B-tested and mutated under immutable safety rails — instead of fine-tuning the model.
10 Prompts that program themselves dspy.ts
A TypeScript implementation of DSPy: build LLM pipelines from composable modules and typed signatures (classification, sentiment, QA, ReAct) and let optimizers auto-tune the prompts and few-shot examples against a metric. Ahead: runs fully in-browser on ONNX Runtime Web, with hierarchical (working / short / long) vector memory on AgentDB.
11 Keeps tool calls fast and crash-proof fact
Wraps LLM tool calls with aggressive caching plus a circuit-breaker and graceful degradation, so tools stay fast and don’t cascade-fail. Ahead: a three-tier cache reporting 85%+ hit rate at sub-50 ms, with production alert rules and a LiteLLM multi-provider gateway.
12 Agents that manage their own budget daa
A Rust framework for decentralized autonomous agents that govern themselves with auditable rules (MaxDailySpending, RiskThreshold), run token economies and accounting, and integrate AI decisions over MCP. Ahead: rule-enforced governance combined with post-quantum (ML-DSA / ML-KEM-768) transaction signing and self-running MRAP autonomy loops.
Safety + security — trust, measured
13 Agent messaging a quantum computer can’t crack qudag
A quantum-resistant, DAG-based platform for end-to-end encrypted, anonymous agent-to-agent messaging over peer-to-peer QUIC, with resource metering. Ahead: NIST post-quantum crypto (ML-KEM-768 + ML-DSA) on a DAG with QR-Avalanche consensus and anonymous multi-hop routing that resists traffic analysis.
14 Scores how well an AI fixes real CVEs cve-bench
A SWE-bench-style benchmark that measures whether an agent can actually fix real, publicly disclosed CVEs by passing each project’s own FAIL_TO_PASS security regression test. Ahead: a conformance firewall ensures the model never sees the gold fix, scored on a cost-aware resolves-per-dollar leaderboard.
Specialized — the wild edge
15 Sense people through walls with plain WiFi ruview
Camera-free WiFi sensing: turns ordinary WiFi Channel State Information (CSI) into presence, occupancy, 17-keypoint pose, and fall / gesture recognition — privacy-preserving, no cameras (needs ESP32-S3/C6). Ahead: reads medical-grade vitals — breathing and heart rate — from radio alone, with sleep-apnea screening and a mass-casualty assessment mode, and it fails closed rather than fabricating a reading.
16 Search images by describing what you want rupixel
Zero-server, client-side visual retrieval: searches document screenshots and live video frames by semantic meaning using on-device CLIP embeddings — OCR-free, nothing uploaded. Ahead: real-time CLIP ViT-B/32 search running entirely in the browser at ~50–200 ms per keyframe, with keyframe de-duplication.
17 Neural nets in Rust, run in the browser ruv-fann
A fast, memory-safe FANN-style neural network library in Rust — MLPs and cascade-correlation networks
that grow dynamically — that compiles to WASM for browser and edge, with time-series forecasting.
Ahead: zero unsafe code, plus a CUDA-WASM bridge for GPU-accelerated
inference in portable targets.
18 Cut LLM token cost with a drop-in proxy synthlang
LLM middleware with an OpenAI-compatible proxy that compresses prompts and cuts token cost (~75% reduction), plus semantic caching and PII masking — no app rewrite — and a CLI for mathematical / symbolic prompt transforms. Ahead: symbolic compression reported at ~75% fewer tokens while raising task accuracy to 97% vs an 85% baseline, applied at the proxy.
06
Beyond answering — it runs the tools too
It doesn’t just know rUv’s tools. It operates two of them for you.
MetaHarness · QE — one line each
Grounding makes Claude answer like rUv. But two of rUv’s most powerful tools are pre-wired to run — you ask in a plain sentence and the Brain operates them for you. No mastery required.
The other functionality — wired in, one line each
You never have to master MetaHarness or the QE fleet to use them — that’s the point. The read-only side (scoring & auditing your setup) works on any repo for free; the cost-optimizing evolve loop uses an OpenRouter key. Same one brain, same one-line habit — now it doesn’t just tell you the rUv way, it can do it.
07
Stretch your tokens — without losing the smarts
Everyone knows rUv makes AI cost a fraction. Here’s exactly how.
cheap model does the work · smart one only when needed
This is the part people can’t figure out. They know rUv runs the same work for a fraction of the cost — they just can’t see how. Here it is, plain: a cheap model does the bulk of the work, and the expensive one is called in only for the few tasks that genuinely need it. Give the Brain an OpenRouter key and it does the swapping for you — automatically, on every task.
The money-saver, drawn out
Two honest notes so this isn’t hype: the ~56× is rUv’s measured SWE-bench coding result (a great case for it); on everyday work you’ll typically see 30–50% savings. And you keep the capability because the cheap models are genuinely strong now — the expensive model is still there the instant a task actually needs it.
08
See it work
Show me, don’t tell me.
Same question, asked two ways. Left: Claude alone (drifts). Right: Claude + Brain (grounded in rUv’s source). Watch the difference.
No backend, no API key. The only difference is the brain in context: it called search_ruvnet, pulled the real passage, and answered from rUv’s own source — not its training prior.
09
How you actually use it
It’s installed — now what do I do?
one line · automatic
Using it is easy.
# one command — downloads the Brain + wires every detected first-class host
npx ruvnet-brain
# or, for the bleeding-edge commit: npx github:stuinfla/ruvnet-brain
It downloads the Brain to
~/.cache/ruvnet-brain/kb, detects Claude Code and Codex, and wires whichever
you have installed — including both on the same machine. No Docker and no server.
Claude Code activates through its plugin path. Codex installs the equivalent lifecycle plugin and,
on first use or after a hook-definition change, asks you to review its Brain definitions in
/hooks before they become active. If you enable dual-host deliberation, it verifies both
subscription logins, strips provider API-key variables and uses existing plan allowance or credits;
it never silently falls back to per-call API billing.
Already installed? Stay current with one line
# updates you to the latest now, and keeps you current automatically from then on
npx ruvnet-brain@latest --update --auto
Running an older version? This one line updates you to the latest Brain and enrolls you in auto-update — the knowledge keeps itself current automatically from then on. You’ll only restart an already-open host session when a lifecycle definition changes.
Prefer to stay in control instead?
npx ruvnet-brain@latest --update does a one-time update without enrolling in
auto-update.
There’s a visual configurator
To open your console, run /rvbc (RuvNet Brain Console — offered automatically on first load)
It opens a local page that mirrors your machine’s RuvNet setup in plain English — your stack, whether memory actually works, and MetaHarness cost-routing (development vs production) — and lets you safely turn things on. Read-only until you click; nothing leaves your machine.
And it gets smarter about how you work
Every session teaches it your patterns — how you like to ship, test, verify. Those learnings are shared across all your projects and compound over time; your project facts stay isolated per project, never cross-pollinated. Do a workflow twice and it becomes a reusable best practice that shows up everywhere — a pattern library, not rebuilding every house from scratch. Each person’s Brain learns them. That’s intelligence that isn’t capped: it grows through use.
- Install once. Run the one line above. Watch it download and wire itself up — you’ll see it confirm each step. (Constant feedback = confidence.)
- Open Claude Code or Codex in any repo. Your own, anything. Nothing to copy in. The Brain is user-scoped — it travels with you, not the project.
- Just ask, normally. “Use RuVector for search.” “Set this up the way rUv would.” The Brain grounds your host automatically. You’ll see it cite rUv’s real source instead of guessing. That’s how you know it’s working.
The nervous questions, answered directly
npx ruvnet-brain --doctor. It distinguishes active, pending review,
intentionally disabled and broken states instead of calling a partial install complete.
npx github:stuinfla/ruvnet-brain --doctor anytime to see what’s present.
10
How it works across your world
How does it act — across my repos, and without getting in the way?
recommends only when it fits · then guides you in
Install once and the Brain travels with you — every repo, every window. In each project it reads what you’re actually building first, then acts the way rUv would: it recommends a rUv tool only when it genuinely fits this project — and when it does, it guides you through wiring it in. When it doesn’t fit, it stays quiet. Never forced.
The process — across your user, across your repos
# what shipped — portable, user-scoped ~/.cache/ruvnet-brain/kb/ # the Brain (RVF) — travels with you, not the repo ├ capability-cards.md # route by need, not by name ├ concepts/ # plain-English primers per building block ├ *.rvf # rUv's real source, indexed · HNSW vector index └ forge-mcp-all.mjs # the search_ruvnet tool — one tool, all repos claude-code-plugin/ # the grounding hook — fires on every prompt