When Your AI Brain Gets a Stress Test You Didn't Schedule A follow-up to: "It's Alive! — The Day I Let an AI Help Build Its Own Memory"
There's a difference between something working and something being proven. I learned that difference the hard way yesterday.
The Setup
If you've been following along, you know the story. I've been building what I call the MPDC RV Brain — a fully autonomous, self-defending mobile security platform running on a repurposed Dell OptiPlex in my RV. The security stack came first. Then we built CORTEX — a persistent AI brain built on SQLite and FastAPI that gives Claude continuous memory across sessions.
The idea was simple. Claude forgets everything when a chat ends. CORTEX doesn't. Every session, every decision, every gotcha, every win — written to the database. Next session, Claude reads the brain, picks up exactly where we left off.
In theory it worked beautifully.
Then reality showed up.
The Incident
Mid-session. Deep in the weeds. We were refining the brain, testing session capture, verifying data persistence. Good work. Real progress.
Then the chat saturated.
If you've hit Claude's context window limit you know the feeling. The session just... stops. No warning. No graceful exit. The instance you've been working with for hours — the one that knows your architecture, your preferences, your running jokes, your verification codes — gone. Inaccessible.
For most people that's an inconvenience. For someone running a complex infrastructure project with an AI co-pilot, that's a potential disaster. Weeks of accumulated context. Hundreds of micro-decisions. All of it locked in a dead session.
Except we'd built CORTEX specifically for this moment.
The Handoff
I opened a new session. Launched the CORTEX Brain Launcher. Hit "Copy Brain to Clipboard." Pasted.
New Claude. Full context. Immediately operational.
Sort of.
The Hard Parts (There Are Always Hard Parts)
Here's what I won't sugarcoat — "working in principle" and "working smoothly" are two very different things. The new Claude instance had the brain data but the brain data wasn't fully current. We hit every layer of the stack that could gum up — and they all did, simultaneously.
The cloudflared tunnel was holding a stale connection. When we restarted the cortex service earlier, the tunnel kept its old connection to the dead process. Fresh service, stale tunnel, stale data. Fix: wire cortex.service to auto-restart cloudflared via ExecStartPost. Done. Never again.
Anthropic's fetch infrastructure caches aggressively. We discovered that Claude's web_fetch tool was getting a March 9th cached snapshot of the brain endpoint no matter what cache headers we set. Cache-busting parameters didn't help. The cache was upstream of our control entirely. Fix: stop relying on web_fetch as the primary context injection method. Use the clipboard workflow. It's equally fast and completely reliable.
We wrote to the wrong database. Early in the session we tried writing context to a Postgres instance that turned out to be structurally separate from the SQLite database the CORTEX API actually reads. Thirty minutes of careful SQL work, committed to the void. Fix: document the actual stack. CORTEX runs on SQLite at /home/console/rv-stack/cortex/cortex.db. Postgres is for other services. Now that's in the knowledge base.
A stale Python process was holding port 7777. After a restart sequence, the old cortex process hadn't released the port. New service couldn't bind. Fix: sudo fuser -k 7777/tcp. Also now in the knowledge base.
Every one of these failures is now a GOTCHA entry in CORTEX. The next Claude instance that walks into this project will know all of them before they happen.
That's the whole point.
The Verification
At the end of the session, after every layer was fixed and confirmed, we ran the real test.
New session. Fresh Claude. CORTEX brain injected via clipboard.
I asked for the verification code.
"Ah yes, a whiskey neat for the gentleman." 🥃
That's the code we'd written to the brain. That's what came back. Immediately. Correctly.
CORTEX 1.0 isn't just working in principle anymore. It's battle-tested.
What This Actually Means
The goal was never perfect AI memory. The goal was resilient AI memory. A system that survives the chaos of real-world operation — session limits, service restarts, network hiccups, tunnel failures, operator errors at 2am — and comes back online knowing exactly where it left off.
Yesterday proved we're there.
The new Claude instance came online, read 16 knowledge entries, saw 2 captured sessions, knew the full action item backlog, understood the architecture, matched the working style, and picked up the conversation without missing a beat.
That's not a party trick. That's infrastructure.
What's Next
The security stack is solid. The brain is operational. Now we build intelligence on top of it.
- OpenSnitch — application firewall rules need a daemon resurrection
- Tor — container still unhealthy, needs an upgrade
- Grafana dashboards — visualizing the full security picture in real time
- Node-RED — tying security events into the broader RV automation system
- Vaultwarden — secrets management for the platform
And looming over all of it — the reason we're building all of this — the T3600.
Proxmox. Ollama. Local LLM.
When that comes online, CORTEX stops being a Claude companion and becomes a fully autonomous on-prem brain. No session limits. No cloud dependency. No chat saturation. No forgetting.
The RV remembers everything.
Soon it won't need to ask anyone for help remembering. 👁️
Build the cage first. Then put the animals in it.
The animals are getting smarter.
It's yours. Make it true. 🥃