Training Billy — When Your Local AI Recommends Norton on Your Wazuh Box
Raw intelligence means nothing without context. We built the cage. Now we train the animal.
A follow-up to "You Wouldn't Give a New Employee Root on Day One. Why Should AI Be Different?"
It finally happened.
The T3600 is online. Proxmox running. Ollama GPU-accelerated on a GTX 1050 Ti. DeepSeek R1 — a fully local, air-gappable, zero-cloud-dependency language model — running on hardware I own, inside the network I built, answering to nobody but me.
No token limits. No rate throttling. No subscription. No phoning home.
I built the cage. I put the animal in it.
Then I gave it its first real test.
It told me to install Norton.
The Incident
I sent DeepSeek a screenshot of an OpenSnitch firewall popup. OpenSnitch is an application-layer firewall — every outbound connection gets intercepted and requires explicit approval. It's one of nine independent layers in the MPDC security stack.
DeepSeek looked at this popup — on my hardened Linux machine, behind CrowdSec, Suricata, Zeek, Wazuh, AdGuard with 1.2 million DNS block rules, and a recursive DNS resolver — and responded with the following:
"Use your legitimate, up-to-date antivirus/anti-malware software (like Windows Defender, Norton, McAfee, Bitdefender, Kaspersky) to perform a full system scan."
On Linux.
Running Wazuh.
Behind CrowdSec.
The AI I built to help defend my network just recommended I install the thing my network was architected to not need.
This Is a Billy Madison Problem
Here's the thing: DeepSeek isn't stupid.
It's actually impressively capable. The reasoning traces alone — watching it think through a problem step by step — are worth the build cost.
But capability without context is Billy Madison on day one of 1st grade.
Billy isn't dumb. He's intelligent, enthusiastic, and completely unequipped for the environment he just walked into. He's got raw material. He just has no idea where he is, who he's talking to, or what the rules are. Left to his own devices, he'll give you the statistically most likely answer for the average person in the average situation.
I am not the average person. This is not the average situation.
DeepSeek gave the right answer for a Windows user on a consumer network. That person genuinely should probably run a scan. The model wasn't wrong. It was answering the wrong question for the wrong person — because nobody told it who it was actually talking to.
That's not a model failure. That's a context failure.
And context failures are fixable.
The 30-Second Fix (And Why It's Not Enough)
The immediate patch took one paragraph:
"You are CORTEX, an AI assistant for Chris at MPDC. The user runs a hardened Linux security stack including Suricata, Zeek, Wazuh, CrowdSec, AdGuard, and OpenSnitch. Never suggest commercial antivirus or Windows solutions. All systems are Linux. Treat the user as an expert sysadmin."
Pasted into the system prompt. Model behavior shifted immediately. Competent peer mode. No more Norton.
But system prompts are duct tape. They work. They're fast. They're also stateless — every session DeepSeek wakes up with complete amnesia and has to be re-briefed from scratch.
CORTEX handles the memory layer right now. Brain endpoint loads, context injects, session starts informed. It works. But it's not training. It's not the same as a model that actually knows what it's doing, who it's working for, and why.
That's where we go next.
The Boot Camp Plan
Billy Madison didn't fail school because he was incapable. He failed because he skipped the foundation. The fix wasn't a lobotomy — it was structured, progressive education from the ground up.
Same principle here. Four phases:
Phase 1 — Modelfile (Now) Bake the system prompt permanently into the model via Ollama Modelfile. DeepSeek loads with MPDC context by default. No manual paste, no Open WebUI dependency, no risk of a cold start. This is the floor, not the ceiling.
Phase 2 — RAG (Next) Rather than a static briefing, connect the model to CORTEX's live knowledge base. Stack docs. Network topology. Decision history. Gotchas. The model stops reading from a fixed briefing card and starts querying a living database. Retrieval-Augmented Generation — the model knows what it knows, and knows where to look for what it doesn't.
Phase 3 — Structured Training Data Every conversation that produces a useful, contextually correct response is a labeled example. Every correction is a negative example. The Norton incident itself is training data — here's the wrong answer, here's why it was wrong, here's the right answer for this specific environment. Build the dataset systematically, session by session.
Phase 4 — Fine-Tuning (When Budget Allows) Actually modify the model weights. Stop briefing a contractor and start growing a staff member. A fine-tuned model doesn't need the system prompt anymore — the context is baked in. It knows the stack. It knows the architecture. It knows Chris doesn't want to hear about Windows Defender. Ever. Under any circumstances.
Where This Is Going
The five articles before this one document the build. The security stack. The memory layer. The stress tests. The governance architecture.
This article is about what comes after the cage.
The endgame for the MPDC RV Brain isn't a helpful chatbot in a Docker container. The endgame is KITT — but expanded into a 38-foot mobile data center and SOC. A fully autonomous, self-healing, self-defending intelligent platform that operates whether I'm at the keyboard or in the field.
The LLM isn't a tool I use. It's the operational core of a system that runs itself.
Every layer we've built feeds it:
- Suricata and Zeek give it eyes on the network
- Wazuh gives it forensic awareness of every system event
- CrowdSec and AdGuard give it active threat response
- Node-RED gives it hands — controlled execution, approved workflows, no freeform commands
- CORTEX gives it memory — every decision, every gotcha, every session, persistent across restarts
The governance model we designed holds: LLM is user. CORTEX is sudo. Node-RED is controlled superuser. Chris is root.
Nothing autonomous touches production without a human in the loop. That trust gets earned incrementally — the same way you'd expand permissions for any new team member who shows they know what they're doing.
Billy doesn't get root on day one.
But Billy is in school now.
The Takeaway
The Norton incident is the best thing that could have happened.
It showed exactly where the gap is. It gave us a clear before-and-after. It became an article. It became training data. It became a movie poster.
The gap is bridgeable. The model is capable. The infrastructure is ready.
We just have to do the work.
Build the cage first. Then put the animals in it.
The cage is built.
Time to educate the animal.
P.S. — Yes, we made a movie poster. Training Day meets Billy Madison. Denzel looming over Adam Sandler at a school desk. "Stay in school... or else." If you don't understand why that's the perfect metaphor for deploying a local LLM without a system prompt, re-read from the top.
We had a good laugh. Then we got to work.