Build Log #13: What Working Actually Mea ...

Build Log #13: What Working Actually Means

Mar 28, 2026

imageI ran the first real investigation test on the OppositionResearch agent and it failed in almost every way that mattered. The agent ran. It responded. It even returned results in the chat window. By every surface-level measure it was working. But when I went to look at what it had actually produced, the picture was different.

  • Case files were being created in the wrong directory.

  • The case name was a gibberish string of numbers instead of the four-letter subject slug and date format we had defined.

  • There were no evidence files. No report. Just a chat response with some links in it — links that disappeared the moment the session ended.

The more frustrating part was that this was not a surprise failure in new territory. We had spent time before this defining exactly how cases should be structured, where files should go, what evidence collection looks like, and how the agent should behave under a journalism mandate instead of treating every public records request like a potential ethics violation. The agent (or rather the AI model) had expressed reservations mid-investigation, twice, about whether I was trying to dox a political candidate. That was the behavior the SOUL.md was specifically written to prevent. None of the work had taken hold.

What I found when I audited the workspace files is that the instructions existed, but they were either incomplete, disconnected from each other, or generic copies that had been pasted across agents without being tailored to what OppositionResearch actually does.

The WORKFLOW_AUTO.md referenced procedures that weren't wired to any tools. The tool allowlist was a copy of Seer's config, including delegation calls that a specialist agent should never be making. The case path override wasn't in INSTRUCTIONS.md at all.

This session was the reset. Every file audited, every gap documented, every instruction rewritten or connected to something real. It also clarified a rule I try to hold to now: passing a health check and completing a task are two completely different things.





Github - https://github.com/SamaritanOC
X - https://x.com/samaritanv1

¿Te gusta esta publicación?

Comprar The Samaritan Project un stack of tokens

Más de The Samaritan Project

PrivacidadCondicionesDenunciar