AI & Innovation
An automated job overwrote the same file two nights in a row. What it actually took to trust an AI agent with write access after that.

Daniel Dopler

The File That Got Destroyed Twice
I trust automation with a lot in my system now: research, content extraction, scheduling, review queues. I did not always trust it with the ability to overwrite my own data, and here's exactly why.
I had a background job, an automated process managing a queue of decisions, tasks that needed a yes or no from me eventually. One evening it ran, and instead of adding to the existing record, it replaced it. Nine decisions I'd already made, gone, overwritten by a near-empty file. I caught it, restored the whole thing from the previous night's backup, verified the count, moved on. Lesson learned, I thought.
Twenty-four hours later it happened again. Same job, same failure mode, this time even after I thought I'd fixed the underlying cause.
That second time is the one that actually taught me something, because the first time I treated it as a bug. The second time made clear it was a design problem: I had given an automated process broad write access to a shared file without any structural protection against it clobbering that file, and "fix the bug" wasn't going to be enough. I needed a rule that made the failure mode structurally impossible, not just less likely.
So the rule became: automated jobs append, they never overwrite. If something needs to replace existing data, that's a human decision, made deliberately, not a side effect of a job doing its normal run. And every job that touches shared state gets backed up before it runs, not just on a nightly schedule, so recovery is always one step away instead of a day away.
The uncomfortable part of this story isn't the two failures. It's that I'd built a genuinely capable multi-agent system, and capability without guardrails is exactly how you lose real work twice in two days. The fix wasn't smarter AI. It was the boring kind of discipline: append-only by default, backup before every risky write, and treat any process that can silently destroy data as something that needs a structural constraint, not just a warning label.
If you're building anything with agents that write to shared files, ask the blunt question before it asks you: what happens if this runs twice, badly, back to back? If the honest answer is "it gets worse both times," that's the gap to close first.





