AI & Innovation

An automated job overwrote the same file two nights in a row. What it actually took to trust an AI agent with write access after that.

Daniel Dopler

A fractured document repaired with gold seams, representing data recovery after an automation failure

The File That Got Destroyed Twice

I trust automation with a lot in my system now: research, content extraction, scheduling, review queues. I did not always trust it with the ability to overwrite my own data, and here's exactly why.

I had a background job, an automated process managing a queue of decisions, tasks that needed a yes or no from me eventually. One evening it ran, and instead of adding to the existing record, it replaced it. Nine decisions I'd already made, gone, overwritten by a near-empty file. I caught it, restored the whole thing from the previous night's backup, verified the count, moved on. Lesson learned, I thought.

Twenty-four hours later it happened again. Same job, same failure mode, this time even after I thought I'd fixed the underlying cause.

That second time is the one that actually taught me something, because the first time I treated it as a bug. The second time made clear it was a design problem: I had given an automated process broad write access to a shared file without any structural protection against it clobbering that file, and "fix the bug" wasn't going to be enough. I needed a rule that made the failure mode structurally impossible, not just less likely.

So the rule became: automated jobs append, they never overwrite. If something needs to replace existing data, that's a human decision, made deliberately, not a side effect of a job doing its normal run. And every job that touches shared state gets backed up before it runs, not just on a nightly schedule, so recovery is always one step away instead of a day away.

The uncomfortable part of this story isn't the two failures. It's that I'd built a genuinely capable multi-agent system, and capability without guardrails is exactly how you lose real work twice in two days. The fix wasn't smarter AI. It was the boring kind of discipline: append-only by default, backup before every risky write, and treat any process that can silently destroy data as something that needs a structural constraint, not just a warning label.

If you're building anything with agents that write to shared files, ask the blunt question before it asks you: what happens if this runs twice, badly, back to back? If the honest answer is "it gets worse both times," that's the gap to close first.

MORE INSIGHTS

person hand in a dramatic lighting

LET’S WORK TOGETHER

If you’re ready to bring structure, clarity, and AI-driven leverage to your business, let’s build it.

person hand in a dramatic lighting

LET’S WORK TOGETHER

If you’re ready to bring structure, clarity, and AI-driven leverage to your business, let’s build it.

person hand in a dramatic lighting

LET’S WORK TOGETHER

If you’re ready to bring structure, clarity, and AI-driven leverage to your business, let’s build it.