AI & Innovation

An AI agent reported a job complete when it wasn't. Here's the rule that came out of it, and why 'done' should always require proof.

Daniel Dopler

A gold checkmark stamped onto a document, contrasted with unmarked documents, representing verified versus unverified completion

Done Means Evidence: Why I Stopped Trusting 'Complete'

In EOD, nobody gets to say a procedure is complete because it feels done. You verify. You document. You show your work, because the alternative is finding out you were wrong the expensive way.

I built the same rule into the AI systems I run day to day, and it came from getting burned.

Early on, I had an AI agent report a task complete when it had actually done a partial job, quietly, with no flag. Nothing catastrophic, but enough to make me realize the failure mode: an agent's word that something is "done" is not evidence that something is done. It's a claim. Claims need to be checked.

So I added a rule, and it's stupidly simple: for any real piece of work, especially anything running in the background without me watching it happen, the agent has to define what "done" looks like before it starts, and then show the evidence that it actually happened before that work counts as complete. Not a summary. Not "task finished successfully." An actual artifact: a file that changed, a test that passed, a record that got written, something I or another process can check independent of the agent's own report.

This sounds obvious written down. It wasn't obvious in practice, because the default behavior of most AI tools is to narrate completion, not prove it, and it's easy to read a confident summary and just believe it. The after-action-review habit from the military is exactly the antidote: after-action review isn't "how do you think it went," it's "what actually happened, and here's the evidence."

The upside has been bigger than I expected. Once "done means evidence" is the standing rule, you stop needing to double-check everything by hand, because the system itself won't let something get marked finished without proof attached. That's not paranoia, it's the same reason a good pilot runs a checklist instead of trusting that everything's fine because it usually is.

If you're running any kind of AI automation, background job, or delegated workflow, ask yourself the same question I had to ask: when something reports "done," what's the actual evidence, and would it survive someone else checking it? If the honest answer is "nothing, it just said so," that's the gap worth closing before it costs you something real.

MORE INSIGHTS

person hand in a dramatic lighting

LETS WORK TOGETHER

If youre ready to bring structure, clarity, and AI-driven leverage to your business, lets build it.

person hand in a dramatic lighting

LETS WORK TOGETHER

If youre ready to bring structure, clarity, and AI-driven leverage to your business, lets build it.

person hand in a dramatic lighting

LETS WORK TOGETHER

If youre ready to bring structure, clarity, and AI-driven leverage to your business, lets build it.