All essays / June 28, 2026

Reliable AI vs. Coin Flip

Ask the same AI the same question twice and you may get two different answers. The difference isn't always the model. It's often how much guessing your data forces it to do. Better data doesn't just improve AI. It makes it reliable.

Reliable AI vs. Coin Flip

Ask the same model the same question twice. Word for word, settings untouched, everything lined up so it should give you the same answer every time. You'll often get two different answers.

Most people assume that's their mistake. A stray setting, a typo, something they'll go fix. It isn't. The thing you're renting changes shape between one call and the next, and the reason is almost insultingly ordinary. That model is handling thousands of requests at the same time, and how many it happens to be juggling the instant yours lands depends on how busy it is. Nothing of yours mixes with anyone else's. No data leaks, nothing like that.

It's simpler and stranger than that. The machine is just doing a different amount of work each time, and that alone nudges a number somewhere deep in the math, and a nudged number becomes a different word, and a different word becomes a different answer. You didn't do a thing. It's a slightly different machine every time you knock.

I keep coming back to this, because of what it does to everything built on top.

We talk about AI as though it were software. Software is the same every time you run it. That's the whole miracle of it: write it once, trust it forever. This isn't that. This is closer to asking a very smart, very tired employee the same question on two different afternoons. Mostly you get the same answer. Sometimes you don't. And the times you don't are not announced.

That last part is the one that should worry you.

Sit through enough demos and you notice the agent always finishes. It always returns something, confident and done. Then you check the work. A good share of the time it's wrong, and not loudly wrong. It's quietly wrong, in a way that looks exactly like being right. The dashboard says complete. The status says resolved. The only way to catch it is to already know the answer. On the standard case it's fine. On the messy one, the patient who doesn't fit the template, the claim with the odd modifier, the edge that was the whole reason you wanted help, that's where it hands you the wrong thing with a straight face and moves on.

Now set that next to the pitch everyone's selling. Pay for outcomes.

It sounds fair. Pay for results, not software. I understand why an owner likes the sound of it. But you can't price a steady outcome on top of an unsteady process, and you can't even agree on whose work the outcome was. Did the AI clear the claim, or did your biller change the workflow last month? Two people will argue that invoice forever. The vendors know it. Watch closely this year and you'll catch them backing out of pure outcome deals, dressing them up again as service contracts, bolting on minimum fees, anything to put a floor under a number that won't hold still. We called this a non-starter over a year ago. The market is finding out the slow way.

The usual fixes don't fix it. A minimum fee plus a slice of the upside just splits the gamble between you and the vendor. The dice are the same dice.

So what actually narrows the swing?

The model wobbles most when it has the most to work out on the spot. A vague task over messy inputs, and it improvises, and improvisation is where the variance lives. Tighten the task and clean what you feed it, and the spread collapses. The same model that was a tossup over a pile of unstructured notes turns boringly reliable over a clean, well-shaped record. Same model. Different odds. The difference is the data, and the data is yours to fix.

This is why the order matters, and why we keep insisting on the unglamorous part. Own your data. Organize it. Shape it so the model barely has to guess. Do that first and outcomes stop being a marketing word and start being something you could actually stand behind, because the thing producing them finally holds still. Skip it, and no pricing scheme on earth saves you. You're only choosing how to split the cost of the wobble.

You don't have to take anyone's word for any of this. You don't have to win an argument with a salesperson either. Run it yourself. Pull ten of your messiest cases, the ones that never go cleanly, and put each one through twice. Count how often the machine disagrees with itself. That's the number no demo will ever put on the screen. And it isn't really a verdict on the model. It's a reading of how much guessing your own data is forcing on it. Clean that up and the disagreements fade. Leave it as it is, and you're paying full price for a coin flip.