Same Model, Different Results: Why the Harness Decides If Your AI Pays Off
Two vendors pitch you the same week. Both say their product is built on the same well-known AI model. One product ends up quietly clearing your invoice pile every morning. The other produces confident mistakes you spend Fridays cleaning up.
Same engine. Opposite outcomes. If you've read parts one and two of this series, you already know where the difference lives: in the harness, the machinery of instructions, tools, loops, and guardrails wrapped around the model. This final post covers the evidence for that claim, and what it means for how you buy.
The Evidence Keeps Pointing Away From the Model
Start with the builders themselves. Anthropic, which makes the Claude models, states plainly that an agent's performance "can vary significantly based on this scaffolding, even when using the same underlying AI model", and credits outside developers with achieving large gains by improving only the machinery around a model they didn't change. When the people who build the engines tell you the car matters this much, believe them.
Then look at outcomes in business. MIT's Project NANDA found that 95% of enterprise generative AI pilots produce no measurable P&L impact, and attributed the failures to the systems around the models: no feedback retention, no workflow fit, no connection to the data where work happens. Those are harness failures, item by item.
And the market is starting to price this in. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. Risk controls are guardrails. Business value comes from workflow fit. The models were never the bottleneck.
of agentic AI projects expected to be canceled by end of 2027, for reasons that map to harness quality rather than model quality (Gartner, 2025)
How Bad Harnesses Actually Fail
Weak harnesses don't announce themselves. The model underneath is articulate either way, so the product demos well and fails later, in patterns worth recognizing:
- Silent failure. A step errors out, nothing escalates, and the agent reports success anyway. You discover the gap at month-end. A good harness treats every failed step as something to retry, route around, or hand to a person.
- Starved context. The product never connects deeply to your documents, calendar, or books, so the model reasons without your data and fills gaps with plausible guesses. Plausible is the expensive kind of wrong.
- Missing approval gates. Consequential actions (payments, customer emails, price changes) flow straight through the loop with nobody checking. Everything works until the day it very much doesn't.
- No memory design. Corrections you made last week never reach this week's briefing, so the agent repeats mistakes you already paid to find. Feedback retention was MIT's top-cited failure, and it's pure harness.
Every one of these failures produces a great demo first. Demos are short, scripted, and run on clean data, which hides exactly the four weaknesses above. Judge harnesses on their worst day, never their best.
The Buyer's Checklist
You can surface harness quality in under fifteen minutes of vendor conversation. Ask these, and expect specific answers:
- "Walk me through one task end to end." You want to hear the loop from part two: briefing, proposed action, execution, checkpoint. Vendors who can only describe outcomes haven't built much.
- "What exact tools does the agent have?" A real answer is a short, concrete list with limits attached. "It integrates with everything" is the wrong answer dressed as the right one.
- "What happens when a step fails?" Listen for retries, escalation rules, and human handoff. Silence here predicts silent failure later.
- "Which actions require my approval, and can I change the list?" No approval queue, no purchase. This one rule filters out most of the products that would have hurt you.
- "How does it remember corrections?" If fixing the same mistake twice is your job, the harness has no memory design.
- "Can I run it on my data before I sign?" The MIT study found purchased tools that embed into real workflows succeed far more often than pilots that stay generic. A vendor confident in their harness will prove it on your invoices, your bookings, your mess.
The Good News in All This
The harness being decisive is the best news a small business could ask for. It means you don't need to win a technology race or predict which frontier model comes out on top next quarter. The engines are already better than most everyday business tasks require, and they keep improving on their own schedule, to every buyer's equal benefit.
What separates winners from the 95% is fit: a harness built around your actual workflow, connected to your actual data, with guardrails matched to your actual risks. Fit is knowable in advance. You find it by understanding your own processes first and making vendors prove themselves against those processes, using the checklist above, before money changes hands.
Next Steps
To put this series to work this week:
- Pick your one highest-friction workflow (for most of our readers it's document handling or scheduling) and write down what a good harness would need: the data it must reach, the tools it would use, and the actions you'd insist on approving.
- Run the six checklist questions against any vendor already in your inbox. Score them on specificity, and drop the ones who answer with adjectives.
- Start from where the money leaks. The best harness in the world only pays off when it's pointed at a workflow that's actually costing you something.
That last step is exactly what our free AI Readiness Assessment does: 4 minutes, 10 questions, and you get a personalized report showing where your business is losing money to manual work, with dollar amounts and a prioritized action plan. Find your leaks first, and every harness decision after that gets easier.
Written by
Michael Sweeting
Is your business leaking revenue?
Take our free 4-minute assessment to find out exactly how much you're losing to manual processes, and get a personalized action plan to fix it.
Start Your Free Assessment