It happens often enough now to be a pattern. An agentic AI pilot looks flawless in the demo. The agent sits on the ERP, picks up purchase orders, spots the exceptions, drafts the reconciliation and closes the loop without anyone touching it. The budget gets signed off in the room. Then, a few quarters later, someone quietly switches it off, and the write-up reads the same as all the others: the model wasn't good enough, let's revisit when the next one lands.
That read is wrong, and it's an expensive kind of wrong, because it sends you looking in the wrong place. The model was almost never the thing that broke.
Gartner is fairly blunt about the scale of this. It expects more than 40% of agentic AI projects to be scrapped by the end of 2027, and the reasons it lists are runaway cost, unclear business value and weak risk controls. Model quality isn't on that list. That gap tells you something. When one of these projects dies, the model is the most visible part of it, so it tends to take the blame. It's rarely what actually failed. The failure sits underneath, in the data the agent had to work with and the systems it was let loose inside.
MIT's NANDA research put some numbers behind the same idea. Give an agent a tight remit — one process, one dataset, an outcome you can measure — and it worked about 67% of the time. Point the same kind of agent at the whole enterprise and success fell to somewhere near 22%. None of that gap is about how clever the model is. It's about how solid the ground is under the agent's feet.
The model is the last 10% of an agentic ERP project. The 90% underneath it — who the agent can be, whether the data holds up, whether there's one version of the truth — is what decides if the thing works. It's also the part nobody wants to fund.
| ● | Access that's properly governed. The agent needs permissions scoped as tightly as you'd scope a new employee's: role-based, logged, easy to pull. An agent with open write access to the ERP isn't a productivity win. It's an unlogged risk sitting on top of your general ledger, waiting. |
|---|---|
| ● | Data the agent can trust. The agent only knows what you feed it. If your master data is full of duplicates, if vendor records disagree from one module to the next, if the same product shows up under three SKUs, the agent will take all of that at face value and automate the mistake faster than any human could. |
| ● | One version of the truth. When the ERP, the warehouse system and the spreadsheet finance actually trusts all give a different answer to “how much stock do we have,” the agent has nothing firm to stand on. It'll pick one of them and be confidently wrong a predictable share of the time. |
| ● | Integration that survives real traffic. A demo runs on tidy, frozen data. Production doesn't. It throws concurrent transactions, half-finished records and edge cases the pilot never saw at the agent all day. That's where most of these projects quietly come apart. |
| 1. | Start smaller than feels comfortable. Pick one process with a real owner and an outcome you can actually measure: invoice-to-reconciliation, the three-way match, a single exception queue. Treat that 67% versus 22% split as a design brief rather than a stat. There'll be pressure to make the first agent do everything. Resist it. |
| 2. | Clean only the data the agent will touch. You don't need enterprise-wide master data nirvana before you can start, and waiting for it is a good way to never start. You need the specific records inside this agent's scope to be clean, deduplicated and tied back to one source. Keep the cleanup bounded to the process you chose. |
| 3. | Sort out access before the agent exists. Define what the agent can do the way you'd define it for a new hire: least privilege, a full audit trail, a kill switch you can actually reach. Doing this first means the agent never operates past a line you can see. |
| 4. | Engineer for production, not the demo. Test against real transaction volumes, concurrent writes and the malformed records your clean pilot data politely left out. The thing that bites you is almost never the happy path. |
| 5. | Keep a human where being wrong is costly, and measure everything. Instrument the agent's decisions from day one. Either the numbers prove it's earning its place, or they let you kill it early and cheaply, before it becomes another line in Gartner's 40%. |
The honest bit
Agentic AI in the ERP is real, and the distance between the companies that get it right and the ones that don't is going to grow. But the winners won't be chosen by their model. They'll be the ones who were willing to pay for the layer that never makes it onto a slide: the governed, trusted, connected data sitting underneath it all.
If your last pilot got switched off, the next foundation model won't bring it back. The engineering under it might.
FAQs
Won't a better foundation model be the quickest way to improve our agentic AI results?
Usually not. Model quality climbs every quarter across every vendor, so it's the one lever you don't really control and don't need to. The levers that decide success — data quality, access governance, integration — only move if you build them. A better model sitting on a broken data foundation just gives you faster, more confident mistakes.
How narrow should our first agentic ERP use case be?
Narrow enough that you can name one process, one owner and one outcome you can measure. MIT NANDA found tightly scoped agents succeeded around 67% of the time against roughly 22% for general-purpose ones. Aim for one job, in one module, done well.
Do we need perfect enterprise master data before we start?
No, and treating it as a prerequisite is how these projects never get off the ground. You need the specific records inside your first agent's scope to be clean and tied to a single source. Keep the remediation bounded, prove the value, then widen it.
What's the most common reason agentic ERP pilots get cancelled?
The data and integration layer buckling under real production conditions it never met in the demo. Gartner puts the projected 40%-plus cancellation rate down to cost, unclear value and weak risk controls rather than model capability. The break sits a layer below the model.
Where should a human stay in the loop?
Anywhere a wrong autonomous action is expensive: financial postings, vendor payments, inventory commitments. Keep human sign-off on the high-stakes decisions, instrument the rest, and only pull the human out once the data shows the agent is reliably right.




