Insights

Hire AI Engineers or Buy Delivery? What the 67% Split Actually Says

5 min read
Two routes to an AI capability compared for teams deciding to hire AI engineers

A retail group approves its first serious AI budget. The board wants to know what it is buying. The CTO is handed a familiar question: build the team, or bring in a partner?

The paper that comes back compares day rates, notice periods, IP ownership and the risk of knowledge walking out of the door. It is a competent paper. It argues the case on cost and control, reaches a defensible conclusion, and gets signed off. Eighteen months later the programme has produced three impressive demos and nothing in the P&L.

The paper asked the wrong question. Nowhere in it is the variable that most reliably predicts whether the work reaches production.

What the research actually measures

MIT's Project NANDA analysed more than 300 public AI deployments alongside 150 leadership interviews and a survey of 350 employees. Its headline finding — that around 95% of enterprise generative AI pilots deliver no measurable profit-and-loss impact — has been quoted to exhaustion. The finding underneath it has not.

The same study found that purchasing AI from specialised vendors succeeds approximately 67% of the time, while internal builds succeed roughly one-third as often. RAND puts overall AI project failure above 80%, around twice the rate of conventional IT projects. S&P Global Market Intelligence measured that the average organization scraps 46% of AI proofs of concept before production, that only 48% of AI projects reach production at all, and that those which do take an average of eight months from prototype to live.

Read quickly, the 67% looks like an argument for outsourcing. That reading is wrong, and acting on it produces its own category of failure.

Vendor engagements do not succeed because vendors are better. They succeed because they arrive with a contract.

Retrieval architecture — designing how a model reaches trustworthy internal context, including chunking strategy, retrieval evaluation and the handling of stale or contradictory sources.
Agentic orchestration — composing multi-step workflows where a model plans, calls tools and recovers from failure, using frameworks that have become standard rather than experimental.
Tool integration via MCP — connecting agents to external systems through the open standard that emerged as the default in 2026, rather than through bespoke glue that nobody else can maintain.
Evaluation design — building the test suite that decides whether the system is good enough to ship, which is the single most frequently skipped step in stalled programmes.
Production hardening — authentication, cost controls, latency budgets, fallback behaviour and audit trails, all of which sit outside the demo and inside the reason the demo never shipped.

Why the 67% is not an argument for outsourcing

A vendor engagement carries structure that an internal mandate does not. There is a defined scope, a stated outcome, acceptance criteria someone has to sign, a fixed end date and a commercial consequence for missing it. The engagement is falsifiable: at a known point it either passed or it did not.

An internal programme frequently carries none of that. It has a budget, a headcount and a direction of travel. Failure has no fixed date attached, so nothing forces the conversation. A pilot with no executable pass-or-fail test cannot succeed or fail; it can only continue to consume budget. That, not the employment status of the engineer, is what the 67% is measuring.

Which leads to a more useful conclusion than build-versus-buy. An internal team held to vendor discipline performs like a vendor. A vendor engaged on a vague discovery retainer performs like a stalled internal programmes. Organisations reporting real financial returns are around twice as likely to have redesigned the workflow before selecting the tooling — a sequencing decision, not a sourcing one.

How to make the decision properly

1. Write the acceptance criteria before the requisition. State what the system must do, on which data, at what accuracy, by when, for the engagement to count as successful. If this cannot be written, the problem is not yet ready for either a hire or a partner.

2. Choose the boundary, then the sourcing. Narrowly scoped, externally bounded work is what the 67% describes. Decide the boundary first; the sourcing question becomes far easier once the scope is fixed.

3. Buy the first use case, build the second. A bounded partner engagement compresses time to a falsifiable answer. Permanent hiring makes sense once one use case has proven value and the work becomes platform work.

4. Screen for evaluation, not model trivia. Ask a candidate how they would prove the system is good enough to release. Strong AI engineers describe a test suite and a threshold. Weaker ones describe a model and a framework.

5. Attach a kill date. Every AI engagement, internal or external, should have a date on which it either graduates to production or stops. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027; the organisations that come out of that well will be the ones that cancelled early and deliberately.

The honest position

There is no sourcing model that rescues a badly scoped AI programme. Hiring excellent engineers into an open-ended mandate produces expensive demos. Engaging an excellent partner on an open-ended retainer produces the same demos with a different invoice.

If the acceptance criteria are written and the boundary is fixed, either route can work. If they are not, neither will — and the hiring decision the board is debating is not the decision that matters.

SyncOrigins

SyncOrigins brings expertise from over a decade of enterprise technology leadership. Focusing on bridging the gap between strategic intent and technical delivery for global organizations.

Enterprise AI

Need a Custom AI Strategy?

We help enterprises architect and deploy production-grade AI solutions. Schedule your strategy session today.