Damian Fozard
Damian Fozard

AI interviews well

A robot sits behind a desk reading a sheet of paper, across from a man seen from behind, against an orange wall

Why most attempts to put AI inside a business deliver less than the pilot promised

Damian Fozard

At lunch a few years ago, Jack Daly, the sales trainer who has written more on the discipline of hiring well than most people, was telling me about the importance of getting the first 30 days right. I asked whether he had ever got the hiring decision wrong himself. He laughed, the way someone laughs who has been asked the question often enough, and told me a story.

He had hired a sales leader for a financial services start-up he was running with a business partner. On the morning of the first day, walking the new man through the operations, his well-trained instincts began to tell him something was wrong. The reactions were not the right ones. By lunchtime he was sure. He went to his partner and said, “I have made a mistake. This is not the right person. We have two choices. I can write him a cheque for twenty thousand dollars today and we move on, or we can spend three months paying him to fail his probation, and then fire him.” His partner trusted his judgement. Jack went back to the new sales leader with the cheque in his hand and said this to him:

“I want to congratulate you on how well you interview. You are not right for this company. Good luck in the future.”

He concluded the story by warning me that there were, in his experience, a great many people who interviewed very well.

I was reminded of this conversation when I started to implement AI in my own businesses, several years later. AI interviews extraordinarily well. You give it a sample task, ask it to produce a report, ask it to analyse a set of figures, and it produces, with apparent ease and considerable detail, work that no ordinary employee can match. The pilot dazzles. The board approves. The pilot turns into a project, the project turns into a deployment, and somewhere in the deployment the dazzle quietly disappears.

Part of the gap is the accuracy ceiling of the models, which is the subject of Two Kinds of Not Knowing. The rest of the gap, which is much larger, is what happens when you put a system that can perform at super-human rates into an organisation that was not designed to absorb one. It is the question I raised in The Over-Manage / Under-Manage Cycle: local performance is not the same thing as global performance, and most of what businesses do with AI optimises the wrong level. The plan is wrong for the same reasons most plans are wrong. Improving a step is not the same as improving a system. What follows is what I have learned about that, after three years of getting it wrong, and what the research suggests is the right shape for the investment.

I. The Interview

The pilot is supposed to be illustrative. It is not. It is a curated demonstration of what the model can do in a clean environment, on a clean task, with no organisational friction. The literature on AI productivity confirms that, under those conditions, the gains are real. A field experiment run across nearly 2,000 software developers at Microsoft and Accenture found a 26% increase in weekly tasks completed when developers had access to GitHub Copilot. An earlier Copilot study reported that developers finished a fixed task 55.8% faster than the control group. Boston Consulting Group, in a controlled experiment with their own consultants, found day-to-day work improved by 12% to 25%. Customer-support agents using a conversational assistant resolved 14% more issues per hour. These numbers are real. They are also tested against tasks shaped to suit the technology, in studies designed to prove the technology works.

What happens at the boundary of those conditions is a different story. The MIT NANDA report The GenAI Divide: State of AI in Business, published in 2025 from a research base of 150 leader interviews, 350 employee surveys, and analysis of 300 public AI deployments, found that 95% of generative-AI pilots produced no measurable return on the company’s profit and loss. Somewhere between $30 billion and $40 billion of enterprise spending had moved no needle. The figure is softer than it looks, since the interview base is small and the report never says what counts as a return, so I take it as a direction rather than a measurement. S&P Global Market Intelligence, surveying over 1,000 enterprises across North America and Europe the same year, reported that 42% of companies had abandoned most of their AI initiatives in 2025, up from 17% in 2024. The average organisation scrapped 46% of its proofs of concept before they reached production. The pattern is consistent across the studies that have looked at it, which is to say that it is not noise.

The diagnosis is the part that matters. The reports converge on the same point. The failures are not, in the main, failures of local model capability. The capability is visible in the pilot. They emerge when organisations attempt to extend that capability across long, stateful, real-world processes, where small uncertainties compound and the controls absent from the pilot turn out to have been doing more of the work than anyone realised. In that sense the failures are deployment failures: brittle workflows, missing data discipline, mismatched expectations, no governance, no integration with the systems where the work actually lives. Vendor-led implementations succeed roughly twice as often as internal builds. The most consistent ingredient in the successful deployments is not the cleverness of the model. It is the operating discipline of the team doing the integration.

This is what I mean when I say AI interviews well. The model is the candidate. The candidate is brilliant in the room. The work is somewhere else.

II. The Plateau

The clearest case I have of this from my own businesses is at CoreAVI, the avionics software company I ran for 20 years, where we put generative AI to work authoring low-level requirements, cut the cost of the step by a quarter within months, and then watched the curve flatten while the model kept improving by every measure we could apply to it. I have told that story in full in The Sixty Per Cent. What belongs here is the diagnosis.

I did not break the plateau at CoreAVI. The diagnosis took me much longer than it should have, and by the time I had it, I had run out of room to act on it. The plateau was not in the AI; the plateau was in the people process around the AI. We had implemented a fast tool inside a slow system, and the slow system had absorbed the speed. People processes adapt to whatever sits in the middle of them. They route around it, validate it, check it, hand it off, batch it, schedule it, and the routing and validating and checking and handing off and batching and scheduling are the rate-limit. The fact that one step is now ten times faster matters very little if the other seven steps are not. This is what Eli Goldratt set out in The Goal in 1984, and what I leaned on in The Over-Manage / Under-Manage Cycle: an hour saved at a non-bottleneck is a mirage, and an hour spent fixing a non-bottleneck is worse than a mirage, because it consumes attention that should have been spent on the constraint. We were optimising a local step inside a global pipeline. The system was telling us, faithfully, that the local step was no longer the constraint, and we had not understood the message.

There was a second pattern at work beneath the first, which took me longer to see. The more autonomy we gave the system, and the more accuracy the certification environment demanded, the larger the surrounding containment infrastructure became. Human review layers, validation stages, sign-offs, reconciliation checks, exception handling, and governance processes expanded to absorb the model’s uncertainty. The AI had accelerated the generation of outputs, and the organisation still carried the burden of deciding whether those outputs were trustworthy. The bottleneck was not just adjacent process friction; it was the containment burden that grew in step with the autonomy we tried to give the system.

The lesson arrived too late for CoreAVI and in time for Squint Cognition, the AI company I went on to build, where we treated the whole eight-step certification pipeline as the thing to be governed rather than inserting a fast model into one step of it, on the assumption that coherence, validation and state had to live inside the system or the old human process would re-form around the model and absorb the gains again. The product that came out of it, Squint AeroCert, is described in The Sixty Per Cent; the cost reduction CoreAVI had been chasing for years arrived only there, and only when the last of the eight steps was inside the system.

The lesson, more general than my own business, is that putting AI into a people process is rarely the highest-leverage move available. The people process is the thing the AI is up against. The question for any operator considering an AI investment is not “can the AI do this task.” Of course it can; that is what the pilot proved. The question is “what is the rest of the pipeline doing while the AI does this task.” If the answer involves a human reviewer, a batched export, a manual validation, a downstream form, a regulatory sign-off, or a meeting on Tuesdays, the gains will compound only if every one of those is absorbed too. Otherwise the pipeline will deliver its old throughput, the AI will deliver a clean dashboard, and the two will diverge until somebody notices that the dashboard is the only thing being optimised.

The objection I hear most often from my own industry is that in certification the human reviewer is not friction. DO-178C requires that requirements be reviewed, that the review be independent of the author at the higher design assurance levels, and that evidence of the review exist; a regulator will not accept a process in which the model checked itself. That is right, and it is the reason the objection does not reach the argument. What the standard requires is a review with the properties it names: independence, coverage, evidence. It does not require the requirement to be re-typed into a second tool, batched for a fortnightly meeting, exported into a spreadsheet for a sign-off the standard never asked for, or routed through the other steps that had grown up around the reviewer because the reviewer was there. At CoreAVI the review the standard asks for was a small part of what the eight steps did with a requirement once the model had written it. The rest was ours, and it was the rest that absorbed the speed. The redesign at Squint kept the decision points the standard requires and put everything around them inside the system. The standard, read carefully, was never the bottleneck. The process built in its name was.

III. The Three Modes

The failure modes I have just described attach to one shape of AI inside a business much more than to others, so the shapes need to be separated. Three modes are in active use, and the returns from each are quite different.

The first mode is AI embedded inside the software you already use. My business bank account, by way of an example I encounter weekly, now reads a supplier invoice and pre-fills the electronic payment with the vendor details, the amount, and the reference number. I check the result and approve. The time saved is real. The error rate is, if anything, lower than mine. This mode is useful, it accumulates across hundreds of small interactions in a working week, and it is not a paradigm shift. The vendors of the underlying software have done the integration work. The user does not have to. The returns are modest, additive, and broad. There is no transformation in this mode, only a steady reduction in friction.

The second mode is AI as a force multiplier. One capable person doing what previously required a team. The widely-circulated stories of an individual shipping a working iPhone application in a weekend, or building a serviceable company website single-handed, are real instances of this mode at consumer scale. The studies I cited earlier (Copilot, BCG, customer support) are this mode at organisational scale: gains of 20 to 55 per cent on well-defined tasks done by people who already know how to do them. The most consequential example, however, is the one playing out in Ukraine. Small Ukrainian units, often six or seven soldiers, equipped with low-cost first-person-view drone technology, have repeatedly halted the advance of Russian armoured columns that on paper outmatched them in every dimension. By 2025, between 70% and 80% of Russian casualties at the front were estimated to come from drone strikes, and Russian columns had been forced to disperse into groups of five to ten vehicles, an order of magnitude smaller than the formations at the start of the war. That is the force multiplier in its starkest form: technology in the hands of trained operators against opponents who have not yet adapted, with results that would be unbelievable if they were not in the daily news. In a business, the equivalent is less dramatic but the shape is the same, which is one operator doing the work of a team because the operator has the right tool and the team did not.

The third mode is the one most enterprise AI ambition claims to be doing but rarely is: AI that owns the pipeline. Agentic AI, with strong governance, well-defined tasks, and clean sources of data, executing an end-to-end process from input to outcome with humans involved only at the decision points where their judgement matters. This is what Squint AeroCert is, and it is, in my experience, the only mode that produces the order-of-magnitude shifts that pilots imply and deployments rarely deliver. It is also the hardest of the three to do well, because it requires the operator to commit to replacing rather than augmenting the process, which means accepting that the people who have been doing the work will either do something else or not be employed by you at all. This is not a comfortable place for many operators to sit. It is, however, where the structural improvement lives.

The point that took me longest to absorb is that governance cannot remain external once the process horizon becomes long enough. A system that fully owns a pipeline has to maintain coherence, state integrity, and validation internally, or the organisation simply reconstitutes the old process around the model in the form of review layers and operational controls. The moment those controls reappear, the gains begin to clip again, in exactly the way they had clipped at CoreAVI.

The reason for separating the modes is that most enterprise AI investment sits in mode one or mode two and is reported to the board as though it were mode three. The dashboards show the per-task productivity gains; they do not show the throughput of the pipeline as a whole. The mistake repeats itself across companies of every size, and it accounts for a substantial fraction of the 95% the MIT report identified.

IV. The Cost of Augmentation

There is a popular vision of the AI workplace that has the technology functioning as a tireless assistant to every employee. Each person, in this vision, becomes more productive in their existing role. The aggregate gain is the per-person gain multiplied by the headcount. The published productivity studies, taken naively, support the arithmetic. The behavioural research that has emerged in the last year suggests the arithmetic is unstable.

The Microsoft and Carnegie Mellon study published in early 2025, The Impact of Generative AI on Critical Thinking, surveyed 319 knowledge workers and found that the more they trusted the tool the less critical thinking they reported doing, not because they had become lazy but because their cognitive role had shifted; an MIT Media Lab study later that year, which did measure brain activity, found reduced engagement in the regions associated with critical thinking among heavy users. The work of generating an answer had been replaced by the work of evaluating one, and evaluation, contrary to the casual assumption, is more cognitively expensive than generation in many contexts, not less. Microsoft’s 2025 Work Trend Index reported that 80% of the global workforce already lacked the time or the energy to do their jobs, and nearly half described their work as chaotic and fragmented. The Harvard Business Review articles of early 2026, When Using AI Leads to Brain Fry and AI Doesn’t Reduce Work, It Intensifies It, describe a consistent picture from heavy users: a buzzing mental fog, slower decision-making, headaches, and the sense that the tool has not lightened the work but accelerated the rate at which work arrives.

The mechanism is not mysterious. AI performs at super-human rates. The person sitting next to it does not. When you connect a fast process to a slow one through a human reviewer, you have not augmented the human. You have created the bottleneck Goldratt warned about, and made it tired as well as overworked. The pattern is visible in the field reports and in the productivity surveys. The companies reporting the largest disappointments are not the ones that bought the wrong models. They are the ones that asked their employees to absorb the speed of the model without redesigning the work around them.

This is not an argument against giving employees access to AI tools. I use them every day. It is an argument for being honest about what they are. A force multiplier in the hands of a small number of capable operators is a different thing from a productivity tool distributed across a workforce. The first compounds. The second saturates, often within months, and the saturation is paid for in the cognitive bandwidth of the people you most need to keep their judgement sharp.

Where to Spend

For an operator deciding where to spend in AI, my own view, refined by three years of getting some of this wrong, is that two investments pay back quickly enough to be worth making before the strategic shape is fully clear.

The first is the conversion of data into information. Most companies sit on years of operational data they have never had the bandwidth to ask questions of. AI is, among other things, an extraordinarily cheap way of asking those questions. The proposal pipeline is the most accessible example: a small team can now produce proposals with the thoroughness, presentation, and competitive positioning that previously belonged to firms with research departments. The depth of expertise that AI can be made to bring to bear on a single document, a single negotiation, a single market analysis, is materially greater than what an individual operator could previously assemble. It does not transform the business. It does meaningfully shift the quality of the inputs to every important decision, and that compounds.

The second is force-multiplier investment, made personally or with a small number of high-leverage individuals in the company, before the competition arrives at the same place. This is the asymmetric bet. The cost is small, the time horizon is short, and the upside is that you are the first operator in your market with the new capability. The cost of not making it is that someone else does. The shape of competitive markets, in my experience, is that the operators who notice the new tool early end up with structural advantages that compound for several cycles before the rest of the field catches up.

The longer-term investment, which is harder and slower, is to look for entire process pipelines in the business that can be folded into a single AI-driven flow. This is mode three, and it is where the order-of-magnitude shifts live. The cost is significant. The political difficulty is greater. The reward is that the business creates extraordinary results and is not just a faster version of the business you had. It is a different business.

The research summarised here argues against distributing AI tools widely across a workforce as the primary AI strategy. The benefits are real but uneven, and the costs (cognitive saturation, brittle workflows, accumulated frustration) are easy to underestimate. The strategy is not wrong; it is incomplete. It works only if the work has been redesigned around the tool, not around the employee.

Much of current enterprise AI strategy is still built around accelerating fragments of workflows while leaving coherence, validation, and judgement outside the system. The businesses that achieve structural advantage will, on the evidence so far, not be the ones that deploy the most AI features, but the ones that redesign processes around systems that can hold a process together with less external governance.

Jack wrote the cheque on the first morning. He did not wait for the probation period to confirm what the interview had hidden and the first hour had shown him; he trusted what he saw in the operation over what he had heard in the room. Most of the AI deployments I have watched fail did the opposite. They trusted the interview for a year. The pilot is the interview. The first months inside the process are the first morning.

More essays