Damian Fozard
Damian Fozard

The 360 review

A lighthouse sweeps a beam of light through concentric coloured orbits carrying small planets, above soft grey clouds, with a tiny figure climbing the steps below

Why an AI employee is easy to hire and hard to manage

Damian Fozard

In 2009 I visited one of my largest customers, in Arizona, and walked into a production hall that had gone quiet. It had built the equipment the company sold. The equipment was now built in Mexico and Malaysia, and the hall stood empty. Six months later I came back to find the same hall full again, this time with temporary tables, soldering stations, oscilloscopes and a small army of people whose job was to find and correct the defects in the systems the outsourced suppliers had built. The engineer walking me through it summed up the arrangement in a sentence: “The person who thought outsourcing this was a good idea got promoted, and left us to deal with the quality issues.” The hall was never empty again on any later visit. For more than ten years, what everyone still called the temporary lab sat there, checking for and repairing the gap between what had been ordered and what had arrived.

I have thought about that hall more than any other room I have visited in business, and it is at the front of this essay because it is the most honest picture I have of what an AI employee costs an organisation. In a companion essay, AI Interviews Well, I described how readily an AI system clears the interview. You set it a sample task, it returns work no ordinary candidate could match, the pilot dazzles, and the board approves. My argument there was that most of the distance between the pilot and the deployment is organisational, the friction of dropping a super-human worker into a process never built to absorb one. This essay is about what happens after the candidate has been hired. It is about the review, the ongoing assessment of an employee whose work you have to live with month after month, and my contention is plain: with these systems, the review is harder than the interview, and correcting course is harder still.

Most businesses have a ritual for this. The 360 review gathers judgements from every direction, the manager above, the peers alongside, the people below, the clients outside, and assembles them into a picture of how someone is actually performing once the novelty has worn off. Run that exercise on an AI employee and you get a strange result. The output is fast, fluent, well presented, and frequently excellent. Yet not one of the people asked to vouch for it can tell you whether it is correct. They can tell you it looks right. That is a different thing, and the difference is the whole of what follows.

The Immaculate Colleague

Let me start with what I mean by Subjective AI, because it is the property that defines every system you are likely to have used. By subjective I do not mean opinionated. I mean that the system is not responsible for the correctness or the consistency of its own output. Look at the small line of text beneath the box where you type your prompt. It says, in one wording or another, that the system can make mistakes and that you should check important information. That disclaimer is not boilerplate. It is the business model. The accountability for whether the answer is true has been handed back to you.

Imagine a colleague who worked this way. Their documents are immaculate. The spelling is perfect, the slides are beautifully composed, the prose is clean, the tone is exactly right. What you never know is whether the figures in the table are correct, whether the sources they cite actually exist, or whether the numbers on the final slide add up. Sometimes they do. Sometimes they do not, and nothing about the surface of the work tells you which case you are in. This is the curse of subjective output: plausible and incorrect, produced at random, with no reliable signal separating the two.

It is that last part, the randomness, that makes the situation so awkward from a management point of view, and without real precedent. Anyone who has managed people is used to employees whose work is poor. There are systems for that, probation, supervision, training, dismissal. What I have almost no experience of is an employee whose work is brilliant and substandard at the same time, in no fixed proportion, varying from one task to the next. There is no developed management practice for the person who is reliably unreliable, because such a person does not usually survive long enough in an organisation to require one.

The Limiting Factor

The instinctive response is to assume the problem will dissolve as the models improve. It will not, and the reason matters. The models will keep getting more capable. They are astonishing now and they will be more astonishing in a year. Capability is not the limiting factor. Subjectivity is. As long as the system remains unaccountable for the correctness of what it produces, the limiting factor is the relationship between the size of the task you hand over and the cost of being sure the result is right. The bigger the task, the more you face one of two bills: a larger management overhead spent checking the output, or a larger number of mistakes left embedded inside it, undetected, to be discovered later by someone who trusted the clean surface.

A business can, of course, choose to live with the mistakes. Some do, deliberately, and the most instructive example I know sits in American health insurance. In March 2023 ProPublica and the Capitol Forum reported that the insurer Cigna had been running an internal system known as PxDx, shorthand for procedure to diagnosis, which matched submitted claims against a list and flagged those that did not fit for denial. The system, they reported, let company doctors reject batches of claims with a single electronic signature without opening the patient files. Over a two-month period in 2022 it was used to deny more than 300,000 claims, with physicians spending an average of about 1.2 seconds on each one. Cigna disputed the characterisation, saying the review happened after treatment, concerned billing and coding rather than the care itself, and was a routine, industry-standard step. What is telling is not the legal question, which the courts have been working through, but the structure of the bet. The company reportedly estimated that only around five per cent of policyholders would ever appeal a denial; one analysis cited in the litigation put the appeal rate as low as 0.2 per cent of denied claims. Across the Medicare Advantage industry, where appeals are actually filed roughly four in five denials are overturned; that figure is not Cigna’s alone, but it describes the market Cigna was pricing. Read those numbers together and you see a process designed around the assumption that wrong outputs would mostly go uncontested. The mistakes were not an accident the system tolerated. They were a cost it had priced, against the near-certainty that most people would pay the bill rather than fight it.

PxDx is a rules-based automation that predates the current generation of models, not a large language model in the modern sense, but it illustrates the management posture exactly. A business that adopts subjective systems at scale, and cannot afford to check everything, is implicitly making Cigna’s wager: that the embedded errors will cost less than the effort of preventing them. For an insurer banking on customer fatigue, that arithmetic may hold. For most businesses, and most tasks, it does not. You cannot ship avionics, file accounts, or advise a client on the assumption that only one customer in 20 will notice the part you got wrong.

There is a version of the wager that is not cynical, and it deserves stating. If a model’s error rate on a task falls below the error rate of the people who would check it, then checking makes the output worse rather than better, and the rational thing is to stop checking; on that arithmetic Cigna’s posture is not fatigue but efficiency. I accept the arithmetic and not the conclusion, for two reasons. The first is that the comparison is between rates and the harm is in cases: a checker who is wrong more often but wrong at random does less damage than a system that is right more often and wrong in the same place every time, on the same kind of claimant, because the second error is invisible until it is a pattern. The second is that the model’s rate is not known. It is estimated, on a test set that was not the world, and the burden of this essay is that a subjective system cannot tell you when the world has moved away from the test set. The wager is rational when the rate is known and the errors are random. Neither holds, which is why the temporary lab fills up.

There is an obvious retort: people make mistakes too. It is true, and it is not a small point. No human employee is free of error, and any honest manager has signed off work that turned out to be wrong. Two differences hold the human case apart. The first is that human mistakes can be course-corrected: a person who makes an error can be shown the error, can understand why it was an error, and can carry that understanding into the next task so that the same mistake is less likely. The second is accountability. We hold people responsible for their work, and that responsibility shapes how they do it. A subjective system offers neither. It does not learn from the correction you give it today in any way that survives to tomorrow, and there is no one inside it to hold responsible. The signature on the work is yours.

Information and Reasoning

Those who know how these systems are built will raise a sharper objection at this point. Does prompt engineering not solve this? Does fine-tuning, or reinforcement learning from human feedback, not exist precisely to make the output reliable? The answer is yes and no, and the distinction is worth drawing slowly, because I got it wrong myself for longer than I should have.

Consider any task as having two components, information and reasoning. The information is the material the task draws on, the structures and correlations and context that bear on it. The reasoning is the logic applied to that material to reach a result. The techniques everyone reaches for, careful prompting, fine-tuning, reinforcement learning, work principally on the first component. They sharpen the system’s grip on the information, steering it towards the kinds of material most useful to the task and the relationships between them. This is real and it is valuable. The system attends to better inputs and draws on more relevant correlations. The reasoning the task requires is, on my reading, largely unchanged. I hold that as a hypothesis rather than a finding, and the part I am sure of is narrower. The errors I am describing are the ones that survive training and optimisation: not the system reaching for the wrong information, but the system reasoning unreliably over the right information. You can feed it a perfectly curated context and still receive a confident conclusion that does not follow from it.

I spent too long treating this as a technical software problem, something the next round of optimisation would clear. It is not, or not only. It is a business management and structure problem, and it has a close historical relative that businesses have been wrestling with for 40 years. The relative is outsourcing.

What Outsourcing Taught Us First

The record of outsourcing is not a single story but two, and the line between them is sharp. Where businesses have handed out functions that are necessary but not core, the results have generally been good. Payroll, human resources administration, bookkeeping, accounts processing: these are skill-centric, task-centric services, well defined, governed by external rules, the same in their essentials from one company to the next. A specialist provider does them at lower cost, with deeper expertise, and with better compliance than most firms could manage in house, and the company gets back the attention it was spending on work that never distinguished it from a competitor. The growth of business process outsourcing from the 1980s onward rests on this logic, and for non-core functions the logic holds.

The other story is what happens when a business outsources its core, the product, the engineering, the manufacturing, the thing it is actually for. Here the failures pile up. Boeing’s 787 Dreamliner is the example most people in my industry reach for. Boeing pushed an unusually large share of the aircraft’s design and manufacture, well over half of it, out to a global network of suppliers, and the programme was hit by years of delay, integration problems, and quality defects that traced back to the difficulty of coordinating core work the company no longer directly controlled. The pattern recurs across surveys of failed outsourcing arrangements: the activity that should not have been outsourced, control lost over the thing that mattered most, hidden costs that surfaced only once the work was somewhere else. The function was gone, but the responsibility for it never left.

The Temporary Lab

I have lived both sides of this. For 20 years I ran CoreAVI, a company that provided outsourced development of software and hardware for avionics systems. When the model works, the benefit to both parties is real and large. My business was incentivised to invest in improving a component to a depth that none of my individual customers could ever have justified on their own; they got the benefit of that investment without bearing its full cost. That is outsourcing at its best, and I saw it produce results neither side could have reached alone.

The trade-off is control, and it falls on the customer. The skill of managing several outsourced processes at once, of holding ultimate responsibility for quality and correctness and timeliness without holding all the levers that determine them, is one of the harder things a business can be asked to do. The hall in Arizona is what that skill costs when it is missing. The work moved, the speed and the saving were real, and the responsibility for correctness came straight back through the door and set up tables in the empty hall. The engineer’s sentence applies to more than outsourcing.

Outsourced Minds

An AI employee is an outsourced employee, in a stricter sense than the metaphor first suggests. You cannot change its fundamental abilities. You did not train it and you cannot retrain it at the root; you can prompt it, frame it, and fine-tune at the margins, but the reasoning underneath is fixed by someone else. More than that, you cannot control how it is improved over time. A newer version of a frontier model will arrive with greater capability, and it will also arrive with different properties, which is to say it will make its mistakes in new and unfamiliar places. The colleague you spent a year learning to check is quietly replaced by a more talented one whose particular failure modes you do not yet know. With a human team you would call this turnover and manage it deliberately. Here it happens on a release schedule you do not set.

This is why an AI employee needs continuous assessment rather than a one-time decision to hire. The system is not adaptable in the way a person is, so the adapting has to be done by everything around it. The processes, the checks, the points at which a human judgement is inserted, all of it has to be built and rebuilt to bridge the standing gap between plausible output and correct results. The 360 review is not an annual event in this world. It is the permanent condition of employment.

The Specialist

Where I stand shapes everything I have written. My background is in high-reliability systems, the kind of engineering where an undetected error is not an inconvenience but a hazard, and where the entire discipline exists to make correctness something you can depend on rather than something you hope for. Subjective AI is, in that light, close to the antithesis of how I was trained to think. That tension is why I have spent the last stretch of my working life building the other kind, objective AI, the system that answers for its own correctness within a stated boundary.

The distinction is simple. Subjective AI is a generalist. It is, as I have described, inconsistent and unaccountable for its output. Objective AI is a specialist: a system that understands the particular task it has been given, can recognise and correct its own mistakes, and in doing so lowers the burden of management and oversight rather than adding to it. The frontier models of subjective AI have already delivered capabilities that astonish the people who know how to use them well. The shift in markets, the redrawing of what a business can do and what it costs to do it, will come not from the generalist that interviews brilliantly and cannot be relied upon, but from the specialist that can. That is my wager, and it should be read as one: Ken Wenger and I have built a company on it, and the evidence so far is a handful of certification programmes rather than a law.

The companion essay was about hiring. This one has been about the review that follows, and about how much harder the second thing is than the first. The subjective employee interviews well and reviews badly, and no amount of capability removes the gap so long as the accountability for being right stays on your side of the desk. The specialist that closes that gap is the harder thing to build and the only hire worth making for work that has to be right, and I have seen what the alternative looks like after ten years: a hall full of temporary tables.

More essays