AI · Judgement · AI Literacy · 15 min read
The Assurance Gap
AI can make professional output almost effortless. Understanding, challenging and taking responsibility for what we produce still takes time, knowledge and judgement.
For most of my career, producing professional work took time. A report had to be researched, written, challenged and revised, software had to be designed, built and tested, and articles or research had to go through some form of review before anyone was comfortable putting their name on the finished product. None of those processes guaranteed quality, but there was usually a certain amount of human effort between having an idea and presenting the finished work to somebody else.
AI is changing that equation remarkably quickly. Work that once took days can now sometimes be completed in hours or minutes, and the amount we can produce has increased with it. Someone who previously wrote one 3,000-word paper a week could theoretically produce ten 10,000-word papers, while organisations can produce far more analysis, documentation, software and written material without anything close to the increase in time or people that would previously have been required.
We normally describe this as productivity, and in many respects it is. If I can achieve the same outcome in two hours instead of two days, there is an obvious benefit. I have written before about using AI to remove repetitive effort so people have more time for judgement, relationships and difficult decisions, rather than simply producing more work.
But I am beginning to wonder what happens when the capacity AI creates is simply used to produce vastly more.
Because someone still has to understand what we have produced.
Or at least, someone used to.
When the mistake is obvious
We have already seen some fairly amusing examples of what happens when that last step disappears. Articles have appeared online where the text itself looks perfectly normal until you reach the bottom and find something that clearly came from the AI conversation used to create it, perhaps an offer to rewrite the article, make it shorter or elaborate on a particular point.
The mistake itself does not really matter. What interests me is everything that must have happened for the mistake to become visible to us. The article was generated, copied into a publishing system and published to the world, yet apparently nobody involved in that process read the finished version closely enough to notice that the AI was still talking to the person who had asked it to write the article.
For an online article that might result in little more than embarrassment, but it exposes a much more interesting problem. The technology did exactly what made it attractive in the first place: it removed much of the time and effort required to produce something. Somewhere along the way, we may also have removed the time and effort required to think properly about what was produced.
That distinction becomes much more important when the output is not an online article, but research, academic work, software, legal analysis, engineering documentation or something contributing to a medical decision. The consequences are very different, but the underlying question remains the same.
Who actually checked it?
From 3,000 words to 100,000
Imagine that I used to produce one 3,000-word report every week. I had enough time to research it, write it, read what I had written, make changes and read it again before sending it to someone else, who might then challenge my assumptions, check the numbers or question the conclusion.
Now give me AI and assume I can produce ten 10,000-word reports in the same week. On paper my productivity has increased enormously, but I have also created 100,000 words that somebody needs to understand. If I am still expected to apply the same level of scrutiny to my own work, I need to read those 100,000 words carefully, check the underlying information, challenge the reasoning, make corrections and then read the revised material again.
Can I actually do that?
More importantly, if the reason for adopting AI was to make me dramatically more productive, will I be given the time to do it?
This is where the conversation about responsible AI becomes uncomfortable. We quite rightly expect people to apply discernment and judgement to AI-generated work rather than blindly trusting the output. That is also consistent with how I have thought about AI for some time: efficient output should not be confused with understanding, and the person using the technology remains responsible for checking the information and challenging the conclusion.
The problem is that this assumes the person has enough time, knowledge and capacity to do it.
If AI allows an organisation to expect significantly more output from the same number of people, while the pressure is simultaneously to deliver faster and at lower cost, careful human scrutiny starts competing directly with the productivity gain we were trying to achieve.
Something eventually has to give, and the easiest thing to remove may be the thing nobody can see.
I can already see the temptation
I am experiencing a very small version of this problem while writing this article.
AI can help me produce a complete draft remarkably quickly, and I could quite easily take the first version, give it a quick read and publish it. Instead, I have read it properly, questioned arguments that did not quite work, rewritten sections, read them again and found other questions that were not even present in the original draft.
The first draft was not bad. There were no spectacular errors and nothing that would have prevented me from publishing it. It was coherent, structured and probably good enough.
But it was not this article.
The distinction matters because assurance is not always about finding something that is factually wrong. Sometimes it is about recognising that an argument is incomplete, questioning an assumption that sounds perfectly reasonable or spending enough time with an idea to discover something that was not there in the first version.
None of this is particularly complicated, but it takes time and attention. Normal life is happening around me while I do it, and with a baby at home I certainly do not have unlimited amounts of either. I am also only producing one article, not 100,000 words of professional output every week.
I can understand the temptation to look at a perfectly reasonable first draft and decide that it is good enough.
In a workplace where AI allows considerably more to be produced and the pressure is simultaneously to deliver faster and at lower cost, that temptation will not disappear. It may become part of the operating model.
So when production becomes almost effortless but scrutiny still takes time, when do we stop looking for excellent and simply accept good enough?
Nobody announces that quality matters less
I doubt many organisations will announce, “From Monday, quality matters less.”
It can happen much more quietly than that. We produce more, review less, trust the tools a little more and gradually redefine what good enough means through behaviour rather than decision.
The first draft looks reasonable, so we send it. An AI review finds nothing concerning, so we trust it. A long report is too much to read properly alongside everything else, so we read the summary. Another deadline arrives, there is more work waiting and nothing went obviously wrong last time.
Eventually, the shortened review process becomes the normal review process, not because anyone deliberately designed it that way, but because it is the only way to keep up with the amount of work we can now produce.
This is not necessarily about people becoming lazy or irresponsible. It may simply be predictable human behaviour under pressure. I have already argued elsewhere that AI governance should not be designed around perfect human attention, because repeated success naturally encourages people to spend less time checking the next output.
The uncomfortable part is that we may never consciously choose between good enough and excellent. The choice may gradually be made for us by volume, cost and time.
The invisible control
Responsible use of AI ultimately relies heavily on two very human capabilities: discernment and judgement.
AI can help me create something, challenge it, check it, rewrite it and improve it, but I am still supposed to understand the output well enough to decide if it makes sense before I put my name on it or pass it to somebody else.
That sounds reasonable, but it creates a surprisingly invisible control.
If I use AI to produce a report and then carefully read the source material, challenge the assumptions, check the important facts and apply my own experience before sending it onwards, I have used AI as a productivity tool while retaining responsibility for the work.
Someone else could use exactly the same technology, generate exactly the same kind of report and send it onwards after reading only the summary.
From the outside, the processes may look identical. Both reports were created with AI, both may have passed an AI review, both have a person's name attached to them and both arrive looking polished and complete.
The difference is that one has been subjected to meaningful human discernment and judgement while the other has not.
Unless something obviously goes wrong, how would the person receiving it know which one they got?
When AI checks AI
This becomes even more complicated when we use AI to solve the assurance problem created by AI.
If AI produces a report, another AI checks the facts, another reviews the reasoning and another summarises everything for the person eventually approving it, we can point to several separate checks in the process. Automated assurance may become extremely capable and could find inconsistencies, errors and risks that human reviewers would have missed.
But it does not remove the need for human judgement. In some ways, it places even more responsibility on the person who eventually decides to trust the result.
If that person understands the subject, scrutinises the output and takes responsibility for it, the process may work extremely well. If they simply accept what the systems tell them, every subsequent person may reasonably assume that meaningful scrutiny already happened.
The creator assumes the reviewing AI checked it, the manager assumes the creator checked it, the director assumes the manager checked it, and the executive sees that the work has passed several stages of review and assumes somebody underneath them understood it.
Everyone can see evidence of a process.
Nobody can see the missing judgement.
Creating more than anyone can consume
There is another consequence of making production almost frictionless. When producing another ten pages, another report or another analysis costs almost nothing, there is very little pressure to ask if we actually need it.
For most of history, effort provided a natural constraint. If a report took three weeks to produce, someone probably asked if it was worth producing before committing three weeks of someone's time to it. If the same report can now be generated in twenty minutes, the easiest answer may simply be to create it.
The problem then moves from the person producing the information to the person expected to consume it.
If someone sends me a ten-page report, I will probably read it. If I start receiving several 60-page reports every morning, I am unlikely to read every page regardless of how quickly they were created, so the obvious solution is to ask AI to summarise them.
We have now created an interesting loop. AI helped someone create information that I do not have enough time to consume, so I use AI again to reduce it into something I do have enough time to consume. Somewhere between the original question and the executive summary, an enormous amount of information may have been generated, analysed and discarded without either of us understanding very much of it.
The cost of creating information has collapsed, but the cost of understanding it hasn't.
What happens when understanding becomes optional?
This is where I think the assurance gap becomes much bigger than having enough people available to review AI-generated work.
People develop judgement by doing things. Researchers learn by researching, writers improve by writing and editing, engineers develop judgement by designing, testing, failing and understanding why something failed, while experienced professionals in almost any field build an instinct for when something does not look quite right because they have spent years working through the underlying problems themselves.
A lot of that work is slow and sometimes frustrating, which makes it an obvious candidate for automation. I use AI for exactly that reason and have no desire to spend hours doing something manually if technology can help me achieve the same outcome in minutes.
The question is where we draw the line between removing unnecessary effort and removing the experience through which we learn the fundamentals.
An experienced professional using AI has something available to them that is difficult to include in a productivity calculation: experience. They can look at an answer that appears perfectly convincing and still question it because something does not fit with what they have seen before. They may not immediately know what is wrong, but they know enough to stop, investigate and challenge the answer rather than simply approve it.
What happens to that ability if the next generation of professionals increasingly learns and works in an environment where AI performs much of the underlying research, analysis, calculation and reasoning for them?
A student who uses AI to create a paper and another AI to check it may produce something impressive, but how much of the subject did they actually learn? A junior employee who asks AI to analyse a problem may produce something considerably better than they could have produced alone, but if they only read the summary, how much experience did they gain from solving the problem themselves? If this continues throughout a career, at what point do we expect them to develop the judgement required to challenge the output they are eventually responsible for approving?
We could unintentionally create a feedback loop where AI reduces the need to perform foundational work, fewer people develop expertise through doing that work, independent scrutiny becomes harder and we consequently become even more dependent on AI to perform the scrutiny.
Today, the problem may be that I know how to scrutinise something but do not have enough time to scrutinise everything I can now produce.
In the future, the bigger problem may be that even if I have the time, I no longer know how.
The high-stakes version
Medicine makes this easier to see because the consequences of getting something wrong can be much more serious, although the same principle applies in many other professions.
Imagine that an AI system produces a complex medical analysis and another system checks it, while a third summarises the findings because the complete analysis contains more information than a doctor realistically has time to review. The doctor reads the summary, considers the recommendation and approves it.
Technically, there is still a human in the loop, but I am not sure that tells us enough.
If the doctor understands the underlying medicine, knows the patient, can challenge the assumptions and has enough experience to recognise something unusual, AI could be an incredibly powerful tool. The doctor is not simply approving the output, they are applying independent judgement to it.
Now take the same model forward a generation and imagine a doctor who was educated, trained and subsequently worked in an environment where AI performed increasing amounts of the foundational reasoning throughout their development. If they have rarely had to work through the underlying problem themselves, how do they know when the answer in front of them deserves to be challenged?
Medicine and academic research already have substantial assurance processes around them for good reason. The question is what happens when the amount of material technology allows us to produce grows much faster than the human attention available to understand it, particularly if the same economic pressure encouraging us to automate production also encourages us to automate the assurance.
Oversight or authorisation?
I think this distinction may become increasingly important.
Human oversight and human authorisation are not the same thing.
Oversight means I understand enough about something to scrutinise it, challenge its assumptions and reject it for reasons of my own. Authorisation can simply mean that my name appears at the end of the process.
It is entirely possible to imagine a future workflow where humans remain at every important approval point while gradually losing the ability to independently evaluate what they are approving. A junior employee might use AI to create an analysis, their manager could use AI to review it, a director could use AI to compare the available options and an executive could use AI to identify the questions they should ask before approving the recommendation.
Everyone participated in the process and everyone may have acted perfectly reasonably, but it is still worth asking if anyone actually understood the complete chain of reasoning behind the final decision.
That concerns me more than the familiar discussion about AI replacing jobs, because keeping a person in a workflow is relatively easy. Maintaining the knowledge, time and experience that allows that person to provide meaningful scrutiny is much harder.
The assurance gap
I don't think the answer is to stop using AI, nor do I think human review is automatically better simply because a human performed it. AI will almost certainly become better at assurance as well as production, and there are many situations where it may find errors, inconsistencies or risks that a person would have missed.
The problem is assuming that increased production and increased assurance are the same thing.
For a long time, creating research, analysis, software and professional work required significant human effort. That effort never guaranteed quality, but it limited how much we could produce and forced people to spend time inside the work they were eventually expected to understand.
AI is removing much of that limitation at exactly the same time that organisations have strong incentives to increase output, reduce cost and move faster. The sheer volume we can create may push the people producing the work towards less scrutiny of their own output, encourage the people consuming it to rely on summaries rather than source material and make automated review increasingly attractive because human review cannot keep pace.
None of this requires anyone to decide that judgement no longer matters.
In fact, judgement may become more important than ever.
The problem is that AI may be making discernment and judgement more important at exactly the same time that the productivity expectations created by AI leave us less time to exercise them, while removing some of the foundational work through which we learned to exercise them in the first place.
The greatest assurance problem may therefore not be that AI produces more than humans can check, but that eventually fewer humans will know how to check it.
We can put a human at the end of every workflow and call it human oversight, but if that person no longer understands how the answer was created, cannot independently assess the reasoning and depends on another AI to tell them if the first AI was correct, I am not sure oversight is the right word.
It may simply be human authorisation.
If producing something becomes almost free while understanding it remains expensive, the most important question may no longer be how much AI allows us to create.
It may be how much of what we create anyone will still have the time, knowledge and judgement to question.
Read Really, Another Leadership Book?
A practical examination of trust, judgement, ownership and leadership for the moments when the situation is more complicated than the advice.