Judgment Happens in the Pause. We Spent Twenty Years Filling It.

Checking an answer and judging it are two different acts, and only one of them gets faster. This is the twenty-year story of how we stopped doing the second one.

A row of tall charcoal blocks across a warm off-white ground, like a filled schedule, with one slot open near the center where a small golden square sits lower than the rest and glows softly onto the paper beneath it
Listen to this article
0:00
0:00
Listen on: Spotify Apple Podcasts YouTube
?Question

Why do executives struggle to overrule AI output they suspect is wrong?

Quick answer

Not because the AI is persuasive. Because the habit of overruling was trained out of them, over about two decades, by a management culture that taught organizations to trust the measurement over the read.

bosio.digital
The Founders’ AItrained on 25 years of our work

Want this made concrete for your company? Ask me, or pick one:

AI did not start that. It is the fastest step in it. The specific thing AI removed is the interval in which a person would have stopped and asked what the answer was actually saying, and a model published this month in the Academy of Management Review argues that interval is where managerial practical wisdom is built in the first place.

The reflex that went quiet

Nobody took your judgment.

You handed it over, one reasonable decision at a time, and every time you did it you were rewarded for it.

That is the uncomfortable part. Not that AI showed up and overwhelmed a generation of executives. That the ground was prepared first, carefully, by people who meant well. A lot of them were consultants. Some of them were us.

An executive reads an AI-generated analysis of a business they have run for a decade. Something in it does not match what they know. They approve it anyway.

Not because the machine was persuasive.

Because they could not produce a defensible reason fast enough, and a view you cannot put on a slide has been the wrong kind of view for a long time now.

That is not a technology failure. That is a trained response.

We do not have a machine problem. We have rooms full of capable people who will not overrule a confident answer.

And the whole industry is looking somewhere else. Hallucination rates. Benchmarks. Evaluation harnesses. Guardrails. Every bit of it aimed at making the output more trustworthy. None of it aimed at the person who already suspected the output was wrong and signed it anyway.

You cannot fix that with a better model.

65local logic errors found in 570 hand-classified code review comments (Bacchelli and Bird, ICSE 2013)
6design-level problems found in those same 570 comments (Bacchelli and Bird, ICSE 2013)
92percent of data and AI leaders who say the barrier is people, not technology (AI and Data Leadership Executive Benchmark Survey, 2025)

Checking an answer and judging it are two different jobs.

Only one of them gets faster when you add a machine.

And the place the second one used to happen has been quietly filled with other work.

Checking and judging are not the same job

Your team is checking more AI output than it has ever checked in its life.

That is true, it is measurable, and it is not the good news it sounds like.

Because “checking” is one word doing the work of two jobs, and the two jobs are nothing alike.

The first one is catching this error, in this output, now. It is local. It is tactical. It gets faster with practice and it gets better with attention, and it is the only one anybody counts.

The second one is seeing why this kind of error keeps happening. What produced it. What it says about the thing generating your work. That one does not get faster. It cannot be done from inside the flow of individual items, because the whole point of it is standing somewhere else.

Call them detection and diagnosis, and notice which one your company measures.

Judgment happens in the pause.

In January, Shelly Palmer (a technology consultant and columnist who writes a daily newsletter on AI) built himself a verification system: an agentic workflow and an AI critic running a thirty-five page rubric that checked accuracy, grammar, structure and sourcing. He pointed it at a draft about Apple, Google and Siri. It verified every claim in the piece.

It also got the most important one backwards. The draft said Apple pays Google roughly twenty billion dollars a year. Google pays Apple.

His own explanation is the sharpest sentence written about AI verification this year. The system had done entity matching. It confirmed that “twenty billion,” “Google,” “Apple” and “search” all appeared together in the sources. It never asked who was paying whom. These tools, he wrote, “excel at confirming that components exist in sources” and “struggle to confirm that relationships between components are correctly represented.”

What caught it was a person who knew about the deal.

Not somebody checking faster. Somebody holding the whole picture.

Verification was turned all the way up. Automated, exhaustive, thirty-five pages of criteria, every claim examined.

And the thing that mattered was still wrong.

Because confirming that a fact appears and knowing what a fact means are two different operations, and only the first one scales.

The point

Checking frequency and checking depth are different variables. Most organizations optimize only the first one.

If this sounds like a new problem, it is not.

A management researcher named Chris Argyris drew this exact line in 1977, and gave it the best image anyone has managed since. A thermostat notices the room is cold and turns the heat on. That is one kind of learning. A thermostat that could ask whether it should be set to sixty-eight degrees at all would be doing something else entirely.

One corrects the error. The other questions the policy that produced it.

His finding is the one that should worry you: organizations are quite good at the first kind. Good at the shallow one. Studying six corporate presidents, he concluded they “were under the illusion that they could learn, when in reality they just kept running around the same track.”

Activity that feels like learning. Measured as learning. Not learning.

And when somebody finally counted the ratio, it was worse than anyone guessed. Two Microsoft researchers hand-classified five hundred and seventy code review comments, inside an organization where forty thousand engineers review each other’s work every day with purpose-built tooling. Of the comments that found defects, sixty-five were local logic errors.

Six were design problems.

Their own summary is that these reviews mostly catch “micro level and superficial concerns,” while the people doing them expected something better. Their interviewees were blunter than that. One senior engineer described watching reviewers flag formatting while missing security holes in the same file. Another admitted that surface commenting is how you avoid doing a real review.

That study is from 2013. It has nothing to do with AI.

AI did not invent this. It raised the volume of things to be reviewed and left the failure precisely where it already was.

The AI Briefing

Tuesdays. 500+ leaders. No hype, just what works.

Twenty years of practice

So where did the second job go?

It was trained out. Deliberately, by people with good reasons, and the reasons were good enough that most of us helped.

Data-driven decision making worked. When researchers looked at a hundred and seventy-nine large public companies in 2011, the ones that adopted it showed output and productivity five to six percent above what their other investments predicted. That is a real result and it deserved to win the argument it won.

What traveled alongside the evidence was a culture, and the culture went considerably further than the evidence did.

Around 2006, product organizations started saying HiPPO. The highest paid person’s opinion. It was not a compliment.

That word did not argue that senior judgment was wrong. It made senior judgment embarrassing, which is far more efficient, because nobody has to win an argument nobody is willing to start. A generation of managers learned that a view the dashboard could not support was a view to keep to yourself.

Then the measurement moved from the room to the person. Accounting researchers gave this its name in 2012 and called it surrogation: managers know a metric is an imperfect stand-in for the thing they actually care about, and then behave as though the metric is the thing. Their experiments found that paying people on a single measure makes it worse.

That is not a reporting problem. That is a person quietly agreeing to want what is measurable instead of what is true.

179large public companies studied in 2011; the data-driven adopters ran five to six percent ahead (Brynjolfsson, Hitt and Kim)
2006the year product organizations started saying HiPPO, the highest paid person's opinion (Kohavi's account)
2012accounting research names surrogation: managing the metric as if it were the thing (Choi, Hecht and Tayler)

I have watched this happen for twenty years. Two decades of it is not a neutral stretch of time. It is practice.

And what it prepared the ground for is not an upgrade. This replacement of oversight, of judgment, of the faculty itself, is sold as an improvement on the thing it displaces. It is not even a full resemblance. It is a shell of what the thing was meant to be.

AI is not asking you to believe something. It is asking you to hand something over. Your judgment. Your critical thinking. Your right to decide what is true.

Read that and the honest reaction is: nobody is asking me that.

Nobody had to ask. It has been handed over in installments for more than two decades, and this is the first thing to take delivery all at once.

That is what we mean by Humans First AI. Not a preference for people over machines. A position on where the deciding happens.

I found the same argument in my own writing before I knew it was this argument. Earlier this year, writing about AI acceptable use policies, I ended up at a line I did not examine closely enough at the time: “A policy is not enforced by a document. It is enforced by people deciding, under pressure, whether the honest path is worth taking.”

That reads like a sentence about policy. It is a sentence about the interval.

Every control your company has ever written comes down to one person, at one moment, with time running, choosing to look harder than they had to.

The strongest case against this argument is structural

There is a competing account of all this, and it is the one most consultants are currently selling. It is also good. Good enough that it deserves its strongest form before anyone says what it misses.

It says the problem is the shape of the organization, not the state of the people inside it.

Palmer made that case in June, in another essay, “The Vice President of Electricity.” Electricity became commercially available in 1881. It did not show up in manufacturing productivity until the early 1920s. The technology worked the entire time. The organizations around it never changed shape.

He is right about the shape and right about the remedy. If your org chart, meeting cadence, decision rights, approval loops and performance reviews were all built for people who think at human speed and run one thread at a time, then adding a model to that is adding horsepower to a machine designed for less of it.

The history is more specific than the retelling usually is, and better. What factories actually did, for roughly twenty-five years, was buy electric motors and use them to turn the same overhead line shafts and leather belts the steam engine had turned.

The power source changed. The building did not.

The gain arrived only when somebody put a motor on each machine. That meant the ceiling no longer had to hold up the drivetrain, which meant the factory could be one story instead of several, laid out in a line, rearranged without shutting the plant down.

Twenty-five years of installing a new capability into an unchanged idea of what a factory was for.

So grant the structural case completely. Then notice what it cannot reach.

You can redesign every workflow in the building and still have a room full of people who will not overrule a confident wrong answer.

Every constraint on that list is one you could photograph. Who reports to whom. Who signs off. How often people meet. All of it real, all of it worth fixing. Not one item on it is about what happens inside the person making the call, in the four seconds before they decide the output looks fine.

The leather belts were a design problem.

This is not.

The point

You can redesign every workflow in the building and still have a room full of people who will not overrule a confident wrong answer.

Three objections worth taking seriously

These cut against everything above. The piece is not worth much if it dodges them.

Maybe executive judgment was never any good. Daniel Kahneman spent a career showing expert intuition is overconfident. Gary Klein spent a career showing it is real. In 2009 they wrote a joint paper working out where each of them was right, and landed on a test: intuition can be trusted only in environments predictable enough to learn, where the person has genuinely had the chance to learn them.

A CEO’s gut about a new market fails that test badly. So on this reading, data culture was a correction, it worked, and this article is nostalgia.

Apply the test honestly and it says something different. An executive guessing at an unfamiliar market should defer. An executive reading an AI analysis of an operation they have run for eleven years, and had feedback from every week of it, is in exactly the environment where the test says trust the read.

That is the judgment that got trained away. A correction aimed at the first case was applied to the second, and nobody noticed the difference.

Maybe more checking is a healthy sign. Amy Edmondson found that hospital nursing units with better management and stronger relationships reported more errors, not fewer, and concluded this reflected their reporting and not their error rate. Read that way, verification going up is a workforce getting more vigilant, and calling it decay is reading a good signal backwards.

What separates the two is where the checking goes. Edmondson’s units reported upward, into a system that could see the pattern and act on it. Correction that moves sideways and informally, between colleagues, under deadline, never reaches anyone who could see a class of error at all.

Reporting builds a record. Absorption destroys one. Same behavior, opposite consequence.

And the one that changed what this article claims. Diagnosis does not always need a pause. Toyota engineered it into the moment: deviations made instantly visible, responses run as real-time experiments by the person at the workstation, supervisors whose actual job is asking why.

Their best illustration is small. Workers trying to cut a changeover from fifteen minutes to five got it to seven and a half, and a manager asked why they had missed. The question surfaced that the five-minute target had been a guess with no reasoning behind it, which is exactly why nobody could work out what went wrong.

That objection held for a long time, and then it turned into the best evidence in the article.

Because the andon cord is a pause. It is a signal-triggered stop, wired into a factory floor, available to anyone who sees something going wrong. Toyota did not eliminate the interval. Toyota gave it infrastructure.

Which means the question was never whether the pause is necessary. It is whether anybody builds for it.

Almost nobody has. And Argyris, working from roughly three thousand cases, found the deeper loop was already rare in 1977.

This is not a story about a golden age. It is a story about something thin that is now gone.

Where this goes next

Not sure where you stand?

Take the 90-second AI readiness read: five dimensions, a scored result, and a clear next step.

Take the readiness read →

Two pauses, and only one can be installed

There are two of these and they are not the same thing, which is why most of the advice about them is useless.

The first is triggered by the decision. It sits in front of anything consequential, it is scheduled, and it is a control. Nobody signs until somebody has produced an objection. You can write that into a process this quarter and it will hold.

The second is triggered by a signal. The work is coming apart. The same problem keeps arriving in slightly different clothes. You have run at it four times and it is no sharper than it was on the first pass.

That is when you stop. Not to rest. To find out why the thing will not resolve, which is a different question from the one you have been asking, and you cannot reach it from inside the fourth attempt.

You cannot schedule that one. Nobody knows in advance which Tuesday it lands on.

And here is the part that should worry you, because it is the whole argument in one turn: noticing that a problem has stopped clarifying is itself an act of judgment.

The capacity you need to recognize you need the second pause is the capacity the second pause was going to restore.

Noticing that a problem has stopped clarifying is itself an act of judgment.

That is circular and it is also exactly what is happening. It is why this degrades quietly instead of announcing itself, and it is why nothing in your reporting will catch it. A team that has lost the second pause does not file a ticket. It just keeps going, productively, into the same wall.

The first pause is what a process can give you. The second is what the first one is buying time to rebuild.

The gap is still on your calendar

Nobody deleted the interval. Look at any executive calendar and the gaps are still there.

They are just occupied now.

But the gap is the symptom, not the problem. What actually changed is your relationship with time. The gap is only where you can see it.

An eight-month Harvard Business Review field study inside a two-hundred-person technology company found workers filling exactly the moments that used to be empty. Prompting over lunch. Prompting in meetings. Prompting while a file loaded. One more query before leaving the desk. Some of them noticed, in hindsight, that downtime had stopped restoring them.

The day did not lose its gaps. Its gaps acquired work.

And there is a meta-analysis that says precisely what that does. Pooling a hundred and fourteen effect sizes, researchers found the incubation effect is real, with one hard condition: fill the gap with a cognitively demanding task and most of the benefit disappears. Fill it with something undemanding and it survives intact.

The interval stayed on the calendar and stopped working in the head.

Which is why the answer is not “take more breaks,” and why anyone who tells you it is has not read the research they are gesturing at. An undesigned pause is not the intervention. A protected one is.

There is a word for what is being lost here, and it is considerably older than any of this.

Phronesis. Aristotle’s term for practical wisdom: not knowing what is true in general, but knowing what this particular situation calls for. It is the thing you cannot get from a book and cannot download, because it is built out of having been wrong before, in this business, with these people.

A model published this month in the Academy of Management Review put that word back into management research, and it is the first serious treatment of AI deskilling aimed at managers instead of at radiologists. Its authors define it as the practical wisdom managers develop through real-world experience, reflection and human interaction. Their argument is that generative AI can erode it, and their condition is the one this whole article has been circling: it happens most when managers are under intense time pressure and reach for AI as a shortcut.

The risk, in their framing, is that managers stop asking important questions, stop seeking other perspectives, and stop learning from the room.

That is not a productivity concern. That is a faculty going quiet.

It is a theoretical model with no sample and no method, which is what that journal publishes and what it should be called. But its prescription is not what you would expect from a paper about wisdom. It is structural. Organizations, the authors write, “need to carefully design roles, responsibilities and workflows to ensure employees continue developing the human skills that AI cannot replicate.”

Redesign the workflow. Not so the machine runs faster.

So the person stays capable.

Putting the friction back

The idea of deliberately adding friction back is not new and I am not going to pretend it is. Forbes covered “friction-maxxing” in February. There is a workshop on frictional AI now in its third year, running alongside an international conference, with organizers at Cambridge, Radboud, Villanova and VU Amsterdam. Their premise is deliberate design choices that create moments of reflection at the expense of speed.

The conversation exists. What is thin is the reason people give for it.

Almost everyone arguing for friction is arguing from overload. People are drowning, give them air. That is true and it will not survive contact with a difficult quarter, because anything filed under employee experience sits on a cost line.

Friction defended as the place judgment happens is a control.

Controls survive bad quarters. Bad quarters are exactly when somebody ships a confident wrong answer.

So the design looks like this.

Five ways to rebuild the interval
1
Answer first, then lookFor any decision that matters, write your own read before you see the model’s. Not after. People accept a confident recommendation far more readily when they have not yet formed a position, including when it is wrong.
2
Require the counterargumentNo major decision is final until somebody has produced one genuine objection and one explicit link to the goal. Not a devil’s advocate ritual. A named person, a written objection, in the record.
3
Log what kind of wrongWhen AI-assisted work comes back wrong, record the type. Reversed relationships. Invented specifics. Confident gaps. Plausible but stale. After a month you are no longer fixing outputs. You are looking at a pattern, which is the only vantage point from which anything gets fixed at the source.
4
Do not respond to urgencyA subtraction, and the hardest one. A protected gap only works if nothing demanding goes into it, which means the pause before a decision cannot also be the moment somebody clears their messages. Urgency will take it back the week after you install it unless somebody is accountable for defending it.
5
Stop when it stops clarifyingThe signal that you need the second kind of pause is that repetition has stopped producing sharpness. Same problem, fourth attempt, no clearer. That is the moment to stop and ask why it will not resolve, and act again when it does.
↻  Review what the checks caught, quarterly. If the answer is only small things, the loop is running shallow and the design needs changing.

Somebody will read that list as bureaucracy. More gates, more process, slower everything.

It is the opposite. Structure is what makes flexibility possible. A team with no settled ground cannot move quickly, it can only react quickly, and those look identical right up until the moment they do not. Take the structure away and what you get is not freedom. It is churn.

None of that is contemplative practice dressed up for the office. It is closer to something we published a while ago, in the mindful prompting framework: “AI’s greatest value comes not from doing our thinking for us, but from creating space for more profound human thought.” The third pillar of that framework is reflection after the interaction. We wrote it down before we had the evidence.

The evidence has now arrived from four directions at once.

There is also a machine half to this, and my colleague documented it. AI systems are built, structurally, to agree with you, and the research on sycophancy and executive judgment shows how invisible that pressure is to ordinary monitoring.

That is the machine leaning toward you. This article is about the person who has stopped leaning back.

Both have to be true for the failure to happen. They usually are.

Where an outside partner fits

There is a reason this is difficult to do from inside, and it is not a reason about capability.

Diagnosis requires someone who is not in the stream. Everyone inside your company is in the stream. The person best placed to notice that your decision architecture has quietly stopped producing decisions is the person whose calendar is the reason it stopped, and they are the least able to see it, because from inside it does not look like a missing interval. It looks like a full week.

There is a second reason, and it is less comfortable.

Every company has a short list of decisions nobody questions. Not because they are correct. Because questioning them costs something, and everyone learned what it costs a long time ago.

Those are exactly the decisions where the interval went missing first.

Nobody on your payroll is going to raise them. Not out of cowardice. Because they are the ones who have to keep working there on Monday.

That protection does not apply to someone who does not work for you. It is most of what you are actually buying, and it is the work we do in the CEO AI program, which is about decision architecture and only incidentally about tooling.

Ask what you own when the engagement ends. Ask us, ask anyone else you are considering.

If the answer is a report about your decisions, you have bought a diagnosis and you will be buying it again next year. If the answer is a working loop your own people run, with named owners and a review cadence that survives our leaving, you have bought the capability.

Ask it in the first conversation. The answer is usually available in about a minute, and it tells you more than the case studies will.

If you would rather not have that conversation at all, the honest alternative is not nothing. It is to run the five steps above yourself, badly at first, and pay attention to what the first month of error logs tells you. Most of the value in this is in looking, and looking is free.

Start Building

Find Out Where Your Pauses Went

Paste this into your AI assistant. It makes you commit a prediction first, then interviews you about one real decision, then tells you where your prediction and the evidence disagreed.

Prompt · paste into your AI

Context: I want to find out where the deliberation has gone out of our decisions. Interview me one question at a time and wait for my reply before asking the next. No more than three questions per step. Push back when I describe a process we aspire to instead of one we actually run, and when I answer with a policy instead of a behavior. If an answer of mine is vague, ask once more in a more concrete form, then move on and note that I could not answer it.

First, your prediction: Before any questions, ask me to write two or three sentences answering this: when an AI-assisted decision goes wrong in my company, what kind of wrong is it usually, and what do I think is causing that. Do not comment on it, do not correct it, do not let it steer your questions. Set it aside until the end.

Step 1. One real decision: Ask me to name one consequential decision from the last quarter where getting it wrong would have cost real money, time or a relationship. Then ask whether it went right or wrong in the end, and what specifically was wrong with the first version. Then ask how many times anyone sat down with it as the only thing they were doing. If the honest answer is once, accept once.

Step 2. The gate: Ask whether there was a point where someone other than the author had to approve it before it went out, and whether that gate actually fired on this decision. If it did, ask what it cost the person to use it. Then ask what would have caught the problem if that person had missed it.

Step 3. What the checking catches: Ask me for three real things our reviews caught recently. Specific instances, not categories. Then ask me what those three have in common. Do not compute a ratio from them, because three remembered examples are not a count and recall favors the memorable ones. Say that to me plainly if I reach for a number.

Step 4. The problem that will not resolve: Ask me to name something we have gone at more than three times that is no clearer now than the first time. Give me this test rather than an example: it qualifies if I attacked it at least three separate times, believed each time that I had addressed it, and it came back in a different shape. Tell me the tell is that I have stopped calling it a problem and started calling it a fact about my business. If I ask you for an example, do not give me one from my company, because I will evaluate yours instead of finding mine. Give me shapes from unrelated businesses instead: a report that keeps getting rebuilt and still goes unused, a handoff that keeps breaking regardless of who runs it. Then ask what question I have been asking about my problem, and whether anyone has stopped long enough to ask a genuinely different one.

Output, in four parts: One, a map of that decision path marking every point where deliberation could have happened, whether it did, and what occupies that space now. Two, what my three examples have in common, stated as the kind of failure it is. Three, the unresolved problem, the question I have been asking about it, and whether my alternative questions were genuinely different or the same question wearing different clothes. Four, my prediction from the beginning, quoted back, with a specific account of where it matched what surfaced and where it did not.

That fourth part is the one to read twice. The gap between what you predicted and what the interview turned up is the most accurate measure you will get of how much of your own judgment is still in contact with your business. See where you stand →

The stand

You do not have a technology problem. You have a room full of capable people who have been quietly taught, over twenty years, that the confident answer on the screen is more trustworthy than the read in their own head. AI did not do that to them. AI is just the first thing to make the cost of it visible within a single quarter.

The interval where your best people would have caught it is not missing from the calendar. It is full. And nothing in your reporting is going to tell you that, because a full week is what a healthy week is supposed to look like.

So find one decision that matters this month. Put a real gap in front of it. Make somebody responsible for keeping that gap empty and somebody else responsible for producing one honest objection before it closes.

Then watch what your people catch.

And when the work starts coming apart, when the fourth attempt is no sharper than the first, stop. Not to rest. To find out why it will not resolve. Move again when it does.

That is not a productivity initiative and it is not a wellness benefit. It is the difference between an organization that fixes what went wrong and an organization that understands why it keeps going wrong. Only one of those compounds.

Intelligence is not a set menu. You are not choosing from what it offers. You are shaping what it becomes, every time you accept something you should have sent back.

Which means you are not using this system. You are training it.

So build from what you can stand behind, not from what the quarter is doing to you. Trust what you know over what the system returns. Then use the system to give what you know a shape.

Your people are not the constraint. The missing moment where they get to be useful is the constraint. Give it back to them and pay attention to what happens.

If you want help finding where it went, that conversation is worth having.

Sources

Frequently Asked Questions

Does using AI actually make executives worse at judgment?

The strongest evidence available is a theoretical model published in the Academy of Management Review in 2026, which argues that generative AI can produce what its authors call epistemic deskilling in managers, most likely when they are under intense time pressure and use AI as a shortcut. It is a conceptual model with no sample. The empirical evidence for skill loss after AI exposure comes from other professions, notably a 2025 observational study of endoscopists whose unassisted detection rates fell after regular AI use, and that study’s own authors describe deskilling as a partial explanation and not a proven mechanism.

Our team checks AI output constantly. Isn't that enough?

It depends what the checking catches. Frequency of review and depth of review are different variables. In a 2013 study of 570 code review comments at Microsoft, the reviews that identified defects found 65 local logic errors and 6 design-level problems, in an organization with mature review tooling and strong review culture. High volume, almost no structural diagnosis. The useful question is not how often anyone checks. It is whether a single named person is looking at what the errors have in common, and whether that person is you.

What is a decision pause, in practical terms?

The version proposed by the Harvard Business Review researchers who studied AI intensification is specific: before a major decision is finalized, require one genuine counterargument and one explicit statement of how the decision connects to an organizational goal. It is not delay. It forces the frame wider than the item in front of you, briefly, at the moment where getting it wrong is expensive. That is the scheduled kind. There is a second kind you cannot schedule, triggered by a signal instead of a decision: the work has stopped clarifying, the same problem keeps returning, and the useful move is to stop and ask why it will not resolve.

Isn't this just an argument for trusting your gut over the data?

No, and the distinction matters. Kahneman and Klein established in 2009 that intuition is reliable only in environments predictable enough to learn, where the person has had real opportunity to learn them. An executive’s instinct about an unfamiliar market fails that test. An executive’s read on an AI-generated analysis of an operation they have run for years passes it. The argument is not that judgment beats data. It is that a correction aimed at the first case was applied to the second.

Why can't we fix this internally?

Some of it you can, and the five steps in this article are all things a company can run without help. What is hard from inside is seeing the norms that made the pause disappear, because those norms are usually invisible to the people inside them. Chris Argyris described the mechanism in 1977: organizations develop a rule against questioning leadership’s chosen direction, and then a second rule against questioning the first. That protection does not bind someone who does not work there, which is the actual reason outside help is useful on this particular problem.

How is this different from adding friction for wellness reasons?

The design can look identical. The justification determines whether it survives. Friction defended as an employee-experience measure sits on a cost line and gets removed in a difficult quarter. Friction defended as a control on decision quality is treated like any other control, which means it survives exactly the period when it is most needed.

What should we measure to know if this is working?

Not the number of reviews and not the volume of AI use, both of which will rise regardless. Look instead at the ratio between problems caught in an individual item and problems caught in the pattern across items. If after a quarter every finding is still item-level, the loop is running shallow and the design needs changing and not more effort.

Where this goes next

Want this scored against your business?

AI Strategy turns this into a prioritized roadmap: where AI pays off for you, and in what order. It grows into the CEO AI Program.

See AI Strategy → Not sure where to start? Take the 90-second readiness read →
Sascha Laura

Say hello.

A 30-minute conversation. If we're not the right fit for where you are, we'll tell you, and point you somewhere better.

Join 500+ leaders The AI Briefing · Tuesdays · no hype
bosio.digital · AI Transformation That Elevates Human Talent · © 2026 Bosio Inc. · SF · Lake Arrowhead