AI in business

Why Teams Trust Artificial Intelligence Too Much or Not at All

Some employees believe every answer the tool gives, others give up after the first mistake, and a four-point sheet of paper corrects both behaviours.

Baltimore, 2001. Peter Pronovost, an intensive care physician at Johns Hopkins Hospital, spends a month walking around the ward with a sheet of paper, recording whether the doctors who insert a central line, a thin tube threaded close to the heart to deliver drugs and fluids to the patient, perform five steps that protect against infection. They include washing hands and disinfecting the skin with chlorhexidine, things every doctor has known since the first year of medical school. After a month Pronovost counts, and finds that in more than a third of patients at least one step was skipped.

The solution he proposed looked like presumption in the hospital hierarchy: nurses were given the right to halt a procedure when a doctor skipped an item on the list. A person lower in the pecking order could now tell a specialist with years of experience that they had to start again. After fifteen months the rate of infections within the first ten days of a line at that hospital fell from 11 percent to zero, which by the hospital's estimate meant 43 infections avoided, 8 deaths avoided and about 2 million dollars saved.

In 2003 the same scheme was introduced in the state of Michigan. Across 103 intensive care units the median infection rate per thousand line-days fell from 2.7 to zero within three months, as reported in December 2006 by the New England Journal of Medicine, and over the first eighteen months of the programme it is estimated to have saved more than fifteen hundred lives. The most effective remedy turned out to be a sheet of paper together with permission for someone lower in rank to say "stop". Does a team that knows everything it needs to know still need permission to say it out loud?

Artificial intelligence (AI) tools, that is, programs that produce answers and recommendations based on patterns learned from enormous collections of text, set the same mechanism in motion in companies. People know the result has to be checked, yet some of them do not check it, while others drop the tool for good after the first mistake. Both behaviours have a psychological source, and both are corrected by a short procedure with a right to object, built on the model of Pronovost's list.

Why a fluent answer seems true

A language model, a program that adds one word after another based on patterns from billions of pages it has read, produces sentences that are smooth and confident whether or not their content is true. The human mind reads fluency as a sign of competence (psychologists call this the processing fluency effect), so a polished paragraph containing an error arouses less suspicion than a ragged note containing the truth. Add automation bias, described in the psychology of work: the more often a machine is right, the less attention we devote to checking its output, and after a few weeks of flawless operation the checking disappears altogether.

Strong evidence comes from a 2023 study in which researchers from Harvard Business School and other universities worked with Boston Consulting Group, one of the largest consulting firms in the world. It involved 758 consultants, about 7 percent of the firm's individual-contributor consultants, and they worked with GPT-4, a model from OpenAI. On tasks within the model's capabilities, the consultants completed 12.2 percent more tasks and did them 25.1 percent faster, and independent evaluators rated the quality more than 40 percent higher.

On a task deliberately set just beyond the model's capabilities the result reversed: consultants with the tool were 19 percentage points less likely to give a correct answer than colleagues who worked without it. The authors called this boundary jagged, because it runs through places the user cannot see: the model solves a task that looks hard without effort and stumbles on one that looks easy. For a manager the conclusion is very concrete, because an employee has no way of telling which side of the boundary they are on when the answer sounds equally convincing in both cases.

Why the same people stop using the tool after the first mistake

The opposite behaviour was described in 2015 by Berkeley Dietvorst, Joseph Simmons and Cade Massey in a paper titled "Algorithm Aversion". In their experiments, participants who saw a forecasting algorithm make a mistake relied on it less afterwards, even though its forecasts were on average more accurate than human ones. We explain a person's mistake by the circumstances, while a machine's mistake is remembered as proof of a design flaw.

The same group of researchers later published in the journal Management Science a result with practical value for companies: people are more willing to use an imperfect algorithm if they can modify its output, even slightly. A sense of agency rebuilds trust, and the ability to correct an answer changes the tool's role from an imposed supervisor into an assistant whose output is subject to editing.

What a checklist does and what it cannot do

Both mechanisms share the absence of a common, explicit rule for handling the tool's output. A checklist supplies one: it turns an individual feeling of "this looks fine" into an action that can be carried out and ticked off. It costs a half-hour meeting and a few extra minutes per document, and it requires no licence.

The limits are clear. A checklist will not tell an employee whether a task lies on the right side of the jagged boundary, because that can only be learned by testing on examples where the correct answer is known. Ritual is a risk as well: after a month people tick the boxes without looking, exactly as Pronovost's doctors skipped steps they knew by heart. That is why the second element of the solution is the right to object, meaning permission for every team member, interns included, to halt a document in which a check was skipped.

The boundary can, however, be probed in-house, and cheaply. All it takes is a set of a dozen or so tasks from the team's past work for which the answers are known (old complaints with their resolutions, contracts with marked clauses, spreadsheets with a calculated result), run through the tool, preferably in the version the team actually uses, because after every model update the boundary shifts in an unknown direction. Two hours of such testing give more than any vendor presentation, because they show the errors on material that employees know and can judge for themselves.

Four questions for every document prepared with the tool

The procedure below fits on one sheet and applies to every text or table that leaves the team:

  1. Sources: every figure, name, date and regulation has a source indicated, and two randomly chosen items have been checked against the original.
  2. Scope of the task: the responsible person has written in one sentence whether the tool was doing something the team had already tested on an example with a known answer.
  3. Input data: no personal data of customers or employees and no confidential contract provisions were entered into the tool.
  4. Signature: the name of the person who read the whole document, not only a summary, and who can stop it from being sent.

The first step for Monday morning

Choose one process in which the tool's output reaches a customer or feeds a decision about people: a sales offer or a candidate assessment. Gather five people who work with it, including the most junior one, and spend half an hour rewriting the four questions above in your own words, adapted to that process. Name the person with the lowest rank in the group as the one with an explicit, written right to halt a dispatch, and ask the manager to confirm it publicly in front of the whole team.

For the following month keep a shared spreadsheet in which every error of the tool gets one row: what was wrong, whether it was caught before dispatch and how many minutes the correction took. The spreadsheet answers a question that would otherwise remain a guess: does the tool get it wrong in two cases out of a hundred, or in twenty? The figure from the spreadsheet then replaces the memory of one loud failure, which after being seen usually decides the assessment of the entire tool.

After a month the team will know more than most departments know today: in which kinds of tasks the tool works without reservation, and in which it always needs a second pair of eyes. The list stops being a general procedure and becomes a map that can be shown to a new employee on the first day, together with a record of who has the right to say "stop". Does a person with that right exist in your company today, and does their manager know about it?