EvolutionLabBlog

Can Judgment Be Trained? What the Research Actually Says

8 min read
Can Judgment Be Trained? What the Research Actually Says

Can Judgment Be Trained? What the Research Actually Says

Yes, within limits that matter. Controlled experiments show a single training session reducing specific decision errors, with effects that hold for months and carry into unrelated real tasks weeks later. The training that works involves practice on realistic decisions with feedback. Being told about your biases does close to nothing.

The evidence at a glance

StudyWhat it measuredWhoResultLimitation
Morewedge et al., Policy Insights Behav. Brain Sci., 2015Confirmation bias, bias blind spot, fundamental attribution error and others, before and after one sessionTwo longitudinal lab experimentsBias down at least 31.9% immediately and at least 23.6% at two months; interactive practice beat videoScale-based measures rather than workplace outcomes
Sellier, Scopelliti & Morewedge, Psychological Science, 2019Choice on an unannounced business case modeled on the Challenger launch290 graduate students across three professional programsTrained participants 19% less likely to choose the inferior hypothesis-confirming solutionOne case, one bias, student population
Heerma van Voss et al., Scientific Reports, 2025Confirmation bias and bias blind spot before and after a 40-minute session56 national risk analysts and 116 master's studentsBoth groups improved, with no significant difference between experts and novicesSmall expert sample, immediate retest only
Kahneman & Klein, American Psychologist, 2009When expert intuition can be trustedReview and synthesisTwo conditions: valid cues in the environment, and prolonged practice with rapid feedbackTheoretical synthesis rather than an experiment

What do researchers mean by "training judgment"?

In this research, judgment means specific, repeatable decision errors with validated measurement scales: testing a hypothesis with evidence that could only confirm it, reading a person's behavior as character while ignoring the situation, rating yourself as less biased than your colleagues. Those errors show up in a pretest and can be measured again after an intervention. When a paper reports that judgment improved, that is what moved. Nobody has run an experiment on wisdom.

That narrowness is a feature. It makes the claim falsifiable, and it is the reason the evidence below is worth more than any amount of confident writing about better thinking.

Does a single training session change anything?

A single session produced measurable change that lasted two months. Morewedge and colleagues, in two longitudinal experiments published in Policy Insights from the Behavioral and Brain Sciences in 2015, found that one interactive training session cut measured bias by at least 31.9% immediately and at least 23.6% at the two-month retest.

The design was simple. Participants took a bias pretest, then received one intervention: either a 30-minute instructional video or a computer game that made them commit to judgments and gave personalized feedback on each one. Then they were tested again, immediately and two months later.

The video did less, at 18.6% immediately and 19.2% two months on. A third laboratory experiment isolated why: the feedback and practice the game provided accounted for its advantage.

01 practice vs explanation

The effects were domain-general. Bias dropped on problems set in contexts the training never touched and in formats it never used. That last point is the one worth holding on to, because transfer is where most training research goes to die.

Does it hold up outside a lab?

The effect survived contact with a real task in one carefully built test. Sellier, Scopelliti, and Morewedge, writing in Psychological Science in 2019, gave 290 graduate students a one-shot debiasing intervention, and those trained beforehand were 19% less likely to choose the inferior hypothesis-confirming solution to an unannounced business case.

The design is what makes it worth quoting. Course scheduling, not the researchers, decided who got the training before the case and who got it after. The case was modeled on the decision to launch the Space Shuttle Challenger, it arrived without warning inside a normal class, and nobody connected it to the training. Written case analyses from trained participants contained less confirmatory hypothesis testing, which is the mechanism the authors credit.

A note on that number, since you will find a different one in circulation. The article as first published reported 29%. A corrigendum followed in 2020, and the figure on the journal's current abstract page is 19%. Most secondary coverage quotes the original.

Does it work on experienced professionals?

Experience neither protects people nor makes the training redundant. An experiment published in Scientific Reports in November 2025 tested more than half of the national risk analysts of a European country alongside a matched sample of master's students, and both groups showed less confirmation bias after a 40-minute intervention, with no significant difference between professionals and novices.

These are people whose job is forecasting war, pandemics, and infrastructure failure for their government, which is why the sample took years to assemble. The intervention was a scripted explanation of decision biases followed by a group case. Both groups showed confirmation bias at pretest. Both showed less of it after. The authors had predicted the professionals would benefit less, and they were wrong.

03 experience gap

Two findings from that paper deserve their own line. The analysts entered with less confirmation bias than students, inside their domain and outside it, which cuts against the older literature on expertise. And bias blind spot, the belief that you are less biased than average, showed up in both groups and correlated with nothing. Not with intelligence, not with education, not with how biased participants actually were.

The sample of analysts was small, 56 valid responses, and the measures were scales rather than live forecasts. The authors say so.

When does experience produce judgment, and when does it produce confidence?

Two conditions decide it, and they come from the paper Kahneman and Klein wrote together in American Psychologist in 2009 after years of disagreeing in print. Skilled intuition requires an environment regular enough for valid cues to exist, and it requires prolonged practice with feedback that is rapid and unambiguous. Where both hold, experience builds real pattern recognition. Where either fails, experience builds the feeling of expertise without the accuracy.

Apply that to reviewing AI output and the problem gets sharp. The environment gives cues that are hard to read, because fluent text carries no signal about whether it is right. Feedback is slow when it arrives at all: a flawed analysis that gets approved rarely comes back with a label on it. Those are precisely the conditions under which years on the job fail to produce calibrated judgment, and under which people get more confident anyway.

02 two conditions

What this means for your team

The research points at a short list.

  • Deliberate practice beats explanation. The intervention with feedback and repeated decisions outperformed the one that only explained the concepts, in the same experiment, on the same measures.
  • Awareness is not the mechanism. People who rate themselves as unbiased are not less biased. Training people to recognize a bias in a slide deck does not mean they will catch it in their own work.
  • Test transfer, not recall. If your training measures whether people remember the material, you have measured the wrong thing. The studies above are credible because they measured performance on tasks nobody connected to the training.
  • Effects decay. The two-month follow-up held, at a lower level than day one. Anything with a certificate at the end and no reinforcement afterward is unlikely to survive the quarter.

Two honest limits. The intervention research covers specific biases, mostly confirmation bias and bias blind spot, and specific tasks. It does not license a claim that any program makes anyone a better decision-maker in general. And the effect sizes are meaningful without being transformative. What the evidence supports is that the thing is trainable, that practice with feedback is the active ingredient, and that the alternative most companies buy is the one that performed worse in a head-to-head test.

Evolution Lab builds on that base. Our training puts people through realistic decisions with feedback rather than through explanations of what good thinking looks like, because that is what the transfer evidence supports. The research behind our method goes into more detail on what we take from each study, and where the evidence runs out.

FAQ

Can judgment really be trained, or is it fixed? Experimental evidence shows measurable improvement on specific decision errors from a single training session, holding at two months and transferring to unrelated tasks. The evidence covers particular biases rather than judgment in general.

Does bias training work on experts, or only students? A 2025 experiment in Scientific Reports on national risk analysts found the same size of improvement in professionals as in students. The analyst sample was small.

Why do most corporate training programs fail to change behavior? The comparison inside the 2015 experiments is instructive: explanation alone produced smaller effects than practice with personalized feedback. Programs built entirely on watching and listening are running the weaker arm of that study.

Does knowing about cognitive biases help? Not by itself. Bias blind spot, the tendency to see yourself as less biased than others, appears in experts and novices alike and shows no correlation with how biased people actually are.

How long do the effects last? Bias reduction of at least 23.6% held two months after a single session in the 2015 experiments. Longer horizons without reinforcement have not been established.