A cognitive bias is a systematic, predictable error in judgement produced by the mental shortcuts people use to make decisions under uncertainty. The modern research programme begins with Amos Tversky and Daniel Kahneman's 1974 paper in Science, which described three heuristics — representativeness, availability, and adjustment from an anchor — and concluded that these heuristics are highly economical and usually effective, but that they lead to systematic and predictable errors. Kahneman received the 2002 Sveriges Riksbank Prize in Economic Sciences for integrating psychological research into economic science; Tversky had died in 1996 and the prize is not awarded posthumously. Three terms are routinely conflated. A heuristic is the shortcut itself, and is usually adaptive. A cognitive bias is the systematic error the shortcut produces. A logical fallacy is a defect in the structure of an argument rather than in judgement. And noise, the subject of Kahneman, Sibony and Sunstein's 2021 book, is undesirable variability in judgements of the same problem — scatter rather than directional error — which their insurance study found produced median premiums for identical fictive customers varying by 55 per cent. The biases that matter most for organisational decisions are confirmation bias, anchoring, availability, overconfidence and overprecision, the planning fallacy, status quo bias, framing, hindsight bias, correspondence bias or the fundamental attribution error, sunk cost, base rate neglect, and the Dunning-Kruger pattern. The evidence for these varies substantially in strength, and honest treatment matters: incidental anchoring from random numbers failed to replicate in Many Labs 2, the letter-frequency demonstration of availability was undercut by Sedlmeier, Hertwig and Gigerenzer in 1998, base rate neglect is strongly format-dependent, and the Dunning-Kruger effect is largely a statistical artifact of regression to the mean rather than a metacognitive deficit. Ego depletion, facial feedback via the pen-in-mouth paradigm, social priming and power posing have all failed large replications. On the remedy side, awareness alone does not work: Pronin, Lin and Ross found participants insisted their self-assessments were accurate after reading a description of the bias affecting them. What does work is structural — mechanical aggregation of judgements outperformed clinical judgement in Grove and colleagues' meta-analysis of 136 studies, debiasing training with practice and feedback persists for months and transfers to professional settings, considering the opposite reduces anchoring even among experts, and reference-class forecasting counters the planning fallacy well enough that HM Treasury requires explicit optimism-bias adjustment. Argumentree addresses these structurally rather than by exhortation: independent submission before discussion removes the anchor and the audience, every argument carries its evidence and can be rated on merit rather than on its author, and a timestamped decision record allows a decision to be judged later on what was known at the time.

A cognitive bias is a systematic, predictable error in judgement — not random sloppiness, but a mistake that leans the same way every time. That predictability is the good news: an error you can anticipate is an error you can design around.
Last updated: 2026-08-25
Biases come from heuristics — shortcuts that Tversky and Kahneman described as "highly economical and usually effective" precisely because they work most of the time. Below are the twelve that matter most for organisational decisions, each with its canonical study, a business example, and an honest replication status — because several famous findings did not survive the last decade, and a reference page that hides that is not a reference page. The fixes that work are structural, not motivational: telling people about bias is the single weakest intervention in the literature.
These get used interchangeably and they are not interchangeable. Getting them apart is most of the practical value on this page.
The research programme starts with Amos Tversky and Daniel Kahneman's 1974 paper in Science, which named three heuristics — representativeness, availability, and adjustment from an anchor — and drew the conclusion that still frames the field: "These heuristics are highly economical and usually effective, but they lead to systematic and predictable errors." Kahneman received the 2002 Nobel Memorial Prize in Economic Sciences for this work; Tversky had died in 1996, and the prize is not awarded posthumously.
Each entry: what it is, a decision where it bites, what the evidence actually supports, and what changes when the reasoning is structured. Replication status is stated plainly — several of these are weaker than their reputations, and two of the most famous findings in adjacent psychology failed outright.
Seeking and weighting evidence that supports what you already think. Where it bites: the vendor shortlist assembled after someone has privately decided, where every reference call confirms rather than tests.
Evidence: Wason's 1960 rule-discovery task is the origin, though his own conclusion applied to the subgroup that persisted after two or more errors. Klayman and Ha (1987) reframed it as a positive test strategy that is normatively sound in most environments — you are not irrational for looking where you expect to find things. Hart and colleagues' 2009 meta-analysis in Psychological Bulletin puts congeniality bias at d = 0.36, and it reverses when disconfirming information is genuinely useful.
In Argumentree: the con branch is not optional. A claim without a recorded counter-case is visibly incomplete rather than quietly unchallenged, and a question chain attaches the challenge to the exact claim it disputes.
The first number mentioned sets the range everyone argues within. Where it bites: whoever says a budget figure, a valuation or a delivery date first has moved every subsequent estimate.
Evidence: Tversky and Kahneman's 1974 demonstration had median estimates of 25 versus 45 for the same quantity after anchors of 10 and 65. Many Labs 1 replicated anchoring robustly — but using provided anchors. Incidental anchoring from random numbers failed in Many Labs 2 (d = 0.04, p = .09, N = 6,826). So: deliberate anchors move people; arbitrary numbers in the environment largely do not.
In Argumentree: positions are submitted independently before discussion, so there is no first number to anchor on. The anchor needs an audience, and asynchronous input removes it.
Judging how likely something is by how easily an example comes to mind. Where it bites: last quarter's outage dominates the risk register while the slow, larger risk nobody has a story about goes unlisted.
Evidence: split, and worth knowing which half. The famous letter-frequency demonstration was undercut — Sedlmeier, Hertwig and Gigerenzer (1998) found judgements generally followed actual proportions. The ease-of-retrieval mechanism survives: Weingarten and Hutchinson's 2018 meta-analysis across 263 studies finds a medium effect, though publication bias shrinks it by up to a third.
In Argumentree: arguments carry their evidence, so a vivid anecdote and a base rate sit side by side and get rated separately — and the timeline shows when each entered the discussion.
Moore and Healy's 2008 paper in Psychological Review splits this into three: overestimating your score, overplacing yourself against others, and overprecision — being too certain your estimate is right. Where it bites: the delivery range that is far too narrow, so every plan downstream inherits false precision.
Evidence: overprecision is the robust one. Soll and Klayman found confidence intervals "sometimes only 40% as large as necessary to be well calibrated." Overestimation and overplacement flip with task difficulty — hard tasks produce overestimation but underplacement — and the hard-easy effect itself is partly a regression artifact (Juslin et al. 2000).
In Argumentree: ratings are on the argument, not the assertion, so a confident claim with thin evidence scores like thin evidence.
Underestimating your own completion times while judging others' realistically. Where it bites: every migration, integration and launch date you have ever set.
Evidence: strong in the field. Flyvbjerg and colleagues (2002) examined 258 transport projects worth about US$90bn and found costs underestimated in roughly nine of ten — rail +45%, fixed links +34%, roads +20%. Buehler, Griffin and Ross (1994) is the lab origin, and notably the bias was eliminated when participants were forced to connect the estimate to past experience. Caveat worth carrying: Halkjelsvik and Jørgensen's 2012 review found underestimation dominant in engineering and management literature but not in the psychology literature.
In Argumentree: the estimate becomes a claim with evidence attached — including the reference class — rather than a number asserted in a planning meeting.
Preferring the current arrangement, and treating "do nothing" as neutral rather than as a choice with consequences. Where it bites: the incumbent vendor renewed because switching requires justification and staying does not.
Evidence: among the strongest here. Samuelson and Zeckhauser (1988) ran 486 subjects and found that only 28% of TIAA-CREF participants had ever changed their allocation. Jachimowicz and colleagues' 2019 meta-analysis of default effects across 58 datasets (n = 73,675) reports d = 0.68 [0.53, 0.83], and Madrian and Shea found 401(k) participation moving from 37% to 86% on a default change alone.
In Argumentree: "keep what we have" is entered as an option with its own pros and cons, so it competes on evidence instead of winning by not being on the agenda.
The same facts described as a gain or a loss produce different choices. Where it bites: "90% success rate" and "10% failure rate" are the same number and do not land the same way.
Evidence: robust in direction, roughly half the original magnitude. Many Labs 1 replicated Tversky and Kahneman's Asian disease item: original d = 1.13 → replication d = 0.62 (N = 6,271). Many Labs 2 replicated a different item from the same 1981 paper — original OR 4.96 → OR 2.06 (N = 7,228). Those two are constantly conflated; they are different studies. Kühberger's meta-analysis of 136 papers calls framing reliable and of small-to-moderate size.
In Argumentree: decision criteria are written before the options are framed, and a claim can be stated once and challenged in both frames rather than surviving on presentation.
Once you know the outcome, you believe you knew it all along. Where it bites: the post-mortem that becomes a search for who should have seen it, which teaches everyone to be quieter next time.
Evidence: robust, magnitude disputed. Fischhoff (1975) showed outcome knowledge raises judged likelihood and that people are largely unaware of the shift. Meta-analyses disagree on size: Christensen-Szalanski and Willham (122 studies) report r = .17; Guilbault and colleagues (95 studies) report a median around .39 — and found debiasing manipulations largely ineffective.
In Argumentree: the decision record is timestamped, so a review can judge the decision on the evidence that existed at the time instead of on the outcome that arrived later.
Explaining behaviour by character while discounting the situation. Where it bites: the project failed because the lead was weak — rather than because the process gave them a deadline set by anchoring and a plan approved by consensus.
Evidence: solid in paradigm. Jones and Harris (1967) found attitude attributions persisted even when the essay position was assigned rather than chosen. Many Labs 2 replicated a correspondence-bias study successfully — original d = 1.75, replication d = 1.82, one of the sturdier results in the set. Cross-culturally it attenuates rather than disappears.
In Argumentree: arguments are rated, people are not. That is the whole design principle, and it is the direct structural answer to this bias.
Counting unrecoverable past spending in a present decision. Where it bites: the project defended because of what it has already consumed.
Evidence: real but moderate, and the measures are contested. Arkes and Blumer (1985) found 85% chose to finish a plane that was already known to be inferior. Sleesman and colleagues' meta-analysis puts sunk cost at ρ = .243 — comparable to, not larger than, personal responsibility. A 2025 analysis found classic sunk-cost vignettes have poor internal consistency (ω = 0.14–0.57), so treat single-vignette findings with care. Note also that in Arkes and Blumer's own theatre study the effect was present in the first half of the season and essentially gone in the second.
In Argumentree: exit criteria can be recorded when the project is approved, so the review reads a document rather than relitigating a colleague's judgement — see escalation of commitment.
Judging by how well a case matches a stereotype while ignoring how common the thing actually is. Where it bites: "this deal feels like the ones we win" — without the win rate.
Evidence: robust as a demonstration, and strongly format-dependent, which is the actionable part. Casscells and colleagues (1978) found only 18% of medical staff gave the correct answer while 45% were off by a factor of nearly fifty; Manrai and colleagues repeated it 36 years later with almost identical results (23% correct, 44% giving the same wrong answer). But Gigerenzer and Hoffrage showed correct responses rising from about 16% to 46–50% when the same problem is posed in natural frequencies rather than probabilities.
In Argumentree: evidence is attached to the claim it supports, so a base rate is present as an argument rather than remembered — or not — by whoever happens to know it.
The claim that the least competent are the most confident. Where it bites: the loudest certainty in the room belonging to the person with the least basis for it.
Evidence — and this one needs care. The descriptive pattern replicates: low performers do overestimate themselves. The explanation largely does not. Nuhfer and colleagues reproduced the classic graph shape from random numbers; Gignac and Zajenkowski (2020, N = 929) found the relationship essentially linear and the effect much smaller than reported; McIntosh and colleagues found that when task difficulty was equalised, the pattern was eliminated — performance level, not metacognitive deficit, was the driver. In fairness, Pennycook and Jansen and colleagues do find low performers less able to evaluate their own answers.
Honest summary: people overestimate themselves, but most of the famous graph is regression to the mean plus better-than-average. Do not cite it as a law.
In Argumentree: confidence is not a scoring input. A rating attaches to an argument's evidence, so certainty carries no weight it has not earned.
This is where most corporate bias work goes wrong, because the intuitive intervention is the weakest one in the literature. The evidence splits cleanly.
The most-purchased intervention has the least support. Pronin, Lin and Ross (2002) found participants "insisted that their self-assessments were accurate and objective even after reading a description of how they could have been affected by the relevant bias." People see bias in others far more readily than in themselves — that is the bias blind spot, and awareness training runs straight into it.
Nemeth and colleagues (2001) found assigned devil's advocacy fostered thinking "primarily aimed at cognitive bolstering of the initial viewpoint," and that an authentic minority was superior to all three role-played variants. Real dissent cannot be cloned by assigning someone to perform it — though giving genuine dissenters cover is a different and better move.
Collect positions in writing, separately, before anyone hears anyone else. This is the single cheapest structural fix on the page, and it defuses anchoring, social proof and authority deference in one step — each of which needs an audience and an order of speaking to operate.
The best-evidenced item here. Grove and colleagues' meta-analysis of 136 studies found mechanical prediction about 10% more accurate than clinical judgement on average, and superiority was consistent "regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined." Combining structured ratings beats a discussion that ends in a feeling.
Actively arguing the other side beats exhorting people to be fair — Lord, Lepper and Preston (1984) found it outperformed instructions to be as unbiased as possible, and Mussweiler and colleagues reduced anchoring in a real-world setting with experts by listing arguments against the anchor. Structurally, that is what a con branch is.
Unlike awareness alone, actual training holds. Morewedge and colleagues (2015) found single-session interventions produced improvements that persisted at two months, and Sellier, Scopelliti and Morewedge (2019) showed transfer to a real professional setting — trained participants were 19% less likely to choose the confirming-but-inferior option.
Against the planning fallacy, base the forecast on comparable completed projects rather than on this project's plan. HM Treasury's Green Book guidance requires it explicitly: appraisers "have the tendency to be over optimistic," so adjustments "should be based on data from past or similar projects."
If two of your reviewers score the same proposal differently, no amount of debiasing will help — that is noise, and it needs shared scales and independent ratings rather than better intentions. It is also invisible until someone runs the same case past two people on purpose.
Notice the shape of the working list: none of it asks anyone to be more objective. Every effective item changes the order in which things happen, what gets written down, or what gets combined — which is precisely what a structured argument tree is. Argumentree's contribution is not a debiasing seminar; it is that independent submission, evidence attached to claims, rating the argument rather than the arguer, and a timestamped record are the four structural interventions the evidence actually supports, running by default instead of by discipline.
What these biases look like once a group is discussing — the same errors with meeting-shaped names and warning signs you can call out live.
The long-form treatment: twelve biases at organisational scale, the counterargument, and the process fixes.
The third layer — what happens when someone uses these mechanisms on purpose.
The five-element discipline that operationalises the structural fixes on this page.
The standard that judges a decision by its process at decision time rather than by how it turned out.
Sunk cost's organisational form, and why authorship of a failure predicts doubling down on it.
A systematic, predictable error in judgement produced by the mental shortcuts people use under uncertainty. The defining feature is that it leans consistently in one direction rather than scattering randomly, which is what makes it possible to design a process that compensates for it. Tversky and Kahneman's 1974 paper in Science established the framing: heuristics are highly economical and usually effective, but they lead to systematic and predictable errors.
A heuristic is the mental shortcut itself and is usually adaptive. A cognitive bias is the systematic error that shortcut produces. A logical fallacy is a defect in the structure of an argument rather than in judgement — a false dichotomy is a fallacy, while anchoring is a bias. A fourth term, noise, is variability rather than error: undesirable inconsistency between judgements of the same problem.
Bias is directional error — your shots land consistently left of the target. Noise is scatter — they land everywhere. Kahneman, Sibony and Sunstein's 2021 book argues organisations spend heavily on bias while ignoring noise, which is often larger and easier to fix. In their insurance study, median premiums set independently for the same five fictive customers varied by 55%, about five times what the underwriters and executives expected.
The twelve on this page: confirmation bias, anchoring, availability, overconfidence and overprecision, the planning fallacy, status quo bias, framing, hindsight bias, correspondence bias, sunk cost, base rate neglect, and the Dunning-Kruger pattern. They earn their place because each maps onto a recurring organisational failure — the shortlist that only confirms, the first number that sets the range, the renewal nobody justified, the post-mortem that hunts for a culprit.
The pattern is real; the popular explanation is not. Low performers do overestimate themselves, but much of the famous chart is reproduced by regression to the mean and the better-than-average effect — Nuhfer and colleagues generated the same graph shape from random numbers, Gignac and Zajenkowski found the relationship essentially linear with a much smaller effect, and McIntosh and colleagues eliminated the pattern by equalising task difficulty. Treat it as a descriptive tendency, not as a metacognitive law.
Yes, and any honest reference page should say so. Ego depletion failed a registered replication (d = 0.04 across 23 labs). The pen-in-mouth facial-feedback paradigm failed a 17-lab replication. Social priming and power posing both failed. Incidental anchoring from random numbers failed in Many Labs 2, and the letter-frequency demonstration of availability was undercut in 1998. More broadly, the Open Science Collaboration replicated 36% of 100 psychology studies, and Many Labs 2 replicated 15 of 28 with median effect sizes falling from 0.60 to 0.15.
Barely, on its own — and this is the most important practical finding here. Pronin, Lin and Ross found participants insisted their self-assessments were accurate even after reading a description of the exact bias affecting them. Training that includes practice and feedback is different: it persists for months and transfers to real professional settings. But a slide deck listing biases is close to the weakest intervention available.
Structural changes rather than motivational ones. Collect judgements independently before discussion, so anchoring and social proof have no audience. Combine structured ratings mechanically — Grove's meta-analysis of 136 studies found mechanical prediction about 10% more accurate than clinical judgement. Require the opposite case to be argued explicitly. Use reference classes rather than fresh optimism for forecasts. And keep a timestamped record so decisions are judged on what was known at the time.
By running the four evidence-backed structural interventions by default. Arguments can be submitted independently before discussion, which removes the anchor and the audience. Every claim carries its evidence, so base rates and vivid anecdotes are visible side by side. Ratings attach to arguments rather than to people, so confidence and seniority carry no weight they have not earned. And the decision record is timestamped, so a later review can judge the process on the information available at the time rather than on the outcome.
Tversky, A., & Kahneman, D. (1974). Judgment under Uncertainty: Heuristics and Biases. Science, 185(4157), 1124–1131.
The founding paper: representativeness, availability, and adjustment from an anchor.
View source →Kahneman, D., & Tversky, A. (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica, 47(2), 263–291.
Value assigned to gains and losses rather than final assets; probabilities replaced by decision weights.
View source →Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A Flaw in Human Judgment.
Noise as undesirable variability in judgements of the same problem — the distinction most bias work omits.
Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517.
The three-way split: overestimation, overplacement, overprecision.
View source →Flyvbjerg, B., Holm, M. S., & Buhl, S. (2002). Underestimating Costs in Public Works Projects. Journal of the American Planning Association.
258 projects, roughly US$90bn: costs underestimated in about nine of ten.
Jachimowicz, J. M., Duncan, S., Weber, E. U., & Johnson, E. J. (2019). When and why defaults influence decisions: a meta-analysis of default effects. Behavioural Public Policy.
d = 0.68 [0.53, 0.83] across 58 datasets, n = 73,675.
Klein, R. A., et al. (2014). Investigating Variation in Replicability: A 'Many Labs' Replication Project. Social Psychology.
Replicated the Asian disease framing item: original d = 1.13 → replication d = 0.62 (N = 6,271).
View source →Klein, R. A., et al. (2018). Many Labs 2. Advances in Methods and Practices in Psychological Science, 1(4), 443–490.
15 of 28 findings replicated; median d fell from 0.60 to 0.15. Incidental anchoring failed (d = 0.04).
View source →Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251).
36% of 100 replications reached significance against 97% of originals.
View source →Hagger, M. S., et al. (2016). A Multilab Preregistered Replication of the Ego-Depletion Effect. Perspectives on Psychological Science.
d = 0.04, 95% CI [−0.07, 0.15] across 23 labs.
Gignac, G. E., & Zajenkowski, M. (2020). The Dunning-Kruger effect is (mostly) a statistical artefact. Intelligence, 80, 101449.
Essentially linear relationship; effect much smaller than reported.
View source →Pronin, E., Lin, D. Y., & Ross, L. (2002). The Bias Blind Spot. Personality and Social Psychology Bulletin, 28(3), 369–381.
Self-assessments defended even after reading a description of the relevant bias.
View source →Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: a meta-analysis. Psychological Assessment, 12(1), 19–30.
136 studies; mechanical prediction about 10% more accurate, consistently across task and judge type.
View source →Sellier, A.-L., Scopelliti, I., & Morewedge, C. K. (2019). Debiasing Training Improves Decision Making in the Field. Psychological Science.
Trained participants 19% less likely to choose the confirming-but-inferior option.
View source →Nemeth, C., Brown, K., & Rogers, J. (2001). Devil's advocate versus authentic dissent. European Journal of Social Psychology.
Assigned devil's advocacy bolstered the initial view; authentic dissent beat all three variants.
View source →HM Treasury. Green Book supplementary guidance: optimism bias.
Requires explicit adjustment based on data from past or similar projects.
View source →Independent submission before discussion, evidence attached to every claim, ratings on arguments rather than authors, and a timestamped record — the four structural interventions the evidence supports, running by default.
Start Free — No Credit Card