Custodes Futurifor the advancement of humanity

Psychological factor, barrier 1 of 8

Cognitive biases

Tversky and Kahneman (1974) documented systematic errors such as anchoring and the availability heuristic. Confirmation bias leads people to favor information that fits existing beliefs.

Evidence

  • Replication. The Open Science Collaboration (2015) repeated 100 published psychology studies. 97% of the originals had significant results but only 36% of the replications did, and replication effects averaged half the size of the originals.
  • Selective exposure. A meta-analysis of 91 studies with nearly 8,000 participants found a moderate preference for information that supports existing views (d = 0.36). The bias was stronger in more closed-minded people and reversed when challenging information was useful for a current goal (Hart et al., 2009).

What can be done

Actions rated on the strength-rating scale (A strong and replicated, B solid but limited, C weak or debated, D contested or failed), applied to the specific claim made.

  • Use validated formulas instead of unaided judgment where outcomes can be scored (A). Across studies, mechanical prediction was about 10% more accurate than clinical judgment on average (Grove et al., 2000), and in hiring and admissions, combining information by formula rather than by holistic judgment improved prediction of job performance by more than 50% (Kuncel et al., 2013). The formula must be validated: a commercial recidivism tool with 137 features was no more accurate than people without criminal justice expertise in one study (B for that setting) (Dressel and Farid, 2018), and people tend to abandon algorithms after seeing them err (Dietvorst et al., 2015).
  • Structured interviews with fixed scoring in hiring (A that they are valid, C that they are the best method). A reanalysis of selection research ranked structured interviews first, though validity estimates for most methods fell by .10 to .20 (Sackett et al., 2022). The ranking itself is disputed.
  • Test policies at scale rather than relying on expert intuition (A for the small size of real effects). Across 126 trials with 23 million people, interventions run by government nudge units raised outcomes by 1.4 percentage points on average, against 8.7 points in academic journal studies (DellaVigna and Linos, 2022), and expert judges failed to predict which of 54 interventions would work best (Milkman et al., 2021a).
  • "Consider the opposite" before deciding (B). Asking people to consider the opposite reduced biased assimilation of evidence more than asking them to be fair (Lord, Lepper, and Preston, 1984), and the same strategy reduced anchoring, including among experts (Mussweiler, Strack, and Pfeiffer, 2000).
  • Contested: forecasting training and teams (C). Training of under an hour was reported to improve forecast accuracy by 6 to 11% (Mellers et al., 2014, Chang et al., 2016), but an independent reanalysis found the effects largely disappeared after controlling for which questions forecasters chose and when they forecast (Hauenstein et al., 2025).
  • Weak or not shown to work: Debiasing games reduced bias in the lab (Morewedge et al., 2015), but evidence of better real-world decisions is weak (C) (Sellier, Scopelliti, and Morewedge, 2019, Korteling, Gerritsma, and Toet, 2021). A surgical checklist cut deaths from 1.5% to 0.8% in eight hospitals (Haynes et al., 2009), but mandated use in 101 Ontario hospitals brought no significant reduction (C) (Urbach et al., 2014). Evidence for pre-mortems is thin and developer-linked (C) (Mitchell, Russo, and Pennington, 1989, Veinott, Klein, and Wiggins, 2010). Expecting large effects from nudges in general is not supported (D): a meta-analysis found an average effect of d = 0.45 (Mertens et al., 2022), but after correcting for publication bias the estimate fell to about 0.04 (Maier et al., 2022), and effects vary widely (Szaszi et al., 2022).

Proposed and experimental methods

Methods that are proposed, under trial, approved in some places, or tried and then failed. Each shows a stage label and an evidence rating. A stage label shows how far a method has progressed, not whether it works. The stage labels are explained on the psychological factor page.

  • Citizens' assemblies and deliberative polls (Approved but not scaled, B). Randomly selected, demographically balanced citizens hear from balanced experts, deliberate in facilitated small groups, and make recommendations. The OECD collected close to 300 such processes and describes a "deliberative wave" gaining momentum since around 2010 (OECD, 2020). In "America in One Room", 526 registered U.S. voters deliberated over a weekend, Republicans and Democrats moved closer together on 22 of 26 highly polarized proposals, and each party's rating of the other rose by about 14 to 15 points more than in a control group of 844 (Fishkin et al., 2021). Ireland's Citizens' Assembly (2016 to 2018) moved toward a more liberal position on abortion as it deliberated, and that position was largely reflected in the later referendum (Farrell et al., 2023), which passed on May 25, 2018 with 66.40% voting yes (ElectionGuide, n.d.). Few studies test whether assemblies lead to better policy outcomes.
  • AI-mediated deliberation (Early trial, B). An AI system drafts and revises group statements from participants' opinions and critiques to find common ground. In a Google DeepMind study with 5,734 participants, people preferred the AI's statements to those written by human mediators, rating them more informative, clear and unbiased, and discussants often moved toward a shared view, a result repeated in a virtual citizens' assembly with a demographically representative UK sample (Tessler et al., 2024). All authors were from the developing company, and effects on real decisions are untested.
  • AI decision assistants and AI "devil's advocates" (Early trial, C). AI systems advise people on decisions, either with a recommendation or by arguing against the current view. A preregistered meta-analysis of 106 experiments (370 effect sizes, 2020 to June 2023) found that human and AI teams on average did worse than the better of the human or the AI alone (Hedges' g = -0.23), with losses in decision tasks and gains in content creation tasks (Vaccaro et al., 2024). In single studies, large language model assistants improved forecasting accuracy by 24% to 28% over a control among 991 participants, with one outlier question affecting the results (Schoenegger et al., 2025), and an interactive AI devil's advocate modestly raised group decision accuracy among 350 people in 120 groups (p = 0.047) but did not significantly reduce reliance on wrong AI advice (Chiang et al., 2024).
  • Prediction markets and futarchy (Early trial, C). In prediction markets people trade contracts that pay out on future events, and prices act as forecasts. Across four U.S. presidential elections, Iowa Electronic Markets prices in the week before the vote had an average absolute error of about 1.5 percentage points versus 2.1 for the final Gallup poll, but the authors expect markets to perform poorly on small probability events (Wolfers and Zitzewitz, 2004). Futarchy is a proposal that voters choose goals while betting markets choose the policies expected to meet them (Hanson, 2013). On the blockchain platform MetaDAO, 62 futarchy decisions across nine organizations drew $2.26 million in cumulative volume with low participation in most, and decision quality was not measured (Pokorny, 2025). The UK government forecasting tournament Cosmic Bazaar, launched in April 2020, had 1,300 forecasters from 41 departments and more than 10,000 forecasts by April 2021, with no published accuracy or policy results (Perry World House, n.d.). This builds on the contested forecasting training item above.
  • Noise audits (Early trial, C). An organization gives identical cases to many of its professionals, measures how much their judgments differ ("noise"), and then reduces unwanted variation with structured guidelines or by averaging independent judgments. In an insurance company audit described by Kahneman, Sibony and Sunstein, the median difference in prices that underwriters set for identical policies was 55% (43% for claims adjusters' payouts), while 828 senior executives had expected about 10% (Kinni, 2021). The audit data are unpublished, and no study was found showing that audits improve later decisions.

Sources cited on this page

  1. Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: heuristics and biases. Science, 185, 1124-1131. DOI A Strong: for anchoring, B for the full set of heuristics
  2. Maier, S. F., & Seligman, M. E. P. (2016). Learned helplessness at fifty. Psychological Review, 123(4), 349-367. DOI B Moderate: for the animal mechanism, C for applying it to human motivation
  3. Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. DOI A Strong
  4. Hart, W., Albarracín, D., Eagly, A. H., Brechan, I., Lindberg, M. J., & Merrill, L. (2009). Feeling validated versus being correct: a meta-analysis of selective exposure to information. Psychological Bulletin, 135(4), 555-588. DOI B Moderate
  5. Chang, W., Chen, E., Mellers, B., & Tetlock, P. (2016). Developing expert political judgment: The impact of training and practice on judgmental accuracy in geopolitical forecasting tournaments. Judgment and Decision Making, 11(5), 509-526. link C Limited
  6. DellaVigna, S., & Linos, E. (2022). RCTs to scale: Comprehensive evidence from two nudge units. Econometrica, 90(1), 81-116. link A Strong: for small average at-scale effects, B for attributing most of the gap to publication bias
  7. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114-126. link B Moderate: for aversion after seeing errors, C as a general claim that people prefer human judgment
  8. Dressel, J., & Farid, H. (2018). The accuracy, fairness, and limits of predicting recidivism. Science Advances, 4(1), eaao5580. link B Moderate: for its setting, C as a general claim that lay people match algorithms
  9. Gal, D., & Rucker, D. D. (2018). The loss of loss aversion: Will it loom larger than its gain? Journal of Consumer Psychology, 28(3), 497-516. link C Limited
  10. Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., & Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12(1), 19-30. link A Strong
  11. Hauenstein, C. E., Thomas, R. P., Illingworth, D. A., & Dougherty, M. R. (2025). Rethinking the role of teams and training in geopolitical forecasting: The effect of uncontrolled method variance on statistical conclusions. Psychological Science, 36(1). link B Moderate
  12. Haynes, A. B., Weiser, T. G., Berry, W. R., Lipsitz, S. R., Breizat, A.-H. S., Dellinger, E. P., et al. (2009). A surgical safety checklist to reduce morbidity and mortality in a global population. The New England Journal of Medicine, 360(5), 491-499. link C Limited
  13. Korteling, J. E., Gerritsma, J. Y. J., & Toet, A. (2021). Retention and transfer of cognitive bias mitigation interventions: A systematic literature study. Frontiers in Psychology, 12, 629354. link B Moderate
  14. Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060-1072. link A Strong
  15. Li, Y., & Bates, T. C. (2019). You can't change your basic ability, but you work at things, and that's how we get hard things done: Testing the role of growth mindset on response to setbacks, educational attainment, and cognitive ability. Journal of Experimental Psychology: General, 148(9), 1640-1655. link B Moderate
  16. Lord, C. G., Lepper, M. R., & Preston, E. (1984). Considering the opposite: A corrective strategy for social judgment. Journal of Personality and Social Psychology, 47(6), 1231-1243. link B Moderate
  17. Maier, M., Bartoš, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L., & Wagenmakers, E.-J. (2022). No evidence for nudging after adjusting for publication bias. Proceedings of the National Academy of Sciences, 119(31), e2200300119. link B Moderate
  18. Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106-1115. link C Limited
  19. Mertens, S., Herberz, M., Hahnel, U. J. J., & Brosch, T. (2022). The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains. Proceedings of the National Academy of Sciences, 119(1), e2107346118. link D Contested: for the headline average effect size
  20. Milkman, K. L., Gromet, D., Ho, H., Kay, J. S., Lee, T. W., Pandiloski, P., et al. (2021a). Megastudies improve the impact of applied behavioural science. Nature, 600(7889), 478-483. link B Moderate
  21. Milkman, K. L., Gandhi, L., Patel, M. S., Graci, H. N., Gromet, D. M., Ho, H., et al. (2022). A 680,000-person megastudy of nudges to encourage vaccination in pharmacies. Proceedings of the National Academy of Sciences, 119(6), e2115126119. link A Strong: for raising vaccination at the measured provider, B for net population uptake
  22. Mitchell, D. J., Russo, J. E., & Pennington, N. (1989). Back to the future: Temporal perspective in the explanation of events. Journal of Behavioral Decision Making, 2(1), 25-38. link C Limited
  23. Morewedge, C. K., Yoon, H., Scopelliti, I., Symborski, C. W., Korris, J. H., & Kassam, K. S. (2015). Debiasing decisions: Improved decision making with a single training intervention. Policy Insights from the Behavioral and Brain Sciences, 2(1), 129-140. link B Moderate: for short-term lab bias reduction, C for better real-world decisions
  24. Mussweiler, T., Strack, F., & Pfeiffer, T. (2000). Overcoming the inevitable anchoring effect: Considering the opposite compensates for selective accessibility. Personality and Social Psychology Bulletin, 26(9), 1142-1150. link B Moderate
  25. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. link A Strong: for structured interviews being valid predictors, C for their top ranking
  26. Sellier, A.-L., Scopelliti, I., & Morewedge, C. K. (2019). Debiasing training improves decision making in the field. Psychological Science, 30(9), 1371-1379. link C Limited
  27. Szaszi, B., Higney, A., Charlton, A., Gelman, A., Ziano, I., Aczel, B., et al. (2022). No reason to expect large and consistent effects of nudge interventions. Proceedings of the National Academy of Sciences, 119(31), e2200732119. link B Moderate
  28. Urbach, D. R., Govindarajan, A., Saskin, R., Wilton, A. S., & Baxter, N. N. (2014). Introduction of surgical safety checklists in Ontario, Canada. The New England Journal of Medicine, 370(11), 1029-1038. link B Moderate
  29. Veinott, B., Klein, G. A., & Wiggins, S. (2010). Evaluating the effectiveness of the PreMortem technique on plan confidence. In Proceedings of the 7th International ISCRAM Conference. Seattle, WA. link C Limited
  30. Chiang, C.-W., Lu, Z., Li, Z., & Yin, M. (2024). Enhancing AI-assisted group decision making through LLM-powered devil's advocate. In Proceedings of the 29th International Conference on Intelligent User Interfaces (pp. 103-117). ACM. link C Limited
  31. Farrell, D. M., Suiter, J., Cunningham, K., & Harris, C. (2023). When mini-publics and maxi-publics coincide: Ireland's national debate on abortion. Representation, 59(1), 55-73. link B Moderate
  32. Fishkin, J., Siu, A., Diamond, L., & Bradburn, N. (2021). Is deliberation an antidote to extreme partisan polarization? Reflections on "America in One Room." American Political Science Review, 115(4), 1464-1481. link B Moderate
  33. Hanson, R. (2013). Shall we vote on values, but bet on beliefs? Journal of Political Philosophy. link C Limited
  34. Kinni, T. (2021, May 19). How noisy is your company? strategy+business. link C Limited
  35. Pokorny, Z. (2025, March 17). The state of onchain futarchy. Galaxy Research. link B Moderate
  36. Schoenegger, P., Park, P. S., Karger, E., Trott, S., & Tetlock, P. E. (2025). AI-augmented predictions: LLM assistants improve human forecasting accuracy. ACM Transactions on Interactive Intelligent Systems, 15(1). link B Moderate
  37. Tessler, M. H., Bakker, M. A., Jarrett, D., Sheahan, H., Chadwick, M. J., Koster, R., et al. (2024). AI can help humans find common ground in democratic deliberation. Science, 386(6719), eadq2852. link B Moderate
  38. Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293-2303. link B Moderate
  39. Wolfers, J., & Zitzewitz, E. (2004). Prediction markets. Journal of Economic Perspectives, 18(2), 107-126. link B Moderate
  40. Sunstein, C. R. (2022). Sludge audits. Behavioural Public Policy, 6(4), 654-673. link C Limited
  41. Mitchell, J. M., Ot'alora G., M., van der Kolk, B., Shannon, S., Bogenschutz, M., Gelfand, Y., et al. (2023). MDMA-assisted therapy for moderate to severe PTSD: A randomized, placebo-controlled phase 3 trial. Nature Medicine, 29(10), 2473-2480. link D Contested

Every source for this factor is listed on the psychological factor page.