Wisdom of Crowds: Galton's Ox — The 1906 Experiment That Proved Crowds Beat Experts
Decision Science

Wisdom of Crowds: Galton's Ox — The 1906 Experiment That Proved Crowds Beat Experts

AT
Argumentree Team
Decision Science
July 25, 2026
14 min read
Galton’s ox experiment is the founding demonstration of the wisdom of crowds. In autumn 1906, at the West of England Fat Stock and Poultry Exhibition in Plymouth, the 84-year-old statistician Francis Galton analyzed a weight-judging contest in which fairgoers paid sixpence to estimate the dressed weight of a live ox. Of roughly 800 cards, 787 were legible; the median guess was 1,207 pounds against an actual weight of 1,198 pounds — within about 0.8% — and, as Galton later acknowledged in follow-up correspondence, the mean was 1,197 pounds, one pound off. Galton, a skeptic of democratic judgment, had expected to prove the average voter incompetent; instead he published the result in Nature (1907) as “Vox Populi.” The experiment works because independent, diverse errors partially cancel when aggregated — a principle formalized in the Condorcet Jury Theorem (1785), codified by James Surowiecki’s four conditions (diversity, independence, decentralization, aggregation, 2004), and applied today in prediction markets and ensemble machine learning. The crowd’s advantage disappears when independence breaks down through information cascades or social influence. In the decision-making lifecycle, Galton’s lesson applies at the Aggregate stage: collect judgments independently before discussion. Argumentree implements this with independent asynchronous argument submission, multi-dimensional rating aggregated into consensus scores, and a documented reasoning trail.
Share:
TL;DR

In 1906, Francis Galton — a brilliant statistician and an open skeptic of democratic judgment — analyzed 787 guesses at an ox's dressed weight, expecting to prove the crowd hopeless. The median guess was 1,207 lb. The ox weighed 1,198 lb. The crowd beat the experts, and the science of collective judgment was born.

  • The setup: a sixpenny weight-judging contest at a Plymouth livestock fair; ~800 cards, 787 valid
  • The result: median 1,207 lb vs. actual 1,198 lb — within ~0.8%; the mean, noted later, was 1,197 lb
  • Why it works: independent, diverse errors partially cancel when aggregated
  • The catch: break independence — let people copy each other — and the magic reverses
  • The legacy: prediction markets, ensemble machine learning, and structured group decision-making all descend from this one afternoon

The scientist who wanted to prove crowds foolish

The best part of the most famous experiment in collective intelligence is that it was designed — in spirit, at least — to prove the opposite of what it found.

By 1906, Francis Galton was one of the most accomplished scientists in Britain. A cousin of Charles Darwin, he had pioneered the statistical concepts of correlation and regression to the mean, invented the weather map, and put fingerprint identification on a scientific footing. He was also — and this matters for the story — deeply skeptical of ordinary people’s judgment. Galton believed that intelligence was concentrated in a small hereditary elite; he coined the term “eugenics” and spent decades promoting it, a legacy rightly condemned today and inseparable from how he approached this question. When he turned his statistical tools on a crowd, he expected to document its incompetence.

So when the 84-year-old Galton wandered through the West of England Fat Stock and Poultry Exhibition at Plymouth in the autumn of 1906 and came across a weight-judging competition, he saw an irresistible data set. Here was a crowd — some expert, most not — making hundreds of independent numerical judgments about a matter of fact, with a small entry fee to keep them honest and a prize to keep them motivated. As he later wrote in Nature, the average competitor was probably as well fitted for judging the ox as an average voter is for judging the merits of most political issues on which he votes. The analogy to democracy was the whole point. Galton expected the vox populi — the voice of the people — to embarrass itself.

A country fair, a fat ox, and 787 sixpenny tickets

The mechanics of the contest were simple, and — as we now know — accidentally close to ideal experimental design. A live ox was on display. For sixpence, anyone could buy a stamped and numbered card, write their name and address on it, and record their estimate of what the ox would weigh once slaughtered and dressed — the commercially meaningful figure, and a harder question than the animal’s live weight. The most accurate guesses won prizes.

Three features of this setup deserve attention, because each one maps onto a condition that later research would show is essential for collective accuracy. First, the sixpenny fee filtered out pure practical jokers — people had skin in the game. Second, each card was filled in privately: there was no show of hands, no auctioneer calling out a running consensus, no visible frontrunner to copy. The judgments were independent. Third, the crowd was genuinely mixed: butchers and farmers with professional expertise stood alongside clerks and families who were, in Galton’s words, guided by the fancy of the moment. The judgments were diverse.

After the contest, Galton borrowed the cards. Of roughly 800, thirteen were illegible or otherwise defective, leaving 787 valid estimates. He sorted them, tabulated them, plotted their distribution, and computed their quantiles — a full statistical workup, published in Nature on 7 March 1907 under the title “Vox Populi” (Galton, F., Nature, 75, 450–451).

The result that surprised him

The middlemost estimate — what we now call the median — was 1,207 pounds. The ox, slaughtered and dressed, weighed 1,198 pounds. The crowd’s collective verdict was off by nine pounds: an error of about 0.8%, closer than the estimates of the cattle experts in attendance.

Galton — to his genuine credit as a scientist — reported the result straight. “This result is, I think, more creditable to the trust-worthiness of a democratic judgment than might have been expected,” he wrote. Coming from a man who had spent his career arguing that sound judgment was a rare hereditary gift, the sentence is a quiet scientific concession: the data had beaten his prior.

The story got better in the correspondence that followed. Readers of Nature wrote in, and in his reply (“The Ballot-Box,” Nature, 75, 509) Galton acknowledged that the arithmetic mean of all 787 guesses was 1,197 pounds — one pound off the true weight. A century later, the econometrician Kenneth Wallis (“Revisiting Francis Galton’s forecasting competition,” 2014) went back over the original data and the surrounding exchange, confirming the essentials of the story that James Surowiecki had made famous. The one-pound miss of the mean is the version usually quoted today; the nine-pound miss of the median is what Galton actually led with. Both are astonishing.

Median or mean? Galton’s statistical choice still matters

Why did Galton prefer the middlemost estimate? His reasoning was explicitly political: the median is the value that a majority of the crowd would ratify in a vote. Every other value would be outvoted by those who thought it too high or too low. The median, in other words, is the crowd’s democratic verdict — which was exactly the question about vox populi he had set out to test.

But there is a statistical argument hiding inside the political one, and it is still taught today. The median is a robust statistic: a handful of absurd guesses — someone writing 10,000 pounds for a laugh — barely moves it, whereas the mean can be dragged arbitrarily far by a single outlier. The mean, on the other hand, uses more of the information in the data and, when errors are roughly symmetric and independent, cancels them more efficiently. Galton’s crowd was well-behaved enough that both worked: the median missed by 9 lb, the mean by 1 lb. In messier crowds — online polls, unmoderated forums — the robustness of the median often wins. Choosing your aggregation mechanism is a decision about how much you trust the tails of your crowd.

Why averaging works: the arithmetic of error cancellation

There is no mysticism in Galton’s result. Each guess can be thought of as the true value plus an error. If the errors are independent and scattered on both sides of the truth — some butchers overestimate, some underestimate — then adding the guesses together cancels much of the error while the shared signal accumulates. The more diverse and independent the errors, the more aggressively they cancel.

Scott Page later formalized the intuition as the diversity prediction theorem (Page, The Difference, 2007): for squared error, the collective’s error equals the average individual error minus the diversity of the predictions. It is a mathematical identity, not an empirical claim — which is what makes it so useful. A crowd beats its average member by exactly as much as its members disagree with each other. Diversity is not a nice-to-have; it is, arithmetically, the entire advantage. A crowd of clones gains nothing from aggregation.

The same logic had been discovered, for yes/no questions rather than quantities, more than a century before Galton. The Marquis de Condorcet proved in 1785 that if each voter is even slightly more likely than chance to be right, and voters judge independently, the probability that the majority is right climbs toward certainty as the group grows. We tell that story — and the dramatic ways the theorem fails when its assumptions break — in our companion piece on the Condorcet Jury Theorem.

From footnote to framework: Surowiecki’s four conditions

For most of the twentieth century, Galton’s ox was a statistical curiosity. That changed in 2004, when James Surowiecki opened The Wisdom of Crowds with the Plymouth story and assembled a century of evidence — from stock markets to search engines — that crowds can be systematically smart. Crucially, Surowiecki did not claim crowds are always wise. He identified four conditions that Galton’s fair satisfied by accident, and that any wise crowd needs on purpose:

Diversity of opinion

Each person brings some private information or a different way of interpreting the question. Butchers, farmers, and casual fairgoers all read the ox differently — and their different errors partially cancelled.

Independence

Each ticket was filled in privately, without a show of hands or a loud expert anchoring the room. Nobody could copy the frontrunner, because there was no visible frontrunner.

Decentralization

No committee decided what the "official" estimate should be. Every entrant drew on their own local knowledge — years of judging cattle, or just a good eye.

Aggregation

A mechanism turned 787 private judgments into one collective answer. Without the pile of tickets and someone willing to tally them, the wisdom stays trapped in individual heads.

Remove any one condition and the crowd doesn’t merely lose its edge — it can become actively worse than its members, because aggregation then amplifies a shared bias instead of cancelling independent errors. For the full science of the four conditions and their failure modes, see our guide to the wisdom of crowds.

The fine print: what the fair got right that meetings get wrong

Here is the uncomfortable question for anyone who runs meetings: how many group decisions in your organization are structured like Galton’s fair — and how many are structured like the exact opposite?

At the fair, every judgment was committed privately, in writing, before anyone saw anyone else’s number. In a typical meeting, the most senior or most confident person speaks first, and every subsequent “estimate” is anchored on theirs. What looks like thirty people agreeing is often one person’s guess with twenty-nine echoes. Economists call the mechanism an information cascade: once a few early opinions are visible, it becomes individually rational for each subsequent person to discount their own information and copy — and the crowd’s independence, the load-bearing condition of the entire phenomenon, quietly collapses (Bikhchandani, Hirshleifer & Welch, 1992). We cover the mechanics in our guide to information cascades.

This is also why replications of Galton’s experiment sometimes disappoint: run the ox-guessing game with a visible running average, or let participants shout their guesses, and accuracy degrades. The magic was never in the crowd — it was in the protocol. Sixpence, a private card, and a locked ballot box turn out to be a better decision architecture than most conference rooms.

Modern descendants of the ox

The aggregation principle Galton documented now runs through a remarkable range of modern systems:

Prediction markets

The Iowa Electronic Markets, founded in 1988, aggregate traders’ beliefs into prices. Berg, Nelson and Rietz (2008) compared the market against 964 national polls across the 1988–2004 presidential elections: the market was closer to the outcome 74% of the time.

Ensemble machine learning

Random forests and other ensemble methods are Galton’s ox in silicon: many imperfect, partially independent models are aggregated — by voting or averaging — into a prediction more accurate than any single model.

Forecast averaging

Combining forecasts — whether from economists, weather models, or election forecasters — routinely beats most individual forecasts, for exactly the error-cancellation reasons Galton stumbled onto.

Estimation exercises

The experiment is easy to replicate: ask a group to estimate a quantity independently and in writing, then aggregate. Classroom and workshop replications remain a standard demonstration of statistical aggregation.

Each descendant aggregates something different — markets aggregate beliefs into prices, ensembles aggregate model outputs into predictions, voting systems aggregate preferences into outcomes — and each inherits the same dependency: the aggregate is only as good as the independence and diversity of what feeds it. That family resemblance is the through-line of the whole research tradition, from Condorcet’s 1785 theorem to the collective intelligence research of MIT’s Center for Collective Intelligence today — a 240-year arc we trace in Collective Intelligence: 240 Years of Research.

What Galton’s ox means for your team

The practical translation of the 1906 experiment is a short checklist, and it applies at one specific point in the life of a decision — the moment you gather and combine judgments, before the group starts deliberating:

  • Collect estimates independently, in writing, first. Before the meeting, before the senior person speaks. The moment judgments become visible, independence — and the crowd’s statistical advantage — starts to decay.
  • Actively recruit diverse perspectives. The diversity prediction theorem is blunt: your collective advantage over your average member is your diversity. A panel of people with the same background and the same information is one guess wearing many name tags.
  • Choose your aggregation mechanism deliberately. Median for messy crowds, mean for well-behaved ones, structured rating when judgments have multiple dimensions. “We discussed it and reached consensus” is not an aggregation mechanism — it is usually an anchoring mechanism.
  • Keep the reasons, not just the numbers. Galton’s cards told him what the crowd thought, never why. For an ox, that’s fine. For a strategy decision, the reasoning is the part you’ll need when circumstances change.

That last point is where modern collaborative decision making goes beyond Galton. Argumentree is built to give a team the fairground protocol with the reasoning attached: participants contribute arguments independently and asynchronously — before the room converges — into a structured pro/con tree; the group then rates each argument on explicit dimensions (helpfulness, clarity, accuracy, completeness), and the ratings aggregate into consensus scores the way 787 cards aggregated into 1,207 pounds. Anonymous contribution options protect independence from hierarchy, and the full audit trail preserves what the ox contest never could: the why behind the number.

Galton went to a fair to prove that ordinary judgment couldn’t be trusted. He left with the strongest evidence anyone had yet produced that — aggregated correctly — it can. The instrument has changed from a ballot box to software. The lesson hasn’t.

Frequently Asked Questions

What was Galton’s ox experiment?

In autumn 1906, the statistician Francis Galton analyzed a weight-judging contest at the West of England Fat Stock and Poultry Exhibition in Plymouth. Fairgoers paid sixpence to guess the weight of a live ox after it had been "slaughtered and dressed." Galton collected the cards — 787 were legible and valid — and found that the middlemost (median) guess was 1,207 pounds against an actual dressed weight of 1,198 pounds: within about 0.8%, and better than the estimates of the cattle experts present. He published the analysis as "Vox Populi" in Nature in March 1907.

Why is Galton’s ox experiment important?

It is the founding empirical demonstration of the wisdom of crowds: under the right conditions, the aggregate of many independent, imperfect judgments can be more accurate than the judgment of individual experts. The result surprised Galton himself, who had expected to demonstrate the incompetence of the average voter. James Surowiecki opened his 2004 book The Wisdom of Crowds with the story, and the underlying statistics now power prediction markets and ensemble machine learning.

Did Galton use the mean or the median?

In the original Nature article, Galton reported the "middlemost estimate" — the median — of 1,207 lb, arguing that it was the democratically fair choice because half the crowd thought the answer higher and half lower. In follow-up correspondence he acknowledged that the arithmetic mean of the guesses was 1,197 lb — just one pound off the true 1,198 lb. The exchange is an early argument about robust statistics: the median resists distortion by wild outliers, while the mean extracts more information when errors are roughly symmetric.

What conditions does a crowd need to be wise?

James Surowiecki (2004) identified four: diversity of opinion (people hold different information and interpretations), independence (judgments are formed without copying others), decentralization (people draw on their own local knowledge), and aggregation (a mechanism combines the private judgments into a collective answer). Galton’s contest satisfied all four — which is precisely why it worked. Remove any one, and the crowd’s advantage erodes or reverses.

When do crowds fail?

Crowds fail when independence breaks down — when people can see and copy each other’s answers, early guesses cascade through the group, errors become correlated, and the aggregate drifts toward whatever the first loud voice said. Information cascades, groupthink, and echo chambers are all versions of this failure. The lesson for teams: collect judgments independently and in writing before anyone discusses them out loud.

Is the wisdom of crowds the same as majority voting?

They are related but not identical. Galton’s ox involved aggregating numerical estimates, where averaging cancels errors. Majority voting on yes/no questions is governed by the Condorcet Jury Theorem (1785), which shows that majorities become dramatically more reliable as group size grows — provided each voter is better than chance and votes independently. Both are aggregation mechanisms; voting aggregates discrete choices, while averaging aggregates quantities.

AT

Argumentree Team

Decision Science

The Argumentree team explores the science of better decisions—from 18th-century mathematics to modern AI.

Give your team the fairground protocol.

Argumentree collects judgments independently, aggregates them into consensus scores, and keeps the reasoning on the record — the conditions that made Galton’s crowd wise, built into your decisions.

Start Free 14-Day Trial →

Related Articles