The same species, opposite results
Plymouth, England, 1906: Francis Galton collects 787 guesses of an ox's dressed weight from a county-fair crowd and finds the middlemost estimate within one percent of the truth — the founding demonstration of crowd wisdom. A century later, crowds with vastly better information technology produced the housing-bubble consensus of 2008, meme-stock manias, and the confidently wrong pile-ons of part one of this series. Same species, same statistical machinery available — opposite results.
The resolution of the paradox is that crowd wisdom was never a property of crowds. It is a property of conditions, and the conditions are specific, known, and fragile. When they hold, aggregation is a superpower; when they break, the same aggregation amplifies error with the confidence of numbers. This post is about the three conditions — and about the uncomfortable fact that modern discussion platforms are condition-breaking machines.
The three conditions
The research line runs from Condorcet's 1785 jury theorem — majorities of better-than-chance, independent voters approach certainty as they grow — through Galton's demonstration to Surowiecki's modern synthesis. Three requirements recur:
- 1Independence. Each judgment forms before seeing the others'. This is what makes errors uncorrelated — one person's mistake doesn't become everyone's mistake. Galton's fair-goers wrote sealed guesses; they couldn't anchor on a leaderboard. Independence is the condition everything else depends on, and the first one platforms destroy.
- 2Diversity. Different information sources, backgrounds, and models — so errors point in different directions and cancel in aggregate. A thousand copies of the same perspective is one perspective with confidence; diversity beats ability precisely because it buys error-cancellation ability can't.
- 3Proper aggregation. A mechanism that actually combines the information: a median, a market, a structured score. Aggregation is not neutral — a mechanism that weights early, loud, or emotionally resonant contributions isn't aggregating information; it's amplifying momentum.
How platforms break all three at once
Social platforms are often described as harnessing crowd wisdom. Mechanically, they do the opposite — each core feature attacks one condition:
- ✗Visible counts kill independence. Every like, upvote and score total is social proof displayed before you judge — the anchoring and cascade machinery that turns crowds into herds. Judgments correlate; errors compound instead of canceling.
- ✗Feeds kill diversity. Engagement-optimized ranking shows you more of what you responded to — narrowing each person's information diet and homogenizing the crowd's inputs. The diversity that should cancel errors gets algorithmically selected out.
- ✗Vote-sorting is improper aggregation. As part one showed, upvote totals blend agreement, timing and entertainment. Sorting by that number doesn't combine the crowd's information — it crowns whatever moved fastest emotionally.
Before trusting any 'the crowd says' number, ask three questions: Did the judgments form independently? Did they draw on different information? Does the aggregation mechanism combine information or amplify momentum? Three yeses are rare — and anything less isn't crowd wisdom; it's crowd volume.
The twist: small structured debates can beat huge crowds
For decades the practical advice was 'protect independence at all costs' — deliberation was suspect because talking correlates errors. Then came a finding that sharpened the picture. Navajas and colleagues, in a study with 5,180 participants, had people answer general-knowledge questions individually, then deliberate in small groups of five and produce consensus answers. Averaging just four of those small-group consensus answers outperformed aggregating thousands of independent individuals.
The lesson is not that independence was wrong — it's that deliberation adds information that averaging can't reach, when it's structured. In the groups, people exchanged reasons: outright errors got caught, implausible estimates got challenged, and the consensus encoded argument quality, not just position. Unstructured virality correlates errors without exchanging reasons; structured deliberation exchanges reasons without (if run well) correlating errors. The variable that matters isn't whether people interact — it's what structure the interaction has.
The honest counterargument: isn't any deliberation contamination?
The steelman of the purist position: once people talk, errors correlate — cascades, conformity, and anchoring are exactly what the failure literature documents, so shouldn't serious aggregation stay silent and independent, prediction-market style?
The answer the evidence supports: independence is the right rule for the estimation step, not for the whole process. The failure cases correlate errors before individuals have committed to a judgment — visible counts, live discussion, anchors. The success cases (Navajas's groups, structured deliberation generally) collect independent positions first, then let reasons collide under structure, then aggregate. Sequence and structure, not silence, are what protect the crowd. That is a design specification — and it's buildable.
How Argumentree implements the three conditions
An argument tree is, mechanically, the three conditions turned into software:
Independence: asynchronous written entry
Positions and ratings form individually — contribution doesn't require reading the room first, and rating happens per-user on each argument, not by following a visible momentum number.
Diversity: pro AND con are structural
The tree displays supporting and attacking branches side by side, so the diversity of views isn't just permitted — its absence is visible. An empty con branch is a warning, not a victory.
Aggregation: recursive merit scoring
An argument's total combines its own ratings with the totals of its supporting children minus its attacking children — so scores encode how well each claim survived examination, propagated up the tree. Failed rebuttals strengthen what they attacked; that's argument quality being aggregated, not applause.
Deliberation: structured, not viral
Challenges run as bounded exchanges attached to specific claims — the Navajas ingredient (reasons colliding) without the cascade ingredient (momentum recruiting votes).
Upvote platforms correlate judgments and amplify momentum. An argument tree collects positions independently, displays disagreement structurally, and aggregates by how claims survive challenge — the three conditions, built in. See collective intelligence and The Argumentree Method.

