Deliberation

Sharing Opinions Changed Nothing: What a Science Study Found About Structured Deliberation

AT
Argumentree Team
Deliberation & Decision Science
September 2, 2026
12 min read
Sharing Opinions Changed Nothing: What a Science Study Found About Structured Deliberation

Does structured deliberation actually change what a group thinks?

A preregistered experiment published in Science in October 2024 (Tessler et al., Science 386, eadq2852) tested this directly with 5,734 UK participants. One group of participants read each other's written opinions on divisive questions; group agreement did not move at all (b = -0.005, P = 0.64). Comparable groups whose opinions were instead synthesised into a candidate group statement, ranked, critiqued and redrafted converged by about eight percentage points, and the proportion of groups reaching unanimity rose from 22.8% to 38.6%. The difference between the two conditions was statistically significant (P = 0.0058). The synthesis was performed by an AI system the authors called the Habermas Machine, whose statements were preferred to those of trained, financially incentivised human mediators 56% of the time. Statements produced after the critique round gave minority positions more weight than their share of the group, not less. The authors state clear limits: participants were all UK residents, groups contained only five people, the system cannot fact-check or moderate, and an exploratory analysis found similar convergence from human mediators, suggesting the effect may come from mediated deliberation as a process rather than from the AI specifically. The practical finding for anyone running a discussion is that exchanging views is not the mechanism that produces agreement; structuring them, and running a round of explicit objection, is.

Share:
TL;DR

In October 2024, Science published a preregistered experiment with 5,734 participants that included a control condition almost nobody has noticed. One set of groups simply read each other's opinions on divisive political questions. Their level of agreement did not change at all. Groups whose same opinions were instead structured into a candidate group statement, ranked, critiqued and redrafted converged by eight percentage points, and unanimity nearly doubled. Free exchange of views is not what produces common ground. Structure is.

  • The control result is the finding. Reading each other's opinions moved group agreement by b = -0.005 (P = 0.64). Structuring the same opinions moved it eight percentage points (P < 0.001 in all three cohorts).
  • Unanimity rose from 22.8% to 38.6%, and 66.6% of participants said the group statement expressed their view better than their own original wording did.
  • The critique round is causal, not an artefact. In a separate cohort of 245 people, redrafts that genuinely used the critiques beat redrafts that secretly ignored them (P = 0.0039).
  • Structure protected the minority. Post-critique statements weighted minority views at 0.36 against a true share of 0.29. This was not majority rule.
  • The honest limits: UK participants only, groups of five, no fact-checking, and human mediators produced similar convergence on the rounds they won. The mechanism may be mediated structure itself rather than the AI.

Five strangers were given a divisive question. Should the National Health Service be privatised? Should the voting age drop to 16? Each of them wrote their honest position in private, in their own words, at reasonable length. Then they read what the other four had written. Their stated positions were measured before, and again after.

Nothing moved. Not slightly, not in the wrong direction, not by an amount too small to matter. The measured change in group agreement was minus 0.005, with a p-value of 0.64, which is as close to literally nothing as an experiment can report.

This was the control condition in a study published in Science on 18 October 2024. Comparable groups, given the same questions and the same amount of time, whose written opinions were instead synthesised into a candidate group statement, ranked, critiqued and redrafted, converged by about eight percentage points. The proportion of groups reaching unanimity rose from 22.8% to 38.6%. The study is usually reported as a story about an AI beating human mediators. The far more useful result is the one that got no headlines: exchanging views does not produce common ground. Structuring them does.

Letting people read each other's opinions moved group agreement by minus half a percentage point.
Structuring those same opinions moved it by eight.

Tessler et al., Science 386, eadq2852 (2024) — opinion-exposure cohort vs. main task cohorts

What the study actually did

The paper is AI can help humans find common ground in democratic deliberation by Michael Henry Tessler, Michiel Bakker and colleagues at Google DeepMind, Oxford and Harvard. It reports seven experiments run between January and August 2023 with a combined 5,734 UK participants. Sample sizes, exclusion criteria and primary analyses were preregistered before data collection, and the code was released publicly. That combination of a preregistration, a genuine control condition, publication in Science and open code is rare enough in this area to be worth stating plainly.

Participants worked in groups of about five and discussed three questions in roughly an hour. For each question the procedure was fixed. Everyone wrote their own position privately first, between 10 and 200 words, averaging about 65. Those written opinions were passed to a system the authors named the Habermas Machine, after the philosopher who argued that agreement emerges when people reason together under fair conditions. The system drafted several candidate statements of the group's common ground. Participants ranked them, and the winner was chosen by a ranked-choice election in which every participant's vote counted equally. The winning statement went back to the group, each person privately wrote a critique of it, and the system produced a revised draft from those critiques. Participants ranked again and finally chose between the original and the revised version.

Two details are worth holding onto. Participants were told the statements were machine-generated, so no deception was involved. And the questions themselves were deliberately contentious, covering immigration, the retirement age, prison numbers, Brexit, climate targets and the minimum wage. These were not warm-up topics chosen to make agreement easy.

The four things the procedure did

Strip out the machine learning and the procedure contains four moves. Each of them is available to any group with a shared document and some discipline.

Everyone commits to a position before hearing anyone else

Opinions were written privately, in full sentences, with reasoning attached. Nobody spoke first and anchored the room, and nobody had the option of quietly agreeing with whoever sounded most confident. A written record of what each person actually thought existed before any social pressure could act on it.

The positions are synthesised into one artefact, not stacked in a list

This is the step the control group did not get. They received exactly the same information, all five opinions in full, as a pile of separate texts. The treatment groups received it as a single drafted statement of where the group actually stood. Same information, different structure, and only one of the two produced movement.

Selection is a ranked vote with equal weight

The winning statement was chosen by aggregating every participant's ranking, one vote each. Nobody's preference counted more for being louder, faster, more senior or more articulate. This is the part that is trivially easy to describe and consistently hard to do in a live meeting.

There is a formal round for objection, and the draft has to change

Each participant privately critiqued the winning statement, and those critiques went back into a revision. Objection was not an interruption or a social risk; it was a scheduled step that everyone performed. Across three cohorts, 68%, 64% and 66% of participants preferred the revised statement to the original.

Why the control condition is the important result

A reasonable objection to any study like this is that people converge simply because they finally heard what everyone else thought. The authors anticipated it and ran the test. An additional cohort of 372 participants in 75 groups did the identical exercise, except that instead of rating a drafted group statement they rated each other's individual opinions. Full exposure to every view in the group, no synthesis.

Group agreement in that condition did not change (b = -0.005, t(230) = -0.46, P = 0.64), and the difference between the two conditions was itself statistically significant (P = 0.0058). This is a controlled, preregistered result, published in Science, showing that circulating everyone's opinions does not move a group toward agreement. It is the closest thing the literature has to a direct test of the comment thread, and the comment thread failed it.

The other half of the mechanism was tested just as carefully. Revised statements might be preferred simply because people saw them second, after investing a little more effort. So the authors ran a further cohort of 245 participants who wrote critiques and were then shown redrafts that either did or did not actually incorporate them. The genuinely critique-informed versions still won (P = 0.0039). The objection round does real work.

Structure did not flatten the minority. It amplified it.

The obvious fear about any consensus-seeking process is that it becomes majority rule with better manners. The authors measured this rather than asserting it, by embedding every opinion and every statement in a vector space and locating them along the axis running from agreement to disagreement on each question.

In roughly 77% of rounds the group started out genuinely divided. The initial group statements weighted minority views at 0.29, against a true average minority share of 0.29, which is proportional representation almost exactly. The revised statements, written after the critique round, weighted them at 0.36. The minority was over-represented relative to its size, and the group preferred those revised statements anyway. Ninety-six percent of group statements fell inside the range actually spanned by the group's own opinions, so the system was not inventing positions nobody held.

Support did still shift toward majority positions overall, and the authors say so. Majorities grew in 29% to 32% of rounds and minorities grew in 20% to 26%. But when they tested whether people moved toward the majority merely because they had been shown majority-leaning statements, they found no relationship at all (P = 0.37). Whatever moved people was the content of the argument, not repeated exposure to it.

The machine's first draft matched the majority exactly.
After one round of objection, it over-represented the minority.

Minority weighting rose from 0.29 to 0.36 against a true minority share of 0.29

It beat the human mediators, and that matters least

The headline result was a head-to-head against people. In 75 groups, one member was assigned to be the mediator, trained to write group statements and offered a bonus of up to 50% of their base pay for every statement the group preferred. Across those rounds the machine's statements were preferred 56% of the time to the humans' 44%, and were rated higher on both endorsement and quality. A separate panel of 155 independent raters judged them clearer, more informative, less illogical, and less likely to smuggle in the writer's own opinion.

This is the finding that travelled, and it is the one to lean on least. The authors themselves are careful here: they call the real opportunity not superhuman mediation but deliberation that is time-efficient, fair and scalable. The system produced statements in seconds where humans took minutes. And in an exploratory analysis, the authors found similar convergence from human mediators on the rounds where the humans won, which suggests the effect may come from mediated structure as a process rather than from anything specific to the AI. If that reading is right it makes the study more useful rather than less, because the benefit is then available to any group willing to adopt the procedure, with or without a model in the loop.

It replicated with a representative sample, and it did not work on everything

The main cohorts were recruited by convenience through a crowdsourcing platform, which is a fair thing to be sceptical of. So the authors partnered with the Sortition Foundation, which runs real citizens' assemblies, to recruit 200 participants stratified to represent the UK on age, gender identity, ethnicity, socioeconomic status and region. They deliberated across three weekly one-hour sessions in April and May 2023, over nine questions. Endorsement levels and the increase in group agreement both replicated.

The interesting part is where it did not work. Across the nine questions, positions converged on five. Support for reducing the prison population rose from 60% to 75%, and for easing asylum entry from 39% to 51%. On Brexit there was no movement whatsoever. On universal free childcare, opinion moved the other way, with opposition rising from 33% to 41%. A process that produced agreement on every question regardless of subject would be evidence of something going wrong. This one left entrenched positions entrenched, which is roughly what an honest deliberation should do.

The limits, stated plainly

Any claim built on a single study should carry that study's own caveats, and this one has real ones. The authors list most of them themselves.

Everyone was British, discussing British questions

All participants were UK residents deliberating on UK policy. The authors say they have no particular reason to expect the findings are UK-specific, but they did not test it. Nothing here has been shown to hold in a different political culture.

Groups of five, not fifty

Every group was roughly five people. The authors argue the method scales in principle with longer-context models, but scale was not tested. Small-group results do not automatically survive contact with a hundred participants, and the problem of verifying that a synthesis is faithful gets much harder as the input grows.

It cannot fact-check, moderate, or keep anyone on topic

In the authors' own words, if the human opinions are ill-informed or harmful then the output may be ill-informed or harmful. The system synthesises what it is given. It has no view on whether any of it is true, and the study's questions were pre-vetted to reduce the risk of harmful discussions.

The training sample was not representative

The model was tuned on crowdsourced participants among whom over-65s were underrepresented. The assembly replication addresses the evaluation sample, not the data the system learned preferences from.

It assumes everyone is arguing sincerely

The procedure takes each written position at face value. Someone who deliberately misstates their view to drag the synthesis toward a preferred outcome is not modelled anywhere, and what counts as rational strategic behaviour under this procedure is an open research question.

The AI may not be the active ingredient

The authors' own exploratory analysis found comparable convergence from human mediators on the rounds they won. The honest reading is that the study demonstrates the value of structured, mediated deliberation, and demonstrates that a model can perform the mediating role competently and quickly. It does not establish that the model is why it worked.

Where the critics have a point

The paper attracted serious criticism from deliberative democracy scholars, and some of it lands. The sharpest comes from Alvaro Oleart and Nicolás Palomo Hernández in the Journal of Deliberative Democracy, who point out that in this procedure the participants never actually argue with each other. They write in private, then rank machine output. There is no rebuttal, no exchange of reasons between people, no moment where one person's argument changes another person's mind directly. They call it deliberation without citizens, and they note the irony of the name: Habermas treated agreement as a precondition of genuine discourse, not as the target to optimise for.

Beth Simone Noveck raises a different objection: consensus is not the hard part of governance. Groups more often fail to agree on what the problem is than on what to do about it, and a machine that is excellent at producing agreement on a well-posed question does not help with a badly posed one. Researchers at the Knight First Amendment Institute add the oversight problem, which is that once a synthesis is drawn from a thousand opinions, no human can verify that it represents them faithfully.

We think the first criticism is the most important one, and we agree with it. Optimising a process toward agreement is not the same as helping people reason together, and a system that produces consensus by routing around human argument has removed the thing that made deliberation valuable in the first place. The finding to carry forward is not let a model write your group's position. It is narrower and more durable: positions have to be explicit, structure beats circulation, and a formal round of objection changes the outcome. Those hold whether the synthesis is done by a model, a facilitator, or the group itself.

What to do with this on Monday

None of the four mechanisms requires an AI, and all four are things most meetings get wrong. Collect written positions before the discussion, not during it. The study's participants wrote roughly 65 words each with their reasoning attached, which is about five minutes of work and is the single cheapest change available. Synthesise rather than circulate. Forwarding everyone's comments is precisely the condition that produced no movement at all; someone has to draft what the group actually holds in common and put it up for correction. Decide by an equal-weight ranking rather than by who is still talking at the end. And schedule the objection round, so that disagreeing is a task everyone performs rather than a social risk somebody has to take.

If you run online discussions, the control result deserves particular attention, because a discussion board or comment thread is the control condition. It circulates opinions and synthesises nothing. That is the structural case made in discussion board alternatives, and it now has a preregistered experiment behind it. The four-phase structure in how to structure a debate enforces the same moves by hand, and the distinction between agreeing and merely not objecting is covered in consent vs consensus.

What this does and does not say about Argumentree

We should be precise about this, because the temptation to overclaim is obvious. The Habermas Machine is not what we built. It generates candidate consensus statements using a personalised reward model that predicts how each participant will rank each draft, and it selects between them with a simulated election. Argumentree does not do that, and this study does not test our product.

What it does test is the mechanism our product is built on. Positions get written down explicitly, with their reasoning, before the discussion converges. Arguments are structured into a shared artefact rather than circulated as a stack of posts. Objection is a first-class action with somewhere to attach, rather than an interruption. Minority positions stay visible and attributable instead of dissolving into the majority view. Those are the four things the study isolated, and they are the four things a structured argument map does by default. The difference is that our synthesis stays inspectable: every claim keeps its author, its objections and its support, so nobody has to trust a summary they cannot check, which is exactly the oversight problem the critics raise.

The test

After your next discussion, ask whether anyone's position actually changed. If everybody read everything and nobody moved, you did not run a deliberation. You ran the control condition.

The result worth remembering

Strip away the model, the citizens' assembly and the argument about whether an AI should be anywhere near a political process, and one number survives all of it. Five people, full access to each other's written reasoning on a question they cared about, measured before and after: P = 0.64. The most common way groups try to reach agreement, which is to put all the views in front of everyone and let people read, has now been tested against a control and produced no measurable effect.

The condition that worked added no new information. Every participant already had every opinion. What changed was that the opinions were given a shape, ranked with equal weight, and subjected to one formal round of objection. That is a procedural finding rather than a technological one, and it is available to anyone willing to run their discussions differently.

The group already had all the arguments. What it lacked was a structure to put them in.

Run the structure, not the thread

Argumentree gives every position an explicit place, every objection somewhere to attach, and every minority argument a record that survives the discussion.

Sources

Frequently Asked Questions

What did the 2024 Science study on AI-mediated deliberation actually find?

It found that structuring a group's written opinions into a candidate group statement, ranking it by equal-weight vote, critiquing it and redrafting it increased group agreement by about eight percentage points, while simply letting the same participants read each other's opinions produced no change at all (P = 0.64). Statements produced this way were preferred to those written by trained, financially incentivised human mediators 56% of the time. The study involved 5,734 UK participants and was preregistered.

Does the study prove that AI is better than people at running a discussion?

No, and the authors do not claim it. The AI produced statements in seconds rather than minutes and was preferred 56% to 44%, but an exploratory analysis in the same paper found similar convergence from human mediators on the rounds they won. The most defensible reading is that mediated structure is what works, and that a model can perform the mediating role competently and fast. The procedure itself does not require an AI.

Did the process just impose the majority view on everyone?

No. Initial group statements weighted minority positions at 0.29 against a true minority share of 0.29, and post-critique statements weighted them at 0.36, over-representing the minority relative to its size. Ninety-six percent of statements fell within the range of opinions the group actually held. Support did shift toward majority positions overall, but there was no relationship between how often people saw majority-leaning statements and how much they moved (P = 0.37), so exposure alone was not driving it.

What are the study's main limitations?

Participants were all UK residents discussing UK policy questions; groups contained only about five people, so nothing was tested at assembly scale; the system cannot fact-check, moderate or keep a discussion on topic, and the authors note that ill-informed input produces ill-informed output; the training sample underrepresented over-65s; and the procedure assumes participants state their views sincerely. The authors also concede that the AI may not be the active ingredient, since human mediators produced similar convergence on the rounds they won.

What is the Habermas Machine?

It is the name the researchers gave their system, after Jürgen Habermas, who argued that agreement emerges when people reason together under fair conditions. It combines a generative model that drafts candidate group statements from participants' written opinions with a personalised reward model that predicts how each participant would rank each draft, then selects a winner through a simulated equal-weight election. It was built on a fine-tuned Chinchilla 70B model, which outperformed a prompted Gemini 1.5 Pro at the same task.

How do I apply this without any AI at all?

Run the four moves the procedure isolated. Have everyone write their position and reasoning privately before any discussion; roughly 65 words is enough. Have someone synthesise those into a single draft of what the group holds in common, rather than circulating the opinions themselves, since circulating them is the condition that produced no effect. Choose between drafts by an equal-weight ranking rather than by who speaks last. Then run a formal round where everyone writes an objection, and revise the draft against it.

Related reading