How to Run a Delphi Study Where the Reasoning Survives the Rounds
The Delphi method gathers expert judgment through iterative anonymous rounds with controlled feedback: independent Round-1 responses, an aggregate fed back, revision in Round 2, repeated until convergence or stable disagreement. Incumbent Delphi tooling is survey software — it carries the numbers between rounds and loses the reasoning, so panelists see that the group moved without seeing why. To run a Delphi where the reasoning survives: frame a falsifiable proposition as the root claim; collect Round-1 positions as independent pro/con arguments (nobody sees others' before submitting); optionally seed additional positions from multiple AI models, visibly labelled; run Round-1 ratings as the quantitative record; make the controlled feedback the positions themselves with their rating spread — not a facilitator's paraphrase; let panelists interrogate positions in Round 2 through Q&A chains (four-turn dialogues with the position's author — AI-authored positions answer automatically, so hybrid panels are practical) and challenge them through parallel Review chains; reconcile near-duplicates through Compromise chains; measure convergence by comparing rating distributions across rounds; and report the consensus position AND the surviving dissent with its reasoning — stable disagreement is a finding, not a failure. Load-bearing limitation stated up front: Argumentree provides INDEPENDENCE (no one sees others' input before submitting), not ANONYMITY — arguments carry their author on a private tenant. If true anonymity is essential to your panel, run it on an Argumentree.AI public tenant, where unauthenticated contribution is supported. There is also no round manager (rounds are a facilitation convention you time-box) and no built-in convergence statistic (export ratings and compute Kendall's W or report the distributions).
The Delphi method's design goals — independence, iteration, measured convergence — describe an argument platform better than a survey tool. Run it so the reasoning survives:
- Decide the anonymity question first: a private tenant gives you independence (attributed authorship, no anchoring); true anonymity needs a public Argumentree.AI tenant
- Round 1: independent positions as arguments, rated by every panelist — the spread is the round's quantitative record
- Controlled feedback is the arguments themselves with their distribution — no facilitator paraphrase in between
- Round 2: interrogate through Q&A chains, challenge through parallel Review chains, reconcile duplicates — then measure convergence across the two distributions, and report the dissent
The panel that converged on nothing anyone could explain
Round 3 of the Delphi study closes and the facilitator emails the panel: consensus reached. The median moved from 6.1 to 7.8 on the ten-year adoption question; standard deviation halved; the report writes itself. Then one panelist — a careful one — replies with the question that unravels it: "I revised my estimate because the Round-2 summary said the group had strong reasons for optimism. Can I see those reasons? I'd like to know what I agreed with."
There is no good answer. The rounds were surveys. Between them, the panel saw medians, quartiles and a facilitator's two-paragraph summary. The reasons — whatever specific arguments moved eleven experts to revise upward — were typed into free-text boxes, compressed into the summary, and lost. The panel converged; on what reasoning, nobody can say.
This is not a broken study; it is standard practice with standard tooling, and it wastes the method's best property. Delphi was designed around a precise insight: anonymity and independence remove the dominance and status effects that distort face-to-face panels, and controlled feedback lets genuine judgment shift without social pressure. Survey software delivers the independence and loses the substance of the feedback. This tutorial runs the same rounds with the reasoning kept.
What survey-based Delphi loses between rounds
A Delphi round produces two things: a distribution (where the panel stands) and a body of reasoning (why). Survey tools are built for the first. The second gets flattened into free-text fields and then into a facilitator's summary — which quietly inserts an unaccountable editor between the experts and each other. Panelists in Round 2 are not responding to their peers' arguments; they are responding to a paraphrase of them. For a method whose entire epistemology is structured collective judgment, that is a strange thing to lose — the history of collective methods is largely a history of fights over exactly this feedback step.
What you need
A panel (classically 8–20 experts), a facilitator, and one Argumentree discussion per Delphi question. Rounds are time-boxed submission windows you announce; the tutorial assumes two rounds plus a report, extendable.
Setup: the question, the panel — and the anonymity decision
Before anything else, one decision that a methodologist will check first: what your panel actually requires — independence or anonymity — because they are not the same thing, and the platform gives you one of them by default.
- 1Frame the proposition as the root claim — falsifiable, not open-ended: "X will be standard clinical practice by 2032", not "what is the future of X?". A vague root produces unrateable positions. Checkpoint: the root can be argued for and against.
Independence ≠ anonymity — choose your tenant accordingly
On a private Argumentree tenant, arguments carry their author: what you get is independence — nobody sees anyone else's input before submitting, which removes anchoring and lets the panel write freely. It does not remove status effects in later rounds, because names are visible. Classic Delphi asks for anonymity. If your panel needs it — senior and junior experts mixed, politically exposed topic — run the study on an Argumentree.AI public tenant, where unauthenticated contribution is supported and positions arrive unattributed. State in your methods section which property you used; the trade-off is real and reviewers will ask.
Round 1: independent positions, then the numbers
- 1Independent submission window. Each panelist adds supporting or opposing positions under the root — without seeing others' until submitted. This is the independence the method requires, enforced by the medium rather than by trust in the facilitator's process. Checkpoint: every panelist has ≥1 position; nobody has seen or edited anyone else's.
- 2Optional: AI panelists, labelled. Seed additional positions from several models. Provenance is stamped on every argument, so machine positions are always distinguishable from expert ones. This widens the position space; it adds no expertise — the distinction matters and belongs in your report (see the hybrid-panel section below). Checkpoint: AI positions visibly labelled.
- 3Round-1 ratings. Every panelist rates every position — one value plus a label, with the anchors you defined ("1 = implausible … 9 = near-certain"). This is the round's quantitative record. Checkpoint: complete rating matrix; anchors stated once, used consistently.
Controlled feedback that carries the reasoning
Now the step survey tools cannot do. The feedback artifact for Round 2 is not a summary the facilitator wrote — it is the tree itself: every position, in its author's (or model's) own words, with its rating distribution attached. Panelists see exactly which arguments the group found strong, which split the panel, and which sank — with no editorial paraphrase in between. The facilitator's job shrinks to logistics: announce the window, point at the tree.
Survey tool, between rounds
Medians and quartiles per item; free-text comments compressed into a facilitator summary. The group's movement is visible; its reasoning is not.
Argument tree, between rounds
The positions themselves, each with its rating spread. A panelist who revises can point at the argument that moved them — and so can your methods section.
Round 2: interrogate, challenge, reconcile
Round 2 is where classic Delphi says "revise in light of the feedback." With the reasoning visible, revision becomes active — three structured moves, each a four-turn dialogue with a position's author:
- 1Interrogate — Q&A chain. Before revising toward a position, understand it: your question, the author's answer, follow-up, answer — complete. For AI positions, the model answers automatically, so the panel is never blocked waiting on one member. Checkpoint: positions that shifted opinion carry chains explaining why.
- 2Challenge — Review chain. Evaluate a position as unsupported; the author responds; follow-up; response. N panelists challenging one position = N parallel chains — which is exactly the structure a Delphi wants: parallel independent critique, never one group thread for status to re-enter through. Checkpoint: contested positions carry completed challenges.
- 3Reconcile near-duplicates — Compromise chain. Two experts holding almost the same position double-counts one view in the ratings. One proposes the merged formulation to the other, on the record. Checkpoint: no duplicate pair left unmerged or unexplained.
- 4Round-2 ratings. Re-rate everything under the same anchors. Checkpoint: second complete matrix, ready for comparison.
Hybrid panels: adding models without adding expertise
Because AI-authored positions answer their own chains — routed back to the model that wrote them — a hybrid human/AI panel is genuinely practical: use models to broaden Round 1's position space, then let the human experts cross-examine the machine positions in Round 2 and get substantive replies without anyone staffing the machine's side. Be precise about what this buys: wider input, not added expertise. A model's position is a well-phrased pattern, not a career of judgment, and the labelling exists so your convergence numbers can be computed with and without the machine rows. When widening input helps and when it merely dilutes is its own literature — when crowds go wrong covers the failure modes; the flip side, why independent judgments aggregate so well when they are genuinely independent, is the Galton and Condorcet story.
The question for your last panel
When it converged — could anyone say why? If the answer is a median that moved, you measured agreement without ever capturing what was agreed to.
Measuring convergence, reporting dissent
Convergence is the narrowing of rating distributions between rounds, and you measure it by comparing them — honestly, by hand or export: there is no built-in convergence statistic, no automatic Kendall's W, no stability test. Export the two matrices and compute what your field expects, or report the distributions qualitatively ("twelve of fourteen positions narrowed; two split further").
And report the dissent. A position where the panel stably disagrees across rounds — same spread, both poles defended in completed chains — is a finding, arguably the most valuable kind: it marks where expert judgment genuinely divides, with the reasoning for both poles attached. Averaging it away is the one unforgivable Delphi sin. The tree makes the dissent citable: the surviving objection, its author (or its anonymized public-tenant stand-in), and the challenges it survived.
Honest limitations
- ✗Independence is not anonymity — stated up front and worth restating: private-tenant arguments are attributed. If anonymity is methodologically essential, use the Argumentree.AI public tenant and say so in your write-up.
- ✗There is no round manager. No close-Round-1 button, no per-round visibility gating. Rounds are a facilitation convention you enforce with time-boxes and instructions — say so, and do not imply automation that is not there.
- ✗No built-in convergence statistic. Kendall's W, stability tests and the like are yours to compute from exported ratings.
- ✗Ratings are a single labelled value, not a Likert instrument with validated anchors. Define your anchors in the label and keep them constant across rounds.
- ✗A chain is four turns, then complete. An expert dispute that needs more gets a call — the chain records what was already asked and answered.
Practical lessons
- ✓Write the anchors before Round 1 and paste them into every window announcement. Half of all messy Delphi data is anchor drift.
- ✓Time-box windows generously across time zones — 72 hours per round works for international panels; the chains are async by construction.
- ✓Cap Round-1 positions per panelist (three is plenty). Delphi rewards considered positions, not volume.
- ✓Freeze the tree between rounds. Announce that submissions outside windows will be removed — the round convention only holds if the facilitator holds it.
The reply you can finally send
Back to the careful panelist's email: "Can I see the reasons? I'd like to know what I agreed with." In this version of the study, the answer is a link. Here are the fourteen positions; here is the one whose rating spread collapsed after two Review chains failed to dent it; here is the exchange that moved you and eleven others. Consensus with the reasoning attached — which was the method's promise all along.
Sources & further reading
- Dalkey, N., & Helmer, O. (1963). An Experimental Application of the Delphi Method to the Use of Experts. Management Science, 9(3).The original RAND formulation: iterative anonymous rounds with controlled feedback.
- Linstone, H. A., & Turoff, M. (eds.) (1975). The Delphi Method: Techniques and Applications. Addison-Wesley.The standard methodological reference, including the feedback-design variations this tutorial's controlled-feedback section addresses.
- Rowe, G., & Wright, G. (1999). The Delphi technique as a forecasting tool: issues and analysis. International Journal of Forecasting, 15(4).The critical review of what makes Delphi work — independence, iteration, feedback quality — and where implementations commonly fail.
Frequently Asked Questions
How do you run a Delphi study step by step?
Define a falsifiable proposition and recruit a panel (classically 8–20 experts). Round 1: each panelist independently submits positions for and against, without seeing others', then everyone rates every position under stated anchors. Controlled feedback: the panel sees the positions themselves with their rating distributions. Round 2: panelists interrogate positions through Q&A chains, challenge them through Review chains, reconcile near-duplicates, and re-rate. Compare the two rating distributions for convergence; repeat a round if needed; then report the consensus positions and the surviving dissent with its reasoning. Time-box each round's submission window — rounds are a facilitation convention, not an automated state.
What is the difference between independence and anonymity in a Delphi?
Independence means nobody sees others' input before submitting their own — it removes anchoring and first-speaker dominance. Anonymity means contributions are unattributed throughout — it additionally removes status effects when the panel reads and reacts to each other's positions. Classic Delphi asks for anonymity. A private Argumentree tenant provides independence with attributed authorship; if anonymity is methodologically essential (mixed seniority, politically exposed topics), run the study on an Argumentree.AI public tenant, where unauthenticated contribution is supported — and state which property you used in your methods section.
What is controlled feedback and why does it matter?
Controlled feedback is what the panel sees between rounds — classically an aggregate (medians, quartiles) plus a summary of reasons. It matters because it is the mechanism by which judgment legitimately shifts: experts revise in light of the group's view without social pressure. The standard failure is that survey tools feed back numbers plus a facilitator's paraphrase, so panelists respond to an editor's compression rather than to their peers' actual arguments. Feeding back the positions themselves, with their rating spread and no paraphrase in between, is the single biggest upgrade this tutorial makes to the standard workflow.
How do you measure consensus or convergence in a Delphi study?
By comparing rating distributions across rounds: convergence is the narrowing of the spread on a position, not a declaration by the facilitator. Export the round matrices and compute what your field expects — Kendall's W for concordance, quartile-deviation thresholds, or a stability test across rounds — because there is no built-in convergence statistic. Equally important: report stable disagreement as a finding. A position with the same wide spread in both rounds, with both poles defended in completed chains, marks a genuine divide in expert judgment, and averaging it away misrepresents the panel.
Can AI models be part of a Delphi panel?
As labelled input-wideners, yes. Seed Round 1 with positions from several models — provenance is stamped, so machine positions are always distinguishable — and because AI-authored arguments answer their own Q&A chains automatically, human experts can cross-examine a machine position in Round 2 and get a substantive reply without anyone staffing it. Two disciplines keep this honest: report convergence with and without the machine rows, and never describe the models as adding expertise. They broaden the space of candidate positions; the judgment remains the humans'.
How many experts and rounds does a Delphi need?
Classic guidance puts panels at roughly 8–20 for a focused question — small enough that every panelist can genuinely engage every position, large enough that the distribution means something. Two rounds resolve most questions once the feedback carries reasoning rather than numbers alone; add a third when the second round's distributions are still moving. More rounds than that usually signals a vague root proposition rather than an undecided panel — sharpen the claim before extending the study.
Run a Delphi your methods section can defend.
Independent rounds, feedback that carries the arguments, convergence you can measure, and dissent you can cite.
Start Free 14-Day Trial