What Makes a Delphi Study Defensible
Four design features separate a defensible Delphi from a survey with extra steps.

Reviewers may be interested in the topic, but the methods still have to answer a simple question: how do we know the result was produced by the panel rather than steered by the investigators? Three parts of the answer are the classic Delphi features described in early RAND work—anonymous response, iteration with controlled feedback, and statistical aggregation. The fourth is the record needed to demonstrate that those features were executed as planned. RAND traces the method to work begun in the 1950s, with foundational publications appearing in the following decade.
Anonymity of response, not secrecy about the panel
The manuscript should disclose how panelists were selected and describe the panel well enough for readers to assess its credibility. That does not mean every panelist must be named. Identity disclosure depends on consent, ethics approval, and the study design; what matters methodologically is stating exactly who was anonymous to whom, at which stages, and how anonymity was protected.
Blinded ratings reduce the chance that hierarchy, reputation, or a forceful personality anchors other panelists. A modified Delphi can include a meeting or discussion, but direct interaction changes the influence structure. Report the modification, who attended, whether discussion occurred, and whether the final rating remained anonymous. The ACCORD explanation and elaboration gives practical examples of these different anonymity arrangements.
Iteration with controlled feedback
Panelists must be able to reconsider their judgments in light of a controlled summary of the group's previous responses. What they see is a design decision, not an afterthought: median and interquartile range, a full distribution, de-identified rationales, their own prior rating, or a combination.
Each choice can shape convergence. Central tendency alone can hide polarization; a distribution can preserve visible disagreement. The feedback policy should be specified in the protocol, applied consistently, and reported clearly. If feedback is personalized—for example, by including each panelist's own prior rating—the rule for personalization should be the same for everyone.
Statistical aggregation with pre-specified rules
Consensus is a rule you declare, not a pattern you notice. The rating scale, statistic, threshold, treatment of missing or unable-to-rate responses, and stopping rule should be defined together before the results are known. We return to this in the next post, because this is where apparently reasonable post hoc decisions can make a study impossible to falsify.
A complete record of what changed and why
Items get reworded between rounds. Panelists drop out. A domain may be split after free-text comments reveal two distinct concepts. These changes can be legitimate, but readers need to be able to reconstruct them. The failure is not change itself; it is losing the link between the original item, the rationale, the revised version, the round in which it appeared, and the analysis that used it.
This auditability requirement is not a fourth historical definition of Delphi. It is the evidence layer that makes the other three features defensible in a protocol, manuscript, or review.
Where fragmented tooling breaks down
None of these features is difficult in isolation. They break down when the study lives across a survey platform, several spreadsheet versions, and an email thread. The item text of record exists in one place, ratings in another, and the reason for a change only in someone's memory. By manuscript time, reconstructing the study record becomes archaeology.
Delphi Studio treats methodology as structured study data. The launch snapshot preserves the declared protocol and consensus configuration; per-round rule snapshots preserve the rules actually applied; item versions retain lineage; and controlled-feedback packets are frozen per panelist. The audit trail and reporting exports then draw from the same records that ran the study.
Next post: why consensus thresholds should account for both inclusion and exclusion, and why you should declare them before Round 1 opens.