Design a defensible modified Delphi study
A defensible Delphi study is not simply a survey sent to people described as experts. It is a governed sequence of independent judgment, controlled feedback, re-rating, and transparent aggregation. The method is most useful when evidence is incomplete, inconsistent, or difficult to apply and a structured account of informed judgment will answer a meaningful research question.
This tutorial connects methodological decisions to the places where Delphi Studio records and enforces them. It is a design guide, not a substitute for a study-specific protocol, ethics review, statistical analysis plan, or applicable reporting guideline.
The four defining features
| Feature | Methodological purpose | Delphi Studio implementation |
|---|---|---|
| Independent responding | Limits status, dominance, and conformity pressure | Panelists respond privately and never see peers' identities or individual responses |
| Iteration | Gives panelists an opportunity to reconsider judgments | Versioned, multi-round rating workflow |
| Controlled feedback | Supports reflection without unmoderated debate | Researcher-reviewed group summaries and locked feedback snapshots |
| Structured aggregation | Makes the group result explicit and reproducible | Pre-specified item-level rules, descriptive statistics, and supplementary concordance measures |
These features are consistent with the classic Delphi model of anonymous response, iteration with controlled feedback, and an aggregated group response described in the RAND literature.
:::note Terminology is not a protocol “Modified Delphi” is an umbrella label, not a complete method specification. Published implementations vary. State exactly what was modified, how initial items were generated, what feedback was provided, how consensus and stability were defined, and why those choices fit the research question. Nasa, Jain, and Juneja provide a useful healthcare-methodology overview. :::
1. Confirm methodological fit
Begin with the decision the study must support—not with a questionnaire.
Good reasons to use Delphi
- Develop recommendations where evidence is limited, inconsistent, or difficult to operationalize.
- Prioritize outcomes, competencies, indicators, research questions, or implementation strategies.
- Refine definitions or classification criteria in an emerging field.
- Identify both areas of agreement and persistent uncertainty.
- Combine evidence with the judgment of people who hold relevant professional or lived expertise.
Reasons to choose another method
| Need | Better starting point |
|---|---|
| Estimate the prevalence of an opinion in a population | Representative survey or probability sample |
| Answer a question adequately resolved by empirical evidence | Evidence synthesis or primary empirical study |
| Negotiate interests or jointly solve a problem in real time | Facilitated workshop, nominal group, or deliberative process |
| Validate a predetermined conclusion | Reframe the project; consensus must not be used as ceremonial endorsement |
Write down the intended output before creating items. Examples include “define a minimum outcome set,” “prioritize implementation barriers,” and “map unresolved disagreement.” Delphi Studio records the objective, intended audience, consensus question, and scope in the Study Design Wizard.
2. Define what “modified” means
Traditional and modified Delphi studies often share the same core safeguards but differ in how the first item set is produced.
| Design feature | Traditional Delphi | Modified Delphi |
|---|---|---|
| Starting point | Often begins with open idea generation | Often begins with candidate items prepared before Round 1 |
| Common item sources | Panelist proposals | Reviews, prior studies, interviews, frameworks, or steering-group work |
| Typical emphasis | Exploration and concept generation | Refinement, prioritization, validation, or appropriateness |
| Reporting obligation | Explain how open responses became rateable items | Explain item sources, steering-group influence, and every other modification |
Starting from an evidence-informed item bank can reduce burden, but efficiency alone is not a methodological justification. Report how the item set was assembled, piloted, edited, and checked for missing perspectives.
3. Engineer the protocol before recruitment
The protocol should make consequential decisions before anyone sees results.
Define the construct and scale
State what every rating means: importance, agreement, feasibility, appropriateness, inclusion, or another construct. Define and report the anchors for every scale point or category. A 1–9 importance scale and a 5-point agreement scale are different measurement designs, even if both ultimately produce an “include” decision.
Use the Rating and consensus step to configure the primary dimension, scale, anchors, zones, and classification predicates.
Pre-specify consensus and item handling
There is no universal Delphi consensus threshold. Select a rule that fits the construct, panel, and intended use; justify it; and state how missing or unable-to-rate responses enter the denominator.
For example, a protocol might define inclusion on a 9-point scale as:
- at least 75% of valid ratings in the high zone (7–9); and
- no more than 15% in the low zone (1–3).
:::caution Example, not Delphi Studio's default The 75/15 rule is illustrative. Delphi Studio's current Modified e-Delphi preset uses a 5-point agreement scale and requires at least 80% in the agree zone, median at least 4, and IQR at most 1. Presets are starting points, not universal standards. :::
Also specify what happens to each result class:
- consensus to include or retain;
- consensus to exclude or drop;
- no consensus or indeterminate;
- polarized or stakeholder-divergent;
- insufficient data; and
- revised wording requiring another rating round.
Pre-specify stopping and stability
Maximum rounds and consensus are not the same as stability. A panel may remain stably divided, which is a meaningful result rather than a failed study.
Define stability operationally. Delphi Studio currently compares item medians and IQRs between rounds, classifies each eligible item-dimension as stable or unstable, and calculates the proportion that remains unstable. The resulting stopping recommendation is advisory; the research team remains responsible for the final decision and its rationale. See Consensus analytics.
Govern changes and overrides
The launch snapshot freezes the declared methodology. Per-round rule snapshots preserve the exact rules used for classification. Where the protocol permits a later amendment or disposition override, Delphi Studio records a new version or audited rationale rather than rewriting historical results.
:::info Guarding against rule drift Snapshots and audit events make post-result changes visible and preserve the rules that were actually applied. They reduce the risk of moving thresholds after seeing results; they do not replace investigator judgment, protocol discipline, or transparent reporting. :::
4. Engineer the panel around relevant expertise
Reputation is not a substitute for fit. Define the knowledge, experience, and perspectives required to answer the consensus question, then recruit against those criteria rather than against known positions on the topic.
An expertise-and-perspective matrix can make those choices reviewable. This example is illustrative rather than a platform default:
| Perspective | Why it is needed | Illustrative eligibility evidence | Illustrative target |
|---|---|---|---|
| Patient or caregiver | Identify outcomes that matter in daily recovery | Direct relevant lived experience | At least 20% of completing panel |
| Clinician | Judge clinical relevance and interpretability | At least three years of relevant practice | Balance across relevant disciplines |
| Outcomes researcher | Assess construct and measurement clarity | Relevant methods or core-outcome work | At least three completing members |
| Implementation leader | Judge operational feasibility | Deployment or quality-design responsibility | At least two completing members |
A narrowly technical question may appropriately require a more homogeneous panel. What matters is that the planned composition follows from the question and that omissions are acknowledged.
Track:
- how candidates were identified;
- eligibility and exclusion criteria;
- numbers screened, invited, enrolled, and completing each round;
- relevant disciplines, regions, roles, and experience;
- reasons for non-participation or withdrawal when known and appropriate to report; and
- any recruitment deviation or steering-group judgment.
Delphi Studio's Panel management tools record study participation, while the Expert directory and recruitment supports organization-scoped, consent-aware reuse of eligible experts.
:::info Response anonymity Panelists never see peer identities or peer-level responses. Researchers work with pseudonyms in routine analytic views. Authorized panel managers may reveal identity only through a deliberate, reasoned, audited workflow for operational needs. :::
5. Build judgeable items
Each item should express one proposition in language that the intended panel can interpret consistently.
| Weak item | Problem | Stronger alternative |
|---|---|---|
| “Remote monitoring is effective, acceptable, and should be routinely implemented.” | Three claims in one rating | “Patient acceptability should be included in the minimum outcome set.” |
Document the source of every initial item. Pilot the instructions, anchors, item wording, unable-to-rate option, and expected completion time with people similar to the intended panel. Piloting is a distinct methodological activity and should not be silently merged into the formal first round.
The item bank preserves source citations and version lineage. Substantive revisions create a new version linked to its parent so later reports can explain what changed, when, and why.
6. Run iterative rounds without steering the panel
The purpose of feedback is reflection, not forced convergence.
For Round 2 and later, a useful feedback packet may contain:
- the panelist's own previous rating;
- the group distribution, median, and IQR;
- zone percentages or another protocol-specified summary;
- approved, de-identified comments or balanced themes; and
- a clear indication when wording changed between rounds.
Delphi Studio locks the approved group-level feedback before the next round opens. Each panelist receives the same group evidence plus their own prior response. See Running rounds and Panelist experience.
Interpret more than the median
| Concept | Question answered | Typical evidence |
|---|---|---|
| Location | Where are ratings concentrated? | Median and zone percentages |
| Dispersion | How spread out are ratings? | IQR, MAD, histogram |
| Item-level consensus | Did the pre-specified classification rule pass? | Frozen rule predicates |
| Concordance | Do panelists rank the set of items similarly? | Kendall's W, when applicable |
| Stability | Are item summaries still changing between rounds? | Change in median and IQR |
| Polarization | Is a central summary hiding opposed groups? | Bimodality and subgroup checks |
Kendall's W is supplementary. Delphi Studio calculates it only when at least two panelists have complete ratings across at least two comparable items, with tie correction. It must not replace the item-level consensus definition.
Use AI as a review-gated assistant
AI-assisted workflows can propose themes, apply an approved codebook, and draft balanced feedback sections. Delphi Studio's feedback schema explicitly separates reasons supporting a statement, reasons against it, conditional positions, and open questions. No AI-generated content reaches panelists or final reporting without human review and approval.
AI can reduce synthesis burden, but it cannot prove that a summary is unbiased. Researchers remain responsible for checking coverage, correcting omissions, preserving meaningful dissent, and documenting their decisions. See Qualitative analysis.
7. Stop transparently and report the full process
Do not report only the final list. Readers need enough information to understand who contributed, how the items and rules evolved, what feedback was provided, and where uncertainty remained.
For biomedical consensus research, use ACCORD as the primary reporting checklist. ACCORD applies across biomedical consensus methods but is explicitly a reporting guideline, not a conduct standard. CREDES offers complementary Delphi-specific conduct and reporting guidance developed in palliative care. See the ACCORD and CREDES crosswalk.
At minimum, report:
- the objective, intended users, and reason for choosing Delphi;
- the exact modification and initial item sources;
- ethics review or determination;
- steering-group roles and conflicts of interest;
- panel eligibility, recruitment, composition, and attrition;
- piloting and survey materials;
- every round's dates, participation, feedback, and results;
- consensus, item-handling, stability, and stopping rules;
- item revisions, overrides, amendments, and deviations; and
- limitations, unresolved disagreement, funding, and applicability.
Delphi Studio's reports and exports are generated from the governed study records, rule snapshots, item lineage, and audit history. The platform currently provides an in-app report, DOCX report, CSV results, qualitative and trace bundles, a CREDES-aligned Markdown methods export, and browser Print/Save PDF. AI-assisted manuscript sections remain review-gated. See Reporting and exports.
Failure-mode diagnostic
| Failure | Why it threatens quality | Remedy |
|---|---|---|
| Post-hoc thresholds | Makes classification vulnerable to outcome-driven interpretation | Pre-specify rules and preserve per-round snapshots |
| Reporting only medians | Can hide dispersion and polarization | Report distributions, IQR, and subgroup or bimodality checks where planned |
| Forced convergence | Confuses fatigue or conformity with agreement | Use stopping rules and report stable disagreement |
| Overwriting revised items | Erases what panelists actually rated | Create versioned items with rationale and lineage |
| Selective feedback | Can steer later-round ratings | Use a declared feedback policy and preserve opposing or conditional positions |
| Undefined expertise | Prevents readers from judging panel relevance | Pre-specify eligibility and report composition |
| Silent attrition | Makes later-round results difficult to interpret | Report denominators and participation flow for every round |
| Unreported steering decisions | Conceals investigator influence | Record rationale, timing, and actor for overrides and revisions |
Pre-launch review
- The objective requires structured judgment rather than a representative population estimate.
- The exact Delphi modification and item sources are documented.
- Panel eligibility and perspective targets follow from the research question.
- Every scale has explicit anchors and an unable-to-rate policy.
- Consensus, polarization, item-handling, stability, and stopping rules are pre-specified.
- Piloting is distinguished from formal rounds.
- Feedback content and approval procedures are declared.
- Planned amendments, overrides, and deviations require rationale and audit evidence.
- ACCORD reporting fields and applicable CREDES elements have an evidence source.
- The research team has reviewed the final protocol snapshot before launch.
Further reading
- Nasa P, Jain R, Juneja D. Delphi methodology in healthcare research: How to decide its appropriateness. World Journal of Methodology. 2021;11(4):116–129.
- Gattrell WT, et al. ACCORD: A reporting guideline for consensus methods in biomedicine. PLOS Medicine. 2024;21(1):e1004326.
- Logullo P, et al. ACCORD explanation and elaboration. PLOS Medicine. 2024;21(5):e1004390.
- Jünger S, et al. Guidance on Conducting and REporting DElphi Studies (CREDES). Palliative Medicine. 2017;31(8):684–706.