Skip to main content

Design a defensible modified Delphi study

A defensible Delphi study is not simply a survey sent to people described as experts. It is a governed sequence of independent judgment, controlled feedback, re-rating, and transparent aggregation. The method is most useful when evidence is incomplete, inconsistent, or difficult to apply and a structured account of informed judgment will answer a meaningful research question.

This tutorial connects methodological decisions to the places where Delphi Studio records and enforces them. It is a design guide, not a substitute for a study-specific protocol, ethics review, statistical analysis plan, or applicable reporting guideline.

The four defining features

FeatureMethodological purposeDelphi Studio implementation
Independent respondingLimits status, dominance, and conformity pressurePanelists respond privately and never see peers' identities or individual responses
IterationGives panelists an opportunity to reconsider judgmentsVersioned, multi-round rating workflow
Controlled feedbackSupports reflection without unmoderated debateResearcher-reviewed group summaries and locked feedback snapshots
Structured aggregationMakes the group result explicit and reproduciblePre-specified item-level rules, descriptive statistics, and supplementary concordance measures

These features are consistent with the classic Delphi model of anonymous response, iteration with controlled feedback, and an aggregated group response described in the RAND literature.

:::note Terminology is not a protocol “Modified Delphi” is an umbrella label, not a complete method specification. Published implementations vary. State exactly what was modified, how initial items were generated, what feedback was provided, how consensus and stability were defined, and why those choices fit the research question. Nasa, Jain, and Juneja provide a useful healthcare-methodology overview. :::

1. Confirm methodological fit

Begin with the decision the study must support—not with a questionnaire.

Good reasons to use Delphi

  • Develop recommendations where evidence is limited, inconsistent, or difficult to operationalize.
  • Prioritize outcomes, competencies, indicators, research questions, or implementation strategies.
  • Refine definitions or classification criteria in an emerging field.
  • Identify both areas of agreement and persistent uncertainty.
  • Combine evidence with the judgment of people who hold relevant professional or lived expertise.

Reasons to choose another method

NeedBetter starting point
Estimate the prevalence of an opinion in a populationRepresentative survey or probability sample
Answer a question adequately resolved by empirical evidenceEvidence synthesis or primary empirical study
Negotiate interests or jointly solve a problem in real timeFacilitated workshop, nominal group, or deliberative process
Validate a predetermined conclusionReframe the project; consensus must not be used as ceremonial endorsement

Write down the intended output before creating items. Examples include “define a minimum outcome set,” “prioritize implementation barriers,” and “map unresolved disagreement.” Delphi Studio records the objective, intended audience, consensus question, and scope in the Study Design Wizard.

2. Define what “modified” means

Traditional and modified Delphi studies often share the same core safeguards but differ in how the first item set is produced.

Design featureTraditional DelphiModified Delphi
Starting pointOften begins with open idea generationOften begins with candidate items prepared before Round 1
Common item sourcesPanelist proposalsReviews, prior studies, interviews, frameworks, or steering-group work
Typical emphasisExploration and concept generationRefinement, prioritization, validation, or appropriateness
Reporting obligationExplain how open responses became rateable itemsExplain item sources, steering-group influence, and every other modification

Starting from an evidence-informed item bank can reduce burden, but efficiency alone is not a methodological justification. Report how the item set was assembled, piloted, edited, and checked for missing perspectives.

3. Engineer the protocol before recruitment

The protocol should make consequential decisions before anyone sees results.

Define the construct and scale

State what every rating means: importance, agreement, feasibility, appropriateness, inclusion, or another construct. Define and report the anchors for every scale point or category. A 1–9 importance scale and a 5-point agreement scale are different measurement designs, even if both ultimately produce an “include” decision.

Use the Rating and consensus step to configure the primary dimension, scale, anchors, zones, and classification predicates.

Pre-specify consensus and item handling

There is no universal Delphi consensus threshold. Select a rule that fits the construct, panel, and intended use; justify it; and state how missing or unable-to-rate responses enter the denominator.

For example, a protocol might define inclusion on a 9-point scale as:

  • at least 75% of valid ratings in the high zone (7–9); and
  • no more than 15% in the low zone (1–3).

:::caution Example, not Delphi Studio's default The 75/15 rule is illustrative. Delphi Studio's current Modified e-Delphi preset uses a 5-point agreement scale and requires at least 80% in the agree zone, median at least 4, and IQR at most 1. Presets are starting points, not universal standards. :::

Also specify what happens to each result class:

  • consensus to include or retain;
  • consensus to exclude or drop;
  • no consensus or indeterminate;
  • polarized or stakeholder-divergent;
  • insufficient data; and
  • revised wording requiring another rating round.

Pre-specify stopping and stability

Maximum rounds and consensus are not the same as stability. A panel may remain stably divided, which is a meaningful result rather than a failed study.

Define stability operationally. Delphi Studio currently compares item medians and IQRs between rounds, classifies each eligible item-dimension as stable or unstable, and calculates the proportion that remains unstable. The resulting stopping recommendation is advisory; the research team remains responsible for the final decision and its rationale. See Consensus analytics.

Govern changes and overrides

The launch snapshot freezes the declared methodology. Per-round rule snapshots preserve the exact rules used for classification. Where the protocol permits a later amendment or disposition override, Delphi Studio records a new version or audited rationale rather than rewriting historical results.

:::info Guarding against rule drift Snapshots and audit events make post-result changes visible and preserve the rules that were actually applied. They reduce the risk of moving thresholds after seeing results; they do not replace investigator judgment, protocol discipline, or transparent reporting. :::

4. Engineer the panel around relevant expertise

Reputation is not a substitute for fit. Define the knowledge, experience, and perspectives required to answer the consensus question, then recruit against those criteria rather than against known positions on the topic.

An expertise-and-perspective matrix can make those choices reviewable. This example is illustrative rather than a platform default:

PerspectiveWhy it is neededIllustrative eligibility evidenceIllustrative target
Patient or caregiverIdentify outcomes that matter in daily recoveryDirect relevant lived experienceAt least 20% of completing panel
ClinicianJudge clinical relevance and interpretabilityAt least three years of relevant practiceBalance across relevant disciplines
Outcomes researcherAssess construct and measurement clarityRelevant methods or core-outcome workAt least three completing members
Implementation leaderJudge operational feasibilityDeployment or quality-design responsibilityAt least two completing members

A narrowly technical question may appropriately require a more homogeneous panel. What matters is that the planned composition follows from the question and that omissions are acknowledged.

Track:

  • how candidates were identified;
  • eligibility and exclusion criteria;
  • numbers screened, invited, enrolled, and completing each round;
  • relevant disciplines, regions, roles, and experience;
  • reasons for non-participation or withdrawal when known and appropriate to report; and
  • any recruitment deviation or steering-group judgment.

Delphi Studio's Panel management tools record study participation, while the Expert directory and recruitment supports organization-scoped, consent-aware reuse of eligible experts.

:::info Response anonymity Panelists never see peer identities or peer-level responses. Researchers work with pseudonyms in routine analytic views. Authorized panel managers may reveal identity only through a deliberate, reasoned, audited workflow for operational needs. :::

5. Build judgeable items

Each item should express one proposition in language that the intended panel can interpret consistently.

Weak itemProblemStronger alternative
“Remote monitoring is effective, acceptable, and should be routinely implemented.”Three claims in one rating“Patient acceptability should be included in the minimum outcome set.”

Document the source of every initial item. Pilot the instructions, anchors, item wording, unable-to-rate option, and expected completion time with people similar to the intended panel. Piloting is a distinct methodological activity and should not be silently merged into the formal first round.

The item bank preserves source citations and version lineage. Substantive revisions create a new version linked to its parent so later reports can explain what changed, when, and why.

6. Run iterative rounds without steering the panel

The purpose of feedback is reflection, not forced convergence.

For Round 2 and later, a useful feedback packet may contain:

  • the panelist's own previous rating;
  • the group distribution, median, and IQR;
  • zone percentages or another protocol-specified summary;
  • approved, de-identified comments or balanced themes; and
  • a clear indication when wording changed between rounds.

Delphi Studio locks the approved group-level feedback before the next round opens. Each panelist receives the same group evidence plus their own prior response. See Running rounds and Panelist experience.

Interpret more than the median

ConceptQuestion answeredTypical evidence
LocationWhere are ratings concentrated?Median and zone percentages
DispersionHow spread out are ratings?IQR, MAD, histogram
Item-level consensusDid the pre-specified classification rule pass?Frozen rule predicates
ConcordanceDo panelists rank the set of items similarly?Kendall's W, when applicable
StabilityAre item summaries still changing between rounds?Change in median and IQR
PolarizationIs a central summary hiding opposed groups?Bimodality and subgroup checks

Kendall's W is supplementary. Delphi Studio calculates it only when at least two panelists have complete ratings across at least two comparable items, with tie correction. It must not replace the item-level consensus definition.

Use AI as a review-gated assistant

AI-assisted workflows can propose themes, apply an approved codebook, and draft balanced feedback sections. Delphi Studio's feedback schema explicitly separates reasons supporting a statement, reasons against it, conditional positions, and open questions. No AI-generated content reaches panelists or final reporting without human review and approval.

AI can reduce synthesis burden, but it cannot prove that a summary is unbiased. Researchers remain responsible for checking coverage, correcting omissions, preserving meaningful dissent, and documenting their decisions. See Qualitative analysis.

7. Stop transparently and report the full process

Do not report only the final list. Readers need enough information to understand who contributed, how the items and rules evolved, what feedback was provided, and where uncertainty remained.

For biomedical consensus research, use ACCORD as the primary reporting checklist. ACCORD applies across biomedical consensus methods but is explicitly a reporting guideline, not a conduct standard. CREDES offers complementary Delphi-specific conduct and reporting guidance developed in palliative care. See the ACCORD and CREDES crosswalk.

At minimum, report:

  • the objective, intended users, and reason for choosing Delphi;
  • the exact modification and initial item sources;
  • ethics review or determination;
  • steering-group roles and conflicts of interest;
  • panel eligibility, recruitment, composition, and attrition;
  • piloting and survey materials;
  • every round's dates, participation, feedback, and results;
  • consensus, item-handling, stability, and stopping rules;
  • item revisions, overrides, amendments, and deviations; and
  • limitations, unresolved disagreement, funding, and applicability.

Delphi Studio's reports and exports are generated from the governed study records, rule snapshots, item lineage, and audit history. The platform currently provides an in-app report, DOCX report, CSV results, qualitative and trace bundles, a CREDES-aligned Markdown methods export, and browser Print/Save PDF. AI-assisted manuscript sections remain review-gated. See Reporting and exports.

Failure-mode diagnostic

FailureWhy it threatens qualityRemedy
Post-hoc thresholdsMakes classification vulnerable to outcome-driven interpretationPre-specify rules and preserve per-round snapshots
Reporting only mediansCan hide dispersion and polarizationReport distributions, IQR, and subgroup or bimodality checks where planned
Forced convergenceConfuses fatigue or conformity with agreementUse stopping rules and report stable disagreement
Overwriting revised itemsErases what panelists actually ratedCreate versioned items with rationale and lineage
Selective feedbackCan steer later-round ratingsUse a declared feedback policy and preserve opposing or conditional positions
Undefined expertisePrevents readers from judging panel relevancePre-specify eligibility and report composition
Silent attritionMakes later-round results difficult to interpretReport denominators and participation flow for every round
Unreported steering decisionsConceals investigator influenceRecord rationale, timing, and actor for overrides and revisions

Pre-launch review

  • The objective requires structured judgment rather than a representative population estimate.
  • The exact Delphi modification and item sources are documented.
  • Panel eligibility and perspective targets follow from the research question.
  • Every scale has explicit anchors and an unable-to-rate policy.
  • Consensus, polarization, item-handling, stability, and stopping rules are pre-specified.
  • Piloting is distinguished from formal rounds.
  • Feedback content and approval procedures are declared.
  • Planned amendments, overrides, and deviations require rationale and audit evidence.
  • ACCORD reporting fields and applicable CREDES elements have an evidence source.
  • The research team has reviewed the final protocol snapshot before launch.

Further reading