Skip to main content

Set Your Consensus Threshold Before You See the Data

· 4 min read
Building infrastructure for defensible consensus research

A threshold chosen after Round 1 is not a threshold. It is a result.

A fixed threshold gate sorting rating responses into symmetric channels

Systematic reviews of Delphi methodology repeatedly find inconsistent definitions and reporting of consensus. In a review of 100 studies, only 42 of the 86 studies that concluded consensus had been achieved had specified an a priori threshold. The usual problem is not fraud but drift: an item lands just below the planned cut point, the investigators believe it belongs, and the rule moves with a plausible-sounding justification. Each decision can sound reasonable while the overall process becomes unfalsifiable. See Diamond and colleagues' systematic review of consensus definitions.

Declare the scale, statistic, and cut point together

A complete threshold specification has at least three connected parts:

  1. The scale. Define its points, anchors, and any agreement zones in the panelist instructions. The RAND/UCLA Appropriateness Method, for example, uses a 9-point scale with 1–3, 4–6, and 7–9 representing inappropriate, uncertain, and appropriate regions; its final classification also considers the panel median and disagreement, not merely the percentage in one region. See the RAND/UCLA user's manual.
  2. The statistic. A percentage in the top region is not equivalent to a median in that region with an interquartile-range ceiling. Both can be legitimate, but they will differ on borderline or polarized items.
  3. The cut point. Justify the value against methodological reasoning and relevant precedent rather than convenience. There is no universal percentage: published definitions vary widely.

Also specify the denominator, the handling of missing and unable-to-rate responses, and any minimum valid-rating requirement. Otherwise, identical observed ratings can receive different classifications under unstated denominator choices.

Account for both inclusion and exclusion

Many studies define retention—such as at least 70% rating an item essential—but leave exclusion undefined. Items that fail retention then return in later rounds without a principled disposition.

For designs intended to classify both positive and negative consensus, a symmetric structure closes that gap:

  • Retain if at least 70% rate the item essential.
  • Exclude if at least 70% rate the item non-essential.
  • Re-rate or review if neither condition is met.

Symmetry is not mandatory for every Delphi question, and some studies appropriately use more stringent exclusion rules. The requirement is that every possible result has a pre-specified interpretation and next action. An item should not linger simply because the protocol never said what failure to retain meant.

Specify the stopping rule too

Consensus thresholds tell you what happens to an item. A stopping rule tells you what happens to the study. Options include a fixed maximum number of rounds, a numerical stability criterion, a target proportion of classified items, or a defined rule for unresolved items.

Fixed-round designs are straightforward to specify and report. Stability-based designs can be informative, but “stable” must be defined numerically in advance—for example, through permitted changes in median, IQR, or classification between adjacent rounds. Consensus and stability are different: an item can remain stably disputed.

Guard the floor

Attrition can make a percentage look more precise than the underlying count warrants. With four valid raters, one response moves a percentage by 25 points; with ten raters, it moves it by 10. Set a minimum valid-rating count or response-rate floor below which an item or round is flagged rather than classified without qualification. State that rule alongside the threshold in the analysis plan.

What the platform preserves

In Delphi Studio, scales, zones, consensus rules, and stopping settings are configured during study setup and captured in the launch snapshot. Each analyzed round also stores the exact rule snapshot used for its classifications. Later permitted changes are audited with before-and-after values, while historical results continue to read their frozen per-round rules. See Consensus analytics for the classification model and Interpreting results for non-destructive re-analysis.

Next post: panel composition, strata, and why the number that matters is not only the total.