Skip to main content

Consensus analytics

Consensus analytics in Delphi Studio computes a configurable set of item-level and round-level statistics against the rules declared during study design. Calculations at round close improve consistency and reduce manual transcription error, but investigators remain responsible for selecting appropriate measures and interpreting them in context.

Analytics pipeline

When a round closes, Delphi Studio processes the data through the following sequence:

Ratings → Per-item statistics → Zone classification → Disposition assignment

Additional metrics are also generated, including:

  • Kendall's W (panel-wide concordance)
  • Stability measures (round-over-round change)
  • Bimodality flags
  • Subgroup breakdowns
  • Stopping recommendations

All computations are performed against the rules frozen in the consensus rule snapshot for that round.

Per-item statistics

For each item and rating dimension, the system calculates:

StatisticRole in Delphi analysisNotes
MedianPrimary measure of central tendencyMost important for consensus decisions on ordinal scales
MeanSupplementary measureReported but rarely used as the basis for consensus
IQRPrimary measure of dispersionCore component of most consensus rules
Standard deviationParametric spreadAvailable but less commonly used in Delphi
MADRobust alternative to IQRUseful when outliers are a concern
HistogramVisual distribution of responsesShown to panelists in controlled feedback
% at or above thresholdPercentage of ratings at or above a score K (pct_at_or_above)One-sided statistic; directly referenced by consensus rules

These statistics form the foundation for classifying items and generating feedback.

Consensus classification

Zones are bands of the rating scale (for example, scores 4–5 as the agree zone on a 5-point Likert scale); they feed the percentage statistics above. Classification then applies the study's consensus rules to the per-item statistics and assigns each item one of the following statuses:

StatusMeaning
consensus_inItem meets the pre-specified criteria for inclusion
consensus_outItem meets the pre-specified criteria for exclusion
no_consensusThresholds not fully met; item requires further rounds or researcher review
consensus_with_subgroup_divergencePanel-wide agreement with substantial opposition from a prespecified stakeholder group (see cross-stakeholder consensus); non-terminal

Two guardrail outcomes protect against thin data: items with fewer than the minimum number of valid ratings (default 3) are marked insufficient_data rather than classified, and items failing a configured global floor are marked excluded_floor.

The new-study screen defaults to the Modified e-Delphi preset: a 5-point agreement scale with an 80% agreement threshold, median and IQR conditions. The separate RAND/UCLA preset uses a 9-point appropriateness scale, 75% zone thresholds, and a Disagreement Index condition. Both are configurable protocol starting points rather than universal rules. The table below shows the Modified e-Delphi rule set:

StatusExample rule (5-point Likert)Meaning
consensus_inMedian ≥ 4, IQR ≤ 1, and ≥80% of ratings are 4 or 5Item meets consensus criteria for inclusion
consensus_outMedian ≤ 2, IQR ≤ 1, and ≥80% of ratings are 1 or 2Item meets consensus criteria for exclusion
no_consensusThresholds not fully metItem requires further rounds or researcher review

Rules are expressed as declarative predicates (e.g., median >= 4, iqr <= 1, pct_at_or_above 4 >= 0.8). The configuration format supports separate rule sets for different item classes (for example, stricter rules for clinical safety items), but studies currently run with a single core class end to end.

Because rules are frozen in a snapshot at the time each round closes, classification is fully reproducible.

RAND/UCLA appropriateness

When using a 9-point appropriateness scale (RAND/UCLA method), the engine additionally calculates the Disagreement Index (DI):

DI valueInterpretation
≤ 1.0Low polarization — panel is relatively united
> 1.0High polarization — panel is divided
nullFewer than 3 valid ratings

A high Disagreement Index can result in an item being classified as indeterminate even if the median falls in an acceptable range.

Kendall's W

Kendall's W (coefficient of concordance) is a supplementary measure of whether panelists rank a comparable set of items similarly within a round:

W valueInterpretation
0.0No concordance
0.1 – 0.3Weak
0.3 – 0.5Moderate
0.5 – 0.7Substantial
0.7 – 1.0Strong

Interpretive labels for W vary across disciplines and should not be treated as universal consensus thresholds. Delphi Studio calculates W with tie correction using panelists who supplied complete ratings for all included items; it returns no value when fewer than two items or two complete panelists are available. Report the number of items and complete panelists with W, and do not use it in place of the pre-specified item-level consensus rule.

Stability tracking

Delphi Studio tracks how responses change between rounds for each item:

MetricWhat it measuresUsed for
Δ MedianChange in median from previous roundDetect movement toward or away from consensus
Δ IQRChange in dispersionAssess whether opinions are becoming more or less polarized
Unstable %Percentage of items whose median or IQR moved beyond the configured deltas (default Δ median ≤ 1, Δ IQR ≤ 1)Compared against the next-round trigger (default 20%); the wizard's zone-switch setting maps to this trigger

Stability is evaluated only for item-dimensions present in both rounds and included in the configured stability dimensions. High instability across many eligible rows can influence the stopping recommendation. Stable but contested items are reported separately from consensus: unchanged disagreement is a substantive finding, not convergence.

Bimodality detection

A median in the "Agree" zone can sometimes hide strong polarization (for example, when roughly half the panel strongly agrees and half strongly disagrees).

Delphi Studio flags an item as bimodal when the share of ratings in each of the agree and disagree zones reaches the study's configured bimodality threshold (set in the wizard, default 20%), at most 20% of ratings fall in the middle (uncertain) zone, and at least 6 ratings are present. Studies that disable bimodality detection get no flag; legacy studies created before the setting existed use a 30% per-tail floor.

The flag is advisory: it routes items for qualitative review but never changes the computed consensus status by default. Hand-written rule configurations may include a bimodal predicate to make it binding.

Subgroup analysis

Results can be broken down by demographic or expertise fields collected during panel enrollment (for example, specialty, region, or years of experience). This helps identify areas of agreement or disagreement across different groups of experts.

Subgroup breakdowns with fewer respondents than the minimum reporting size (default 5) are suppressed rather than reported, so unstable percentages never appear in results. Subgroup divergence is flagged for an item when the spread between the highest and lowest subgroup medians reaches the divergence threshold (default 2 scale points).

Primary stakeholder roles (when configured) are the preferred stratum for composition reporting and for the cross-stakeholder rule below.

Cross-stakeholder consensus

Optionally, consensus rules can include a cross-stakeholder condition: panel-wide agreement is not enough if a prespecified stakeholder group of sufficient size shows substantial opposition.

SettingMeaning
maxOppositionPctMaximum fraction of a stakeholder group that may rate in the disagree zone
minGroupNGroups smaller than this are skipped (unstable percentages are not treated as a veto)

When the condition fires, the item is classified as agreement with stakeholder disagreement (consensus_with_subgroup_divergence) — a non-terminal status that typically rolls the item forward for re-rating and is reported distinctly from plain consensus. This prevents a numerically dominant group from producing apparent consensus over a smaller group's objection.

Data-quality guardrails

Several safeguards run automatically alongside classification:

GuardrailBehavior
Minimum respondersItems with fewer than 3 valid ratings (configurable) are marked insufficient_data rather than classified
Unable-to-rate accountingUnable-to-rate responses are excluded from medians, IQRs, and zone percentages but tallied and reported per item as % unable to rate
Response-rate thresholdsA round response rate below 70% attaches a low-response caveat to the stopping recommendation; below 50%, round close is held until the researcher explicitly acknowledges the low response
Stopping recommendationOne of six values per round: consensus_reached, tentative_consensus_low_response, stable_contested, still_moving, max_rounds_reached, or insufficient_history (first round, no comparison possible)
Predicate traceEvery classification stores the evaluated predicates and their results, so you can see exactly which condition passed or failed for each item
Consensus roundEach item records the round in which it first reached consensus (consensus_reached_round) for reporting
Kendall's W preconditionsW is computed on complete blocks and reported as null when fewer than 2 raters or fewer than 2 commonly rated items are available

Sensitivity analysis

You can test how sensitive the final dispositions are to your chosen thresholds. For example, you can compare results using an 80% agreement threshold versus a 75% threshold. Sensitivity runs are labeled post hoc so they are not confused with the pre-specified rules frozen at round close. This analysis is useful when writing robustness or limitations paragraphs in manuscripts.

Derived re-analysis

When a panelist withdraws after a round is closed, or a data correction changes who should count, Delphi Studio can recompute classifications for a closed/analyzed round with selected panelists excluded — without mutating the original frozen dispositions or the audit chain.

Each re-analysis stores:

  • Reason (withdrawal or data correction)
  • Excluded panelist set
  • Per-item original vs new consensus status
  • Count of items whose classification changed

You may designate a re-analysis as the authoritative record of reference for reporting while the original remains preserved. See Interpreting results.

Attrition caveats and rule snapshots

When response rates drop significantly between rounds, the system adds attrition caveats to stopping recommendations and reports.

At the close of every round, the exact consensus rules in effect are saved in a consensusRuleSnapshot. This ensures that even if rules are later adjusted, each round's classifications remain interpretable in their original context. Any mid-study changes to rules are also recorded in the audit log with before-and-after values.

Next steps