Consensus analytics
Consensus analytics in Delphi Studio computes a configurable set of item-level and round-level statistics against the rules declared during study design. Calculations at round close improve consistency and reduce manual transcription error, but investigators remain responsible for selecting appropriate measures and interpreting them in context.
Analytics pipeline
When a round closes, Delphi Studio processes the data through the following sequence:
Ratings → Per-item statistics → Zone classification → Disposition assignment
Additional metrics are also generated, including:
- Kendall's W (panel-wide concordance)
- Stability measures (round-over-round change)
- Bimodality flags
- Subgroup breakdowns
- Stopping recommendations
All computations are performed against the rules frozen in the consensus rule snapshot for that round.
Per-item statistics
For each item and rating dimension, the system calculates:
| Statistic | Role in Delphi analysis | Notes |
|---|---|---|
| Median | Primary measure of central tendency | Most important for consensus decisions on ordinal scales |
| Mean | Supplementary measure | Reported but rarely used as the basis for consensus |
| IQR | Primary measure of dispersion | Core component of most consensus rules |
| Standard deviation | Parametric spread | Available but less commonly used in Delphi |
| MAD | Robust alternative to IQR | Useful when outliers are a concern |
| Histogram | Visual distribution of responses | Shown to panelists in controlled feedback |
| % at or above threshold | Percentage of ratings at or above a score K (pct_at_or_above) | One-sided statistic; directly referenced by consensus rules |
These statistics form the foundation for classifying items and generating feedback.
Consensus classification
Zones are bands of the rating scale (for example, scores 4–5 as the agree zone on a 5-point Likert scale); they feed the percentage statistics above. Classification then applies the study's consensus rules to the per-item statistics and assigns each item one of the following statuses:
| Status | Meaning |
|---|---|
consensus_in | Item meets the pre-specified criteria for inclusion |
consensus_out | Item meets the pre-specified criteria for exclusion |
no_consensus | Thresholds not fully met; item requires further rounds or researcher review |
consensus_with_subgroup_divergence | Panel-wide agreement with substantial opposition from a prespecified stakeholder group (see cross-stakeholder consensus); non-terminal |
Two guardrail outcomes protect against thin data: items with fewer than the minimum number of valid ratings (default 3) are marked insufficient_data rather than classified, and items failing a configured global floor are marked excluded_floor.
The new-study screen defaults to the Modified e-Delphi preset: a 5-point agreement scale with an 80% agreement threshold, median and IQR conditions. The separate RAND/UCLA preset uses a 9-point appropriateness scale, 75% zone thresholds, and a Disagreement Index condition. Both are configurable protocol starting points rather than universal rules. The table below shows the Modified e-Delphi rule set:
| Status | Example rule (5-point Likert) | Meaning |
|---|---|---|
consensus_in | Median ≥ 4, IQR ≤ 1, and ≥80% of ratings are 4 or 5 | Item meets consensus criteria for inclusion |
consensus_out | Median ≤ 2, IQR ≤ 1, and ≥80% of ratings are 1 or 2 | Item meets consensus criteria for exclusion |
no_consensus | Thresholds not fully met | Item requires further rounds or researcher review |
Rules are expressed as declarative predicates (e.g., median >= 4, iqr <= 1, pct_at_or_above 4 >= 0.8). The configuration format supports separate rule sets for different item classes (for example, stricter rules for clinical safety items), but studies currently run with a single core class end to end.
Because rules are frozen in a snapshot at the time each round closes, classification is fully reproducible.
RAND/UCLA appropriateness
When using a 9-point appropriateness scale (RAND/UCLA method), the engine additionally calculates the Disagreement Index (DI):
| DI value | Interpretation |
|---|---|
| ≤ 1.0 | Low polarization — panel is relatively united |
| > 1.0 | High polarization — panel is divided |
| null | Fewer than 3 valid ratings |
A high Disagreement Index can result in an item being classified as indeterminate even if the median falls in an acceptable range.
Kendall's W
Kendall's W (coefficient of concordance) is a supplementary measure of whether panelists rank a comparable set of items similarly within a round:
| W value | Interpretation |
|---|---|
| 0.0 | No concordance |
| 0.1 – 0.3 | Weak |
| 0.3 – 0.5 | Moderate |
| 0.5 – 0.7 | Substantial |
| 0.7 – 1.0 | Strong |
Interpretive labels for W vary across disciplines and should not be treated as universal consensus thresholds. Delphi Studio calculates W with tie correction using panelists who supplied complete ratings for all included items; it returns no value when fewer than two items or two complete panelists are available. Report the number of items and complete panelists with W, and do not use it in place of the pre-specified item-level consensus rule.
Stability tracking
Delphi Studio tracks how responses change between rounds for each item:
| Metric | What it measures | Used for |
|---|---|---|
| Δ Median | Change in median from previous round | Detect movement toward or away from consensus |
| Δ IQR | Change in dispersion | Assess whether opinions are becoming more or less polarized |
| Unstable % | Percentage of items whose median or IQR moved beyond the configured deltas (default Δ median ≤ 1, Δ IQR ≤ 1) | Compared against the next-round trigger (default 20%); the wizard's zone-switch setting maps to this trigger |
Stability is evaluated only for item-dimensions present in both rounds and included in the configured stability dimensions. High instability across many eligible rows can influence the stopping recommendation. Stable but contested items are reported separately from consensus: unchanged disagreement is a substantive finding, not convergence.
Bimodality detection
A median in the "Agree" zone can sometimes hide strong polarization (for example, when roughly half the panel strongly agrees and half strongly disagrees).
Delphi Studio flags an item as bimodal when the share of ratings in each of the agree and disagree zones reaches the study's configured bimodality threshold (set in the wizard, default 20%), at most 20% of ratings fall in the middle (uncertain) zone, and at least 6 ratings are present. Studies that disable bimodality detection get no flag; legacy studies created before the setting existed use a 30% per-tail floor.
The flag is advisory: it routes items for qualitative review but never changes the computed consensus status by default. Hand-written rule configurations may include a bimodal predicate to make it binding.
Subgroup analysis
Results can be broken down by demographic or expertise fields collected during panel enrollment (for example, specialty, region, or years of experience). This helps identify areas of agreement or disagreement across different groups of experts.
Subgroup breakdowns with fewer respondents than the minimum reporting size (default 5) are suppressed rather than reported, so unstable percentages never appear in results. Subgroup divergence is flagged for an item when the spread between the highest and lowest subgroup medians reaches the divergence threshold (default 2 scale points).
Primary stakeholder roles (when configured) are the preferred stratum for composition reporting and for the cross-stakeholder rule below.
Cross-stakeholder consensus
Optionally, consensus rules can include a cross-stakeholder condition: panel-wide agreement is not enough if a prespecified stakeholder group of sufficient size shows substantial opposition.
| Setting | Meaning |
|---|---|
| maxOppositionPct | Maximum fraction of a stakeholder group that may rate in the disagree zone |
| minGroupN | Groups smaller than this are skipped (unstable percentages are not treated as a veto) |
When the condition fires, the item is classified as agreement with stakeholder disagreement (consensus_with_subgroup_divergence) — a non-terminal status that typically rolls the item forward for re-rating and is reported distinctly from plain consensus. This prevents a numerically dominant group from producing apparent consensus over a smaller group's objection.
Data-quality guardrails
Several safeguards run automatically alongside classification:
| Guardrail | Behavior |
|---|---|
| Minimum responders | Items with fewer than 3 valid ratings (configurable) are marked insufficient_data rather than classified |
| Unable-to-rate accounting | Unable-to-rate responses are excluded from medians, IQRs, and zone percentages but tallied and reported per item as % unable to rate |
| Response-rate thresholds | A round response rate below 70% attaches a low-response caveat to the stopping recommendation; below 50%, round close is held until the researcher explicitly acknowledges the low response |
| Stopping recommendation | One of six values per round: consensus_reached, tentative_consensus_low_response, stable_contested, still_moving, max_rounds_reached, or insufficient_history (first round, no comparison possible) |
| Predicate trace | Every classification stores the evaluated predicates and their results, so you can see exactly which condition passed or failed for each item |
| Consensus round | Each item records the round in which it first reached consensus (consensus_reached_round) for reporting |
| Kendall's W preconditions | W is computed on complete blocks and reported as null when fewer than 2 raters or fewer than 2 commonly rated items are available |
Sensitivity analysis
You can test how sensitive the final dispositions are to your chosen thresholds. For example, you can compare results using an 80% agreement threshold versus a 75% threshold. Sensitivity runs are labeled post hoc so they are not confused with the pre-specified rules frozen at round close. This analysis is useful when writing robustness or limitations paragraphs in manuscripts.
Derived re-analysis
When a panelist withdraws after a round is closed, or a data correction changes who should count, Delphi Studio can recompute classifications for a closed/analyzed round with selected panelists excluded — without mutating the original frozen dispositions or the audit chain.
Each re-analysis stores:
- Reason (withdrawal or data correction)
- Excluded panelist set
- Per-item original vs new consensus status
- Count of items whose classification changed
You may designate a re-analysis as the authoritative record of reference for reporting while the original remains preserved. See Interpreting results.
Attrition caveats and rule snapshots
When response rates drop significantly between rounds, the system adds attrition caveats to stopping recommendations and reports.
At the close of every round, the exact consensus rules in effect are saved in a consensusRuleSnapshot. This ensures that even if rules are later adjusted, each round's classifications remain interpretable in their original context. Any mid-study changes to rules are also recorded in the audit log with before-and-after values.
Next steps
- Interpreting results — heatmaps, overrides, re-analysis, stopping
- Audit trail — reproducibility and study event log
- Reporting & exports — consensus tables