Cognitive debriefing (CD) is foundational to patient-reported outcome (PRO) measure development and content validation as it provides direct evidence that patients understand instructions, items, and response options of the measure as intended. Regulators and methodological organizations clearly recognize the value of documenting respondents’ understanding of PRO measures, but they do not consistently specify clear quantitative thresholds for interpreting CD results. This can create ambiguity when comprehension issues are raised by a minority of participants and can make cross-study learning difficult.

At ISPOR 2026 in Philadelphia, Lumanity presented a targeted review examining how CD results are reported and whether studies use quantitative thresholds when interpreting participant understanding. The findings suggest limited transparency and inconsistent reporting practices, with thresholds (when used) varying widely and often lacking justification. In response to these findings, we proposed pragmatic, easy-to-apply interpretability thresholds (to be used alongside qualitative judgment) and a strong recommendation to improve transparency in CD reporting – at minimum, by clearly stating how many participants experienced issues for each item/attribute.

This thought leadership piece summarizes the rationale, key observations from our review, and a practical set of proposed thresholds intended to support more consistent interpretation and reporting, while explicitly recognizing that CD evidence remains inherently contextual and qualitative in nature.

CD sits at the intersection of qualitative insight and decision-making: teams must decide whether wording should change, whether a concept lands as intended, and whether “minor” issues are meaningful enough to revise.

Two pressures are increasing the need for clearer, more comparable CD reporting:

Contact us

For more information about how Lumanity can support you with your Cognitive Debriefing strategy, please contact us.

Regulatory expectations for evidence of understanding
Regulatory agencies emphasize documenting whether respondents understand PRO measure content as intended, but do not necessarily provide quantitative cutoffs that would standardize interpretation across programs
A growing need for cross-study comparability and shared learning
Without transparent reporting of CD issue rates (e.g. the number of participants with comprehension problems per item), it is difficult for the field to compare results, identify best practices, or understand what “good enough” looks like in real-world CD datasets

In accordance with the Food and Drug Administration (FDA) guidance1 emphasizing documentation of respondent understanding, Lumanity’s Patient-Centered Outcomes (PCO) team supports clients in designing and executing CD as part of:

  • De novo PRO development (e.g. refining instructions/items/response options through iterative rounds), and
  • Content validation of existing measures (e.g. evaluating whether an instrument remains fit-for-purpose in a new context of use, population, or setting)

This includes helping teams plan CD to produce evidence that is (a) credible for decision-making and (b) reportable with sufficient transparency for internal governance, publication expectations, and regulatory submission needs.

CD is widely valued, but interpretation cutoffs are not standardized

Our poster highlighted a practical gap: although CD is routinely used to assess whether patients interpret PRO measures as intended, there are no clear recommendations for interpretation cutoffs across key guidance, which creates ambiguity, particularly when only a minority of participants report issues.

We also noted why this is hard in practice: if comprehension challenges are reported by a small number of participants, the implications may be less straightforward and require nuanced judgment.

Our objective

We aimed to examine how CD results have been categorized in published research and in regulatory/organizational guidance, and whether standard quantitative thresholds are used.

High-level method

We conducted a targeted review of ISPOR and ISOQOL conference abstract databases using “cognitive debriefing” and “content validation,” then assessed whether abstracts reported CD results quantitatively and whether they specified predefined thresholds/cutoffs.

1) Quantitative reporting is uncommon; explicit thresholds are rare

We found that few abstracts reported CD results quantitatively (e.g. proportion interpreting items as intended), and among those that did, explicit thresholds were rare.

2) Where thresholds exist, they vary and are often not linked to clear decision rules

When thresholds were mentioned, they varied considerably and were often applied descriptively rather than as formal decision rules.

In our extracted sample, thresholds ranged from 60% to 92%, with mean 74% and median75%.

3) No broadly endorsed standard emerges from guidance sources

Our review of regulatory and organizational guidance did not identify published quantitative thresholds for interpreting CD results.

We summarized that FDA guidance emphasizes documentation of respondent understanding, while other organizations (e.g. ISPOR, ISOQOL, COSMIN) emphasize iteration, documentation, and context-driven interpretation, without specifying quantitative acceptability cutoffs.

4) The field has a transparency problem (not just a thresholds problem) Our conclusion was not that qualitative judgment should be replaced, but that limited quantitative reporting reduces transparency and comparability across studies.

We encourage the research community to, at a minimum, be transparent about CD issue frequency so results can be interpreted, compared, and learned from. If researchers are transparent about their data, we can more effectively evaluate the evidence and, where appropriate, reuse existing clinical outcome assessments instead of reinventing the wheel – saving time and cost and reducing unnecessary burden on patient communities. Concretely, that means for each item (and other assessed elements), the research community should report how many participants had an issue (and, where possible, what type of issue).

Recommended minimum reporting elements (practical and submission-ready)

For each assessed component (instructions, recall period, response scale, each item):

  1. N assessed (how many participants provided interpretable feedback for that element)
  2. Number of participants with comprehension issues
  3. Nature of issue (e.g. misunderstood concept, ambiguous term, misread recall period, response option mismatch)
  4. Severity/decision impact (minor edit vs. must-fix vs. monitor)
  5. Action taken (no change; wording tweak; revise concept; remove item; retest)
  6. Evidence of resolution if iterative rounds were conducted (what changed, what improved)

This level of granularity supports regulatory-grade traceability even if a journal publication ultimately condenses the detail.

Why propose thresholds at all?

Thresholds should be proposed because teams repeatedly face the same decision: Is the item “good enough,” or should we revise and retest? Without shared conventions, different programs may apply very different standards to similar results, making outcomes hard to compare and harder to defend.

As presented at ISPOR, we proposed the following thresholds for item comprehension, with explicit cautions about small sample sizes and the need for qualitative interpretation. These categories are intended for interpretability/comprehension; relevance may behave differently and can be more heterogeneous by nature.

Proposed thresholds for item comprehension (N ≥ 10)

CategoryThresholdSuggested interpretation/action
Acceptable≥ 80% interpret as intendedNo modification required
Marginal60–79% interpret as intendedReview qualitative feedback; minor wording adjustments may be warranted
Problematic< 60% interpret as intendedItem revision strongly recommended; re-testing in subsequent CD round
These were presented in Table 3 of the poster.

Small sample guidance (N < 10): use counts, not percentages

Percentages can be misleading in small samples. As presented, we recommend that two or more participants with comprehension issues should trigger review, regardless of percentage.

We also highlighted that even a single participant concern may warrant revision depending on the nature and severity of the issue.

Thresholds should be treated as signals, not automated decision rules. We recommend applying them alongside:

Issue type

  • Terminology confusion may be fixable with a small edit
  • Conceptual mismatch may require substantive revision or item removal
  • Recall period misunderstanding can invalidate the item even if only a few participants flag it

Population context

  • Heterogeneity in literacy, symptom experience, disease severity, or treatment journey can legitimately affect comprehension and relevance patterns

Item complexity

  • Multi-clause items, conditional phrasing, and abstract concepts may require higher scrutiny even when they “pass” a numerical cutoff

Intended context of use

  • The tolerance for ambiguity differs by context of use (e.g. endpoint in a pivotal trial vs. exploratory data collection)

To enable the field to learn faster and reduce preventable inconsistencies, we encourage the community to adopt the following near-term improvements:

  1. Publish counts of issues by item/element (not just narrative summaries)
    Even if space is limited, reporting “x/y participants misunderstood item 4” is far more comparable than “some participants found item 4 unclear”
  2. State whether thresholds were used, and how
    If a threshold guided decisions, name it and describe the decision rule (e.g. revise and retest; revise without retest; monitor only)
  3. Report what happened when items failed
    Our review noted that many studies did not specify what occurred if item interpretability did not meet the threshold; the field benefits when authors report the downstream action
  4. Evolve thresholds collaboratively as transparency improves
    The thresholds proposed here are intentionally pragmatic and preliminary. They should improve with broader reporting and shared evidence – especially across therapeutic areas, populations, and instrument types

For sponsors developing or adapting PRO measures, increased transparency and pragmatic thresholds can deliver:

  • More defensible instrument decisions (internally and externally)
  • Clearer rationale for iteration and retesting
  • Better alignment between publication narratives and regulatory-ready evidence packages
  • Improved cross-program learning, reducing repeated “reinvention” of CD interpretation approaches

The proposed thresholds warrant additional validation. Examining how the 80% and 60% cutoffs perform across a broader set of CD datasets, and whether they align with downstream psychometric outcomes, would strengthen their evidential basis and support adoption in practice.

Another area of research is to develop analogous guidance for item relevancy. Unlike comprehension, relevancy is expected to vary across participants as a function of disease severity and symptom heterogeneity, which complicates threshold-setting; nevertheless, the field would benefit from clearer frameworks for distinguishing acceptable variability from a signal that an item is poorly targeted to its intended population.

Future work might also examine whether thresholds should be differentiated by item type (e.g. instructions vs. response options vs. item stems) or by study context, such as early-phase concept elicitation vs. late-stage content validation of a near-final instrument.

1.         U.S. Department of Health and Human Services F and DA. Patient-Focused Drug Development: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments [Internet]. 2022. Available from: https://www.fda.gov/media/159500/download