The Center of CX
Subscribe
Open the tool
Published method

QA Scorecard Builder: how forms are checked and evaluators calibrated

Whether a QA (quality assurance) form is built to produce scores you can defend, and whether your evaluators, scoring the same calls without seeing each other's marks, agree closely enough that a score means the same thing whoever gives it.

Model version 1.0, published 2026-09-24. The Center of CX Calibration Method 1.0.

How a form is checked

Before anyone scores against it, the form is checked for how it is built. A criterion's swing is its category's weight divided by the number of criteria in the category: the points one mark moves the score. What the form measures is shown as a share of its weight on customer outcome, compliance, process and behavior, as a fact with no threshold, because no source says what the right mix is. An auto-fail must name its reason: legal, regulatory, security, customer harm.

FindingRaised whenSeverity
Weights do not total 100Category weights must total exactly 100.Critical
Category with no criteriaA category carries weight but has no criteria to score.Critical
Criterion with no definitionEvery criterion needs a written definition of what earns a yes, so two evaluators mark it the same way.High
Auto-fail with no stated reasonEvery critical-fail criterion names its reason: legal, regulatory, security or customer harm. An auto-fail that cannot name one is hard to defend when it leads to discipline.High
One judgment call swings the scoreA non-critical criterion moves the score by more than 15 points (its category weight divided by the criteria in the category). Heuristic threshold.Medium
Criterion with no focus tagA criterion is not tagged as customer outcome, compliance or process, so the form's mix cannot be shown in full.Note

Blind calibration

Each evaluator scores the same calls alone and sees only their own score. They send the QA lead a submission code, which carries their initials, the call ID, their marks and a fingerprint of the form, so a code scored on a different form is set aside. The tool compares nothing, and shows no score, bias or agreement, until every evaluator has scored every call. Until then it shows only how many evaluators have scored each call. A session needs at least 2 evaluators besides any reference.

The Center of CX Calibration Method, version 1.0

Our own combination of published reliability statistics, chosen for how QA calibration actually runs: few calls, several evaluators, and rare critical fails. Krippendorff's alpha grades total scores and item marks. Gwet's AC1 grades agreement on critical fails, because when fails are rare, older statistics such as kappa report good agreement as poor. Percent agreement sits beside both, and a bootstrap interval on each shows how far a small session can be trusted. The method is our own combination; every statistic inside it is published and cited below.

MeasureStatistic
Total score reliabilityKrippendorff's alpha, interval metric, on each evaluator's weighted score before critical fails
Item reliabilityKrippendorff's alpha, nominal metric, on every criterion mark on every call
Critical-fail agreementGwet's AC1 on every critical criterion mark on every call

Every measure carries a 95% bootstrap interval: the calls are resampled 1,000 times with a fixed seed, so the same session always gives the same interval. A measure is graded by where its whole interval falls. An interval that crosses a band line is inconclusive, and a session of fewer than 3 calls is shown but not graded. When every evaluator gives every mark the same value, alpha has no value to compute, because there is no variation to measure; the session reads as unanimous and is never reported as a failure. Percent agreement sits beside each measure because it is the number a supervisor reads first.

Rules and bands

BandCut pointMeaning
Reliable0.800 and aboveEvaluators agree closely enough to rely on the scores.
Tentative0.667 to below 0.800Agreement supports tentative conclusions only. Reword the flagged criteria and calibrate again before scores feed coaching or pay.
Not reliableBelow 0.667Evaluators do not agree closely enough for the scores to mean the same thing whoever gives them.

The cut points are Krippendorff's (2004). Gwet publishes no fixed cut points for AC1, so the method applies the same two to it.

FindingRaised whenSeverity
Scores not reliableA measure's interval lies below 0.667, where Krippendorff advises discarding the data.High
Reliability inconclusiveA measure's interval crosses a band line, so the session is too small to say which band it is in.Medium
Reliability tentativeA measure's interval lies between 0.667 and 0.800.Medium
Evaluators read a criterion differentlyFewer than 80% of evaluator pairs mark a criterion the same way across the calibration calls. Heuristic threshold.Medium
Evaluator scores harder or softerAn evaluator scores more than 5 points above or below the others on the same calls, on average. Heuristic threshold.Medium
Wide spread on a callAn evaluator's score sits more than 5 points from the call's average. Heuristic threshold.Medium
Evaluator differs from the referenceWhen a reference evaluator is named, an evaluator matches fewer than 80% of the reference's marks. Heuristic threshold.Medium
Critical fails decide most scoresMore than 50% of calibration evaluations end in a critical fail, so the weighted score rarely matters. Heuristic threshold.Medium
Too few calls to gradeA session with fewer than 3 calls reports its measures but grades none of them. Heuristic threshold.Note

Why kappa is not used for critical fails

Two evaluators score 100 calls on one auto-fail criterion. Both pass 90; each fails 5 that the other passed; neither fails the same call. They agree on 90% of calls. Cohen's kappa for this table is -0.05, which reads as no agreement at all, because rare fails make chance agreement look almost certain. Gwet's AC1 for the same table is 0.89. Critical fails are rare by design, so the method grades them with AC1 (Feinstein and Cicchetti 1990; Gwet 2008).

Thresholds

ThresholdValueBasis
Points one non-critical criterion can move the score15Heuristic, no published source
Share of calibration evaluations decided by a critical fail50%Heuristic, no published source
Points an evaluator scores above or below the others, on average5Heuristic, no published source
Points an evaluator's score sits from the call's average5Heuristic, no published source
Share of evaluator pairs that mark a criterion the same way80%Heuristic, no published source
Calls a session needs before reliability is graded3Heuristic, no published source
Evaluators, besides any reference, a session needs2Rule of the method

The next step

While the form has a critical or high finding, the next step is to fix the form. Then to calibrate; then to calibrate again while any measure is not reliable, inconclusive, tentative or ungraded. Once the form passes and evaluators agree, the next step is FCR Leakage, to test whether the scores track the repeat contacts they should prevent.

Checked against

  • Krippendorff's alpha, the published reliability example (published value). Four observers, twelve units, values 1 to 5 with missing data, as printed in the source. Result: nominal alpha 0.743, interval alpha 0.849. Krippendorff, K. (2011). Computing Krippendorff's Alpha-Reliability. Annenberg School for Communication, University of Pennsylvania.

Sources

  • Krippendorff, K. (2004). Content Analysis: An Introduction to Its Methodology, 2nd edition. Sage. Source of alpha's cut points, 0.800 and 0.667.
  • Krippendorff, K. (2011). Computing Krippendorff's Alpha-Reliability. Annenberg School for Communication, University of Pennsylvania. The computation used here; its worked example is pinned in the test suite.
  • Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29 to 48. Source of AC1.
  • Feinstein, A. R. and Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543 to 549. Why kappa misreads rare critical fails.
  • Efron, B. and Tibshirani, R. J. (1993). An Introduction to the Bootstrap. Chapman and Hall. The percentile interval used on every measure.

What this tool cannot tell you

  • Form checks read how the form is built. Whether a criterion measures what matters to your customers is a question for outcome data, which is why the next step is FCR Leakage (first contact resolution: the share of issues solved on the first contact).
  • Calibration measures agreement between the evaluators who took part, on the calls they scored. Evaluators can agree and all be wrong. Name a reference evaluator, someone whose marks you treat as correct, to measure accuracy as well.
  • A small session gives wide intervals. The method shows that width openly and grades nothing below the minimum number of calls.
  • Submission codes are not encrypted. Blind scoring holds because the tool compares nothing until every evaluator has scored every call, and because each evaluator sees only their own code.
  • The thresholds marked heuristic have no published source. They are starting points. Set your own once your program has history.
Build and calibrate a QA form
How to cite

The Center of CX, "QA Scorecard Builder method", version 1.0, 24 September 2026, https://www.contactcentercx.com/methodology/qa-scorecard