How a form is checked
Before anyone scores against it, the form is checked for how it is built. A criterion's swing is its category's weight divided by the number of criteria in the category: the points one mark moves the score. What the form measures is shown as a share of its weight on customer outcome, compliance, process and behavior, as a fact with no threshold, because no source says what the right mix is. An auto-fail must name its reason: legal, regulatory, security, customer harm.
| Finding | Raised when | Severity |
|---|---|---|
| Weights do not total 100 | Category weights must total exactly 100. | Critical |
| Category with no criteria | A category carries weight but has no criteria to score. | Critical |
| Criterion with no definition | Every criterion needs a written definition of what earns a yes, so two evaluators mark it the same way. | High |
| Auto-fail with no stated reason | Every critical-fail criterion names its reason: legal, regulatory, security or customer harm. An auto-fail that cannot name one is hard to defend when it leads to discipline. | High |
| One judgment call swings the score | A non-critical criterion moves the score by more than 15 points (its category weight divided by the criteria in the category). Heuristic threshold. | Medium |
| Criterion with no focus tag | A criterion is not tagged as customer outcome, compliance or process, so the form's mix cannot be shown in full. | Note |
Blind calibration
Each evaluator scores the same calls alone and sees only their own score. They send the QA lead a submission code, which carries their initials, the call ID, their marks and a fingerprint of the form, so a code scored on a different form is set aside. The tool compares nothing, and shows no score, bias or agreement, until every evaluator has scored every call. Until then it shows only how many evaluators have scored each call. A session needs at least 2 evaluators besides any reference.
The Center of CX Calibration Method, version 1.0
Our own combination of published reliability statistics, chosen for how QA calibration actually runs: few calls, several evaluators, and rare critical fails. Krippendorff's alpha grades total scores and item marks. Gwet's AC1 grades agreement on critical fails, because when fails are rare, older statistics such as kappa report good agreement as poor. Percent agreement sits beside both, and a bootstrap interval on each shows how far a small session can be trusted. The method is our own combination; every statistic inside it is published and cited below.
| Measure | Statistic |
|---|---|
| Total score reliability | Krippendorff's alpha, interval metric, on each evaluator's weighted score before critical fails |
| Item reliability | Krippendorff's alpha, nominal metric, on every criterion mark on every call |
| Critical-fail agreement | Gwet's AC1 on every critical criterion mark on every call |
Every measure carries a 95% bootstrap interval: the calls are resampled 1,000 times with a fixed seed, so the same session always gives the same interval. A measure is graded by where its whole interval falls. An interval that crosses a band line is inconclusive, and a session of fewer than 3 calls is shown but not graded. When every evaluator gives every mark the same value, alpha has no value to compute, because there is no variation to measure; the session reads as unanimous and is never reported as a failure. Percent agreement sits beside each measure because it is the number a supervisor reads first.
Rules and bands
| Band | Cut point | Meaning |
|---|---|---|
| Reliable | 0.800 and above | Evaluators agree closely enough to rely on the scores. |
| Tentative | 0.667 to below 0.800 | Agreement supports tentative conclusions only. Reword the flagged criteria and calibrate again before scores feed coaching or pay. |
| Not reliable | Below 0.667 | Evaluators do not agree closely enough for the scores to mean the same thing whoever gives them. |
The cut points are Krippendorff's (2004). Gwet publishes no fixed cut points for AC1, so the method applies the same two to it.
| Finding | Raised when | Severity |
|---|---|---|
| Scores not reliable | A measure's interval lies below 0.667, where Krippendorff advises discarding the data. | High |
| Reliability inconclusive | A measure's interval crosses a band line, so the session is too small to say which band it is in. | Medium |
| Reliability tentative | A measure's interval lies between 0.667 and 0.800. | Medium |
| Evaluators read a criterion differently | Fewer than 80% of evaluator pairs mark a criterion the same way across the calibration calls. Heuristic threshold. | Medium |
| Evaluator scores harder or softer | An evaluator scores more than 5 points above or below the others on the same calls, on average. Heuristic threshold. | Medium |
| Wide spread on a call | An evaluator's score sits more than 5 points from the call's average. Heuristic threshold. | Medium |
| Evaluator differs from the reference | When a reference evaluator is named, an evaluator matches fewer than 80% of the reference's marks. Heuristic threshold. | Medium |
| Critical fails decide most scores | More than 50% of calibration evaluations end in a critical fail, so the weighted score rarely matters. Heuristic threshold. | Medium |
| Too few calls to grade | A session with fewer than 3 calls reports its measures but grades none of them. Heuristic threshold. | Note |
Why kappa is not used for critical fails
Two evaluators score 100 calls on one auto-fail criterion. Both pass 90; each fails 5 that the other passed; neither fails the same call. They agree on 90% of calls. Cohen's kappa for this table is -0.05, which reads as no agreement at all, because rare fails make chance agreement look almost certain. Gwet's AC1 for the same table is 0.89. Critical fails are rare by design, so the method grades them with AC1 (Feinstein and Cicchetti 1990; Gwet 2008).
Thresholds
| Threshold | Value | Basis |
|---|---|---|
| Points one non-critical criterion can move the score | 15 | Heuristic, no published source |
| Share of calibration evaluations decided by a critical fail | 50% | Heuristic, no published source |
| Points an evaluator scores above or below the others, on average | 5 | Heuristic, no published source |
| Points an evaluator's score sits from the call's average | 5 | Heuristic, no published source |
| Share of evaluator pairs that mark a criterion the same way | 80% | Heuristic, no published source |
| Calls a session needs before reliability is graded | 3 | Heuristic, no published source |
| Evaluators, besides any reference, a session needs | 2 | Rule of the method |
The next step
While the form has a critical or high finding, the next step is to fix the form. Then to calibrate; then to calibrate again while any measure is not reliable, inconclusive, tentative or ungraded. Once the form passes and evaluators agree, the next step is FCR Leakage, to test whether the scores track the repeat contacts they should prevent.
Checked against
- Krippendorff's alpha, the published reliability example (published value). Four observers, twelve units, values 1 to 5 with missing data, as printed in the source. Result: nominal alpha 0.743, interval alpha 0.849. Krippendorff, K. (2011). Computing Krippendorff's Alpha-Reliability. Annenberg School for Communication, University of Pennsylvania.
Sources
- Krippendorff, K. (2004). Content Analysis: An Introduction to Its Methodology, 2nd edition. Sage. Source of alpha's cut points, 0.800 and 0.667.
- Krippendorff, K. (2011). Computing Krippendorff's Alpha-Reliability. Annenberg School for Communication, University of Pennsylvania. The computation used here; its worked example is pinned in the test suite.
- Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29 to 48. Source of AC1.
- Feinstein, A. R. and Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543 to 549. Why kappa misreads rare critical fails.
- Efron, B. and Tibshirani, R. J. (1993). An Introduction to the Bootstrap. Chapman and Hall. The percentile interval used on every measure.
What this tool cannot tell you
- Form checks read how the form is built. Whether a criterion measures what matters to your customers is a question for outcome data, which is why the next step is FCR Leakage (first contact resolution: the share of issues solved on the first contact).
- Calibration measures agreement between the evaluators who took part, on the calls they scored. Evaluators can agree and all be wrong. Name a reference evaluator, someone whose marks you treat as correct, to measure accuracy as well.
- A small session gives wide intervals. The method shows that width openly and grades nothing below the minimum number of calls.
- Submission codes are not encrypted. Blind scoring holds because the tool compares nothing until every evaluator has scored every call, and because each evaluator sees only their own code.
- The thresholds marked heuristic have no published source. They are starting points. Set your own once your program has history.