← All guides

Evaluation & operations · Explore this field ↗ · Practice · 4 min read

Conversation quality review: use an anchored rubric, not a vibes score

A human review rubric can make task correctness, evidence, action safety, and repair quality discussable—and calibration can expose where the policy itself is vague.

A model to inspect

Calibration loop

  1. 01Sample cases
  2. 02Score independently
  3. 03Compare evidence
  4. 04Resolve policy gap
  5. 05Update rubric or test

Disagreement is evidence to investigate, not noise to erase.

Original conceptual diagram · not a live trace or measured result.
01

Score the service promise, not the prose polish

A fluent reply can be wrong, and a plain reply can safely complete a task. Start a rubric with dimensions that map to actual obligations: task outcome, evidence or source fit, authorization and action correctness, disclosure, repair after misunderstanding, and handoff quality. Give each dimension a small anchored scale—for example, 0 harmful or absent, 1 incomplete or unverified, 2 acceptable, 3 clear and verified—plus a not-applicable option. Avoid a single “helpfulness” score that lets style hide a severe action failure.

NIST presents risk management as spanning evaluation as well as design and use. The practical implication is that reviewers should capture the evidence behind a score. “2 because the answer cited the current owner-approved page and stated its exception” is useful for a decision; “pretty good” is not. The rubric should refer to the actual task contract, authorized source, and tool trace rather than asking reviewers to guess what the assistant was allowed to do.

02

Make anchors observable

Write an anchor from a turn and system event, not a personality trait. For action correctness, a 3 might mean the tool arguments match confirmed user intent, authorization is present, and the visible result agrees with the system of record. A 0 might mean the assistant claims a cancellation that no tool confirmed. For repair, a 3 might mean it acknowledges the correction, clears stale state, and continues; a 1 repeats the old assumption.

Keep the rubric short enough to use. If evidence review needs an hour for a routine sample, separate a lightweight service review from an investigation queue. Reserve detailed review for high-consequence tasks, disputed cases, and releases that change an action, source, or boundary. Include a “cannot assess” option when necessary evidence is unavailable. Forcing a score in that condition converts missing observability into an invented quality judgment.

03

Calibrate before trusting averages

Give two or more reviewers the same small, masked set of cases. Ask them to score independently, identify the evidence, then compare dimension by dimension. Do not force immediate consensus. A disagreement may reveal ambiguous policy, incomplete traces, different assumptions about the user goal, or a bad anchor. Record the chosen interpretation and assign its owner. Calibration has succeeded when reviewers learn where the contract is underspecified, not when everyone simply converges on the senior person’s opinion.

Worked example, hypothetical: two reviewers disagree on a refund answer. One awards evidence quality because the bot links an approved policy; the other scores it down because the policy has a stated exception for the customer’s product class. The resolution is not to average the scores. Add the exception to the expected evidence, repair retrieval or routing, and retain the case for future calibration. The case becomes a source and evaluation improvement, not a reviewer-performance dispute.

04

Turn findings into a review loop

Sample by task and consequence, not only by random traffic. Include clear requests, corrections, missing entitlements, tool failures, requests for a person, and abandoned turns. Keep a case card with scenario, permitted actions, expected evidence, reviewer scores, rationale, disagreement status, and follow-up owner. Link it to the relevant flow, source, tool trace, or policy revision. Minimize transcripts and restrict access; reviewers need enough context to assess the service, not an unrestricted customer history.

Review outcomes need a taxonomy: source defect, interface defect, routing defect, tool defect, policy ambiguity, reviewer instruction defect, or acceptable variation. A prompt change is only one possible remedy. This prevents a rubric from becoming a factory for cosmetic rewrites while the service remains operationally wrong. Confirm that an owner actually accepted the remediation and add the revised case to a regression or calibration set where it can prevent recurrence.

05

Success criteria and limitations

A useful rubric produces stable rationale for clear cases, makes uncertainty visible for hard cases, and lets a service owner see which dimension changed after a release. Success is not perfect reviewer agreement or a single composite number. Report score distributions, not-applicable rates, disagreement reasons, skipped evidence, and the small sample’s limitations. Review quality by slice, because a total can hide a poor result for one language, accessibility route, or high-consequence task.

Human review has its own bias and cost. Reviewers may infer intent the user never expressed, overvalue polished language, or see private context the assistant lacked. Use restricted access, rotating calibration, and an appeal route for agent or policy owners. This method complements deterministic action checks and regression fixtures; it does not replace them. A rubric can make a decision auditable, but it cannot prove the entire service is correct beyond the cases and evidence actually inspected.

Take it into the review

Anchored review card

Dimension0 anchor2 anchor3 anchorEvidence
Task outcomeWrong or harmful endingPermitted next stepVerified task resultTool status
EvidenceUnsupported claimRelevant approved sourceSource and exception fitCitation and version
RepairRepeats stale assumptionAsks useful correctionClears state and continuesTurn trace

A starting artifact to adapt to your service—not a ready-made policy, compliance certificate or test result.

Primary reading

Sources and limits

These links support the architecture, policy, or product behavior discussed above. Vendor documentation describes vendor features; it is not independent proof of performance. Current details should be rechecked before a production decision.

  1. NIST — AI Risk Management Framework
  2. Google Cloud — Dialogflow CX conversation history

What the sources establish

AI Risk Management Framework

NIST describes the AI RMF as supporting risk consideration in AI system evaluation alongside design, development, and use.

Limits: NIST does not supply the specific 0–3 anchors in this article.

Checked 2026-09-19 · NIST · source publication date not established.

Open original source ↗
Conversation history

Dialogflow CX documents turn-level conversation details, handoff filters, permissions, and configurable retention for its history tool.

Limits: Those observability features are product-specific and do not establish review quality.

Checked 2026-09-19 · Google Cloud · source publication date not established.

Open original source ↗

Procedures and worked examples are editorial synthesis. Preparation/review dates are not claimed historical publication dates.

Find your next good decision.

Start typing to explore the guides.

76 sourced guides · Escape to close