Skip to content
Brent Dunham
← Writing

Measuring before fixing: raising citation quality with evals

A low groundedness score turned out to be mostly the judge. How calibrating the instrument first, then adding a citation verifier, took claim coverage from 48% to 95%.


The research agent in Waypoint writes briefings with numbered citations. Its evaluation scores each briefing on several criteria, and one was far below the rest: grounded, whether every claim is supported by the source it cites, averaged 2.48 out of 4. Everything else scored above 3.3.

The obvious move was to change the agent. I measured the measurement first, and that turned out to be most of the answer.

Read the reasons, not the scores

The judge writes its reasoning before it scores. I read that reasoning for all 21 briefings rather than looking at the mean, and the deductions fell into five groups:

  • summary sentences with no citation;
  • a citation to the right finding but the wrong source within it;
  • claims the judge could not check, because it saw a source's title but not what the source said;
  • deductions for source quality, which the rubric does not ask about;
  • one claim marked down as "future-dated", which was simply a recent source.

Only the first two are faults in the briefing. The other three are faults in the instrument. Changing the agent at that point would have meant scoring the change with something that could not see it.

Fix the instrument

The search API returns, with each citation, the exact text the answer drew on. I kept up to three of those quotes per source and showed them to the judge next to each [n]. I rewrote the rubric to judge support and attribution only, and added a deterministic check, claims cited: at least 80% of a briefing's factual sentences carry a citation. Code cannot drift the way a judge can.

Then I rebuilt the human-labelled calibration set for this criterion, from real paragraphs in groups of three: the text as written, the same text with citations corrected, and a version with one planted error. Planted errors give examples with a known right answer.

Calibration then failed, which was the most useful result of the whole exercise:

JudgeAgreement with labels
Haiku 4.567%, then 75% on a rerun
Sonnet 4.575%
Opus 4.592%

The cheap judge was fine for four criteria but could not tell a citation to the wrong page from a right one, which is the one thing this criterion exists to catch. Grounded now routes to the stronger judge; the rest stay cheap. And because a judge should not grade its own model's work, I did not use the model that writes the briefings.

Re-scored by the calibrated judge, the unchanged briefings scored 3.67, not 2.48. Most of the problem had been the ruler.

Then fix the agent

What remained was real: half the briefings left more than a fifth of their factual sentences uncited, and some cited a neighbouring source. I added a verify step after the writer. A small model lists only the sentences that are not supported as cited. Wrong and missing citations are fixed in code; only sentences that nothing supports go back to the writer, one replacement each, spliced in by position so a rewrite cannot touch anything else.

It took four iterations. The first failed because the small model spent its whole token budget reasoning before it wrote any output. The second fixed that, but asking the writer for the whole revised briefing made it drop everything except the rewritten sentences, and coverage collapsed. Each failure became a design rule, and replaying every unchanged model call kept each iteration to cents.

BeforeAfter
Briefings with 80%+ of claims cited48%95%
Grounded (calibrated judge)3.673.86
Added cost per research runabout $0.02

Check it on new data

The iterations replayed the same research, so I ran the whole pipeline fresh. The gate flagged a drop, but the fresh searches were also weaker. Because the fresh run recorded every call, I replayed those exact briefings with the verifier switched off. The only difference between the two runs was the verifier, and it still helped: grounded 3.43 without it, 3.57 with it. One pattern did repeat: coverage was slightly lower with the verifier, because deleting an unsupported sentence sometimes deletes a point the reader wanted. That is the next thing to fix.

What I would tell a team

  1. When a score looks wrong, read the judge's reasons before touching the system.
  2. A judge is an instrument. Calibrate it against labels, per criterion, and expect a cheap one to fail on the criterion that matters most.
  3. Put everything code can decide in checks, not in the judge.
  4. Make evaluation cheap enough to run on every change. Replay is what made four iterations and a paired comparison cost a few dollars.

The full write-up, with every table and the scorecards behind them, is in the Waypoint repository.