Tagged “metrics”
-
Recall@k and What It Hides
Recall, precision, MRR and nDCG measure different things badly. What each one is blind to, and which to report.
-
Why Offline Gains Vanish Online
Seven reasons an eval-set improvement doesn't reach users, and how to shrink the gap by evaluating through the production path.
-
Was That a Real Improvement or Noise?
Four sources of variance sit between your change and your eval score. Paired comparison and bootstrapping tell you which deltas to believe.
-
The Aggregate Score Hides the Bug
One number over a mixed eval set averages your best and worst behaviour into a figure that describes neither. Which slices to report, always.
-
Scoring Answers Against a Reference
Exact match and embedding similarity both fail on long-form answers. Key-point coverage is the reference metric that survives paraphrase.
-
Component or End-to-End: Which Evaluation Answers Which Question
End-to-end scores tell you whether users are served. Component scores tell you where to fix it. Neither substitutes for the other.