Essays··9 min read
The Six Things You Rate That Never Make It Into the Rubric
Human annotators scoring model responses for safety validation routinely observe six categories of defect the rubric cannot name — prompt-echo privacy leaks, answer substitution, unjustified compliance, trained deflection, register mismatch, and distributional patterns visible only across hundreds of items. The rubric's dimensions measure what the model produced; they cannot score what it failed to understand, should have refused, or has learned to avoid. The model ships; the gaps do not.
ai-evaluationhuman-annotationrlhfmodel-safety
Read