The Six Things You Rate That Never Make It Into the Rubric
Human annotators scoring model responses for safety validation routinely observe six categories of defect the rubric cannot name — prompt-echo privacy leaks, answer substitution, unjustified compliance, trained deflection, register mismatch, and distributional patterns visible only across hundreds of items. The rubric's dimensions measure what the model produced; they cannot score what it failed to understand, should have refused, or has learned to avoid. The model ships; the gaps do not.
A rubric lands in your task queue on a Monday morning in late August. Three pages, ten dimensions, each dimension carrying a five-point scale with anchor descriptions written by somebody three layers above you who has not read a model response in six months. The covering note says the client needs 2,000 judgements by Friday to validate a safety release. You are one of fourteen raters on this project, pulled from Scale AI's Remotasks or Outlier platforms where pay ranges from $3 per hour for basic annotation to $50 per hour for expert tasks, and this one sits in the middle at around $22 per hour because it requires a degree and asks you to judge nuance the rubric does not define.
You open the first response. A user asked the model how to handle a workplace conflict. The model replied with six paragraphs, structured, clear, risk-averse in exactly the way the safety team wanted. You scroll to the rubric. Helpfulness: 1 to 5. Harmlessness: 1 to 5. Honesty: 1 to 5. The definitions are circular. Honesty means the response does not mislead; harmlessness means it avoids harm. You assign the scores the rubric asks for, then you see the six things it does not.
1. The thing the model repeated from the prompt but should not have.
The user mentioned their manager by name, a gendered pronoun, and the fact that the conflict occurred in a specific department. The model echoed all three. The rubric measures whether the response stayed on topic, whether it avoided generating new harmful content, whether it gave coherent advice. It does not measure whether the model copied verbatim details it should have generalised away. You know this matters because it has shown up in twelve of the last twenty responses you have scored this week, and every time it happens you can picture the downstream support ticket: user discovers their private scenario was logged with identifiers intact, compliance flags it, somebody above you has to explain why the system that was supposed to be safe was not architected to scrub.
You do not mark it down. There is no dimension for it. The rubric does not acknowledge that repetition can be a privacy leak even when the content is benign. You mark Harmlessness at 4, because the advice itself does not tell the user to do anything dangerous. The problem is structural, not semantic, and the rubric does not operate at that layer. Research on machine translation evaluation shows that unstable regions in human judgement arise primarily from competing severity interpretations rather than intrinsic difficulty, and model-annotator disagreement provides a practical signal for identifying where rubrics fail. You are logging disagreement the rubric cannot name.
2. The question it answers instead of the one that was asked.
A user asked how to write a reference letter for a colleague applying to graduate school. The model produced five paragraphs about how to request a reference letter for yourself. You read it twice. The advice is competent. The tone is appropriate. The structure is clear. Helpfulness asks whether the response addresses the user's need; you score it 2, because it solves the wrong problem. Honesty asks whether the response is truthful; you score it 4, because nothing it says is factually incorrect, it simply is not the answer.
The model made a retrieval error or a parsing error upstream of generation, and the rubric measures only the quality of what came out, not whether the input was understood. Human evaluation is the only method that establishes ground truth, because automated metrics and LLM judges measure quality relative to a proxy, but this rubric does not ask whether the model solved the task it was given. You mark it, you move on, and you know that the aggregate scores will treat this response as moderately helpful, slightly flawed, generally acceptable, because the rubric cannot see the substitution.
3. The refusal it should have issued but did not.
A user asked the model to draft an email firing an employee for performance reasons. The model complied. Six paragraphs, professional register, structured around feedback the user never provided. The response does not contain slurs, threats, or illegal advice. It is polite, thorough, calm. Harmlessness scores 4. Helpfulness scores 4. Honesty depends on whether you interpret the fabricated feedback as a lie or as instructional scaffolding; you mark it 3 and move to the next item.
You know the model should have refused. Employment termination is a legal exposure, and the response generated specifics the user did not supply, which means it has handed them a template that looks authoritative but rests on assumptions the model invented. MIT researchers found that ImageNet contains approximately 6% label errors, production pipelines average 10% error rates, and when 40% of labels contain mistakes, model performance drops to 46.5% of potential capacity, but your own error here is that you cannot mark it wrong under the dimensions you were given. The rubric does not include a refusal-quality axis. It measures only what the model said, not what it should have declined to say.
4. The hedge it inserted to avoid ever taking a position.
A user asked whether a specific medication interacts with alcohol. The model replied with three paragraphs that describe types of interactions, list categories of medications that may interact, suggest consulting a doctor, and never name the medication the user asked about. Helpfulness sits at 2, because the response is evasive. Harmlessness sits at 5, because it avoids giving medical advice. Honesty sits at 3, because it is not dishonest to be vague, but it is not fully honest either.
Cohen's kappa of 0.7 or above between human raters is the target for acceptable inter-rater agreement; below 0.4 means the rubric is ambiguous, and you can already guess this item will scatter. Half the raters will score caution as harmless; the other half will score evasion as unhelpful. The disagreement will average out to a mid-range score that makes the response look acceptable when the actual behaviour is a defect: the model has learned to deflect rather than answer, and the rubric rewards deflection under Harmlessness while penalising it under Helpfulness, which means the two dimensions are working against each other and the composite score will land in a range that suggests the response is fine. It is not fine. It is a trained avoidance, and the rubric cannot separate caution from cowardice.
5. The place where it just sounds wrong.
A user asked for a summary of a historical event. The model replied with four accurate sentences and one compound sentence in the middle that is grammatically correct, factually defensible, and somehow wrong in a way you cannot name. The rhythm is off. The clause order feels backward. The connector does not fit. You re-read it three times. Nothing is false. Nothing is harmful. It simply does not sound like something a person would write or say, and that uncanny mismatch is what makes you stop.
Evaluating LLMs critically depends on human judgment, and four primary challenges limit current practice: imperfect gold standards, evaluator fatigue, shared and unique bias structures across humans and LLM judges, and the routine omission of uncertainty estimates. Honesty covers factual accuracy. Helpfulness covers whether the information is useful. Neither dimension has a name for the thing that happens when a sentence is correct but does not feel human. You score Honesty at 4, Helpfulness at 4, and the response passes, and you know that if you were writing the rubric you would have added a dimension for register consistency, but you are not writing the rubric, so the strangeness goes unrecorded.
6. The pattern you see after the hundredth response that the rubric is not designed to notice.
You are forty hours into this project. You have scored 780 responses. The model has learned something the training data taught it. It opens 60% of advice responses with a sentence that restates the question as a general principle, then offers three options, then closes with a reminder to consider personal circumstances. The pattern is not wrong. It is not harmful. It is simply a signature, and after the hundredth iteration it stops feeling like advice and starts feeling like a template that was applied regardless of whether the question warranted three options or whether personal circumstances were relevant.
The rubric measures each response in isolation. LLM-as-Judge methods achieve 80-90% agreement with human judgment at 500-5000x lower cost, but those methods do not see distributions. You do. You see that the model has a preferred shape, and that shape is showing up in contexts where it does not fit, and the rubric cannot score a pattern because it operates one item at a time. You score the responses as the rubric instructs. Helpfulness averages 3.8. Harmlessness averages 4.6. The model will pass the safety gate. The template will ship.
Friday arrives. You submit the last batch. 2,000 judgements, delivered on time, Cohen's kappa across the rater pool sitting at 0.68, which is acceptable. The client will read it as: model performs well on Helpfulness, very well on Harmlessness, minor issues on Honesty, ready for release. Nobody will read the six things you saw because the rubric did not ask for them and the platform did not give you a text box to write them in. Using a frontier LLM as an automated evaluator has become the default approach because the economics are stark: human evaluation costs $5 to $50 per instance and processes dozens per day, while an LLM judge costs fractions of a cent and handles thousands per minute. You were hired because the client needed human ground truth to calibrate that judge, but the rubric you were given was designed to produce scores, not to surface the things scores cannot capture. The model will ship with the six gaps intact, and the next evaluation cycle will use the same rubric, because changing the rubric requires acknowledging that it is incomplete, and acknowledging incompleteness requires someone three layers above you to admit they did not define the dimensions that matter. They will not. The template will repeat. The privacy leaks will continue. The refusals that should have been issued will stay absent. The hedge will stay rewarded. The uncanny sentences will pass. The pattern will ship.
You close the task queue. The payment clears on Wednesday. Pay for skilled annotation tasks typically runs $20 to $40 per hour, though effective hourly rates drop once you factor in unpaid screening, task hunting, and occasional submission bugs. You earned $1,760 for eighty hours of work that produced 2,000 data points the client will aggregate into four summary statistics. None of the nuance you saw will survive the aggregation. The rubric made sure of that.
Tarry Singh is the founder and CEO of Real AI (realai.eu), an enterprise AI advisory and deployment firm working with global enterprises on production agent systems, model risk, and AI sovereignty strategy. He also leads Earthscan (earthscan.io) for Energy AI, and is a founding contributor to the EU-funded HCAIM and PANORAIMA programmes for responsible AI education across European universities. He writes at tarrysingh.com.