Field Note
The Human in the Evaluation Loop
Why evaluation cannot collapse judgment into a single metric.
A metric is a decision in disguise
Evaluation is often presented as measurement after the real work is done. In practice, it shapes the work from the beginning. What we decide to score becomes what a system learns to notice.
Accuracy can be counted. Coherence can be approximated. But relevance, dignity, surprise, and taste live inside situations. They resist becoming a single number not because they are vague, but because they contain more than one legitimate value.
The human is not an oracle
Putting a person in the loop does not solve judgment. People disagree, tire, imitate one another, and inherit the assumptions of the interface in front of them. Human evaluation needs design: contrastive examples, visible uncertainty, space for reasons, and ways to preserve meaningful disagreement.
The goal is not to protect a mystical human answer. It is to build a process in which judgment remains inspectable and responsibility remains located.
The unresolved edge
Can an evaluation system become more rigorous without becoming less able to recognize what has never been valued before?