# Human-rater agreement evidence

Retained source:
`docs/reference/ref-hule-research-selection-decision-noise-asymmetry.md`
at revision `c9d8b6ef60c19bbbdb17b9b46243ab4abfce35bb`.

This public extract contains the aggregate reliability result used in the
system overview. It excludes operational details that are irrelevant to the
claim.

## Test and human-reference roles

The ELLIPSE test partition is separate from cross-validation and guarded
against text overlap. Test values are descriptive evidence.

The current pairwise raw-rater result is quadratic weighted kappa
`0.5849381457369485` over `5,728` admitted Overall-score pairs from `5,740`
prepared rows. Twelve rows were unresolved, giving coverage
`0.9979094076655052`.

This is contextual reliability evidence for that verified population and score
pair. It is neither a model-performance target nor a universal ceiling.
