Human-rater agreement evidence
Retained source:
docs/reference/ref-hule-research-selection-decision-noise-asymmetry.md
at revision c9d8b6ef60c19bbbdb17b9b46243ab4abfce35bb.
This public extract contains the aggregate reliability result used in the system overview. It excludes operational details that are irrelevant to the claim.
Test and human-reference roles
The ELLIPSE test partition is separate from cross-validation and guarded against text overlap. Test values are descriptive evidence.
The current pairwise raw-rater result is quadratic weighted kappa
0.5849381457369485 over 5,728 admitted Overall-score pairs from 5,740
prepared rows. Twelve rows were unresolved, giving coverage
0.9979094076655052.
This is contextual reliability evidence for that verified population and score pair. It is neither a model-performance target nor a universal ceiling.