Human-rater agreement evidence

Retained source: docs/reference/ref-hule-research-selection-decision-noise-asymmetry.md at revision c9d8b6ef60c19bbbdb17b9b46243ab4abfce35bb.

This public extract contains the aggregate reliability result used in the system overview. It excludes operational details that are irrelevant to the claim.

Test and human-reference roles

The ELLIPSE test partition is separate from cross-validation and guarded against text overlap. Test values are descriptive evidence.

The current pairwise raw-rater result is quadratic weighted kappa 0.5849381457369485 over 5,728 admitted Overall-score pairs from 5,740 prepared rows. Twelve rows were unresolved, giving coverage 0.9979094076655052.

This is contextual reliability evidence for that verified population and score pair. It is neither a model-performance target nor a universal ceiling.