// Read this first
What the scores mean.
Prediction can be right for the wrong reason. We test interventions and compare their estimated effects with confidence intervals.
Each score compares a reconstructed experiment with the original human study.
Rank order, effect direction, parameter distance, and uncertainty measure different parts of the match.
A benchmark is evidence, not a universal guarantee.