Skip to content

The causal AI
leaderboard.

417 replicated study runs across 5 domains. Mid-range scores and failures included.

Read the paper↗︎

// Read this first

What the scores mean.

Prediction can be right for the wrong reason. We test interventions and compare their estimated effects with confidence intervals.

Each score compares a reconstructed experiment with the original human study.

Rank order, effect direction, parameter distance, and uncertainty measure different parts of the match.

A benchmark is evidence, not a universal guarantee.

417Replicated study runs
5Research domains
OpenFailures included

// Protocol

The method.

Reconstruct the study. Match the respondent profile. Run the experiment. Estimate the original model.

Compare the result with the human data and publish the limitation.

Every score has a source and boundary.

Live rankings

ModelScoreCoverageMatchStudies
Claude Sonnet (Databricks)56.7840.0782.6319
GPT-4 (Azure OpenAI)55.6246.0278.2119
Claude Sonnet (Databricks)54.6345.7970.9523
Gemini Flash (Google)54.6323.8481.5919
GPT-4 (Azure OpenAI)53.8735.2773.222
Claude Haiku (Databricks)53.6239.9276.6019
GPT-3 Instruct (Azure OpenAI)52.7941.5075.9718
GPT-4 (Azure OpenAI)52.6734.9278.749
Gemini (Google)51.0933.4676.839
Claude Sonnet (Databricks)50.8228.5477.559
Similarity
Parameter proximity
Coverage
Confidence overlap
Match
Effect direction
Rank Corr.
Preference ordering
1 / 4

Domain breakdowns

Each table preserves the original domain-level ranking surface from the legacy leaderboard.

Public Health

8 models ranked

#public-health
ModelScoreCoverageMatchStudies
Claude Sonnet (Databricks)54.6345.7970.9523
Claude Haiku (Databricks)50.3044.9365.8323
GPT-4 (Azure OpenAI)50.2343.1267.5923
Gemini Flash (Google)49.4825.3768.9322
o1 (Azure OpenAI)48.4925.6768.7220
Gemini (Google)39.7140.4755.4223
GPT-3 Instruct (Azure OpenAI)38.9441.1252.2722
o3-mini (Azure OpenAI)37.5121.0555.2122
Similarity
Parameter proximity
Coverage
Confidence overlap

Consumer Research

8 models ranked

#consumer-research
ModelScoreCoverageMatchStudies
Claude Sonnet (Databricks)56.7840.0782.6319
GPT-4 (Azure OpenAI)55.6246.0278.2119
Gemini Flash (Google)54.6323.8481.5919
Claude Haiku (Databricks)53.6239.9276.6019
GPT-3 Instruct (Azure OpenAI)52.7941.5075.9718
o1 (Azure OpenAI)50.6525.3375.2019
Gemini (Google)50.0137.1274.8119
o3-mini (Azure OpenAI)33.1117.9156.5919
Similarity
Parameter proximity
Coverage
Confidence overlap

Economics

8 models ranked

#economics
ModelScoreCoverageMatchStudies
GPT-4 (Azure OpenAI)52.6734.9278.749
Gemini (Google)51.0933.4676.839
Claude Sonnet (Databricks)50.8228.5477.559
Gemini Flash (Google)47.4723.8177.879
GPT-3 Instruct (Azure OpenAI)45.0321.5875.799
Claude Haiku (Databricks)43.9231.1773.579
o1 (Azure OpenAI)38.4018.6765.199
o3-mini (Azure OpenAI)33.194.7763.359
Similarity
Parameter proximity
Coverage
Confidence overlap

Agricultural Sciences

4 models ranked

#agricultural-sciences
ModelScoreCoverageMatchStudies
GPT-4 (Azure OpenAI)43.1824.3053.482
Claude Sonnet (Databricks)38.5229.0847.402
Claude Haiku (Databricks)32.5628.3043.322
Gemini (Google)29.3228.9844.802
Similarity
Parameter proximity
Coverage
Confidence overlap

Business Administration

4 models ranked

#business-administration
ModelScoreCoverageMatchStudies
GPT-4 (Azure OpenAI)53.8735.2773.222
Claude Sonnet (Databricks)44.6415.6368.582
Claude Haiku (Databricks)44.0018.5162.302
Gemini (Google)43.1229.1660.102
Similarity
Parameter proximity
Coverage
Confidence overlap