Ranking Methodology

← Back to Leaderboard

Aggregating across tasks: HELM-style win rate

The per-task metrics are on different scales and directions, so directly averaging them is not meaningful. Following HELM, each task is reduced to pairwise comparisons:

WRtask = (models beaten + 0.5 × models tied) / (N − 1)

where N is the number of models with data on that task. A model's mean win rate is the mean of its per-task win rates across the tasks in scope. The Individual and Distributional columns on the leaderboard are this mean computed over the corresponding subset of tasks; the overall column is their average.

ELO Rating (alternative ranking)

Each task generates pairwise matchups (ties counted as draws). Matchups are processed with K = 32 and initial rating 1500, shuffled 200 times with a fixed seed and averaged. Reported separately for Individual and Distributional.

Verbalized vs. simulated distributions

The Verbalized vs. Simulated Distributions page scores the distributional tasks a second way. In the simulated setting, the one the Distributional tab reports, the model answers once per individual and the pooled answers are compared with the human distribution. In the verbalized setting the model is asked once per condition to write out the population distribution itself, with no sampling; each question is asked several times and the distances are averaged, and an answer that cannot be read as a distribution is scored as the uniform distribution over the answer space. Both settings use the same Wasserstein-1 distance on the same scale.

Both mean win rates on that page use the formula above, computed over the models that have a run in both settings. That pool is smaller than the leaderboard's, so the simulated ranks there are not the Distributional ranks here, although the simulated cells are the same numbers.