SOOHAK ยท Math Benchmark

SOOHAK Leaderboard

Avg@3 (%) of language models on SOOHAK, a mathematician-curated benchmark of research-level math โ€” over two subsets: Challenge (340 hard items) and Refusal (99 unanswerable items).

Tracking 23 models on 340 Challenge + 99 Refusal items.
46
Challenge problems remain unsolved by any modelof 340
News
Frontier tracker ยท Challenge problems solved by โ‰คN models over time
no model โ‰ค1 model โ‰ค2 models
Cumulative over every model released to date (pass@3, 340 Challenge items). As the field advanced, problems solved by no model fell 289 โ†’ 46 and solved by โ‰ค2 models fell 340 โ†’ 86. Hover a point for the model(s) released that day.
Uniquely solved ยท Challenge
Challenge problems (of 340) that exactly one model solves โ€” cracks no other model can reproduce, so they show where each model uniquely extends the frontier. Across all models, gpt-5.5 alone gets 15 that every other model misses.

How to cite

@article{son2026soohak,
  title={Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs},
  author={Son, Guijin and Kim, Seungone and Arnett, Catherine and Ko, Hyunwoo and Lee, Hyein and Kang, Hyeonah and Longxi, Jiang and Yun, Jin and Lee, JungYup and Lee, Kyungmin and others},
  journal={arXiv preprint arXiv:2605.09063},
  year={2026}
}