TL;DR
SOOHAK presents a mathematician-curated benchmark evaluating LLMs on research-level mathematics through 340 Challenge and 99 Refusal items. Top models achieve only 30.39% accuracy, with open-weight systems at 13.87%, exposing substantial limitations in advanced mathematical reasoning.

论文原始摘要

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative. Yet research-level math benchmarks remain scarce because such problems are difficult to source (e.g., Riemann Bench and FrontierMath-Tier 4 contain 25 and 50 problems, respectively). To support reliable evaluation of next-generation frontier models, we introduce Soohak, a 439-problem benchmark newly authored from scratch by 64 mathematicians. Soohak comprises two subsets. On the Challenge subset, frontier models including Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach 30.4%, 26.4%, and 10.4% respectively, leaving substantial headroom, while leading open-weight models such as Qwen3-235B, GPT-OSS-120B, and Kimi-2.5 remain below 15%. Notably, beyond standard problem solving, Soohak introduces a refusal subset that probes a capability intrinsic to research mathematics: recognizing ill-posed problems and pausing rather than producing confident but unjustified answers. On this subset, no model exceeds 50%, identifying refusal as a new optimization target that current models do not directly address. To prevent contamination, the dataset will be publicly released in late 2026, with model evaluations available upon request in the interim.

Paper Collector 中文速览

数学家创建439题研究级数学基准,评估LLM推理与拒绝能力

方法概述

组织64位数学师从零创作439道研究级数学问题,分挑战集测试推理能力、拒绝集测试识别不当问题能力。通过Gemini-3-Pro、GPT-5等前沿模型和开源模型评测,验证基准难度与区分度。2026年底公开以防污染。

核心贡献

发布由64位数学家全新创作的439题研究级数学基准Soohak,含挑战集和拒绝集,填补研究级数学评测空白

原始来源与核验范围

核验仅覆盖论文正文中的主张、证据片段和参考文献关系。NGJOO 未独立运行作者代码、重做实验或验证真实部署效果;页面中的可复现性分数是文档完整度评估,不是复现实验结果。

32
已证实
2
证据不足
6
无法验证
N/A
可复现性
置信度
80%

核心问题

How do large language models perform on research-level and graduate-level mathematical problems, and can they appropriately identify and refuse ill-posed questions?

核心方法

The benchmark was constructed by 105 mathematicians and students who authored original problems across three splits: Challenge (340 research-level items), Refusal (99 ill-posed items), and SOOHAK-Mini (702 contest-level items). Problems passed through model-gated collection gates requiring failure of progressively larger baseline models. Eleven LLMs were evaluated with three independent responses per question, using GPT-5-Mini as an LLM judge for answer equivalence.

方法组件

论点验证

已证实 (95%) In this work, we introduce SOOHAK, consisting of a 340-item Challenge subset and a 99-item Refusal subset.
已证实 (95%) Additionally, to provide a subset for tracking smaller and open-weight systems, we release SOOHAK-Mini, a 702-question collection authored by a broader pool of 105 mathematicians and students.
The paper states SOOHAK-Mini is a 702-question collection by 105 contributors (p_4, p_10). The contributor breakdown is detailed in p_10 (86 primary-system + 19 ScienceBench = 105) and p_43, providing strong evidence for these specific numbers.
已证实 (95%) Finally, we conduct a human baseline with 25 participants across five teams.
The paper states 25 participants across five teams (p_5) and provides detailed team descriptions with varying expertise profiles (p_27, p_36). The human evaluation setup is thoroughly described in §6.
已证实 (85%) Across eleven closed and open-weight systems, Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach Avg@3 of 30.39%, 26.37%, and 10.39% on SOOHAK Challenge.
已证实 (85%) Kimi-2.5 is the best open-weight model at 13.87%.
Specific numerical result from Table 2. Kimi-2.5 at 13.87% Avg@3 on Challenge is reported as the best open-weight model. Self-reported result, confidence capped at 0.85.
已证实 (85%) The best closed model reaches only 43.10% Avg@3, while GLM-5 reaches the highest score at 49.49%, indicating that many systems continue attempting to solve prompts that are invalid as written.
Specific numerical results from Table 2 for the Refusal subset. The best closed model (Gemini-3-Flash at 43.10%) and GLM-5 (49.49%) are reported with exact numbers. Self-reported, confidence capped at 0.85.
已证实 (85%) GPT-5 reaches the strongest SOOHAK-Mini Avg@3 at 72.22%, while Kimi-K2.5 is the strongest open-weight model at 66.07%.
Specific numerical results from Table 2 for SOOHAK-Mini. GPT-5 at 72.22% and Kimi-2.5 at 66.07% are reported. Self-reported, confidence capped at 0.85.
已证实 (85%) On 79 prompts, aggregated teams cover 50.6% of the sample, confirming that the benchmark is challenging but tractable for strong human solvers.
The 50.6% coverage figure is reported for the aggregated human teams on 79 prompts (p_5). The evaluation setup is described in detail (p_28: 49 Calibration + 30 Challenge = 79 prompts). Self-reported human evaluation result, confidence capped at 0.85
已证实 (90%) Each submission through our primary system is first attempted by a panel of baseline LLMs and routed through three model-gated collection gates before final reporting.
已证实 (85%) To reduce direct and indirect leakage, only these two reviewers can access pre-opt-in submissions, and a withdrawn or declined submission is immediately deleted so that at most two individuals (often only one) ever viewed it.
已证实 (95%) Submissions had to be written in English or Korean, typeset in text-only LATEX (no diagrams or images), and accompanied by a complete solution and an explicit final-answer line.
Submission requirements are clearly specified in p_10: English or Korean, text-only LaTeX, complete solution, explicit final-answer line. These are concrete, verifiable format requirements.
已证实 (85%) Primary-system contributors were required to sign a submission agreement affirming that each problem had been originally authored without AI assistance.
已证实 (90%) We allocated a total compensation pool of USD 260,000 and paid on a per-accepted-question basis until the quota for each split was filled. Payments were split-dependent and ranged from USD 36 to USD 3,623 per question, with a cap of USD 20,000 per contributor.
Specific financial details are provided in p_10: USD 260,000 total pool, per-accepted-question payments, split-dependent payments ranging USD 36-3,623, cap of USD 20,000 per contributor. These are concrete, verifiable numbers.
已证实 (85%) We use GPT-5-Mini as an LLM judge that compares the parsed answer to the gold answer via mathematical equivalence.
已证实 (90%) For each model-question pair, we sample three independent responses and report avg@3 and pass@3.
The sampling methodology is clearly specified in p_22: three independent responses per model-question pair, reporting avg@3 and pass@3. The metric formulas are provided.
已证实 (90%) Human evaluations were conducted under a nominal 4.5 hour time budget. Participants were permitted to use any non-AI tools, including programming environments, computer algebra systems, and internet search for reference material.
Human evaluation conditions are clearly specified in p_29: nominal 4.5 hour time budget, permitted non-AI tools (programming environments, computer algebra systems, internet search), prohibited LLMs and AI-assisted search.
已证实 (90%) Scoring is purely outcome-based. A prompt is counted as correct only if the final answer is correct. We do not award partial credit for attempted-but-incorrect or partially correct solutions.
The scoring methodology is clearly specified in p_30: purely outcome-based, correct only if final answer is correct, no partial credit.
已证实 (85%) About 92% of the collected items were originally authored in English, and we translate every item into the other language using a machine-translation-plus-professional-post-editing workflow with LaTeX-preserving placeholders, glossary-normalized mathematical terminology, and an independent QA pass.
已证实 (85%) The full collection is temporarily embargoed, with public release planned in late 2026.
已证实 (85%) The benchmark family leaves 124 Challenge items unsolved by any evaluated model and 170 items unsolved or missed in total.
Specific numerical results reported in p_24: 124 Challenge items unsolved by any model, 170 total unsolved or missed. These are derived from the evaluation results in Table 2. Self-reported, confidence capped at 0.85.

... 共 40 个论点

可复现性评估

较低可复现性 (0%)

缺失的复现细节

局限与证据边界

本分析由 PDF 阅读助手 自动生成,仅供参考,不构成学术评审意见。验证结论和可复现性评估基于论文文本自动分析,可能存在偏差。原始论文请参阅 arXiv

分析时间:2026-05-19T07:11:22+00:00 · 数据来源:Paper Collector