HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?

Fangchen Yu, Haiyuan Wan, Qianjia Cheng, Yuchen Zhang, Jiacheng Chen, Fujun Han, Yulun Wu, Junchi Yao, Ruilizhen Hu, Ning Ding, Yu Cheng, Tao Chen, Lei Bai, Dongzhan Zhou, Yun Luo, Ganqu Cui, Peng Ye

Introduction

“Physics is essentially an intuitive and concrete science.” — Albert Einstein

Large language models (LLMs) and Multimodal LLMs (MLLMs) have recently attracted increasing attention for their physical reasoning capabilities. However, existing datasets for physics remain limited in scope, lacking both systematic coverage of physics Olympiads and direct comparison with human contestants. By contrast, mathematical Olympiads have already received extensive attention—models such as GPT-5 have achieved over 90% accuracy on challenging benchmarks like AIME-2025 (Ye et al. 2025), and DeepMind’s Gemini (“Deep Think”) even reached gold-medal level at the 2025 International Mathematical Olympiad (IMO). As the crown jewel of high school physics, Olympiad problems offer a uniquely rigorous testbed that remains underexplored, underscoring the need for a new benchmark with systematic evaluation and human-aligned comparison.

Physics Olympiads, such as the International Physics Olympiad (IPhO), represent the pinnacle of high school physics competitions. Unlike mathematical Olympiads, they require a deep understanding of real-world physical principles, the derivation of abstract physics formulas, and the ability to reason with complex multimodal diagrams. These characteristics make physics Olympiads an ideal testbed for evaluating whether (M)LLMs can perform authentic visual and physical reasoning. Moreover, they offer the unique advantage of enabling direct comparison with human contestants.

However, current Olympiad-related physics datasets exhibit several critical limitations: (1) Outdated coverage: Datasets like OlympiadBench (He et al. 2024) includes IPhO problems only up to 2021 and Asian Physics Olympiad (APhO) problems up to 2015, omitting the most recent two years of exams. (2) Lack of multimodal content: Physics Olympiad problems often involve complex diagrams. However, datasets such as PHYBench (Qiu et al. 2025) and PHYSICS (Zheng et al. 2025) are entirely text-based. (3) Limited evaluation quality: Most datasets stop at coarse answer-level evaluation, without step-level grading aligned with official marking schemes. (4) No human-level comparison: Existing datasets typically report accuracy, without exam-score-based comparisons against human contestants in real-world physics competitions.

To address these limitations, we introduce HiPhO (High School Physics Olympiad), the first benchmark dedicated to recent physics Olympiads with human-aligned evaluation. As illustrated in Fig. 1, HiPhO incorporates four key improvements: (1) Up-to-date coverage: It compiles the latest 13 Olympiad exams from 2024–2025, including both international and regional contests. (2) Mixed-modal content: Problems are collected as complete exams, spanning text-only to diagram-based, with fine-grained categorization of diagram types (see Fig. 4). (3) Professional evaluation: We adopt official marking schemes to perform fine-grained grading at both the answer and step level, fully aligned with human examiners to ensure rigorous and domain-specific assessment. (4) Human-level comparison: We compute full exam scores for models and map them to gold, silver, and bronze medal thresholds, thereby enabling direct comparison with human contestants.

In our large-scale evaluation of 30 state-of-the-art (M)LLMs, we observe a clear performance hierarchy. Closed-source reasoning MLLMs dominate the medal table, winning 6–12 gold medals in 13 Olympiads. However, even the strongest models, such as Gemini-2.5-Pro and GPT-5, still fall short of the very best human contestants, especially in challenging exams like IPhO and EuPhO. In contrast, open-source chat MLLMs failed to secure any gold medals, with most scoring only at or below the bronze threshold. More encouragingly, several open-source reasoning (M)LLMs, including Intern-S1 and DeepSeek-R1, each achieved 4–8 gold medals, particularly in relatively easier exams such as F=MA. Taken together, these results reveal the strong but still non-parity performance of closed-source models, the evident limitations of open-source chat MLLMs, and the promising trajectory of open-source reasoning (M)LLMs in advancing physics problem solving.

In summary, our main contributions are as follows:

∙\bullet Olympiad-focused Benchmark. We present HiPhO, the first benchmark dedicated to high school physics Olympiads, comprising 13 Olympiad exam papers from 2024–2025. All problems are expert-curated, structurally extracted, and manually verified to ensure high quality and consistency.

∙\bullet Human-aligned Scoring. We evaluate performance using exam scores rather than the commonly used accuracy, applying both answer- and step-level grading based on official marking schemes. This enables direct, medal-standard comparisons with actual human contestants.

∙\bullet Human-level Comparison. Compared to human contestants, closed-source reasoning MLLMs reach gold in 6–12 exams, while open-source MLLMs stay mostly at bronze; some open-source LLMs show stronger reasoning with multiple golds, yet all remain far from the very top students.

Related Work

Among Olympiad-related physics datasets (Table 1), PHYBench (Qiu et al. 2025) and PHYSICS (Zheng et al. 2025) are text-only, while OlympiadBench (He et al. 2024) and OlympicArena (Huang et al. 2024) cover outdated exams. Multimodal datasets such as SeePhys (Xiang et al. 2025) and PhysReason (Zhang et al. 2025) remain restricted in scope (e.g., IPhO or CPhO). By contrast, HiPhO compiles 13 recent Olympiads (2024–2025), spanning problems with mixed modalities. More importantly, whereas existing datasets generally stop at coarse answer-level accuracy, HiPhO introduces fine-grained evaluation based on official marking schemes and adopts exam scores as the metric, enabling direct comparison between (M)LLMs and human contestants.

The HiPhO Benchmark for High School Physics Olympiad

We introduce HiPhO, the first high school physics Olympiad benchmark designed to compare the reasoning capabilities of (M)LLMs against human contestants. It contributes along three key dimensions: (1) a comprehensive and up-to-date dataset from real Olympiad problems, (2) professional evaluation using answer- and step-level grading aligned with official marking schemes, and (3) human-level comparison by mapping model scores to official medal thresholds.

Dataset Perspective. As shown in Table 2, HiPhO includes 360 problems and 519 subquestions from 13 Olympiad-level exam papers held in 2024–2025, making it the most up-to-date benchmark. All problems are carefully verified by expert annotators to ensure accuracy. Each problem is categorized along two axes: (1) physics taxonomy—covering five fields (Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics); and (2) modality type—spanning four formats (Text-Only, Text+Illustration Figure, Text+Variable Figure, and Text+Data Figure; see Fig. 4), enabling detailed analysis of both physical and visual reasoning.

Evaluation Perspective. HiPhO introduces a new evaluation method that combines answer-level correctness with step-level assessment based on official marking schemes—the first benchmark to do so. Beyond fine-grained scoring, we use exam score instead of accuracy to provide a more faithful measure of overall exam performance. This enables direct comparison with actual student scores and official medal thresholds, providing the first quantitative analysis of the performance gap between state-of-the-art (M)LLMs and top human contestants in real-world physics competitions.

2 Dataset Construction

Up-to-date Coverage. HiPhO comprises 13 exam papers from seven major Physics Olympiads (PhOs) and competitions in 2024–2025 (see Appendix A.1). These were selected based on their influence and the availability of human scores Notably, CPhO, USAPhO, and APhO-2024 were excluded due to the absence of human scores.. Compared with existing datasets, HiPhO offers superior timeliness and breadth: OlympiadBench includes older problems (e.g., APhO-2015) prone to contamination, while PhysReason covers only IPhOs. In contrast, HiPhO integrates the most recent 2024–2025 Olympiad exams. Supported by robust data processing tools, it can be efficiently maintained and expanded with new problems each year, ensuring a comprehensive and continuously updated benchmark. We further categorize the included exams into three difficulty levels:

Hard: IPhO (International PhO), APhO (Asian PhO), EuPhO (European PhO)

Medium: NBPhO (Nordic-Baltic PhO), PanPhO (Pan Pearl River Delta PhO)

Easy: PanMechanics (Pan Pearl River Delta Mechanics Test), F=MA

High-Quality Extraction Pipeline. We collect official PDF exam papers from Olympiad websites E.g., IPhO 2024–2025: https://ipho.olimpicos.net/; other websites listed in Appendix A.1. and process them through a structured pipeline to ensure high-quality extraction. (1) Data Extraction: PDFs are converted into markdown files using OCR tools, preserving LaTeX formatting for physical expressions. (2) QA Matching: Question indices are used to align each problem with its corresponding answer. (3) Human Verification: Unlike large-scale datasets that rely on automated or sampled checks, HiPhO manually verifies every QA pair. Experts fix OCR errors, correct mismatches, and validate answers to ensure consistency with the original exams. (4) Marking Scheme Structuring: For exams with official marking schemes, we extract step-level criteria and convert them into clear, model-readable rubrics for fine-grained evaluation. (5) Post-processing: Verified QAs are then refined in terms of context, question structure, and unit clarity, as detailed below.

Post-Processing and QA Refinement. To simulate real exam conditions and improve evaluation accuracy, we apply three key post-processing steps, rarely seen in existing datasets. (1) Context Completion: Many Olympiad problems include long stems followed by interdependent subquestions. To ensure contextual coherence, we merge relevant prior content into the context field. (2) Subquestion Structuring: Subquestions are often semantically linked and jointly graded in official marking schemes. We preserve their structure and explicitly label each part (e.g., “Find the speed and acceleration” →\rightarrow “(1) Find the speed. (2) Find the acceleration”) to guide ordered model responses and reduce omissions. (3) Unit Specification: We explicitly clarify required units within the question (e.g., “Find the speed in m/sm/s”) to avoid penalization from trivial unit mismatches, enhancing scoring fairness and accuracy without altering problem difficulty.

Comprehensive Data Annotation. To support fine-grained analysis, we annotate the dataset across three dimensions. (1) Physics Domain Categorization: Problems are labeled by five major fields—Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. (2) Visual Modality Classification: As shown in Fig. 4, we categorize problems into four types based on the role of figures, reflecting varying levels of visual reasoning:

Text-Only (TO): Problems are stated entirely in text with no figures.

Text+Illustration Figure (TI): Figures depict scenarios, while the text provides the description.

Text+Variable Figure (TV): Figures specify key variables or spatial configurations.

Text+Data Figure (TD): Figures present data, plots, or functions that are not given in the text.

Evaluation Method

Comprehensive Evaluation Framework. Our framework integrates both answer-level and step-level scoring, closely mirroring the grading approach of human examiners. The process is as follows: (1) Answer-level (coarse-grained) scoring: We first apply rule-based math-verifier (Kydlíček) to determine whether the extracted model answer matches the ground truth. If the check fails, possibly due to equivalent expressions, we invoke a strong model to assess correctness. (2) Step-level (fine-grained) scoring: We compare the model’s solution with the official marking scheme. A strong model awards partial credit for correctly completed steps. The final score for each problem is

following human grading conventions: a correct final answer receives full credit while an incorrect one is still graded fairly through intermediate steps. For problems with multiple official solutions (e.g., NBPhO), we score all marking schemes and take the highest result to ensure fairness across alternative solving paths. The exam score is then obtained by summing the scores of all problems.

Key Improvements in Evaluation. Our evaluation introduces three major enhancements. (1) Step-level scoring: Most datasets assess only the final answer (Table 1), awarding zero credit if it is incorrect. We incorporate step-level scoring to recognize partial correctness. (2) Marking-scheme alignment: Unlike prior step-level methods (e.g., OlympicArena generates steps with GPT-4), we are the first to extract steps strictly from the official marking scheme, awarding points based on explicit marking criteria for greater rigor and consistency. (3) Stronger grader: We employ Gemini-2.5-Flash as the grading model. Previous works (Xu et al. 2025; Feng et al. 2025) often rely on GPT-4o, whose competition scores are relatively low, making it less suitable for evaluating Olympiad-level problems. Compared to GPT-4o, Gemini-2.5-Flash provides step-level scores more closely aligned with human experts, with under 1-point differences in the examples (see Appendix B.4).

Experiments

To systematically evaluate the performance of state-of-the-art (M)LLMs against top-performing students in physics Olympiads, we selected 30 representative models. The experimental setup is detailed in Appendix C.1, and the medal scorelines are illustrated in Appendix C.2.

11 Closed-source MLLMs: GPT-5 (OpenAI b), o3 (OpenAI c), o4-mini (high) (OpenAI c), o4-mini (OpenAI c), GPT-4o (OpenAI a), Gemini-2.5 Series (Comanici et al. 2025), Grok-4 (xAI), Claude-4-Sonnet(-Thinking) (Anthropic), Mistral-Medium-3 (Mistral).

11 Open-source MLLMs: Intern-S1 (Bai et al. 2025a), InternVL3 Series (Zhu et al. 2025), Qwen2.5-VL Series (Bai et al. 2025b), GLM-4.5V (Team et al. 2025b), DeepSeek-VL2 (Wu et al. 2024), LLaMA4-Scout-17B (Meta), Phi-4-multimodal (Abouelenin et al. 2025).

8 Open-source LLMs: GPT-OSS (Agarwal et al. 2025), Kimi-K2-Instruct (Team et al. 2025a), DeepSeek-R1 (Guo et al. 2025), DeepSeek-V3 (Liu et al. 2024), Qwen3 Series (Yang et al. 2025).

What are the most powerful (M)LLMs in physics Olympiads? The medal table in Table 3 ranks models just like an Olympiad. The top five positions are occupied by Gemini-2.5-Pro, Gemini-2.5-Flash-Thinking, GPT-5, o3, and Grok-4. However, in the challenging IPhO-2025, only three of them reached the gold threshold, with the top two separated by merely 0.4 points. This narrow margin highlights the intense competition at the frontier of closed-source model performance.

How far are (M)LLMs from top-performing students? HiPhO enables human-level comparisons. (1) Closed-source reasoning MLLMs collected 6–12 gold medals across the 13 exams, yet still fell short of the very best students—for instance, in IPhO-2025 the top human scored 29.2/30, while the best model achieved only 22.7/29.4. (2) Open-source MLLMs mostly remained at or below the bronze threshold, with Intern-S1 the only exception to reach gold. (3) Open-source LLMs generally outperformed open-source MLLMs, earning gold in easier contests such as F=MA and even reaching the gold threshold in IPhO-2024, but still lagged well behind the very top students.

Can open-source models catch up with closed-source champions? While closed-source reasoning models still dominate, recent progress in the open-source community is noteworthy. Intern-S1 distinguished itself as the only open-source MLLM to win gold, obtaining four medals. Even more impressively, open-source LLMs including DeepSeek-R1 and Qwen3-235B-A22B both achieved eight gold medals. These advances reveal both the enduring gap with closed-source leaders and the emerging potential of open-source research to narrow it.

2 Can SOTA (M)LLMs Compete with Human Medalists in Physics Olympiads?

SOTA closed-source MLLMs remain unable to rival top-1 human contestants in most exams. As shown in Fig. 5, closed-source MLLMs frequently reached the gold threshold across Olympiads, yet consistently fell short of the very best human medalists in most exams. For example, Gemini-2.5-Pro achieved 22.7 points in IPhO-2025, the highest among all models, but 21 of the 37 human gold medalists still scored higher. In PanPhO-2025, its 60.3 points were far below the top human score of 81. These results show that while closed-source models can reliably “reach the podium,” they remain unable to rival the very top-performing students.

SOTA open-source MLLMs usually reach bronze to silver levels, but still fall short of human gold medalists in most Olympiads. As shown in Fig. 5, open-source models delivered respectable results in several contests, often scoring near the median of human medalists. For instance, in PanPhO-2025 Intern-S1 scored above 40 points, surpassing the median performance of human medalists, yet still below the gold threshold. By contrast, in more challenging Olympiads such as EuPhO-2025, which featured optics problems with variable figures, open-source MLLMs dropped back to the bronze level. These results illustrate both the steady progress of open-source models and the considerable gap that remains between them and the top-performing human contestants.

Mechanics-focused Olympiads demonstrate the potential of open-source (M)LLMs to rival human gold medalists. In PanMechanics-2025, which targets younger high-school students and places strong emphasis on classical mechanics, open-source (M)LLMs not only reached the gold threshold but in some cases even surpassed top human scores. For example, Kimi-K2-Instruct achieved 65.9 points compared to the top human score of 62.0. These results suggest that models are already approaching near-expert competence in mechanics, where reasoning is more deterministic and better aligned with training priors, while continuing to struggle in Olympiads that include optics problems, such as EuPhO-2025 and PanPhO-2024.

3 Do visual inputs pose a real challenge for MLLMs’ physical reasoning?

To investigate the impact of modality on performance, we categorize all problems into four types (Fig. 4): Text-Only (TO), Text+Illustration Figure (TI), Text+Variable Figure (TV), and Text+Data Figure (TD). For each type, we report the Mean Normalized score (MNS), defined as:

where M∈{TO,TI,TV,TD}M\in\{\text{TO},\text{TI},\text{TV},\text{TD}\}, NMN_{M} is the number of questions in MM, and QQ denotes a single question.

As shown in Fig. 6, diagram-based problems are consistently more difficult than text-only ones, leading to a sharp decline in mean normalized scores. For instance, the leading closed-source model, Gemini-2.5-Pro, the score reaches 86% on TO but decreases progressively as visual complexity increases: 81% on TI, 75% on TV, and 67% on TD. Grok-4 drops to 51% on TD, highlighting persistent challenges in extracting numerical values from plots and interpreting functional plots such as peaks. Comparatively, open-source models fall further behind: the gap is visible on TO questions and widens significantly with visual input. On TV problems, GPT-5 achieves 75%, while Qwen2.5-VL-72B-Instruct reaches only 29%, showing its limitation in interpreting figures with complex variables.

Overall, these results highlight three major areas where the visual reasoning of MLLMs in physics problems can be further advanced: (1) interpreting illustration diagrams, (2) reasoning over variable-based graphs, and (3) accurately extracting quantitative information from data figures.

4 Which Physics Fields Challenge MLLMs the Most?

Physics Olympiad problems span five fields: Mechanics (Mech.), Electromagnetism (Elec.), Thermodynamics (Ther.), Optics (Opt.), and Modern Physics (Mode.). We report mean normalized scores (Eq. 1) for each field in Table 4, leading to three key observations. (1) Optics is the most challenging field, with all models scoring below 55%. Its difficulty arises from two aspects: geometrical optics requires diagram interpretation, while wave optics demands precise symbolic derivations—both remain weak points for current MLLMs. (2) Modern Physics tends to yield higher scores than Mechanics, likely because the problems are less visually intensive and the concepts tested do not extend to university-level depth. (3) Mechanics and Electromagnetism show a clear source gap. Closed-source models average above 70%. Most open-source MLLMs remain below 45%, with one exception: Intern-S1 narrows the gap, reaching 63–67% in Mech./Elec. Overall, optics is the most challenging, while closed-source models show steady gains.

5 How Can (M)LLMs Achieve True Human-Level Physics Reasoning?

Beyond theoretical performance, (M)LLMs face three structural limitations compared to human contestants. (1) Multimodal ability: LLMs lack the capacity to interpret diagrams, while MLLMs, though capable of reading figures, still struggle with variable-based and data-intensive figures. (2) Generative ability: physics Olympiads require not only interpreting diagrams but also sketching functional plots. Most (M)LLMs cannot generate such diagrams directly, and evaluating the quality of generated plots remains non-trivial. (3) Embodied ability: IPhO, APhO, and EuPhO include separate experimental exams, and NBPhO integrates theoretical and experimental problems within a single paper. Since most models lack embodiment and generative capabilities, all experimental and diagram-generation problems are excluded to ensure fair evaluation. As a result, the Full Mark (Model) in Table 3 may be lower than the Full Mark (Human). A detailed overview of exam formats and scoring components is provided in Table 6 of Appendix A.1.

Reaching true human-level physics reasoning will thus require advances in three dimensions: (1) multimodality, for robust integration of textual and visual inputs; (2) generation, to produce diagrams and functional plots; and (3) embodiment, to support experimental reasoning and physical interaction. Without progress in these areas, even the most capable (M)LLMs will remain constrained relative to the holistic problem-solving abilities of human contestants.

Conclusion and Future Work

In this work, we presented HiPhO, the first benchmark designed to systematically evaluate (M)LLMs against top-performing students in high school physics Olympiads. Spanning 13 Olympiad exams from 2024–2025, HiPhO covers both international and regional contests. The problems include mixed modalities, spanning text-only to diagram-based, with fine-grained categorization. HiPhO is the first to combine answer-level and step-level grading based on official marking schemes, enabling rigorous, human-aligned evaluation. We compare model exam scores against human contestants and observe significant performance gaps. SOTA closed-source reasoning MLLMs achieve gold-level performance on most exams but still fall short of the very best human contestants. Most open-source MLLMs remain at the bronze level or below, while open-source LLMs obtain multiple golds. These findings highlight both the progress of (M)LLMs in physical reasoning and the substantial gap that remains before achieving human-level mastery.

Looking forward, HiPhO can be extended in two key directions. First, expanding coverage to include additional Olympiads such as the Chinese Physics Olympiad (CPhO) and the USA Physics Olympiad (USAPhO) would provide a broader and more representative benchmark for evaluating physics reasoning. Second, introducing a dynamic update mechanism that incorporates newly released Olympiad problems and human scores would ensure data freshness and benchmark integrity, while reducing risks of training contamination. By both broadening its scope and keeping the dataset continuously refreshed, future iterations of HiPhO can offer an even more comprehensive and reliable standard, further driving progress in multimodal physical reasoning.

This work was supported by a locally commissioned task from the Shanghai Municipal Government. We also gratefully acknowledge the support of the APhO Committee.

References

Appendix A Details of HiPhO Benchmark

HiPhO covers seven types of high school physics Olympiads, from international to regional exams.

IPhO (International Physics Olympiad): The most prestigious global physics Olympiad for high school students since 1967, featuring challenging theoretical and experimental exams.

APhO (Asian Physics Olympiad): A regional contest launched in 2000 for Asian and Oceanian students, structured similarly to IPhO with both theory and laboratory components.

EuPhO (European Physics Olympiad): Established in 2017, a multi-day competition for European students that emphasizes creative problem solving in both theory and experiment.

NBPhO (Nordic-Baltic Physics Olympiad): A regional contest among Nordic and Baltic countries, primarily theoretical but also including experimental problems.

PanPhO (Pan Pearl River Delta Physics Olympiad): An invitational competition for top schools from the Pearl River Delta and neighboring regions in China.

PanMechanics (Pan Pearl River Delta Mechanics Test): A specialized subset of PanPhO focusing solely on mechanics, typically offered as a shorter, single-field exam.

F=MA: A U.S. mechanics contest organized by the American Association of Physics Teachers, serving as the entry test for the U.S. Physics Olympiad (USAPhO).

Physics Olympiads adopt diverse exam formats. As shown in Table 6, IPhO, APhO, and EuPhO include a separate 20-point experimental exam, while NBPhO combines theoretical and experimental components into a single paper (72-point total). PanPhO, PanMechanics, and F=MA include only theoretical problems. Since most (M)LLMs lack embodied experimental and diagram-drawing capabilities, we exclude all experimental and diagram-generation questions for fair and consistent evaluation. The Full Mark (Model) in Table 3 therefore refers only to the theoretical score of the theoretical exam.

Importantly, we introduce for the first time a step-level score based on official marking schemes. Among the 13 exams, 7 provide official marking schemes, and 4 support multiple solutions. This alignment improves fairness and ensures consistency with human grading standards.

A.2 Data Processing Pipeline

We implemented a structured, multi-stage data processing pipeline, as illustrated in Fig. 7 (left).

Data Extraction. Official Olympiad exam papers in PDF format were processed with OCR-based tools. Text blocks were reflowed into markdown for easier parsing and editing.

QA Matching. Questions were aligned with their official answers by matching indices and numbering patterns, ensuring that every problem statement had a corresponding answer entry.

Human Verification. Human experts carefully reviewed all extracted content. This included correcting OCR misrecognitions and ensuring textual consistency with the original exam.

Marking Scheme Structuring. For exams with marking schemes, step-level scoring criteria were systematically extracted, and then reformatted into standardized, model-readable rules.

Post-Processing. Verified QAs underwent additional refinement:

Context Completion: Long problem stems and necessary background information were consolidated to provide self-contained contexts.

Subquestion Structuring: Interdependent subparts were explicitly separated and re-labeled to reflect the intended order of solution.

Unit Specification: Required physical units were clarified in the problem statement to minimize ambiguity and avoid unfair penalization.

As shown in Fig. 7 (right), the finalized dataset is organized into a unified json format:

Information: Physics constants table on the exam cover page (if available), e.g., g=9.8 m/s2g=9.8~m/s^{2}.

Context: Problem stem content, including introductory descriptions or preceding subparts.

Question: Specific question text, with multiple subquestions explicitly split into (1), (2), etc.

Diagram: Path to required figures or diagrams associated with the question.

Marking: Step-level scoring points in the unified form “Award xx pt if the answer …”.

Answer: Ground-truth solution including value, units, answer type, and allocated points.

Physics Field: The physics domain of the problem (Mechanics, Electromagnetism, Thermodynamics, Optics, Modern Physics).

Modality Type: The modality type of the problem (Text-Only, Text+Illustration Figure, Text+Variable Figure, Text+Data Figure).

Source: The corresponding Olympiad exam, e.g., IPhO 2025.

Appendix B Details of Evaluation Framework

To ensure consistent and reproducible evaluation across models, we adopt a compact instruction template that: (i) enforces LaTeX-formatted math, (ii) separates full reasoning from the final answer using and tags, and (iii) standardizes multi-part and multiple-choice outputs with boxed answers.

To align with the original language of each Olympiad exam, we use English prompts for English-language exams and Chinese prompts for Chinese-language exams. Specifically, the language settings are as follows:

English Exams: IPhO, APhO, EuPhO, NBPhO, PanPhO, F=MA

For completeness, we provide both the English and Chinese prompt templates below.

B.2 Answer-Level Coarse-Grained Score

The answer-level coarse-grained evaluation pipeline extends the PHYSICS framework (Zheng et al. 2025) and follows a structured sequence to assess whether the model’s final answers are correct. The process comprises the following key steps:

Answer Extraction. Final answers are extracted from the model’s solution using automatic parsing of boxed expressions (e.g., \boxed{...}). This ensures that evaluation targets the model’s intended outputs, while ignoring intermediate reasoning steps.

Rule-Based Matching. A rule-based math-verifier (Kydlíček) compares the extracted answers with the ground-truth answers. It performs correct matching on either numeric or symbolic expressions, while also accounting for units and answer types.

Model-Based Verification. If no exact match is found, a second-pass verification is performed using a powerful judge model (Gemini-2.5-Flash). This model compares the model’s answer and the ground-truth answer to determine physical or mathematical equivalence, thereby reducing false negatives caused by alternative expressions, equivalent transformations, or formatting inconsistencies.

Multi-Part Matching Logic. For problems consisting of multiple sub-questions (e.g., labeled as (1), (2), etc.), the intended answer order is explicitly defined by the problem statement. To reflect this, we extract boxed answers from the model’s response in reverse order and compare them sequentially with the ground-truth, following the original question structure.

Unlike the original PHYSICS’s evaluation framework, which relied on a fine-tuned lightweight 8B verifier, our pipeline leverages Gemini-2.5-Flash for model-based verification. This substitution substantially improves the reliability of equivalence judgments. The corresponding evaluation prompt is adapted from the PHYSICS design (Zheng et al. 2025), as shown below.

B.3 Step-Level Fine-Grained Score

The step-level fine-grained evaluation measures the model’s reasoning quality by comparing its solution steps against detailed criteria from official marking schemes, with each criterion representing a specific conceptual, physical, or mathematical step. The evaluation follows the steps below:

Parse Marking Criteria. Each problem is accompanied by a list of step-level marking points, such as “Applies energy conservation correctly” or “Derives the correct force expression.” These marking points are parsed from the dataset and include descriptions and assigned partial scores.

Model-Based Step Scoring. For each marking point, the model’s solution is independently assessed using Gemini-2.5-Flash, a strong judge model. A dedicated prompt is constructed to instruct the model to evaluate whether the student’s solution satisfies the given criterion and to return a numerical score (e.g., 1.0, 0.5, or 0.0) accordingly. This approach enables semantic and physical equivalence checking beyond surface-level matching.

Score Aggregation. The scores assigned for each marking point are aggregated to compute the final fine-grained score for the question. These scores are stored alongside their corresponding criteria for detailed feedback and analysis. For EuPhO and NBPhO problems with multiple official marking schemes, we take the highest score across schemes, reflecting the fact that solutions may follow different valid approaches.

Compared to commonly used answer-level evaluation, this fully model-based step-level evaluation provides a more accurate and context-aware assessment of the model’s reasoning. The fine-grained scoring closely aligns with how human graders assign partial credit in real Olympiad exams, making it one of the core innovations of our benchmark and a more faithful reflection of true exam performance. The prompt used for step-level model-based evaluation is shown below.

B.4 Answer-Level vs. Step-Level vs. Human Expert Scores

As shown in Table 7, we compared grading outcomes across different graders. Step-level scores are typically higher than answer-level scores in the final exam results, reflecting the partial credit awarded for intermediate steps. Moreover, when used as a grader, the more powerful model Gemini-2.5-Flash produces results much closer to those of human experts, demonstrating its greater accuracy in evaluating exam performance, particularly at the step level.

Appendix C Details of Experiments and Results

Hyperparameter Settings. We adopt VLMEvalKit (Duan et al. 2024) as the evaluation framework for benchmarking MLLMs. To reduce randomness and improve the reliability of evaluation, each problem is tested using eight independent inference runs at a temperature of 0.6; its score is averaged across runs, and the exam score is the sum of these averages. This setup helps reduce score fluctuations caused by randomness in the model’s responses. In addition, the maximum token limit is set in reference to the largest value permitted by each model, helping prevent output truncation and ensuring that responses are complete and valid.

Evaluated Models. To assess the physics reasoning capabilities of state-of-the-art (M)LLMs relative to top-performing human contestants, we evaluate 30 representative models, including 11 closed-source MLLMs, 11 open-source MLLMs, and 8 open-source LLMs, based on the following criteria: (1) Recency: Most models were released after April 2025, with the newest launched in August 2025 (e.g., GPT-5). (2) Diversity: The selection includes both closed- and open-source models, covering reasoning-specialized and general-purpose architectures from a wide range of developers. (3) Model Scale: A range of model sizes—small, medium, and large—is included to enable performance comparisons by scale. The full list of evaluated models is provided in Table 8.

C.2 Illustration of Medal Scorelines

IPhO, APhO, EuPhO: These include separate theoretical and experimental exams, with medals awarded on total scores. As (M)LLMs cannot do experiments, we use the lowest theoretical exam score of gold medalists as the gold cutoff, and analogously for silver and bronze.

NBPhO: Theory and experiment appear in the same paper, with medals based on total scores. We set the gold line as the lowest total score of gold medalists. Since the paper contains experimental and plotting items, the model’s full mark is lower than the human’s; such items are counted as zero to reflect current limitations of (M)LLMs.

PanPhO, PanMechanics: These are theory-only exams, but the official website reports scores only for Hong Kong SAR contestants. Thus, thresholds are based on the lowest scores of Hong Kong SAR medalists. If scores of all medalists were available, the top-1 human score would likely be higher, while the medal thresholds could be lower.

F=MA: With over 5,000 participants, only score histograms are published and no official medals are awarded. As a USAPhO qualifier, the cutoff score for advancement is treated as the gold line. Silver and bronze thresholds are inferred from Physics Bowl conventions (top 20% and 35%), estimated from the histogram, while the gold line also aligns with the Physics Bowl top-10% rule.