Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

Jump to content

MMMU

From Wikipedia, the free encyclopedia

MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a benchmark for evaluating multimodal large language models on tasks that require college-level subject knowledge and deliberate reasoning.[1]

It was released in November 2023 by a team of 22 researchers led by Xiang Yue at Ohio State University and Wenhu Chen at the University of Waterloo, and was presented as an oral paper at CVPR 2024.[2]

The benchmark comprises 11.5 thousand multimodal questions manually collected from college exams, quizzes, and textbooks, spanning six disciplines — Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering — covering 30 subjects and 183 subfields. The questions incorporate highly heterogeneous image types, including charts, diagrams, maps, tables, music sheets, and chemical structures, and are designed to test expert-level perception, knowledge, and reasoning beyond commonsense visual understanding.At the time of release, the best-performing models achieved only around 56% accuracy, well below the human expert range of 76–89%. MMMU has since become a standard evaluation for frontier multimodal models, with scores routinely reported in the technical reports of systems such as GPT-4o, Gemini, and Qwen-VL.[3][4][5][6]

References

[edit]
  1. "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI". mmmu-benchmark.github.io. Retrieved 2026-08-30.
  2. Yue, Xiang; Ni, Yuansheng; Zheng, Tianyu; Zhang, Kai; Liu, Ruoqi; Zhang, Ge; Stevens, Samuel; Jiang, Dongfu; Ren, Weiming; Sun, Yuxuan; Wei, Cong; Yu, Botao; Yuan, Ruibin; Sun, Renliang; Yin, Ming (2024). "MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI". 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR): 9556–9567. Bibcode:2024cvpr.conf..913Y. doi:10.1109/CVPR52733.2024.00913. ISBN 979-8-3503-5300-6. ISSN 2575-7075.{{cite journal}}: CS1 maint: periodical has ISBN (link)
  3. "MMMU" (PDF).
  4. "dblp: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI". dblp.org. Retrieved 2026-08-30.
  5. Team, Gemini; Georgiev, Petko; Lei, Ving Ian; Burnell, Ryan; Bai, Libin; Gulati, Anmol; Tanzer, Garrett; Vincent, Damien; Pan, Zhufeng (2024-12-16). "Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context". arXiv:2403.05530 [cs.CL].
  6. Yue, Xiang; Zheng, Tianyu; Ni, Yuansheng; Wang, Yubo; Zhang, Kai; Tong, Shengbang; Sun, Yuxuan; Yu, Botao; Zhang, Ge (2025-05-22). "MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark". arXiv:2409.02813 [cs.CL].