Pith. sign in

REVIEW 4 cited by

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10563 v3 pith:P3DCINMW submitted 2024-10-14 cs.CV

classification cs.CV
keywords tasksevaluationmega-benchmultimodalacrosscapabilitiescoverdimensions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation. In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MMBench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks. Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.

  2. MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning

    cs.AI 2025-06 conditional novelty 6.0 of 10

    State-of-the-art multimodal language models perform at or near random chance on MARBLE, a new hard benchmark for spatial reasoning and planning.

  3. VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new benchmark with edited-image question pairs shows that multimodal reasoning models lose accuracy when visual cues change, suggesting their reasoning is often not faithfully tied to the image.

  4. OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 57-task benchmark for multimodal image generation that uses automated visual parsers and an LLM judge to show GPT-4o-Native leads current models.

Pith tools