REVIEW 4 cited by
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation. In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MMBench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks. Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions.
Forward citations
Cited by 4 Pith papers
-
ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.
-
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
State-of-the-art multimodal language models perform at or near random chance on MARBLE, a new hard benchmark for spatial reasoning and planning.
-
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
A new benchmark with edited-image question pairs shows that multimodal reasoning models lose accuracy when visual cues change, suggesting their reasoning is often not faithfully tied to the image.
-
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
A 57-task benchmark for multimodal image generation that uses automated visual parsers and an LLM judge to show GPT-4o-Native leads current models.
Discussion (0). Continue with ORCID to comment.