REVIEW 12 cited by
GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various multi-modal tasks, leveraging GPT-4V as a generalist evaluator for these tasks has not yet been systematically explored. We comprehensively validate GPT-4V's capabilities for evaluation purposes, addressing tasks ranging from foundational image-to-text and text-to-image synthesis to high-level image-to-image translations and multi-images to text alignment. We employ two evaluation methods, single-answer grading and pairwise comparison, using GPT-4V. Notably, GPT-4V shows promising agreement with humans across various tasks and evaluation methods, demonstrating immense potential for multi-modal LLMs as evaluators. Despite limitations like restricted visual clarity grading and real-world complex reasoning, its ability to provide human-aligned scores enriched with detailed explanations is promising for universal automatic evaluator.
Forward citations
Cited by 12 Pith papers
-
MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
MEDit-Bench shows the same source video yields strikingly different professional edits under different narrative messages, and state-of-the-art video models still fall well short of human editors at strict cut-precisi...
-
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.
-
Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models
A grounding-guided attack that concentrates perturbation on text-matched image regions and disrupts global and local semantic alignment consistently improves adversarial transferability across multiple vision-language models.
-
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
SciFig automatically generates editable methodology figures from scientific text and claims state-of-the-art quality on its own SciFig-Eval rubric-based benchmark.
-
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
Q-Ponder is a two-stage pipeline (distill-then-reinforce) that makes a 7B multimodal model both more accurate at image quality scoring and better at explaining its judgments.
-
Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation
ABP evaluates and improves how well text-to-image models render implicit real-world knowledge.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
MultiCompose combines embedding regularization, cross-attention suppression, and mask-guided denoising to compose independently personalized subjects into one image while keeping each subject's attributes exclusive.
-
Local Brushstroke Quality Assessment via Vision-Language Feedback
Multimodal LLMs approximate expert absolute scores for calligraphy brushstrokes but show no significant rank correlation; RAG guidance trades accuracy for ranking.
-
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
A capability-oriented multimodal judge benchmark and MCTS-based preference-data generation improve judge models on some benchmarks, but the paper's SOTA claim is not supported by its own numbers.
-
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.
-
LA-RCS: LLM-Agent-Based Robot Control System
LA-RCS reports that a dual-agent LLM system controls a small car robot to complete 18 of 20 self-designed commands with the GPT-4o variant, but the supporting evaluation is inconsistent and not reproducible.
Discussion (0). Continue with ORCID to comment.