Pith. sign in

REVIEW 12 cited by

GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01361 v1 pith:Z4YGASRU submitted 2023-11-02 cs.CV cs.CL

classification cs.CVcs.CL
keywords gpt-4vtasksevaluationevaluatorpromisinggeneralistgradinglimitations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various multi-modal tasks, leveraging GPT-4V as a generalist evaluator for these tasks has not yet been systematically explored. We comprehensively validate GPT-4V's capabilities for evaluation purposes, addressing tasks ranging from foundational image-to-text and text-to-image synthesis to high-level image-to-image translations and multi-images to text alignment. We employ two evaluation methods, single-answer grading and pairwise comparison, using GPT-4V. Notably, GPT-4V shows promising agreement with humans across various tasks and evaluation methods, demonstrating immense potential for multi-modal LLMs as evaluators. Despite limitations like restricted visual clarity grading and real-world complex reasoning, its ability to provide human-aligned scores enriched with detailed explanations is promising for universal automatic evaluator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    MEDit-Bench shows the same source video yields strikingly different professional edits under different narrative messages, and state-of-the-art video models still fall well short of human editors at strict cut-precisi...

  2. H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.

  3. Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models

    cs.CR 2026-02 conditional novelty 6.0 of 10

    A grounding-guided attack that concentrates perturbation on text-matched image regions and disrupts global and local semantic alignment consistently improves adversarial transferability across multiple vision-language models.

  4. SciFig: Towards Automating Editable Figure Generation for Scientific Papers

    cs.AI 2026-01 conditional novelty 6.0 of 10

    SciFig automatically generates editable methodology figures from scientific text and claims state-of-the-art quality on its own SciFig-Eval rubric-based benchmark.

  5. Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Q-Ponder is a two-stage pipeline (distill-then-reinforce) that makes a 7B multimodal model both more accurate at image quality scoring and better at explaining its judgments.

  6. Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ABP evaluates and improves how well text-to-image models render implicit real-world knowledge.

  7. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.

  8. MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

    cs.CV 2026-08 conditional novelty 5.0 of 10

    MultiCompose combines embedding regularization, cross-attention suppression, and mask-guided denoising to compose independently personalized subjects into one image while keeping each subject's attributes exclusive.

  9. Local Brushstroke Quality Assessment via Vision-Language Feedback

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Multimodal LLMs approximate expert absolute scores for calligraphy brushstrokes but show no significant rank correlation; RAG guidance trades accuracy for ranking.

  10. Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

    cs.AI 2026-02 reject novelty 5.0 of 10

    A capability-oriented multimodal judge benchmark and MCTS-based preference-data generation improve judge models on some benchmarks, but the paper's SOTA claim is not supported by its own numbers.

  11. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

  12. LA-RCS: LLM-Agent-Based Robot Control System

    cs.RO 2025-05 reject novelty 4.0 of 10

    LA-RCS reports that a dual-agent LLM system controls a small car robot to complete 18 of 20 self-designed commands with the GPT-4o variant, but the supporting evaluation is inconsistent and not reproducible.

Pith tools