Pith. sign in

REVIEW 3 cited by

Unifying Vision-and-Language Tasks via Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.02779 v2 pith:TZMQDXTN submitted 2021-02-04 cs.CL cs.AIcs.CVcs.LG

classification cs.CLcs.AIcs.CVcs.LG
keywords singlevision-and-languagevisualarchitecturemodelstaskstextanswering
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for referring expression comprehension, and a language decoder for image captioning, etc. To alleviate these hassles, in this work, we propose a unified framework that learns different tasks in a single architecture with the same language modeling objective, i.e., multimodal conditional text generation, where our models learn to generate labels in text based on the visual and textual inputs. On 7 popular vision-and-language benchmarks, including visual question answering, referring expression comprehension, visual commonsense reasoning, most of which have been previously modeled as discriminative tasks, our generative approach (with a single unified architecture) reaches comparable performance to recent task-specific state-of-the-art vision-and-language models. Moreover, our generative approach shows better generalization ability on questions that have rare answers. Also, we show that our framework allows multi-task learning in a single architecture with a single set of parameters, achieving similar performance to separately optimized single-task models. Our code is publicly available at: https://github.com/j-min/VL-T5

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-modal single-cell foundation models via dynamic token adaptation

    q-bio.GN 2025-04 conditional novelty 7.0 of 10

    The authors introduce dynamic token adaptation to combine DNA language models with single-cell foundation models, and show that mutating GATA4's promoter in silico shifts predicted target gene embeddings in fetal card...

  2. FLIP Reasoning Challenge

    cs.CV 2025-04 conditional novelty 6.0 of 10

    The FLIP benchmark of 11,674 blockchain image-story puzzles shows best open and closed AI models reach 75.5% and 77.9% accuracy, below the 95.3% human consensus baseline.

  3. Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems

    cs.CV 2025-06 reject novelty 3.0 of 10

    A survey that organizes vision-language segmentation methods for intelligent transportation, but its synthesis is undermined by fabricated references and unverifiable benchmarks.

Pith tools