Pith. sign in

REVIEW 4 cited by

Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07356 v2 pith:MZYFFLJY submitted 2024-07-10 cs.CV

classification cs.CV
keywords videomodelsdemonstrationvisualautoregressivecapacityimitationin-context
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

People interact with the real-world largely dependent on visual signal, which are ubiquitous and illustrate detailed demonstrations. In this paper, we explore utilizing visual signals as a new interface for models to interact with the environment. Specifically, we choose videos as a representative visual signal. And by training autoregressive Transformers on video datasets in a self-supervised objective, we find that the model emerges a zero-shot capability to infer the semantics from a demonstration video, and imitate the semantics to an unseen scenario. This allows the models to perform unseen tasks by watching the demonstration video in an in-context manner, without further fine-tuning. To validate the imitation capacity, we design various evaluation metrics including both objective and subjective measures. The results show that our models can generate high-quality video clips that accurately align with the semantic guidance provided by the demonstration videos, and we also show that the imitation capacity follows the scaling law. Code and models have been open-sourced.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Training on synthetic compositional task sequences with sequence-level masking lets a transformer-based in-context learner follow multi-step medical imaging instructions on held-out images, but well below codebook upp...

  2. PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An LLM-guided, iterative prompt-refinement system that captions generated videos, detects physics and semantic mismatches, and rewrites prompts, improving physics-adherence scores on VideoPhy and PhyGenBench.

  3. VidTok: A Versatile and Open-Source Video Tokenizer

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VidTok reports state-of-the-art video reconstruction accuracy among open and published video tokenizers, using FSQ for discrete tokens, 2D+1D convolutions, and a two-stage low-resolution-to-high-resolution training recipe.

  4. Video Diffusion Transformers are In-Context Learners

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.

Pith tools