REVIEW 2 cited by
Transframer: Arbitrary Frame Prediction with Generative Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a general-purpose framework for image modelling and vision tasks based on probabilistic frame prediction. Our approach unifies a broad range of tasks, from image segmentation, to novel view synthesis and video interpolation. We pair this framework with an architecture we term Transframer, which uses U-Net and Transformer components to condition on annotated context frames, and outputs sequences of sparse, compressed image features. Transframer is the state-of-the-art on a variety of video generation benchmarks, is competitive with the strongest models on few-shot view synthesis, and can generate coherent 30 second videos from a single image without any explicit geometric information. A single generalist Transframer simultaneously produces promising results on 8 tasks, including semantic segmentation, image classification and optical flow prediction with no task-specific architectural components, demonstrating that multi-task computer vision can be tackled using probabilistic image models. Our approach can in principle be applied to a wide range of applications that require learning the conditional structure of annotated image-formatted data.
Forward citations
Cited by 2 Pith papers
-
Masked Generative Nested Transformers with Decode Time Scaling
MaGNeTS schedules progressively larger nested transformer sub-models over decode iterations and caches key-value pairs of unmasked tokens, achieving 2.5-3.7x compute reduction with competitive FID/FVD.
-
CAT: Content-Adaptive Image Tokenization
CAT uses LLM-scored captions to assign each image an 8x, 16x, or 32x compression and trains a nested VAE with variable-length latents, improving ImageNet generation FID and throughput over fixed-ratio baselines.
Discussion (0). Continue with ORCID to comment.