Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:08.639585Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 0 inbound Pith citation observations for arXiv:2412.04432.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:08.639585Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 88694565-ba9c-483c-9310-62f1116ee105 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a66d8743-ee4c-4cf2-a342-e0680679f100 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Frozen in time: A joint video and image encoder for end-to- end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3653f49-ce7b-4db9-bca3-17cfb9b98bfe · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Label-Efficient Semantic Segmentation with Diffusion Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1141abc7-39f4-4ee4-8990-e6794b58c03d · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe63e62-cdbc-402e-9126-83c9b5a1a303 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Align your latents: High-resolution video synthesis with la- tent diffusion models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909f728c-c898-4638-8dcf-45839aabe643 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Language models are few-shot learners
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c3547e-c80a-4fff-86e7-bd5e635aa5ee · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Collecting highly parallel data for paraphrase evaluation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b16fd6b-064d-41c6-964e-870eaf58ba7a · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ce776a-30ba-417f-b86f-382fbed87148 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Panda-70m: Captioning 70m videos with multiple cross- modality teachers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b448b3a-e129-4e2f-ae98-5b15608bdc83 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5ce52d-6fdd-4227-813f-db3e6673e0e7 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation PaLM: Scaling Language Modeling with Pathways
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed797b34-139a-4a90-b8a4-5d1692e981bc · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13653f45-ea1f-44b6-84ae-2caaae821200 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6688d341-cf16-46d8-93cb-9d55b7306015 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Preserve your own correlation: A noise prior for video diffusion models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e1d6ab-c3e5-4ca6-b4f9-94f8caa32e1e · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Planting a SEED of Vision in Large Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340650dd-c441-4c0e-abe9-d8a31787b639 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Making LLaMA SEE and Draw with SEED Tokenizer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation becff120-b99a-4655-9fd4-f82b888aec5f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32baaa1-befe-4e5e-aea0-dd386d2109a2 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation The” something something” video database for learning and evaluating visual common sense
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba5d9ee-2ce3-4e7b-819b-55bf959b6088 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Ego4d: Around the world in 3,000 hours of egocentric video
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a968b18-3a95-4fea-81de-f8fef9e5b59f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Denoising diffu- sion probabilistic models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b755da2d-c28e-4cc4-9a73-49602892c054 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22240207-50b5-41fd-8fcf-ec7b6705b946 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LoRA: Low-Rank Adaptation of Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5f184f-f8a5-4363-897c-e257a19ae903 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Soda: Bottle- 9 neck diffusion models for representation learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7702ea-a184-4ec5-bd09-88df05d65a5f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Mistral 7B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29319ca1-0dca-4645-828c-3294497c6842 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified language-vision pretraining with dynamic discrete visual tokenization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75d80948-3a44-42e7-9357-0bdca8bf4644 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29bb5ce6-43a3-407e-b862-90e29103bc8e · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation The Kinetics Human Action Video Dataset
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f4df18-8a6f-4929-89a7-9cd91a6c13e1 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Auto-Encoding Variational Bayes
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b8fc7b-47c5-489e-8ce0-d763fdc692b5 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f919c1c9-babb-4d9a-a1db-0dd7a2363342 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae669222-29c1-4638-a93e-276574568f4a · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa4197d-2cbb-4e51-ab97-6a21c561bf87 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1861889d-2ac4-4230-8f59-d8d84a041b02 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Autoregressive Image Generation without Vector Quantization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e12df6-ef38-45fc-86af-f14c4ddeeb1c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Tgif: A new dataset and benchmark on animated gif description
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af313e59-6bb0-49e6-8aba-a3778c0c5bbf · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634a98ee-5c43-499e-a08b-822ca1cf8a47 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0289fb-3bf5-4700-8762-b53642a822f8 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2babac02-5914-4200-a41a-0875694d62ee · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Llava-next: Improved reason- ing, ocr, and world knowledge, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 810fa34a-c5b2-484b-b674-741ba7105249 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Visual instruction tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e1afc70-109d-4c54-9589-d9d3e2102445 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ef846f-071d-402c-b08e-c4830fcd9b60 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bdb0bea-db6f-4b54-983f-14f23fce4394 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac167ca-53c0-45e6-9f98-d13922eae03e · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7982e3ec-573f-4b6d-b0f6-0bf9b7ce002e · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b83e4cab-1aff-4174-94e9-3686c3919300 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthe- sis
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 584aa6ec-297e-4804-bb3b-1efd5c2e749b · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gpt-4v(ision) system card, 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e460fd8c-2921-4eb6-aa4a-144f7c26046f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gpt-4o system card, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2d65fc94-30e8-43fe-8193-24a3d17670ff · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Per- ception test: A diagnostic benchmark for multimodal video models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b5e5b8c-d601-4d6d-81e5-48b7e6a3ac27 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Learning transferable visual models from natural language supervi- sion
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 627e1c39-512d-46b5-aead-988b3954a783 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation High-resolution image synthesis with latent diffusion models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00f56ba-dccd-4e88-8fa6-4119135a5d2f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Laion-5b: An open large-scale dataset for training next gen- eration image-text models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 53bf96ac-008b-4f6f-86d1-23bb9a64e7cd · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aa65f3de-f780-458d-854a-b55b950e8bc5 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e5760db-8a9d-48ed-a2a0-cae4ef3d152c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Generative Multimodal Models are In-Context Learners
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d37623-318c-40de-bb60-39c1e808cc58 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Emu: Generative Pretraining in Multimodality
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67f96e7-aeec-4e04-bfc4-f8fe67278784 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd95997-71a4-4c28-a41d-d502774f0377 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gemini: A Family of Highly Capable Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d011547-2090-42b1-bca7-07e699c125c0 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e2c36e-18a0-4084-a2f3-ab37912dc0b5 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LLaMA: Open and Efficient Foundation Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d9e5565-71fb-4537-ad98-e25f5de9769c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Givt: Generative infinite-vocabulary transformers
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d44ab93e-0129-40e2-9e46-fa3c620d224e · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e3564c-f760-455b-a901-5daaa7423144 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Neural discrete representation learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fabbf4d-909f-490c-900e-0d38893bd01c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd3ac56-11cc-44f0-af08-c1734937bba7 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Diffusion Feedback Helps CLIP See Better
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c894d03-6918-4594-8b30-10cc637b7bdc · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Videocomposer: Compositional video synthesis with motion controllability
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c79c8ba-9b7d-429c-a1ac-5da90d8ee83b · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Emu3: Next-Token Prediction is All You Need
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04bee03-6122-498b-b595-26ce1cbf779a · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68f103a-d157-45b2-89d8-ff723c089931 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3874d8e5-ffbf-437c-b867-d3e88740b3c1 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Diffusion models as masked autoencoders
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 895b9fd6-7f2c-4ee3-ba7a-fb937343097c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8e2a87-06d8-46c0-8006-f621f08f4bf6 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4aec21-24c1-4aa6-a82b-ac9ff1b200ac · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8353176b-b0b5-496b-be99-8f8a59fae0ab · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2458da6-fa1f-4222-b13f-a3f3ff68e8e4 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Denoising diffusion autoencoders are unified self-supervised 11 learners
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 87b8863b-1442-43e3-b4aa-97c7a56ae7f9 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Next-qa: Next phase of question-answering to explaining tem- poral actions
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f5a051-14fd-4c19-a0d9-39a0c5b42968 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05bebe21-3796-4d4a-b25b-e5b0b4b53e1f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Dynamicrafter: Animating open- domain images with video diffusion priors
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a79945-bb2e-4b85-ad0c-333def12026f · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Msr-vtt: A large video description dataset for bridging video and language
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 424b2904-3b46-4d0d-aa8d-298bcf40faba · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Open-vocabulary panoptic segmentation with text-to-image diffusion models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4a214b0c-b715-44a1-877a-52a85490991c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4506835-9bcd-491f-b237-7ce7b03527e5 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoGPT: Video Generation using VQ-VAE and Transformers
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa14b4d0-9b7e-4349-8337-a3f88209eff9 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d15063-c71c-42ca-afbd-4471e4f63b73 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation SEED-Story: Multimodal Long Story Generation with Large Language Model
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bee4769-a4db-476a-9edd-2644cc8520d7 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e4bc1c-ffd9-4d5f-8291-b8edfe12a830 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95660d3-6d9b-49a9-8628-e96bdc167775 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Capsfu- sion: Rethinking image-text data at scale
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5fc67231-9f23-4423-ae6f-ccfda4171d1c · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e1c248f8-9454-4c7e-9168-e4bc1ffe0389 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MonoFormer: One Transformer for Both Diffusion and Autoregression
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55abc623-112a-43fc-b6d1-e5db92259baf · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unleashing text-to-image diffusion mod- els for visual perception
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aabc4237-ac62-4232-a343-d0d6653d64cd · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326892c7-b028-44cb-9c25-a51330864047 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Towards auto- matic learning of procedures from web instructional videos
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 226b838e-729e-416e-b6dd-31d478af8707 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3206fa4-c4f1-4d61-be4b-b95d74817d99 · outbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Curious George
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.