Pith. sign in

Paper Citation Record · LEDGER

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

As of 13 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 0 inbound Pith citation observations for arXiv:2412.04432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04432 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:08.639585Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88694565-ba9c-483c-9310-62f1116ee105 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.168681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.168681Z digest=sha256:8f3fc06ee34e3292fa16fed1d4d96ff654d5188b56b92d99060affce1d517772

Observation a66d8743-ee4c-4cf2-a342-e0680679f100 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.176057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.176057Z digest=sha256:566f90295d8fd65615ff31a130ebdf45ddc6f865fcccbbc71d13227a34380413

Observation c3653f49-ce7b-4db9-bca3-17cfb9b98bfe · outbound

This paper cites Label-Efficient Semantic Segmentation with Diffusion Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Label-Efficient Semantic Segmentation with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.182315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.182315Z digest=sha256:77c564a06d5e0c3e89912cacfdbbe52499032d105c8a3e8e93f35ff2ed88081b

Observation 1141abc7-39f4-4ee4-8990-e6794b58c03d · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.188344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.188344Z digest=sha256:13b3ee51bdeb1034b4b7b0d6620012c17b0407917b3ce1cf102634928284f346

Observation efe63e62-cdbc-402e-9126-83c9b5a1a303 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.194537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.194537Z digest=sha256:7bef8186c8c08a1b1a898309c7ae4f5d99c9d58725efda7c0f153766c666c1a4

Observation 909f728c-c898-4638-8dcf-45839aabe643 · outbound

This paper cites Language models are few-shot learners.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.200580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.200580Z digest=sha256:f45ff6ea272e44b05197febdb59ae0339450ae9f8dc724e6c547a7788bee776f

Observation 86c3547e-c80a-4fff-86e7-bd5e635aa5ee · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Collecting highly parallel data for paraphrase evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.207912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.207912Z digest=sha256:c854d9e76227d81eda1d84630f0bfa75a148ac1f5de7c080d2d37959bd396cce

Observation 2b16fd6b-064d-41c6-964e-870eaf58ba7a · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.215677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.215677Z digest=sha256:3bf47b9c9348d7511d2e75ab61d838463232874db2fc8b9a4395a7094ced71bd

Observation c0ce776a-30ba-417f-b86f-382fbed87148 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross- modality teachers.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Panda-70m: Captioning 70m videos with multiple cross- modality teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.223266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.223266Z digest=sha256:3b3bd87faa6e611bd432aa63e8e0b819e43ee4d0a2d435f730d8bdc66069d980

Observation 1b448b3a-e129-4e2f-ae98-5b15608bdc83 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.228320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.228320Z digest=sha256:c87c6936292fe0d4c4f963bae2ce1efc2664f552d3bfe4de786f1dbc314632cd

Observation 1e5ce52d-6fdd-4227-813f-db3e6673e0e7 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation PaLM: Scaling Language Modeling with Pathways

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.233464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.233464Z digest=sha256:89ed18539658ec7dd234c18b501e5f2dd825dcc2a52ee5e3f76147e29f5c218d

Observation ed797b34-139a-4a90-b8a4-5d1692e981bc · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.238448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.238448Z digest=sha256:eb85e6c2a71ef3ec59dc72b0acf29f479163262d0c24e0bbcac5c3f34d827fb2

Observation 13653f45-ea1f-44b6-84ae-2caaae821200 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.243491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.243491Z digest=sha256:5b7f40128fa3d74551435d9103ba06ac664471c52beb9a02e56f97ea74b2d9bf

Observation 6688d341-cf16-46d8-93cb-9d55b7306015 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Preserve your own correlation: A noise prior for video diffusion models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.248720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.248720Z digest=sha256:df03fc5280efac251924e28cf6f937c1c09c8ba1e44f33565d1a51c278334aaa

Observation 61e1d6ab-c3e5-4ca6-b4f9-94f8caa32e1e · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Planting a SEED of Vision in Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.253576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.253576Z digest=sha256:79b97a83721a93de2e190a8e6f532ddcf053636ef58587522ae6b72204f39959

Observation 340650dd-c441-4c0e-abe9-d8a31787b639 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.258608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.258608Z digest=sha256:bada0ea8e54d6f7485fe1314692da7f597c3d857b5049d1b2b2c56eaae41a413

Observation becff120-b99a-4655-9fd4-f82b888aec5f · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.264244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.264244Z digest=sha256:5d18f6b42d01d6c18c9f247e4fbaf0c07fc047972167b1972315c7c1212a1f82

Observation e32baaa1-befe-4e5e-aea0-dd386d2109a2 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation The” something something” video database for learning and evaluating visual common sense

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.270201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.270201Z digest=sha256:db983d67bdd87e709f47bf652dce838749d7a4cfad691e932ed98902c0805f4a

Observation 3ba5d9ee-2ce3-4e7b-819b-55bf959b6088 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.275037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.275037Z digest=sha256:f4c717c1f161c96305b9e452e935ee14bc1c8d4df0cd5907a3b12aa9ea27fe58

Observation 0a968b18-3a95-4fea-81de-f8fef9e5b59f · outbound

This paper cites Denoising diffu- sion probabilistic models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Denoising diffu- sion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.279158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.279158Z digest=sha256:1ee5fb9bfe7065d56eef62b02ad1411718c4e6f624c69505e1cb15eb2ca84872

Observation b755da2d-c28e-4cc4-9a73-49602892c054 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.283173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.283173Z digest=sha256:a994ac30cbf543e8b49fdb64731941f6a3a6058a41aae714e930c99c613c0200

Observation 22240207-50b5-41fd-8fcf-ec7b6705b946 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.288095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.288095Z digest=sha256:f1d63ea3ec87ba25181802b11fd23d446537b678a973a78d15ff71865201fa3f

Observation 5d5f184f-f8a5-4363-897c-e257a19ae903 · outbound

This paper cites Soda: Bottle- 9 neck diffusion models for representation learning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Soda: Bottle- 9 neck diffusion models for representation learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.292734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.292734Z digest=sha256:7f2b339bd1ebff84247f4c2ac01bcecac33f27c519f26ac726fde41e51dd9aae

Observation cb7702ea-a184-4ec5-bd09-88df05d65a5f · outbound

This paper cites Mistral 7B.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.296984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.296984Z digest=sha256:447c1a9c0dc8e2af194e60d066f91809b85a6cd47ab2ea7cd116a154015623ef

Observation 29319ca1-0dca-4645-828c-3294497c6842 · outbound

This paper cites Unified language-vision pretraining with dynamic discrete visual tokenization.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified language-vision pretraining with dynamic discrete visual tokenization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.301962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.301962Z digest=sha256:5acfe18eaddf0e590cbe6f84a0474f427f899ddb536062ca9b2476770aef5f62

Observation 75d80948-3a44-42e7-9357-0bdca8bf4644 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.307015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.307015Z digest=sha256:7b3265cc799774b9f2ab2322676da4192a5a7a4ba6e9603e36426d280873a1ad

Observation 29bb5ce6-43a3-407e-b862-90e29103bc8e · outbound

This paper cites The Kinetics Human Action Video Dataset.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation The Kinetics Human Action Video Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.311855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.311855Z digest=sha256:b89b9f4d17e507ace41d97adf7fec126144d6873c8b178ac0d89222ef56aa423

Observation 85f4df18-8a6f-4929-89a7-9cd91a6c13e1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Auto-Encoding Variational Bayes

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.317096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.317096Z digest=sha256:4b88db9c0f2dc338c3b67cdf16fa8ed57948db76e73cc2b922776ad79a2c5e7d

Observation 89b8fc7b-47c5-489e-8ce0-d763fdc692b5 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.321924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.321924Z digest=sha256:2b2264bab214d90388b2dbf2a39446200a34b9aa1606f584f3171af784fc3c79

Observation f919c1c9-babb-4d9a-a1db-0dd7a2363342 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.327003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.327003Z digest=sha256:d0c90122786efcbfa00b73e96e68eec75d5fbe95d4816b58845501fc07fa09af

Observation ae669222-29c1-4638-a93e-276574568f4a · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.333026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.333026Z digest=sha256:cd031a3cd8164d89d5d824f3da899e9fc3abf0bcc54d35a4887624b0adff6821

Observation 9fa4197d-2cbb-4e51-ab97-6a21c561bf87 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.338245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.338245Z digest=sha256:2d559ccaeaec9aaeea9c4e51ee3dca5c90e856d4387b3678d187f6700c3df60d

Observation 1861889d-2ac4-4230-8f59-d8d84a041b02 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Autoregressive Image Generation without Vector Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.343535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.343535Z digest=sha256:e5bc2335f856f795f028da9310ae1ca2f5b3094c075d85020b06a252fd6f1de4

Observation b5e12df6-ef38-45fc-86af-f14c4ddeeb1c · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Tgif: A new dataset and benchmark on animated gif description

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.348486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.348486Z digest=sha256:c5fc379c6041b25dfa64b0c522219f55f2450a5462ba7e92bfeb7444a889426f

Observation af313e59-6bb0-49e6-8aba-a3778c0c5bbf · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.353411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.353411Z digest=sha256:72edd2c168a8022c94f182e78184c76df1f9138e9e39ca104dba68f05452ae29

Observation 634a98ee-5c43-499e-a08b-822ca1cf8a47 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.358731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.358731Z digest=sha256:d73622799fe32af9a537ac5c490f03005f47a5526f47c9821395427db8783008

Observation cf0289fb-3bf5-4700-8762-b53642a822f8 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.364042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.364042Z digest=sha256:da72f04856767e60f69e0cb0443553a1208459b0216620a96d0695f1337d9b92

Observation 2babac02-5914-4200-a41a-0875694d62ee · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.121409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.369328Z digest=sha256:20255cda6fbb642a328a18787662204bc16cfe764c89ed74fedaedec39f0445e

Observation 810fa34a-c5b2-484b-b674-741ba7105249 · outbound

This paper cites Visual instruction tuning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.374239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.374239Z digest=sha256:b42d082691c3d767ed48dfa00bf0512fe82e3cfbadef4712f5a0e8a58c8aff0a

Observation 2e1afc70-109d-4c54-9589-d9d3e2102445 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.379980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.379980Z digest=sha256:39a465f4e11b60c49f13a8e25fe570bed52594d3bc2a57387bc8d623b9d02bd9

Observation e8ef846f-071d-402c-b08e-c4830fcd9b60 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.385354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.385354Z digest=sha256:0ca7c22f59dd662c345e9b3f53c3a66ec6465edbd3cbf00bc5602d321527f489

Observation 3bdb0bea-db6f-4b54-983f-14f23fce4394 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.390529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.390529Z digest=sha256:601fa9ec0c116b6c36f317e4859b65e437bc08f5f825790b40a987b75729cda9

Observation 9ac167ca-53c0-45e6-9f98-d13922eae03e · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.395867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.395867Z digest=sha256:a52aab534bdf62fdf158d5783db5e590f07e3c42e1126b752c65cbebf15cf86c

Observation 7982e3ec-573f-4b6d-b0f6-0bf9b7ce002e · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.092270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.401277Z digest=sha256:99037ade16f0d9f90a9391ef2249836bf97b73efa20f3ce4b56d3c2cadcd9fb9

Observation b83e4cab-1aff-4174-94e9-3686c3919300 · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthe- sis.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthe- sis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.073428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.405298Z digest=sha256:5438686db9b6fd9dbf19f0d892645a33e146747db9675bca9740fe30a554ac1e

Observation 584aa6ec-297e-4804-bb3b-1efd5c2e749b · outbound

This paper cites Gpt-4v(ision) system card, 2023.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gpt-4v(ision) system card, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.056575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.409431Z digest=sha256:79369eff8755ffeafa1bdef031cbb1529b5888b719ae77ab7448ff7fbbf4222c

Observation e460fd8c-2921-4eb6-aa4a-144f7c26046f · outbound

This paper cites Gpt-4o system card, 2024.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gpt-4o system card, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.413406Z digest=sha256:5d8d9aabe2c58243789eca90673ae145036735e9c509221bf73ca93211e5beed

Observation 2d65fc94-30e8-43fe-8193-24a3d17670ff · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Per- ception test: A diagnostic benchmark for multimodal video models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.025216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.417513Z digest=sha256:1d804526fd7475e2192ba238edff95914d5446e580b30e526f09e62438fc8f13

Observation 8b5e5b8c-d601-4d6d-81e5-48b7e6a3ac27 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Learning transferable visual models from natural language supervi- sion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:10.008927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.421944Z digest=sha256:89b343d47686982a6ce1317b579e3b30646ad7ab02d7bb46054f97d04e9af689

Observation 627e1c39-512d-46b5-aead-988b3954a783 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation High-resolution image synthesis with latent diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.426561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.426561Z digest=sha256:9bdd55829819692eabfc6ab3c06499c95fad33c03d05b35a9c869a0b14119cd1

Observation a00f56ba-dccd-4e88-8fa6-4119135a5d2f · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.977976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.431183Z digest=sha256:ccdbcf90a42c9ee2cd286f729d6b41f7b4e53b437cca7f2b58e0497a44c2e962

Observation 53bf96ac-008b-4f6f-86d1-23bb9a64e7cd · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.961577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.435975Z digest=sha256:ffa8804ecd4827c7f4b681b765c692b0170b0ffdb5dd76218a22d408b1a4ef6b

Observation aa65f3de-f780-458d-854a-b55b950e8bc5 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.440449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.440449Z digest=sha256:94a2cf5839e2348afee309b1013bf42766084a7bbb20739507112e0bb9674e83

Observation 3e5760db-8a9d-48ed-a2a0-cae4ef3d152c · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Generative Multimodal Models are In-Context Learners

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.445180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.445180Z digest=sha256:3b2d6f2899842e7f89cd5830bcf26ebc1fda366a37f86f2ca01fa1fbe0e7656c

Observation 23d37623-318c-40de-bb60-39c1e808cc58 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Emu: Generative Pretraining in Multimodality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.455468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.455468Z digest=sha256:218c89fad08b51f1502639467db551a62f96617118a2f1c4d6c84fe8c9618ce5

Observation d67f96e7-aeec-4e04-bfc4-f8fe67278784 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.460392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.460392Z digest=sha256:46ecae8b511b5dd1cf5742fe8d7b2f27040304de78b162381917119a5bd758c8

Observation ffd95997-71a4-4c28-a41d-d502774f0377 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.465177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.465177Z digest=sha256:2ba873108b37e0b104395e7393477a27a51e68edae46cf75ed5722f83792b9bb

Observation 9d011547-2090-42b1-bca7-07e699c125c0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.469723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.469723Z digest=sha256:f6fde26d9ef180160e98a35e6daec0a54fdd35b7f26be3fc33778239e723b784

Observation a2e2c36e-18a0-4084-a2f3-ab37912dc0b5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.474548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.474548Z digest=sha256:21f8801e78d988b85b7b6dc5898f544082dfef36ae21c14aa634ae4a270ed465

Observation 4d9e5565-71fb-4537-ad98-e25f5de9769c · outbound

This paper cites Givt: Generative infinite-vocabulary transformers.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Givt: Generative infinite-vocabulary transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.944108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.479670Z digest=sha256:e2a810caf430210153f5e95b3507d4398158a0b84cbcb9ba1862e0136d6e1118

Observation d44ab93e-0129-40e2-9e46-fa3c620d224e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.484543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.484543Z digest=sha256:639286fb253b594d7362b5474e85c09932574f204d26febcb11ea22810ff90b0

Observation d8e3564c-f760-455b-a901-5daaa7423144 · outbound

This paper cites Neural discrete representation learning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Neural discrete representation learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.489443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.489443Z digest=sha256:98aac2dbdbbd274f15e41dc432d523b834871cc48de468ce7ae7c0b4dd0160e9

Observation 8fabbf4d-909f-490c-900e-0d38893bd01c · outbound

This paper cites LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.494766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.494766Z digest=sha256:cf439e66d90b909d3595454a4c15100daf5a738bf4a8535146e74483b2aae031

Observation acd3ac56-11cc-44f0-af08-c1734937bba7 · outbound

This paper cites Diffusion Feedback Helps CLIP See Better.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Diffusion Feedback Helps CLIP See Better

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.500992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.500992Z digest=sha256:78f2772f33ef23cad7e0cd01585a98f9fa0cd84ed90390bb98ae4efc9893cf1f

Observation 9c894d03-6918-4594-8b30-10cc637b7bdc · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Videocomposer: Compositional video synthesis with motion controllability

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.918380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.506963Z digest=sha256:1faf0afd6db76221cca693396c2953218d21e4890783595bc1f04c24e62255c0

Observation 8c79c8ba-9b7d-429c-a1ac-5da90d8ee83b · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Emu3: Next-Token Prediction is All You Need

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.511826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.511826Z digest=sha256:3ba68763bc35861b0b643bbbe5e0e1a93751b2effd1de58555b794b60cfcb681

Observation d04bee03-6122-498b-b595-26ce1cbf779a · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.516416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.516416Z digest=sha256:5de9616fc8b1ed6c49d9ade71408451bc08c68edf7ec88480811319c96168466

Observation a68f103a-d157-45b2-89d8-ff723c089931 · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.521692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.521692Z digest=sha256:3d84570b309ff03b203a34646212990e76fee32f0949419025ace8645f90c868

Observation 3874d8e5-ffbf-437c-b867-d3e88740b3c1 · outbound

This paper cites Diffusion models as masked autoencoders.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Diffusion models as masked autoencoders

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.903063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.525932Z digest=sha256:eb80a354f006c74d68e6eb002a71e9567440d5e9c096722566e82498f637d32c

Observation 895b9fd6-7f2c-4ee3-ba7a-fb937343097c · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.529926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.529926Z digest=sha256:c6e1d6528e37f0eb96ae20870b3579ed6ab043fbf787f19bd4ea53c077415fa5

Observation 5f8e2a87-06d8-46c0-8006-f621f08f4bf6 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.534148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.534148Z digest=sha256:4886078f6ac1537d2e27a724b71470ca0fd09700cfe8c85fa0aa506c74ce29a7

Observation ae4aec21-24c1-4aa6-a82b-ac9ff1b200ac · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.538409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.538409Z digest=sha256:c8b7684d28ba23da72081f52c2953b33ce8675458ad651b9a1bbd7082a9def5b

Observation 8353176b-b0b5-496b-be99-8f8a59fae0ab · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.543892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.543892Z digest=sha256:c3a2da29f4beca30a2b3045797be1154d0a2b8322e18c91a796f616760169a1f

Observation b2458da6-fa1f-4222-b13f-a3f3ff68e8e4 · outbound

This paper cites Denoising diffusion autoencoders are unified self-supervised 11 learners.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Denoising diffusion autoencoders are unified self-supervised 11 learners

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.886501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.548835Z digest=sha256:afc316fc87931a8bd1ba2d4b3068f02cbb4f1e5592bf2b9704663327917cd80a

Observation 87b8863b-1442-43e3-b4aa-97c7a56ae7f9 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining tem- poral actions.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Next-qa: Next phase of question-answering to explaining tem- poral actions

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.553398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.553398Z digest=sha256:c8dd847032eef8bd647d24716027592255991c121a1c7af729876af4cf9f5c4e

Observation f7f5a051-14fd-4c19-a0d9-39a0c5b42968 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.558166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.558166Z digest=sha256:d59cc854bc4a7c24e5ff12546e31be39a64f60c0c6c393504ceb38287efa31fd

Observation 05bebe21-3796-4d4a-b25b-e5b0b4b53e1f · outbound

This paper cites Dynamicrafter: Animating open- domain images with video diffusion priors.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Dynamicrafter: Animating open- domain images with video diffusion priors

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.563006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.563006Z digest=sha256:9166d94f1e662b2cd54f9e100ee467f51ab80031aed830cef2ab6eb1b5e8785f

Observation a7a79945-bb2e-4b85-ad0c-333def12026f · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Msr-vtt: A large video description dataset for bridging video and language

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.849437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.567869Z digest=sha256:fcaeed4d27b49a0c13fe5c55c344381ef5e2a91fa91ceccb2e6389866995f25b

Observation 424b2904-3b46-4d0d-aa8d-298bcf40faba · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.831409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.572476Z digest=sha256:58e48450576a8d1f0b2e51599420e0a5b2d32eef4fbb149d135199e8088229e0

Observation 4a214b0c-b715-44a1-877a-52a85490991c · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.577527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.577527Z digest=sha256:2e3f726e7a2e50bc6f03af53a219a67799c57bb120112bff110acfbc32f93872

Observation e4506835-9bcd-491f-b237-7ce7b03527e5 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.583416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.583416Z digest=sha256:6fa388dfde28a1c2cf8613be29b5353547052d3da10bee25697ef8db36d43de4

Observation fa14b4d0-9b7e-4349-8337-a3f88209eff9 · outbound

This paper cites MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.588482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.588482Z digest=sha256:6531f0393aa76e5c9201923a2ee0ff07cb75e3c5d7666e3a0cd3cfbca45a3aaf

Observation d6d15063-c71c-42ca-afbd-4471e4f63b73 · outbound

This paper cites SEED-Story: Multimodal Long Story Generation with Large Language Model.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.593471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.593471Z digest=sha256:7abe53e091c29c87336237a0a3073c93ec012800e32f540d5d03f78873cdd92b

Observation 1bee4769-a4db-476a-9edd-2644cc8520d7 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.598608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.598608Z digest=sha256:3fd850fd08225926ba6a9ef899a1472ff519ff2e562db682d469e938c43ef0e1

Observation 14e4bc1c-ffd9-4d5f-8291-b8edfe12a830 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.603875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.603875Z digest=sha256:fa53aee3fa8343510b289c142758c8bfe010f75b29868d9d2774eee7f2a05a26

Observation e95660d3-6d9b-49a9-8628-e96bdc167775 · outbound

This paper cites Capsfu- sion: Rethinking image-text data at scale.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Capsfu- sion: Rethinking image-text data at scale

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.814067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.608451Z digest=sha256:3637ea744c3227cb2568dfcc9ab2c65579e75c7e0ad4a2dc20aa4ee4a98443c8

Observation 5fc67231-9f23-4423-ae6f-ccfda4171d1c · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.797396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.612556Z digest=sha256:ed474d50c1997e9139e6f414b1b00e4543eace1699e42106fc8c99877528bb1b

Observation e1c248f8-9454-4c7e-9168-e4bc1ffe0389 · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.616778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.616778Z digest=sha256:a0933fabf02cc47f16a6dc4d7ff3f147aed0a78f098376057dcea4de871fd57e

Observation 55abc623-112a-43fc-b6d1-e5db92259baf · outbound

This paper cites Unleashing text-to-image diffusion mod- els for visual perception.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unleashing text-to-image diffusion mod- els for visual perception

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.781187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.620784Z digest=sha256:7ddcb087607b6f20f802863b9caa8c00363cf70918f9ae47a808b8cbca8bf832

Observation aabc4237-ac62-4232-a343-d0d6653d64cd · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.624654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.624654Z digest=sha256:f49274cad9a2c1257f510911d79e3e17f3ddf60311404d70d5ae0f9efb1f9f15

Observation 326892c7-b028-44cb-9c25-a51330864047 · outbound

This paper cites Towards auto- matic learning of procedures from web instructional videos.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Towards auto- matic learning of procedures from web instructional videos

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.763515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.628936Z digest=sha256:122b651e8f6ff447fffd51a4224c4fb8c1a209d012f234ba5eb3a9e583436bea

Observation 226b838e-729e-416e-b6dd-31d478af8707 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.634738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.634738Z digest=sha256:c10bec5d53236e730cab3df27ceb896d5710c74d50a387b45ec31492408b561c

Observation c3206fa4-c4f1-4d61-be4b-b95d74817d99 · outbound

This paper cites Curious George.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Curious George

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:09.746735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:30:08.639585Z digest=sha256:0af1516bd1375b4616b3ee5641e4725ddbfd156c6c0448279695a3cbe06af03e

Pith citing papers

No inbound Pith citation observations are available.