Pith. sign in

Paper Citation Record · LEDGER

Everything is a Video: Unifying Modalities through Next-Frame Prediction

As of 14 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2411.10503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10503 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:59:00.269362Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:42.989079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:49:52.967735Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e38793b4-4c0b-413c-a19f-c16a564bf07f · outbound

This paper cites From methods to datasets: A survey on image-caption generators.

Everything is a Video: Unifying Modalities through Next-Frame Prediction From methods to datasets: A survey on image-caption generators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.957912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:58:59.985267Z digest=sha256:d57a5c92f198f12bb29a504917cbaef15bf0963c3ee3725141ff29187ceca857

Observation 07e62f32-e74b-472f-8214-0da90b355604 · outbound

This paper cites data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language.

Everything is a Video: Unifying Modalities through Next-Frame Prediction data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:58:59.991539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:58:59.991539Z digest=sha256:539bf79882d6a0935172357931405f6058612bce86ad18761cf75fb84dc8ad4f

Observation ede4becb-eb9b-4d1d-9244-01c0b8270f6d · outbound

This paper cites Audiomnist: Exploring explainable artificial intelli- gence for audio analysis on a simple benchmark.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Audiomnist: Exploring explainable artificial intelli- gence for audio analysis on a simple benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.938193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:58:59.997316Z digest=sha256:12a2666d8568bdccb9f601784341c9a6f32435c32506da8b2d52739daddecc6d

Observation bb47d286-b6a0-48af-854c-3876671d61c9 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.919706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.009812Z digest=sha256:a75fa1c1229115b24d97cd31c71e208aab8c32e4ab23b397068ce0e74f4ae8b2

Observation 1e4e1f29-19c9-4d6c-869a-00d1968a8a6d · outbound

This paper cites Pcanet: A simple deep learning baseline for image classification? IEEE transactions on image pro- cessing, 24(12):5017–5032, 2015.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Pcanet: A simple deep learning baseline for image classification? IEEE transactions on image pro- cessing, 24(12):5017–5032, 2015

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.899756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.022299Z digest=sha256:c9690a974d9f312d51e58ca93185779f86a8bc59f9f22342ee927b336c865ee9

Observation ebd41328-7dee-49d8-81d4-4ef700d44ea0 · outbound

This paper cites Generative pre- training from pixels.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Generative pre- training from pixels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.031644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.031644Z digest=sha256:f1f2db56fede57828c2eafae57b0e99137de431d3f6719d8418f5458a3299352

Observation be956813-70c8-4907-803b-68d82d645b5f · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction UNITER: UNiversal Image-TExt Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.046717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.046717Z digest=sha256:eb74ef4bbe662ce071938969f15b809f1a0902a3cfae19f703bfae7c09a87809

Observation f116849b-3e73-444b-adb7-5411c45c0cea · outbound

This paper cites High-performance long- term tracking with meta-updater.

Everything is a Video: Unifying Modalities through Next-Frame Prediction High-performance long- term tracking with meta-updater

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.865798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.055661Z digest=sha256:2143e7c322a132b4bf6416539d0c76242a760d81c1e10a3d57d51a70e4c3a082

Observation dea018fd-f403-4c3f-8eac-732c7f57a13f · outbound

This paper cites Tinyvirat: Low-resolution video action recognition.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Tinyvirat: Low-resolution video action recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.848075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.062318Z digest=sha256:0156eb8dfee6c2b2120556cdfebcb7fa0ca33ed7560a9e0bcaa12dd159c5dcc5

Observation 6e13c638-90f6-46ce-b6cd-b42eb5df37a8 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Everything is a Video: Unifying Modalities through Next-Frame Prediction An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.831679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.068835Z digest=sha256:d03f946ebc472360215565787810ae3c8a9f12d9279f342727ee68bd67fd6df0

Observation bded84f4-ef3f-4f61-ba8c-e8b27111a274 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.077421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.077421Z digest=sha256:9757fe5e6a9be8aa48466b26bf517627e9f8de5b381edcd677713a6a4c5da387

Observation 4529372c-b4d8-4775-bb83-368e7da1498d · outbound

This paper cites Improving Language Understanding from Screenshots.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Improving Language Understanding from Screenshots

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.085021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.085021Z digest=sha256:3aa7665652d9e0188797ad9fb9f177177b58685fbc71043682b10deaf4c4b4db

Observation 1edaf6a9-600e-4a08-b89f-46ac57e44bd0 · outbound

This paper cites Photorealistic video generation with diffusion models.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Photorealistic video generation with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.804522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.091353Z digest=sha256:cf642a342af0bba7758f3774e5270b3fb7db39efee04b53713545fa0d5417d27

Observation 32a06e5c-ea02-4501-8931-9ef37cf18346 · outbound

This paper cites Unit: Multimodal mul- titask learning with a unified transformer.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Unit: Multimodal mul- titask learning with a unified transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.787385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.097416Z digest=sha256:33a0271f3c973566ab275e7d3bf15a03c191030b9fb57bc75ee96c602c9183ee

Observation b3568299-f139-4847-8c3b-24d4ba37b3c7 · outbound

This paper cites Video Question Answering with Spatio-Temporal Reasoning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Video Question Answering with Spatio-Temporal Reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.770191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.103573Z digest=sha256:cd7000858a86e0d6be9ff3340200d6a573fdf24c0d4d8c3edef8836b3a26c063

Observation 4b4732d1-b4da-4fe8-b4c3-f5baa2a6ff1d · outbound

This paper cites Unifying Question Answering, Text Classification, and Regression via Span Extraction.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Unifying Question Answering, Text Classification, and Regression via Span Extraction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.112379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.112379Z digest=sha256:3dfdf255aa7a2de768492ba8d66da34d30a6ba5554307ffb9c01276a459a6a47

Observation c6886f5d-86be-438c-afaa-8ee0a23febef · outbound

This paper cites Learning multiple layers of features from tiny images.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Learning multiple layers of features from tiny images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.119758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.119758Z digest=sha256:3b82dcf9148cd1c17e547bdadb6a00626deaefafc1e4c7e9050aa3cb6c4fe844

Observation d7940d5d-b5c9-44f4-b8fe-e9759f30d852 · outbound

This paper cites Temporally consistent video colorization with deep feature propagation and self-regularization learning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Temporally consistent video colorization with deep feature propagation and self-regularization learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.739590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.126027Z digest=sha256:bbf06bd27aafb84f087d90c22cfdbb6687ccd23ccf046d9472badbb9d2513cec

Observation 4360691c-c323-4759-ae1d-7fb0ae0608ab · outbound

This paper cites Decoupled weight de- cay regularization.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Decoupled weight de- cay regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.134549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.134549Z digest=sha256:fe2a5c18d6d80975c2c0fd518de4f9e75eb9d72f5ed982a171d127a42a1f121c

Observation 3c5327fb-a60e-4265-b2d4-6f95c06ba0af · outbound

This paper cites The multi-modal fusion in vi- sual question answering: a review of attention mechanisms.

Everything is a Video: Unifying Modalities through Next-Frame Prediction The multi-modal fusion in vi- sual question answering: a review of attention mechanisms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.710446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.140434Z digest=sha256:66773794f715c751ce8018e7ccd0c0002939e805a99bb4dbfd42acb250470cfb

Observation 36d92b00-531d-4b93-ba7d-06937ba713d6 · outbound

This paper cites The Natural Language Decathlon: Multitask Learning as Question Answering.

Everything is a Video: Unifying Modalities through Next-Frame Prediction The Natural Language Decathlon: Multitask Learning as Question Answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.145573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.145573Z digest=sha256:3b185309c4e1eae91ae4fb9969c7f0bc557ee6d4baa6258df3a61e3ed51d9f7c

Observation d2b3e8e3-6335-4962-8e27-23dcfc6297b2 · outbound

This paper cites Transframer: Arbitrary Frame Prediction with Generative Models.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Transframer: Arbitrary Frame Prediction with Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.151096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.151096Z digest=sha256:ef6d26288ecac4e735af4ed08ad0669329d518829fd79ca9915710c3e5554ce9

Observation e8febd18-221e-44e8-870b-771d3b1c81a4 · outbound

This paper cites Gpt-4 technical report, 2024.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Gpt-4 technical report, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.157591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.157591Z digest=sha256:4bf77119b9544ad0f4587b7ff8174edd148acc93598aa35211987fb36f37c56b

Observation 28a728f5-9c93-45f8-a5c9-c7221187eec2 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Video (language) modeling: a baseline for generative models of natural videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.167425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.167425Z digest=sha256:d71a0d1001bfcab4f610e1e33c50e16553a31de6f1e9d26d63e26989779edff9

Observation b8f9b94e-8d86-4e8c-a9f5-04a2d48dbaeb · outbound

This paper cites Language Modelling with Pixels.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Language Modelling with Pixels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.174488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.174488Z digest=sha256:751f358f06e80932bce11b6fd536dcb6074f14117378142b68fd5d500c7ad255

Observation 84a023d8-5388-4b65-8205-75d150f0180a · outbound

This paper cites Implicit Stacked Autoregressive Model for Video Prediction.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Implicit Stacked Autoregressive Model for Video Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.180810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.180810Z digest=sha256:7f44948f1b5544594cdf887cf82908beee75416f1a732588c5ca847abd163a40

Observation 3645c40e-e2f8-41a7-8b0e-74fd0f97f629 · outbound

This paper cites FLAVA: A Foundational Language And Vision Alignment Model.

Everything is a Video: Unifying Modalities through Next-Frame Prediction FLAVA: A Foundational Language And Vision Alignment Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.186534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.186534Z digest=sha256:fc537fe8e7784819316026c026dd172149a78a5597c4ff2d99a1b6acb3f1476b

Observation 9b5591c1-6ae5-4507-8aba-f7e3079ee87c · outbound

This paper cites Recursive deep models for semantic compositional- ity over a sentiment treebank.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Recursive deep models for semantic compositional- ity over a sentiment treebank

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.684208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.192904Z digest=sha256:a77726d348847439e6450ccd7c8d269211fe77af73f106872e53faa8f1f4b9db

Observation 3bc84ed7-9dfe-438c-a501-a72f9754e8c3 · outbound

This paper cites PIXAR: Auto-Regressive Language Modeling in Pixel Space.

Everything is a Video: Unifying Modalities through Next-Frame Prediction PIXAR: Auto-Regressive Language Modeling in Pixel Space

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.198753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.198753Z digest=sha256:396d40cef4ab83f008ef8f741cac3c225738306add176f8a64eb1c6a47a484ea

Observation 1f7ab565-4240-441f-b0eb-b5330b29b97d · outbound

This paper cites Transformation-Based Models of Video Sequences.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Transformation-Based Models of Video Sequences

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.204388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.204388Z digest=sha256:d5e44aee62ab84d8ce5354f506ac6a69be1e14788ba18311106fd0f7ed449f98

Observation 2e09d4d6-8b24-41be-a5a3-25a80c23e98a · outbound

This paper cites Multimodal llm enhanced cross- lingual cross-modal retrieval.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Multimodal llm enhanced cross- lingual cross-modal retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.667610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.210160Z digest=sha256:f224f9dc8616e39593857cd08b145e543a932a997a28462fa950fcfea51d9ffd

Observation b589a8ee-9745-4773-88b5-6140da2de8c1 · outbound

This paper cites Bovik, H.R.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Bovik, H.R

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.216372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.216372Z digest=sha256:c1cef56a5d862083f34a793f0c5d1d4280181c83ea17439e1ba472cb4ef5085f

Observation 42f29cc1-00e4-4b44-a243-a669c6a0c5eb · outbound

This paper cites Visual question answer- ing: A survey of methods and datasets.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Visual question answer- ing: A survey of methods and datasets

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.222174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.222174Z digest=sha256:6c456dfb46d4140c5d319d041199fb84440abec19563a8df55c6e73523b7c76a

Observation 59f3cd8b-2217-46d2-afab-376787778509 · outbound

This paper cites Pixel Sentence Representation Learning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Pixel Sentence Representation Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.228785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.228785Z digest=sha256:d3c5a53bd2357bb1f2558b5b7e24045a0ade2fd412197784ed2d50faca0aef29

Observation b2a3e8b8-2ffb-492f-a9ff-e7108e387f74 · outbound

This paper cites Videogpt: Video generation using vq-vae and trans- formers, 2021.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Videogpt: Video generation using vq-vae and trans- formers, 2021

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.235445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.235445Z digest=sha256:2ac1577b97c125c14012d23ca30f0d9636de9f82c3f611a5497c7ed03bbb2324

Observation b11476c7-6673-4d84-afec-b5c57ff54520 · outbound

This paper cites Colormnet: A memory-based deep spatial-temporal feature propagation network for video colorization.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Colormnet: A memory-based deep spatial-temporal feature propagation network for video colorization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.621333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.240859Z digest=sha256:52bf8a892978d15944bda58253493376fc5dcd8397610b28fbf954c5d3f0dc76

Observation 8b438b27-b62c-4236-8da5-c7e2ddcf297b · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Everything is a Video: Unifying Modalities through Next-Frame Prediction CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:59:00.245986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:59:00.245986Z digest=sha256:f7d0a26ece465d37eff8b044a76802f695ba360e55a9cc0f4f1940502b526b95

Observation bcc370f7-493e-411a-bdae-cb5f5a8c8bcf · outbound

This paper cites Akin Yilmaz and Ahmet Murat Tekalp.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Akin Yilmaz and Ahmet Murat Tekalp

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.604360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.252101Z digest=sha256:ecd5fe195a9a34acf4b9a230196b88f62c4b11317304017997c98fb613b03853

Observation 6b0c4265-dd2e-4798-bac9-51cb76b4d9b6 · outbound

This paper cites Colorful image colorization.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Colorful image colorization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.587679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.257545Z digest=sha256:4fc338d1b14f56c107548dbe48b0361f52f947420e9272789ab7097b1713c617

Observation 3902d12c-2589-4e0d-8d3d-082882a42ac4 · outbound

This paper cites Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE.

Everything is a Video: Unifying Modalities through Next-Frame Prediction Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:59:00.323391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.263231Z digest=sha256:21f7534ab590bc593dacb425f7ebd06fcf9aebdd229debeced270aeebd22be4a

Observation 2856d329-c2aa-4a9e-be1c-342be7dafe80 · outbound

This paper cites A survey on vqa: Datasets and approaches.

Everything is a Video: Unifying Modalities through Next-Frame Prediction A survey on vqa: Datasets and approaches

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:59:00.569716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:59:00.269362Z digest=sha256:ba18d93661e853bf432164cbc63438db01e694d97783aa089902f10df20879a4

Pith citing papers

Observation 3c62bd83-a8a4-42bc-a3c9-82446053c66c · inbound

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review cites this paper.

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review Everything is a Video: Unifying Modalities through Next-Frame Prediction

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:49:53.110190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:42.989079Z digest=sha256:f6790dd59fb41cd17abdd91983bf9f9bf0524dc606de698eaa033f57a5876e98