Pith. sign in

Paper Citation Record · LEDGER

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

As of 14 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2411.14762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14762 v4

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:00:31.863822Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T08:52:31.686474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:52:31.987543Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fdcf8ded-eaa4-48f4-ac5e-3ef9afc0e199 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.797094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.572869Z digest=sha256:010649e83652edbdf3e4505d04713150d9a3147825ba1a551f162f53ab7a578d

Observation f75b73ef-49af-4791-8269-8e9a2363f1b3 · outbound

This paper cites Video generation models as world simulators.Ope- nAI Blog, 2024.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Video generation models as world simulators.Ope- nAI Blog, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.783433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.577481Z digest=sha256:569c15e5512b8906214c853ade39d1370ebc0263d32f45b644e4f4e0242b7565

Observation eb1be525-3e00-4e47-9835-08ef5b59f461 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Quo vadis, action recognition? a new model and the kinetics dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.769593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.581261Z digest=sha256:117711b72db61f182f14fcb9717651fb8758c7b3c29047e5953d44b79c22faa1

Observation 3d82438e-94fb-4764-95b3-9d558cdb7a53 · outbound

This paper cites Efficient geometry-aware 3d generative adversarial networks.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Efficient geometry-aware 3d generative adversarial networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.754149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.585350Z digest=sha256:bc9dbb63b5a1c5ac62ade018239a3213429730fa88d0bbfac979c6028f0f40d0

Observation aa3b9a8d-a7bd-45b2-a84e-1161229179ad · outbound

This paper cites MaskGIT: Masked generative image transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction MaskGIT: Masked generative image transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.739628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.588957Z digest=sha256:76a77af34a3f269a550a3eb4e9ad75bafa920a680102719933c17f4bc2c81a67

Observation 5e026116-80e1-4e49-bc7d-a2b52585c019 · outbound

This paper cites VideoINR: Learning video implicit neural represen- tation for continuous space-time super-resolution.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction VideoINR: Learning video implicit neural represen- tation for continuous space-time super-resolution

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.726521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.592709Z digest=sha256:3c6dd732be40b71287f1f42588ec5abf0f56fd94fe249fe587573d6ee925c568

Observation eb7fb135-130b-4adc-8f15-7e81c1d3dfcc · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.713298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.596493Z digest=sha256:c1f9a4e66476e0c055374c8186d94e4b10767ea40215529cf5c2566df522cddc

Observation e3104583-958f-464b-8f19-53c50cea674e · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Taming transformers for high-resolution image synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.698626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.600143Z digest=sha256:85a6c5bdd6a8cb51e361fc3a6d658fb5e0b50ba38948b5fb6edd6842b89d1e43

Observation f3793d6f-e1eb-4587-aa6c-6d9df4e143f2 · outbound

This paper cites Cosmos world foundation model platform for physical ai.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Cosmos world foundation model platform for physical ai

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.683996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.603912Z digest=sha256:36dae874db39de3cdbf4f05170a18e5c9c445120fda8d2d1e9d08c2cb8f959da

Observation 6ab4570e-0667-49a7-a293-7d0f1f07e86f · outbound

This paper cites Latte: Latent diffusion transformer for video generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Latte: Latent diffusion transformer for video generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.670460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.607393Z digest=sha256:9e6b6ea087a54bc23d910cfc45ccd59e307528582f4ab43565c0ca6d01027905

Observation 34bcc14b-7ae8-4cd8-8814-30d0d0ec37d6 · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Long video generation with time-agnostic vqgan and time- sensitive transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.658379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.611475Z digest=sha256:e01af69f7cb21ff2b876cf3242c8202f6be37a538c80e573f837566ab0adb408

Observation 97a2ac37-c453-4817-861f-2f0f12886975 · outbound

This paper cites Generative adversarial nets.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Generative adversarial nets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.615363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.615363Z digest=sha256:12563ee13e5d3a51a6ced59103d7d79b27915e525205b2dc4ea5997ce3af8aad

Observation 23499501-afd2-4824-a499-3965b05ee614 · outbound

This paper cites Photorealistic video generation with diffusion models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Photorealistic video generation with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.638853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.619377Z digest=sha256:e0e39321bd6d6104a62af139b57f3e39b765ef97f99269e30708634e9d21b96d

Observation 4db7a237-a658-48e2-87a1-e618b108fc0c · outbound

This paper cites A technical overview of av1.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction A technical overview of av1

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.626160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.623316Z digest=sha256:44738fd3d4865dc24e5b2f289e8559732f10ef7ec8ea760f09a8dd0a0abfc0fc

Observation 9b64092f-be6f-4740-991c-a871edfe2b6f · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.627658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.627658Z digest=sha256:b53f281e0bacd3443c7f97d06db3d5024c668294a8c1fd51947fda4d015a27cc

Observation 7986910f-9192-4604-b608-ee87e563ef7d · outbound

This paper cites Denoising dif- fusion probabilistic models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Denoising dif- fusion probabilistic models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.613170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.632169Z digest=sha256:4d81b993e62fa18ae42ec7d8d9276df4c255045752ae4625f421749da54e0398

Observation 06fba8ce-8705-4ce4-aa4e-c69e33e6896f · outbound

This paper cites CogVideo: Large-scale pretraining for text-to- video generation via transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction CogVideo: Large-scale pretraining for text-to- video generation via transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.600180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.635831Z digest=sha256:de3f0530c3b05f3b08c5ec59936e4626c83d1b5dda3932ea267bc4c750ab51e5

Observation b89f7285-0f2e-4717-92a7-7e192ee4b81c · outbound

This paper cites LRM: Large reconstruction model for single image to 3D.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction LRM: Large reconstruction model for single image to 3D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.587947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.639953Z digest=sha256:dcee72943effd490465451a253fead5f632943c29f9cfefb2d1316ec395ce0a4

Observation ccc72d19-cf63-4fcd-a704-37c7873e9d70 · outbound

This paper cites Shap-E: Generating Conditional 3D Implicit Functions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Shap-E: Generating Conditional 3D Implicit Functions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.643821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.643821Z digest=sha256:aa776f8c58689ca247f76045f84712288691bca387f31a772f300f813ba24c8f

Observation 119ccc18-5727-4a3a-8897-8997a5ba4761 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Analyzing and improving the image quality of stylegan

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.575638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.648125Z digest=sha256:a6a5b8c28a917fa731c94a25595cc2041e0da10290593ba4c8b5d6112383027d

Observation 2583c661-2372-4d60-87fe-c2a7f33577af · outbound

This paper cites Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.564535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.651927Z digest=sha256:f0d21cf98fb04265e0cc3785c2b21296eddabe188101fa831b81714ed0192aff

Observation 9bcbdfec-f263-4587-879f-00f1ef29cf94 · outbound

This paper cites Scal- able neural video representations with learnable positional features.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Scal- able neural video representations with learnable positional features

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.552140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.656101Z digest=sha256:3b68e8d3972ac26ef764f9beb91704fc12f3a3c9fde1b35a4ca854723d22e999

Observation b01e019c-8d81-4a69-91fb-c4a39d026ed2 · outbound

This paper cites 9 Videopoet: A large language model for zero-shot video gen- eration.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction 9 Videopoet: A large language model for zero-shot video gen- eration

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.540537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.660065Z digest=sha256:253a75cd693bc31ccb437f0de3eb9905115fafac44e15ad416d5bedc0e95f930

Observation 89bcc7af-9278-4274-9cf9-8bb033ec06c1 · outbound

This paper cites One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.527621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.663773Z digest=sha256:bb4b55e049a81fb3339c8eab9d0bbeb93892a796ce610b131b38014c2f773fa0

Observation 25a2cb62-04f1-4d72-9ff9-3a1e82af3650 · outbound

This paper cites Decoupled weight de- cay regularization.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Decoupled weight de- cay regularization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.513140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.667725Z digest=sha256:c9cb883931c13fbc707edf1f2d95dc73ed76edb3c95209b1b41cb647302af22d

Observation 4b571676-00d6-4ce0-a8ca-29d01e70efc0 · outbound

This paper cites Vdt: General-purpose video diffusion transformers via mask modeling.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Vdt: General-purpose video diffusion transformers via mask modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.500891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.672238Z digest=sha256:179947eb263b3ef469b527f5f6e30662c3e63679f259c0044c28cf5f726047b1

Observation bc0fb863-4141-40b1-a09b-09d638cb210f · outbound

This paper cites SiT: Explor- ing flow and diffusion-based generative models with scalable interpolant transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction SiT: Explor- ing flow and diffusion-based generative models with scalable interpolant transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.489274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.676893Z digest=sha256:869b22f524cfca6da523b6a86622aa5a088018b295d0c4a5efb40e8245923ce6

Observation d457ce73-e466-4c2e-8381-becbc9a8c05f · outbound

This paper cites an unresolved cited work.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:00:32.477424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.681174Z digest=sha256:aa94622d6371641382093553cf3736935d19d94970d59a1e3042306a2df77aee

Observation 67b4eefb-c27d-444d-9096-c39431ff008a · outbound

This paper cites GTA: A geometry-aware attention mechanism for multi-view transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction GTA: A geometry-aware attention mechanism for multi-view transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.466216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.685014Z digest=sha256:4ef5fa32908ff1a08e77a585e0e49d0c35587b225f061991c70fe84db246d787

Observation 1dbf2046-3ea5-4a77-93b0-f49b07b4714c · outbound

This paper cites A technical overview of vp9—the latest open- source video codec.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction A technical overview of vp9—the latest open- source video codec

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.455366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.689257Z digest=sha256:12399cd1272a58a2319534849bdf83954161d1ed6131161a890350089b313b22

Observation ad6e4414-6f95-47fe-a77c-c24e55c5d805 · outbound

This paper cites Scalable diffusion mod- els with transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Scalable diffusion mod- els with transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.443835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.692949Z digest=sha256:de8a2aeb0d944870a005ddeec02d82186eb4651a2ae4f1d63947e093f6ebcdbf

Observation 57d5776b-bb05-4dcc-a1c5-a1ef3a6604e4 · outbound

This paper cites Searching for Activation Functions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Searching for Activation Functions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.696802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.696802Z digest=sha256:bb7acb51b49331d50ca274cdf7a291039c72452e708e9fd0d8bede23b1efbbc3

Observation 5aa3a3e9-6158-4f0f-9702-dac7df8b7070 · outbound

This paper cites Gen- erating diverse high-fidelity images with VQ-V AE-2.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Gen- erating diverse high-fidelity images with VQ-V AE-2

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.432784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.701691Z digest=sha256:17b9ba875c505825b335cb491c702c88853af7f16e265144bfc2061d18a87f34

Observation f2886f5a-f047-40c4-af65-c37818a9c56c · outbound

This paper cites an unresolved cited work.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:00:32.421758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.705694Z digest=sha256:3afab06ddf24b2bf78b7079ffc456e920ee7e26fb0d0ab160c5e15560250bf35

Observation a18308e8-319a-4fe1-b3e3-757968f84b19 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction High-resolution image syn- thesis with latent diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.409639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.709569Z digest=sha256:180496c266b93c4a0f339e794b9f7236a2692f07c46d8252727b2a22486456ac

Observation 3d34d4a6-566b-48cd-a0bb-1362fd2f2706 · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Very deep convolutional networks for large-scale image recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.396543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.713764Z digest=sha256:c1e484e085335a8bf09ce943910336a6204987d51ece978cf710dd62c513b01d

Observation 91c6350e-d218-4ee8-9233-7f2c3d88a18b · outbound

This paper cites Make-a-video: Text-to-video genera- tion without text-video data.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Make-a-video: Text-to-video genera- tion without text-video data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.384816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.717461Z digest=sha256:15bc6b8f39198c0545863e53dfef46f2f63a627080823d5ad9ba36a75120405b

Observation f3185d23-d447-4bdf-a296-083fe22438fb · outbound

This paper cites Implicit neural representa- tions with periodic activation functions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Implicit neural representa- tions with periodic activation functions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.373542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.721583Z digest=sha256:b835471c8269301f568639b05e41274ff5cef0e75c3f9ab36f276ed38e39e7a5

Observation 13444ad6-3774-42a8-9510-6b8b174bca62 · outbound

This paper cites StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.362332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.725686Z digest=sha256:2de40721eee0e4a9c70e75e7872ca5d7493c7e9903a620133a8741f4a2aef641

Observation 9fff845b-dfb7-496c-b99e-717e285337d9 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Deep unsupervised learning using nonequilibrium thermodynamics

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.350579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.729635Z digest=sha256:f45ae720a64fc9d96bf0cca540056290cdb18155b4b4219b224193aedd7ded40

Observation 1f597a3a-b91d-4043-a645-2e654262f980 · outbound

This paper cites Denois- ing diffusion implicit models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Denois- ing diffusion implicit models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.733649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.733649Z digest=sha256:a628b69913f8a188508faa1e1f19d4de20d8e42be1ebc28c8079308a256a7f62

Observation d5c6504b-2b2a-4f1e-be6f-544205df4572 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.737759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.737759Z digest=sha256:38f70af0ba212a0e011d30ba814edba3d55a797d2846e961f808d4ecf52f7850

Observation 4d800118-b594-4f40-a788-3b185c273d1f · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Overview of the high efficiency video coding (hevc) standard

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.327288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.742133Z digest=sha256:09e4106b49328968d872c8283652ca29c37e50de233793cfede84914d77f642a

Observation 3ef89911-018b-4328-8a64-e81a56f45be6 · outbound

This paper cites High efficiency video coding (hevc).

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction High efficiency video coding (hevc)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.313267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.746289Z digest=sha256:2477ce8d4b931ade6a0debf1b376afc1ebe49417846783bd8900c8959965821f

Observation cb39484a-e85c-428f-8d84-f58fea5c2a79 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.750358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.750358Z digest=sha256:9bba0f585fefec749f1595b5e9cf55474bfe28982db7315f470f952c1ec94ecf

Observation 0215fffb-bf82-4712-85b6-80e83bcdab26 · outbound

This paper cites A good image generator is what you need for high-resolution video synthe- sis.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction A good image generator is what you need for high-resolution video synthe- sis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.300153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.755065Z digest=sha256:9d5e786b8cea73e5fca02e9317340b113123e5357dbac41081d69e692965b68e

Observation 7fb30c56-3d25-4d54-9f8f-83e725facdc9 · outbound

This paper cites MoCoGAN: Decomposing motion and content for video generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction MoCoGAN: Decomposing motion and content for video generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.286251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.759054Z digest=sha256:6dc7bbdb3c4766d0d375cf5d533f48fbecb33bb9031976b67517bb189eae783a

Observation 9097d2a1-1571-4097-897d-24e5046186b7 · outbound

This paper cites FVD: A new metric for video generation, 2019.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction FVD: A new metric for video generation, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.274143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.763337Z digest=sha256:707fe2a63d43732b3a2377b65f2e4a20ec615136ad4564139bb9643702f9dfbc

Observation 4d08fda6-9913-4f49-bf22-d1030e6af355 · outbound

This paper cites Neural discrete representation learning.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Neural discrete representation learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.260738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.767489Z digest=sha256:be482096e8417cd03e8ac7c1a118e349ffd105fd63c4fec036aaf89503f45363

Observation 77a92c12-139b-4592-98c1-d859db907317 · outbound

This paper cites Attention is all you need.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.771230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.771230Z digest=sha256:bbffe1e8a53be97c33245553b6db04a8d55f9daa93db36b7905c7d6d8679b285

Observation 1e750476-372d-4cb5-a05b-aff5c483d654 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Phenaki: Variable length video generation from open domain textual descriptions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.239330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.775210Z digest=sha256:64787f53c65f185fa12a07c4c634ad1a06a7936e4f8f395b3301da5d2bb98e8d

Observation 5e73021d-f44d-4ac8-8af3-0535260537c7 · outbound

This paper cites LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.779652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.779652Z digest=sha256:1c8a3ae01921650f058e6deeb0ef0a997144ebee06d18877dd2fca55285884d2

Observation f7fc426f-fe45-4413-af71-f7bb0955850c · outbound

This paper cites OmniTokenizer: A joint image- video tokenizer for visual generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction OmniTokenizer: A joint image- video tokenizer for visual generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.225131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.784395Z digest=sha256:801f5004d189f406345b58456c523eab194bfa7ba59750375531e11558b66bda

Observation 2d9ba59e-87c2-4426-b397-bd9c44f16421 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Emu3: Next-Token Prediction is All You Need

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.788200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.788200Z digest=sha256:24e3e7615cde3e32c5a26c46e77e2af39a15a8025c5473db956f4ae6473365ec

Observation 62662224-3a6e-4b82-bace-80303f622e74 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Image quality assessment: from error visibility to structural similarity

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.793683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.793683Z digest=sha256:e13bec2d4e9e4f4eb01ed5adc6b917ba09a2474cefd29e669cf226d04716257b

Observation f9676269-962f-46f7-aaa5-e263ee516df8 · outbound

This paper cites Overview of the h.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Overview of the h

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.201195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.798745Z digest=sha256:c2e04c65e274c2b1c88f4601a83a1e6bf7df0ab9002d4013138c7faea077d2bb

Observation 8f20aa74-0595-41db-8739-232e8b4a8c1e · outbound

This paper cites Group normalization.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Group normalization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.181726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.803530Z digest=sha256:092431696a4feccaebb115254521b12c324d825348abd9cfcc575b3df5a12f90

Observation 0a4afa8b-aee9-4585-8e79-e699a29ca685 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.807670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.807670Z digest=sha256:4c6c3c0b86b6965c0eb334c050e67befe553372bf8806ea0532d204bdb67f2d0

Observation d2082f80-eb7f-4a40-9876-f6cc0be50748 · outbound

This paper cites Temporally consistent transformers for video gen- eration.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Temporally consistent transformers for video gen- eration

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.168884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.811547Z digest=sha256:fb24b573528e7a0bf717b2a5cd6bf16eae21cde723a979b0d9ca28b2f5e025b5

Observation 47d1455c-4771-4e51-b3fa-d1ece7871991 · outbound

This paper cites ElasticTok: Adaptive Tokenization for Image and Video.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction ElasticTok: Adaptive Tokenization for Image and Video

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.815756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.815756Z digest=sha256:d84a792f4aadeb8c6cc45eb5b58aada7823df57578d6d448129787744964af8a

Observation 4992d67d-bb4f-46f8-b665-d215be3bcc53 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.819677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.819677Z digest=sha256:dcc47e57de71a6dd12cb2ed0b7bed5ab6ddeae43c7399a8a5e100ba397af07d0

Observation 1dd3f11c-d7c5-4685-83c3-505f9eec3555 · outbound

This paper cites Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.155536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.823951Z digest=sha256:120fa8ea916370f58903203e3a87d046ef0e839cfecabe696152331f18b9eaad

Observation 165bd51e-bf33-4a2c-ae8b-8b7a02a5f4fe · outbound

This paper cites Magvit: Masked generative video transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Magvit: Masked generative video transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.142567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.828818Z digest=sha256:b47855bb29ede768276ba8588f3ec8edb288cb25b2b0c8a1329d117aa6466a25

Observation 63aef45f-b108-4b22-9938-1df582be05ac · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Language model beats diffusion–tokenizer is key to visual generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.129370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.832494Z digest=sha256:e4bb9b610ad6ddbdc1039df9d9985d3d7713cdd11df6bff12c1e0adb76ca8066

Observation 610180e5-aa84-4e69-8a43-d941f533ed4d · outbound

This paper cites Generating videos with dynamics-aware implicit generative adversarial net- works.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Generating videos with dynamics-aware implicit generative adversarial net- works

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.116479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.836383Z digest=sha256:1af4db453334b74fa5c5a687d2a6c04910225ae675c49eb770415d4c7ba37a03

Observation acdafe3f-e627-4e2f-8b1e-01db5e71de82 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Video probabilistic diffusion models in projected latent space

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.102853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.839912Z digest=sha256:cc583679e48ed66f8c996dc955ff3259605a540fcde7c855ad0e3009c085b5ff

Observation b29d40b4-e2b6-46a0-89e3-48e5402a8ac6 · outbound

This paper cites Efficient video diffusion models via content-frame motion-latent decomposition.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Efficient video diffusion models via content-frame motion-latent decomposition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.090647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.843624Z digest=sha256:81753f49318b2153a64681dab82e5278b52ff498cad933cca0f968b405c48a95

Observation a37244c2-9d80-4882-ab5b-a1cc90c788c6 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction The unreasonable effectiveness of deep features as a perceptual metric

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.076180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.848252Z digest=sha256:7ea69facf97319b2740a313dea6d7b59583bc303fc1490d8a3b37b26e9fd19d2

Observation 242ebf3b-1c68-4012-a82f-02b04a5cc823 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.851805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.851805Z digest=sha256:2ea243d80b0cc4631a87aee97c15f2e8a0e40d46b901dbbb8dacedd96189fa48

Observation 6b27ceea-0e36-440c-ac8c-cd605693df84 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.855686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.855686Z digest=sha256:35b8cf5bc4cb37f5f89e08c7ff11a84d7079b081eb19421a0c93a65cd1e4ec0c

Observation 9edbdde8-17b5-41f8-8473-b1604ffba88c · outbound

This paper cites Architecture We use the same structure as SiT, except that our patch embedding and final projection layers are im- plemented separately for each plane.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Architecture We use the same structure as SiT, except that our patch embedding and final projection layers are im- plemented separately for each plane

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.061934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.859918Z digest=sha256:92c7bb51d33001c256c86ed10ea3b9afebb0feaa6aacbe1702393307cc80771d

Observation 4d1fa4bd-360f-484c-8f47-e2142ff03862 · outbound

This paper cites an unresolved cited work.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:00:32.047444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:00:31.863822Z digest=sha256:35add746814ab4aa2372bcf4175d34fa27c3b88d9e04f6920ad360e39dca0d59

Pith citing papers

Observation 70bc0a3e-e3ef-4560-86b2-de67a24c6e60 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.993969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:69a59c0e7e3eb1dde8456d31358f017bca87777e584ff5732fe03ca0dc16706e