Pith. sign in

Paper Citation Record · LEDGER

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

As of 14 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2411.14762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14762 v4

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:00:31.863822Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T08:52:31.686474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:52:31.987543Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fdcf8ded-eaa4-48f4-ac5e-3ef9afc0e199 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.797094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.572869Z digest=sha256:12ae4f0b620da5bf7130e893d6f230ffdb1a5daf4160cc2ac9de725cdd0fc8eb

Observation f75b73ef-49af-4791-8269-8e9a2363f1b3 · outbound

This paper cites Video generation models as world simulators.Ope- nAI Blog, 2024.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Video generation models as world simulators.Ope- nAI Blog, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.783433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.577481Z digest=sha256:25ce5c12f457045ea4255b8afe2a442822604c0d739a355fd37dec289d877bad

Observation eb1be525-3e00-4e47-9835-08ef5b59f461 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Quo vadis, action recognition? a new model and the kinetics dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.769593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.581261Z digest=sha256:15ad23609f4cf056fe67a1a11ba4a6431a686d6592b51fb32ef7b0aafac41cdb

Observation 3d82438e-94fb-4764-95b3-9d558cdb7a53 · outbound

This paper cites Efficient geometry-aware 3d generative adversarial networks.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Efficient geometry-aware 3d generative adversarial networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.754149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.585350Z digest=sha256:6f2bf0ae87d30c36fdb8ef22a0bcf74a5807dc44a796076e538d6989d52966b3

Observation aa3b9a8d-a7bd-45b2-a84e-1161229179ad · outbound

This paper cites MaskGIT: Masked generative image transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction MaskGIT: Masked generative image transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.739628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.588957Z digest=sha256:3c8c2465dc71e91a52ffe786399cc7c3bef7c0e2d9b330a5d439d00940685add

Observation 5e026116-80e1-4e49-bc7d-a2b52585c019 · outbound

This paper cites VideoINR: Learning video implicit neural represen- tation for continuous space-time super-resolution.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction VideoINR: Learning video implicit neural represen- tation for continuous space-time super-resolution

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.726521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.592709Z digest=sha256:d3554d96278aedfef7b3af7f8a54ef78758ef5cd2935575cf730b375262b6732

Observation eb7fb135-130b-4adc-8f15-7e81c1d3dfcc · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.713298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.596493Z digest=sha256:7fcf2476b97ce0569894f72b1403d79a6e035ed5b7f1bbcd02a980661db3ec77

Observation e3104583-958f-464b-8f19-53c50cea674e · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Taming transformers for high-resolution image synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.698626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.600143Z digest=sha256:dba0a76a5dfaa595a56e7a1c11c2d8d9a31832542701f492b7e8af59ed0427a8

Observation f3793d6f-e1eb-4587-aa6c-6d9df4e143f2 · outbound

This paper cites Cosmos world foundation model platform for physical ai.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Cosmos world foundation model platform for physical ai

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.683996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.603912Z digest=sha256:afb06094bb7264a9bd37d0c379616c5ff7e9335559b0828643b7f39891e5a9ca

Observation 6ab4570e-0667-49a7-a293-7d0f1f07e86f · outbound

This paper cites Latte: Latent diffusion transformer for video generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Latte: Latent diffusion transformer for video generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.670460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.607393Z digest=sha256:7017d81e2c1e1c24fcfa04ea472257f5ffecb6f933015a71f76ec7b4d5626d52

Observation 34bcc14b-7ae8-4cd8-8814-30d0d0ec37d6 · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Long video generation with time-agnostic vqgan and time- sensitive transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.658379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.611475Z digest=sha256:3bc0901fa114109366766a2b21a0eca2059d1af78fe0c5bb172e5fe5862f2262

Observation 97a2ac37-c453-4817-861f-2f0f12886975 · outbound

This paper cites Generative adversarial nets.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Generative adversarial nets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.615363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.615363Z digest=sha256:12563ee13e5d3a51a6ced59103d7d79b27915e525205b2dc4ea5997ce3af8aad

Observation 23499501-afd2-4824-a499-3965b05ee614 · outbound

This paper cites Photorealistic video generation with diffusion models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Photorealistic video generation with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.638853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.619377Z digest=sha256:ffd010803078a851502da84825ec0dca3aa1205d3d4cdbb6ae5cac55603be662

Observation 4db7a237-a658-48e2-87a1-e618b108fc0c · outbound

This paper cites A technical overview of av1.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction A technical overview of av1

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.626160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.623316Z digest=sha256:a30b61a1c9db115b9044c5c25de0b5b182df3a316541e681f16c1f7c121b6674

Observation 9b64092f-be6f-4740-991c-a871edfe2b6f · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.627658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.627658Z digest=sha256:b53f281e0bacd3443c7f97d06db3d5024c668294a8c1fd51947fda4d015a27cc

Observation 7986910f-9192-4604-b608-ee87e563ef7d · outbound

This paper cites Denoising dif- fusion probabilistic models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Denoising dif- fusion probabilistic models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.613170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.632169Z digest=sha256:219aa2d16816594a76558af7fbf0a6a719226b724b05239ab28b9d8743969813

Observation 06fba8ce-8705-4ce4-aa4e-c69e33e6896f · outbound

This paper cites CogVideo: Large-scale pretraining for text-to- video generation via transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction CogVideo: Large-scale pretraining for text-to- video generation via transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.600180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.635831Z digest=sha256:03131d3e255f602d2d2fb9d83dee3dd58c4a445514c046759adea206357613d3

Observation b89f7285-0f2e-4717-92a7-7e192ee4b81c · outbound

This paper cites LRM: Large reconstruction model for single image to 3D.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction LRM: Large reconstruction model for single image to 3D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.587947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.639953Z digest=sha256:d9829b90057f56a1d4266538c90a24077b802924691fa156e81260aa711dc89e

Observation ccc72d19-cf63-4fcd-a704-37c7873e9d70 · outbound

This paper cites Shap-E: Generating Conditional 3D Implicit Functions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Shap-E: Generating Conditional 3D Implicit Functions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.643821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.643821Z digest=sha256:aa776f8c58689ca247f76045f84712288691bca387f31a772f300f813ba24c8f

Observation 119ccc18-5727-4a3a-8897-8997a5ba4761 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Analyzing and improving the image quality of stylegan

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.575638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.648125Z digest=sha256:d05c74c271b065ddb3afe2e402d3adaab994943593fb42add6a86833474bc369

Observation 2583c661-2372-4d60-87fe-c2a7f33577af · outbound

This paper cites Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.564535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.651927Z digest=sha256:f444359fe1c3d652b0794b5fa5ff55e9723c6001913d67c01afa2e051bfe6e23

Observation 9bcbdfec-f263-4587-879f-00f1ef29cf94 · outbound

This paper cites Scal- able neural video representations with learnable positional features.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Scal- able neural video representations with learnable positional features

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.552140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.656101Z digest=sha256:6e5ad3ff81578142e34f507c40a76a6ad9b5a57e1e6861a7a40bdbb3058889ed

Observation b01e019c-8d81-4a69-91fb-c4a39d026ed2 · outbound

This paper cites 9 Videopoet: A large language model for zero-shot video gen- eration.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction 9 Videopoet: A large language model for zero-shot video gen- eration

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.540537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.660065Z digest=sha256:b463cf4223c0072fa20e8a3da486f8d3246462ea292b3a0e27a7bc5107f22c3a

Observation 89bcc7af-9278-4274-9cf9-8bb033ec06c1 · outbound

This paper cites One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.527621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.663773Z digest=sha256:3889b2c0009c9d22708974603c72a7d6b7096a0782eb889a4c5927042310910d

Observation 25a2cb62-04f1-4d72-9ff9-3a1e82af3650 · outbound

This paper cites Decoupled weight de- cay regularization.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Decoupled weight de- cay regularization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.513140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.667725Z digest=sha256:6d4b4895be52ecec0c526477330749886210bd55ab6b317d6cd33054fdb02590

Observation 4b571676-00d6-4ce0-a8ca-29d01e70efc0 · outbound

This paper cites Vdt: General-purpose video diffusion transformers via mask modeling.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Vdt: General-purpose video diffusion transformers via mask modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.500891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.672238Z digest=sha256:dd7dca8e646e96fec5f96ecd733fcc95d2bad76910e09369c12dd14bcc1b7226

Observation bc0fb863-4141-40b1-a09b-09d638cb210f · outbound

This paper cites SiT: Explor- ing flow and diffusion-based generative models with scalable interpolant transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction SiT: Explor- ing flow and diffusion-based generative models with scalable interpolant transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.489274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.676893Z digest=sha256:4fe6ce5313809adc93ca18300f2fa5a48df2b7c73177b269dc7e1e16856ac267

Observation d457ce73-e466-4c2e-8381-becbc9a8c05f · outbound

This paper cites an unresolved cited work.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:00:32.477424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.681174Z digest=sha256:393434fb8971236811f47b3cee3921d33f2789f4936015cf97d3acf4ccbe7556

Observation 67b4eefb-c27d-444d-9096-c39431ff008a · outbound

This paper cites GTA: A geometry-aware attention mechanism for multi-view transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction GTA: A geometry-aware attention mechanism for multi-view transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.466216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.685014Z digest=sha256:ab3ab770f185ed67cb9b95a6212b179ed5a3a7cf02540d0459edd49fd21a034f

Observation 1dbf2046-3ea5-4a77-93b0-f49b07b4714c · outbound

This paper cites A technical overview of vp9—the latest open- source video codec.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction A technical overview of vp9—the latest open- source video codec

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.455366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.689257Z digest=sha256:a684ea2e218ae763c0ca6ddf93f7e360c53454781063c6c678ac2f4d7ba8ee47

Observation ad6e4414-6f95-47fe-a77c-c24e55c5d805 · outbound

This paper cites Scalable diffusion mod- els with transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Scalable diffusion mod- els with transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.443835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.692949Z digest=sha256:614e9deba00257f24d3549089537caa7069edc0e9281560b87c79783103d817e

Observation 57d5776b-bb05-4dcc-a1c5-a1ef3a6604e4 · outbound

This paper cites Searching for Activation Functions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Searching for Activation Functions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.696802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.696802Z digest=sha256:bb7acb51b49331d50ca274cdf7a291039c72452e708e9fd0d8bede23b1efbbc3

Observation 5aa3a3e9-6158-4f0f-9702-dac7df8b7070 · outbound

This paper cites Gen- erating diverse high-fidelity images with VQ-V AE-2.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Gen- erating diverse high-fidelity images with VQ-V AE-2

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.432784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.701691Z digest=sha256:9d814915dd1a69b139631a403cb117ce69f2d0376344e1942bf1b415fd3b253c

Observation f2886f5a-f047-40c4-af65-c37818a9c56c · outbound

This paper cites an unresolved cited work.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:00:32.421758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.705694Z digest=sha256:2888b45c6af93202fc6e8dff5476bd0bd087dad7af72991896e1abbcd8ea5065

Observation a18308e8-319a-4fe1-b3e3-757968f84b19 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction High-resolution image syn- thesis with latent diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.409639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.709569Z digest=sha256:64ed7214985e3a06befd6b808b1a6637f1d1b77fddaca79d6daac0c94ae1cdec

Observation 3d34d4a6-566b-48cd-a0bb-1362fd2f2706 · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Very deep convolutional networks for large-scale image recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.396543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.713764Z digest=sha256:245d101587e98d94a66562a3fdf67d7aa1f96827b2b7a6c0499fec477cb6ce66

Observation 91c6350e-d218-4ee8-9233-7f2c3d88a18b · outbound

This paper cites Make-a-video: Text-to-video genera- tion without text-video data.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Make-a-video: Text-to-video genera- tion without text-video data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.384816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.717461Z digest=sha256:ff76cf3be8fe13796b57ff10e712c9d75c5576fe37d9a14332b2bbda62de2219

Observation f3185d23-d447-4bdf-a296-083fe22438fb · outbound

This paper cites Implicit neural representa- tions with periodic activation functions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Implicit neural representa- tions with periodic activation functions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.373542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.721583Z digest=sha256:ed5cb95f4db33ac49ab965bc8f50c8b974f2cf5fd990a6230edefa1cadf0c1a5

Observation 13444ad6-3774-42a8-9510-6b8b174bca62 · outbound

This paper cites StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.362332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.725686Z digest=sha256:54f7655bb34b1de80f0fa60682eb5fad77cda6727a65a5e652321211cfd7d203

Observation 9fff845b-dfb7-496c-b99e-717e285337d9 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Deep unsupervised learning using nonequilibrium thermodynamics

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.350579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.729635Z digest=sha256:c566116dcd335b102f459c4ca8921585bc6bbc1f9da991cfba182d4aa4694633

Observation 1f597a3a-b91d-4043-a645-2e654262f980 · outbound

This paper cites Denois- ing diffusion implicit models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Denois- ing diffusion implicit models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.733649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.733649Z digest=sha256:a628b69913f8a188508faa1e1f19d4de20d8e42be1ebc28c8079308a256a7f62

Observation d5c6504b-2b2a-4f1e-be6f-544205df4572 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.737759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.737759Z digest=sha256:38f70af0ba212a0e011d30ba814edba3d55a797d2846e961f808d4ecf52f7850

Observation 4d800118-b594-4f40-a788-3b185c273d1f · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Overview of the high efficiency video coding (hevc) standard

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.327288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.742133Z digest=sha256:fe1ad8079a4018d5ce6ebde5ed9700fd0ac816db1bcda4494f95ea43087ab444

Observation 3ef89911-018b-4328-8a64-e81a56f45be6 · outbound

This paper cites High efficiency video coding (hevc).

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction High efficiency video coding (hevc)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.313267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.746289Z digest=sha256:c12e384d998a54c499a4d79d20af71ec2e2e5204d88addcb893d225887852f11

Observation cb39484a-e85c-428f-8d84-f58fea5c2a79 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.750358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.750358Z digest=sha256:9bba0f585fefec749f1595b5e9cf55474bfe28982db7315f470f952c1ec94ecf

Observation 0215fffb-bf82-4712-85b6-80e83bcdab26 · outbound

This paper cites A good image generator is what you need for high-resolution video synthe- sis.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction A good image generator is what you need for high-resolution video synthe- sis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.300153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.755065Z digest=sha256:91bb2ccdbcd7581e4827ab1d4e8d3c5f82624bff8141944ff15ad53b69c2b054

Observation 7fb30c56-3d25-4d54-9f8f-83e725facdc9 · outbound

This paper cites MoCoGAN: Decomposing motion and content for video generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction MoCoGAN: Decomposing motion and content for video generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.286251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.759054Z digest=sha256:46042293b9b2db85031a8e2d5343a9899e8ffb961f36fde72edba2a25c986b5a

Observation 9097d2a1-1571-4097-897d-24e5046186b7 · outbound

This paper cites FVD: A new metric for video generation, 2019.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction FVD: A new metric for video generation, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.274143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.763337Z digest=sha256:93c9b23fbb5c3448a059f8c94ae1be37baf5513f0de98e0a16f21c774a10e124

Observation 4d08fda6-9913-4f49-bf22-d1030e6af355 · outbound

This paper cites Neural discrete representation learning.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Neural discrete representation learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.260738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.767489Z digest=sha256:3ce97a8e472b2d139ed207df62358a506304140d063d244388000fbe21393ca3

Observation 77a92c12-139b-4592-98c1-d859db907317 · outbound

This paper cites Attention is all you need.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.771230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.771230Z digest=sha256:bbffe1e8a53be97c33245553b6db04a8d55f9daa93db36b7905c7d6d8679b285

Observation 1e750476-372d-4cb5-a05b-aff5c483d654 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Phenaki: Variable length video generation from open domain textual descriptions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.239330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.775210Z digest=sha256:cbe4d51f5f2879a1896b72d0afaf170a8931d6e5ad0987ad6ef14e58657bb8a8

Observation 5e73021d-f44d-4ac8-8af3-0535260537c7 · outbound

This paper cites LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.779652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.779652Z digest=sha256:1c8a3ae01921650f058e6deeb0ef0a997144ebee06d18877dd2fca55285884d2

Observation f7fc426f-fe45-4413-af71-f7bb0955850c · outbound

This paper cites OmniTokenizer: A joint image- video tokenizer for visual generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction OmniTokenizer: A joint image- video tokenizer for visual generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.225131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.784395Z digest=sha256:083f62c2b8b6232ca4f7daad652d85777968e759eda44c7f867864fce483fe65

Observation 2d9ba59e-87c2-4426-b397-bd9c44f16421 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Emu3: Next-Token Prediction is All You Need

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.788200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.788200Z digest=sha256:24e3e7615cde3e32c5a26c46e77e2af39a15a8025c5473db956f4ae6473365ec

Observation 62662224-3a6e-4b82-bace-80303f622e74 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Image quality assessment: from error visibility to structural similarity

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.793683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.793683Z digest=sha256:e13bec2d4e9e4f4eb01ed5adc6b917ba09a2474cefd29e669cf226d04716257b

Observation f9676269-962f-46f7-aaa5-e263ee516df8 · outbound

This paper cites Overview of the h.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Overview of the h

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.201195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.798745Z digest=sha256:312d9360584303af8d7b0d381ca3bb3d404afa649aae1a3eeebadbdaddea8e8e

Observation 8f20aa74-0595-41db-8739-232e8b4a8c1e · outbound

This paper cites Group normalization.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Group normalization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.181726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.803530Z digest=sha256:d49f0c0826275aa1e6734b81c816508c9af38ab1c1dd827bda752610a721fe69

Observation 0a4afa8b-aee9-4585-8e79-e699a29ca685 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.807670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.807670Z digest=sha256:4c6c3c0b86b6965c0eb334c050e67befe553372bf8806ea0532d204bdb67f2d0

Observation d2082f80-eb7f-4a40-9876-f6cc0be50748 · outbound

This paper cites Temporally consistent transformers for video gen- eration.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Temporally consistent transformers for video gen- eration

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.168884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.811547Z digest=sha256:499480ce234e5db113eeef1a2e6c1c40c4286358f50547ac97cc150187349eb8

Observation 47d1455c-4771-4e51-b3fa-d1ece7871991 · outbound

This paper cites ElasticTok: Adaptive Tokenization for Image and Video.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction ElasticTok: Adaptive Tokenization for Image and Video

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.815756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.815756Z digest=sha256:d84a792f4aadeb8c6cc45eb5b58aada7823df57578d6d448129787744964af8a

Observation 4992d67d-bb4f-46f8-b665-d215be3bcc53 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.819677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.819677Z digest=sha256:dcc47e57de71a6dd12cb2ed0b7bed5ab6ddeae43c7399a8a5e100ba397af07d0

Observation 1dd3f11c-d7c5-4685-83c3-505f9eec3555 · outbound

This paper cites Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.155536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.823951Z digest=sha256:89fbf5c3ae96a0399bdf384d38305481c5f4a21a5580ede67fe1c56db703b7fc

Observation 165bd51e-bf33-4a2c-ae8b-8b7a02a5f4fe · outbound

This paper cites Magvit: Masked generative video transformer.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Magvit: Masked generative video transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.142567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.828818Z digest=sha256:c8e3ffb58321aefb449b0a6f802ffe8d79272de3ef149aa930355cf7938c45e6

Observation 63aef45f-b108-4b22-9938-1df582be05ac · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Language model beats diffusion–tokenizer is key to visual generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.129370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.832494Z digest=sha256:ddfe3d7290cd34310fcdde4942481a9142d63833544ee815cd3378dd914fab05

Observation 610180e5-aa84-4e69-8a43-d941f533ed4d · outbound

This paper cites Generating videos with dynamics-aware implicit generative adversarial net- works.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Generating videos with dynamics-aware implicit generative adversarial net- works

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.116479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.836383Z digest=sha256:897f936a40c516809e8d299c01523724e81fa0582ca38b29a047715c596edc81

Observation acdafe3f-e627-4e2f-8b1e-01db5e71de82 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Video probabilistic diffusion models in projected latent space

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.102853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.839912Z digest=sha256:a87d21e2bdc10bf8ca9769ec1cd25dd04c4ea53034b48254f98d9c5dd386f948

Observation b29d40b4-e2b6-46a0-89e3-48e5402a8ac6 · outbound

This paper cites Efficient video diffusion models via content-frame motion-latent decomposition.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Efficient video diffusion models via content-frame motion-latent decomposition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.090647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.843624Z digest=sha256:960fa13f8526b5be59da0caf676b515593bd00e6793165b736cbc2e6f32c3d54

Observation a37244c2-9d80-4882-ab5b-a1cc90c788c6 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction The unreasonable effectiveness of deep features as a perceptual metric

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.076180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.848252Z digest=sha256:a42f4776c6233581ef8ca766dbe70640c9342ae0875313a7aa5d04101370c1fc

Observation 242ebf3b-1c68-4012-a82f-02b04a5cc823 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.851805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.851805Z digest=sha256:2ea243d80b0cc4631a87aee97c15f2e8a0e40d46b901dbbb8dacedd96189fa48

Observation 6b27ceea-0e36-440c-ac8c-cd605693df84 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:00:31.855686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:00:31.855686Z digest=sha256:35b8cf5bc4cb37f5f89e08c7ff11a84d7079b081eb19421a0c93a65cd1e4ec0c

Observation 9edbdde8-17b5-41f8-8473-b1604ffba88c · outbound

This paper cites Architecture We use the same structure as SiT, except that our patch embedding and final projection layers are im- plemented separately for each plane.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Architecture We use the same structure as SiT, except that our patch embedding and final projection layers are im- plemented separately for each plane

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:00:32.061934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.859918Z digest=sha256:43f81b0d02c72580c14cc477fcb4fb685878b565a44c64172e973e34f55af296

Observation 4d1fa4bd-360f-484c-8f47-e2142ff03862 · outbound

This paper cites an unresolved cited work.

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:00:32.047444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:00:31.863822Z digest=sha256:dac03c5e35c096547fa8531eceb64694d671f1c22b55598642a4dfc0f743ddd3

Pith citing papers

Observation 70bc0a3e-e3ef-4560-86b2-de67a24c6e60 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.993969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:ce370be582aa3def834b3c01c7a1e4c8a41d191aab9efead1b802431c1459a74