Pith. sign in

Paper Citation Record · LEDGER

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization

As of 14 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2412.10443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10443 v3

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:59:36.794196Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:20:32.508409Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:28:55.532932Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f634ee9b-78dd-4f3c-ab01-fd27cce0af3a · outbound

This paper cites Qwen Technical Report.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.344044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.344044Z digest=sha256:922f87b911f11832ebb6597d42007f93e4efc4aaab14cb3720d72a9a633e07d1

Observation 71491f69-a707-4d05-8975-ad2eaf8a759c · outbound

This paper cites Language Models are Few-Shot Learners.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language Models are Few-Shot Learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.350357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.350357Z digest=sha256:f8039832efbb8b337a6c3c50230d73722a966d23182b793f35da53b90ca3c14d

Observation 6a4497d6-f0d1-49da-bde6-b8d708fc57b0 · outbound

This paper cites End- to-end object detection with transformers.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization End- to-end object detection with transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.358325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.358325Z digest=sha256:ebd02860dbef8719924753166f4b3bae4f59ff5fc225040d80ec9af4a5f26bb6

Observation 6a715bda-cc16-4c5d-979d-cea8fd1f9c25 · outbound

This paper cites A short note about kinetics-.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization A short note about kinetics-

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.364476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.364476Z digest=sha256:c9254ff3bd2e3de7ee27a83701e4a30841a2e673cc4cc29dac05c828fde1a35f

Observation a57b8f75-1b49-4082-8570-7131645bab46 · outbound

This paper cites Maskgit: Masked generative image transformer.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Maskgit: Masked generative image transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:38.169986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.380212Z digest=sha256:bd0ef72e13380fd9a157e43f1658fa80ac270f86f17c44d164c19c1bab94767d

Observation d9edbc7e-0e6f-4054-8ba4-bd1424e111f4 · outbound

This paper cites Palm: Scaling language modeling with pathways.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Palm: Scaling language modeling with pathways

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:38.146366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.386691Z digest=sha256:fb3432faaa2ed96a1cf571badefde661ec735052b4eb3f6dfe310ccc6a278304

Observation 8cdea5bc-c0aa-47ac-83e8-53e8f13526d0 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Imagenet: A large-scale hierarchical image database

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:38.119413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.394212Z digest=sha256:83d4bc43d2a4c63b0e446fed0596930af1e7336ff6fe72a58fbeb57ada70a09c

Observation 27031164-1ba2-4614-9208-8eef7c0f54c2 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:38.098205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.402336Z digest=sha256:ef782c2ea1cc4bd4f71839100dc751d77bb917617851c79e847a76906f4e21eb

Observation 0ca42ab5-0596-4bf9-8331-7643f4587299 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:38.075974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.410230Z digest=sha256:5db4b507b4cc6f2b6b4f9ec43d6cc7e58cd600a019c78576a3169603091c6a04

Observation d8e1eb6c-ca63-41a9-9d10-d1d3cabc4ca7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Taming transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.416735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.416735Z digest=sha256:b5492cebe9df0cdf285d453c05e3fe5e262fc81093115f4e8b03fd5818cdd51a

Observation 1076691a-c275-4857-8126-7412d4262aee · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Long video generation with time-agnostic vqgan and time- sensitive transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:38.038739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.423291Z digest=sha256:afdf8be4ea3718a57841b10d6f8d29b9dcc7435d7daca47cf0a6f53be558e377

Observation 51dacad5-eb1a-47a6-b833-7650d0da9238 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.429338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.429338Z digest=sha256:f9aa5d93d8af2acbb2b65f00e2370dc2a3790dde82460b4c9d39dc2604d250ec

Observation 5a81665d-89ab-4346-b607-3b6f6aa400af · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.437373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.437373Z digest=sha256:1b25bebc45e4a28b6481a2df42e4d1555c73f8bf96cd36e12eab19d6af8c5c27

Observation 5f623e22-7c16-48dd-b4f7-228542323d09 · outbound

This paper cites Video-lavit: Unified video- language pre-training with decoupled visual-motional tok- enization.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Video-lavit: Unified video- language pre-training with decoupled visual-motional tok- enization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.998333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.457304Z digest=sha256:90d5f3cd3b70dd8ed6059a32f8bbf6a4bf675bd014761705d2b74b070a9266a1

Observation f4b85abb-128f-48d2-9e92-63ce885fad3e · outbound

This paper cites Unified language-vision pre- training in llm with dynamic discrete visual tokenization.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Unified language-vision pre- training in llm with dynamic discrete visual tokenization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.974956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.465436Z digest=sha256:d24bd04a0eebe20f8d6cba618d2151e647a290356719431321a6cfd0932bead3

Observation e6641644-3148-462f-8aec-51579427366d · outbound

This paper cites The Kinetics Human Action Video Dataset.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization The Kinetics Human Action Video Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.471459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.471459Z digest=sha256:7cb03cadba819494457e620c5700b084a0f140c6cde47e4bbb9728749f1447a0

Observation 1be71f54-55f5-4b17-bfc5-424bfbe85832 · outbound

This paper cites Auto-encoding variational bayes.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Auto-encoding variational bayes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.949263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.486052Z digest=sha256:885e53ac1171c807dd03ce9968fe69ce7ef768af88ca075a2735bfcac77ab301

Observation 945b8942-8d6d-41ac-8fa1-8bca06121c9b · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.496799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.496799Z digest=sha256:f3570d54387f8bd27e4b9c6862d0364dbb3cd1da3f707696723e917d3c21cc0c

Observation b63f6f55-b613-41e1-8ae5-71a27ae45d90 · outbound

This paper cites Semi-supervised classifi- cation with graph convolutional networks.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Semi-supervised classifi- cation with graph convolutional networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.922814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.502676Z digest=sha256:61ac428cbd9dddfe05d4031f172e7f2b2d05e98f7debeeba6f19594a93d7f16e

Observation 6a136bf9-69fe-47a5-8ae4-3b95ef5f41ae · outbound

This paper cites Few Shot Activity Recognition Using Variational Inference.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Few Shot Activity Recognition Using Variational Inference

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:59:37.063800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.508594Z digest=sha256:9333bfc47764da9e5a33c8230eb673a25a5ce63c3dbf50fb3d93141f9fa3e641

Observation 968cd9ef-e880-467c-bb60-10da10c89a25 · outbound

This paper cites Autoregressive image generation using residual quantization.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Autoregressive image generation using residual quantization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.900772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.514927Z digest=sha256:2ca4f86ce017f0c6000cc2ebd6d8a3859bd00cc72f06256cf89a2407dcbdc172

Observation c4e0ab49-b0b7-41f1-8331-03dfe4465fc2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.881820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.522503Z digest=sha256:46b5631ca4c5b55c432ce4dc119245c0461240535095ba9e2af543ba88d61f54

Observation 211ad92b-f31d-4486-a9f4-25a9bf0f52aa · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Video-llava: Learning united visual representation by alignment before projection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.860558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.530575Z digest=sha256:611d39b72b54e14dad6839172943c461ec678bf18f0155bd123b6342d9b0b568

Observation e273481b-cba3-4d46-b71e-8a8c6a482911 · outbound

This paper cites Language quan- tized autoencoders: Towards unsupervised text-image align- ment.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language quan- tized autoencoders: Towards unsupervised text-image align- ment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.836268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.536630Z digest=sha256:591e7bee5bddfd3eb952d0b797984833421b778930da2bb3ea117c0f5276e0d6

Observation 5df510ed-3dba-488d-a51c-b6cb1d70606b · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.547616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.547616Z digest=sha256:c5b1049512db850ffe07da4dbe15aa9221dd544e38339ee510d1f15a4587806c

Observation 9f7c6b2b-26f6-4834-bc1f-318f6f1ccd48 · outbound

This paper cites Language models are unsu- pervised multitask learners.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language models are unsu- pervised multitask learners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.809123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.554655Z digest=sha256:38bf9313036682f38c75113a73eaae97d3eee15a2b02bc1b7224c1e74ff7613a

Observation 9a323101-e69d-48ba-98db-1db98b6dc1e1 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Learn- ing transferable visual models from natural language super- vision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.789162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.562189Z digest=sha256:c2a42234233d6bd0ddc94404825193f880c705475bea0cb73e2ee1e82020b9e3

Observation 0b53fd07-f0ef-46b4-a302-a3d17184de8d · outbound

This paper cites Gen- erating diverse high-fidelity images with vq-vae-2.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Gen- erating diverse high-fidelity images with vq-vae-2

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.771282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.574766Z digest=sha256:2b52d0a8ab24bc971be9dd43411b7d4244917f0ae603ce8931b3fed025940f7b

Observation 3819aea2-0056-4517-af94-f79d33b26bf5 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization High-resolution image syn- thesis with latent diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.586598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.586598Z digest=sha256:6fdf725f9ee1ff7835a29024d834c9bc38df70a8784f12c20cf02cb86519f2b8

Observation 59ee048b-eb94-4b44-a119-36745675eca5 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.597282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.597282Z digest=sha256:363d21068948e657a06180e620147b7c5497ff71cbc547e27ca5db9b5f975dc7

Observation 38cd57ad-b971-4384-95e5-d0349e67821f · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.606993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.606993Z digest=sha256:154656b4851ea0d9cc58b904b20e48f9195907e3aef2d2be34587fb387bd60c6

Observation 8cd2ac95-c35b-44e8-aa51-4b9e854e0c77 · outbound

This paper cites Generative pretraining in multi- modality.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Generative pretraining in multi- modality

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.738569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.614175Z digest=sha256:3c5865e04f8dbc8a7bb0cbab387435abba772842b3fb971f40562a334208b0f9

Observation 9523df4c-8de3-4e26-9389-7f5d77b4c31f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.620381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.620381Z digest=sha256:5d061339f5dee28511d895d180e78a26ddbf6bc948beb43b6d54e3e13e1b9c1b

Observation 127c0691-132b-4380-88be-fa6457fa7560 · outbound

This paper cites Multimodal few-shot learning with frozen language models.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Multimodal few-shot learning with frozen language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.717409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.627771Z digest=sha256:df02e1c310fc522e63729ed819a94404e591a12b447becb72b382de6508d0ebf

Observation cd4fe5d6-f837-4844-b485-01fa2be2b4a9 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.634407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.634407Z digest=sha256:59536d3747e9bb2553d34d403f69bc6070812df7ad3f67ece29e55c233660272

Observation f5774657-632d-44e9-aa56-5ac20279ac83 · outbound

This paper cites Neural discrete representation learning.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Neural discrete representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.700155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.640644Z digest=sha256:c31658d46510a296480a28f8efc7e590058b140f24172f6ee681a70a5a89f1d8

Observation d637f0b4-aa73-49d5-83f1-6d1f5827fde4 · outbound

This paper cites Attention is all you need.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Attention is all you need

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.647046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.647046Z digest=sha256:812d549fd74f5f0eb6ef25978579b5a55d41f323fa25929606191ad66e81e30f

Observation ad5233e0-c6f6-4f7c-95a8-29db850161bc · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Phenaki: Variable length video generation from open domain textual descriptions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.668417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.654064Z digest=sha256:f586eaf1a78c63ab73887df30337521aa73307a956378c83d9f9a757aa25d5f5

Observation 40f906bf-7435-48b5-99db-ea375161fe36 · outbound

This paper cites LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.661509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.661509Z digest=sha256:2dd5724a3f4f66bcc92af816eda081f0932cecedd552afdedd5e04c4bb71ad37

Observation 47a0d732-ed91-4c61-9932-0e8cb2f64b06 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Omnivid: A generative framework for universal video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.641124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.668216Z digest=sha256:ed62df9ef384fdaba19272be3f1441b10ea9a2a215f848ae04f8b799a0342ea4

Observation 43331cb1-d65b-48e0-b710-798d2cb80e4d · outbound

This paper cites Omnitokenizer: A joint image- video tokenizer for visual generation.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Omnitokenizer: A joint image- video tokenizer for visual generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.614591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.675383Z digest=sha256:a2998dd0c7a67fc9971e76dc5bba5f11c480b95c8c232c909640ff63a032ea3f

Observation 085300f3-bda5-4665-99aa-1b86c3deefc9 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.594358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.682041Z digest=sha256:61cf71ef97f12a4ed9649792184196aed0766b32c03eabc5e9b399ca568ebf22

Observation 81439c27-75d2-4c21-8185-8d74e302f3e9 · outbound

This paper cites De-diffusion makes text a strong cross- modal interface.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization De-diffusion makes text a strong cross- modal interface

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.572619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.688675Z digest=sha256:79249c98d2d3d8e4e03b1882baf70d8d617634d5757a6c1efa509b7d9ced5156

Observation cd2a7da0-6e65-4afb-93e8-f4c3b646fef9 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.697826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.697826Z digest=sha256:c4b729268bc183b50c4601e1aceb615479746af3d803e42ab5a4226f70221b1d

Observation 12804a7a-50fd-4d25-b397-e47ba164fd20 · outbound

This paper cites Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.550993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.705007Z digest=sha256:6ad62f952556e01de7cefb909bdddcd670a775f80eb1ff79643794fe21a4b36f

Observation 6086b9fa-b514-431c-ad34-c7ba61ba6bdc · outbound

This paper cites Vector-quantized image modeling with improved vqgan.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Vector-quantized image modeling with improved vqgan

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.522142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.711425Z digest=sha256:a9b3e56b029f1a6bea11edae716802d60e3f40153ac4f1cde43287bec9797265

Observation 2ea36884-0cf5-4716-b389-726e4ac4b9ae · outbound

This paper cites Scaling autoregres- sive models for content-rich text-to-image generation.ICLR,.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Scaling autoregres- sive models for content-rich text-to-image generation.ICLR,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.499444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.718292Z digest=sha256:895933fa0be71a59a794449e53d9a77bafbbefb30721bfec50d14967cc8cf8b1

Observation 3a84aa0f-197d-439a-9d56-a797a50f47bd · outbound

This paper cites MAGVIT: Masked generative video transformer.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization MAGVIT: Masked generative video transformer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.476987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.726515Z digest=sha256:1515a57738c9caabb396354aa4b90864c86cc591867cd7da755f2bad1b5c98dd

Observation 2712b039-6360-480f-8be5-9e698c9bdf6e · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.443170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.733286Z digest=sha256:6635c4004ab4cd275ebf7f399c6952d1daaaf30c70c0f6118c4738456821320e

Observation 5ebfcc74-81b4-48a7-bafd-167b2d5b3256 · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language model beats diffusion–tokenizer is key to visual generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.420600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.744271Z digest=sha256:7426c7b00aef7ff79bbd6ed84908c36224c798cef022f0e50aa5b18a2b1d5a9b

Observation d55f7f7f-eec3-43ee-b530-bfd01eb1349e · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization An image is worth 32 tokens for reconstruction and generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.388111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.751607Z digest=sha256:22edae3f9411d1e20e1b47a1bdc3b4b5f1aecb4765397f8f77d762edee40b252

Observation 3d8b81bb-0891-4f13-9d92-c23b56d3278b · outbound

This paper cites Codebook transfer with part-of-speech for vector-quantized image modeling.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Codebook transfer with part-of-speech for vector-quantized image modeling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.358649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.758094Z digest=sha256:096c826cb5fe6f86875bf8d7c1d1c4b14657797d53b5a6f926a5b358266fb513

Observation b1f89d5c-a185-4574-b5af-94784b189269 · outbound

This paper cites Few-shot action recog- nition with permutation-invariant attention.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Few-shot action recog- nition with permutation-invariant attention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.332058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.764498Z digest=sha256:6087df4b637d51315ca241d2752ed9dc3f69dae151da79041bf7901bd76b4904

Observation 48e76d7f-93fc-47d4-8e9c-e396faaf397d · outbound

This paper cites Beyond text: Frozen large language models in visual signal comprehension.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Beyond text: Frozen large language models in visual signal comprehension

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.304341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.772186Z digest=sha256:633190915e4594e66019c4c0afcca25bda0114246a529b90a2f25b7953e16b97

Observation a5e2f6f3-0c49-42f8-80b7-7b51a6083305 · outbound

This paper cites Model Implementation Details Visual Tokenizer.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Model Implementation Details Visual Tokenizer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.273559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.780363Z digest=sha256:2be0f3116daf2537d61c20cfb5a5bcffe9c886e12718f106fd744c62ca04c94a

Observation cf8fdd52-de64-46e5-bbf3-7775f13e1624 · outbound

This paper cites More evaluation metrics We assess SweetTok using additional metrics: PSNR, SSIM, and LPIPS.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization More evaluation metrics We assess SweetTok using additional metrics: PSNR, SSIM, and LPIPS

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:59:37.252753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.786936Z digest=sha256:51040696578ccfd1cd467168edd951d7f193fe496022c915daea71b245a4b24c

Observation 3a6be956-8f12-4a23-a0fd-31d0cf93c06e · outbound

This paper cites an unresolved cited work.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:59:37.229840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:59:36.794196Z digest=sha256:b7826cbcfe6d93e78fb99b1ef0c18379954f2239914f267726af93bb6670627a

Observation 9e6d8dbe-d8ff-4ad1-becf-fe83f8bc503e · outbound

This paper cites A Short Note about Kinetics-600.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization A Short Note about Kinetics-600

Reference 600

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.373349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.373349Z digest=sha256:4f41c14d6aed17dc53769fe0fa6efa08bc63ac1a99adb5f3ced77942deeec601

Pith citing papers

Observation dcdfdacf-f854-4699-b61b-7559fe8035f0 · inbound

TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization cites this paper.

TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:28:55.534476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T01:20:32.508409Z digest=sha256:fc138489f21d1a22ecc679a0d8f7bec4bfdb50f0fea1ce79336733b7c2163e08