Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:59:36.794196Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2412.10443.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:59:36.794196Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:20:32.508409Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:28:55.532932Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f634ee9b-78dd-4f3c-ab01-fd27cce0af3a · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71491f69-a707-4d05-8975-ad2eaf8a759c · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language Models are Few-Shot Learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4497d6-f0d1-49da-bde6-b8d708fc57b0 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization End- to-end object detection with transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a715bda-cc16-4c5d-979d-cea8fd1f9c25 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization A short note about kinetics-
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57b8f75-1b49-4082-8570-7131645bab46 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Maskgit: Masked generative image transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9edbc7e-0e6f-4054-8ba4-bd1424e111f4 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Palm: Scaling language modeling with pathways
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8cdea5bc-c0aa-47ac-83e8-53e8f13526d0 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Imagenet: A large-scale hierarchical image database
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27031164-1ba2-4614-9208-8eef7c0f54c2 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0ca42ab5-0596-4bf9-8331-7643f4587299 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d8e1eb6c-ca63-41a9-9d10-d1d3cabc4ca7 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Taming transformers for high-resolution image synthesis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1076691a-c275-4857-8126-7412d4262aee · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Long video generation with time-agnostic vqgan and time- sensitive transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51dacad5-eb1a-47a6-b833-7650d0da9238 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a81665d-89ab-4346-b607-3b6f6aa400af · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f623e22-7c16-48dd-b4f7-228542323d09 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Video-lavit: Unified video- language pre-training with decoupled visual-motional tok- enization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f4b85abb-128f-48d2-9e92-63ce885fad3e · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Unified language-vision pre- training in llm with dynamic discrete visual tokenization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6641644-3148-462f-8aec-51579427366d · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization The Kinetics Human Action Video Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be71f54-55f5-4b17-bfc5-424bfbe85832 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Auto-encoding variational bayes
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 945b8942-8d6d-41ac-8fa1-8bca06121c9b · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Adam: A Method for Stochastic Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63f6f55-b613-41e1-8ae5-71a27ae45d90 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Semi-supervised classifi- cation with graph convolutional networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a136bf9-69fe-47a5-8ae4-3b95ef5f41ae · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Few Shot Activity Recognition Using Variational Inference
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 968cd9ef-e880-467c-bb60-10da10c89a25 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Autoregressive image generation using residual quantization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4e0ab49-b0b7-41f1-8331-03dfe4465fc2 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 211ad92b-f31d-4486-a9f4-25a9bf0f52aa · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Video-llava: Learning united visual representation by alignment before projection
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e273481b-cba3-4d46-b71e-8a8c6a482911 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language quan- tized autoencoders: Towards unsupervised text-image align- ment
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5df510ed-3dba-488d-a51c-b6cb1d70606b · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7c6b2b-26f6-4834-bc1f-318f6f1ccd48 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language models are unsu- pervised multitask learners
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9a323101-e69d-48ba-98db-1db98b6dc1e1 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Learn- ing transferable visual models from natural language super- vision
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0b53fd07-f0ef-46b4-a302-a3d17184de8d · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Gen- erating diverse high-fidelity images with vq-vae-2
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3819aea2-0056-4517-af94-f79d33b26bf5 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization High-resolution image syn- thesis with latent diffusion models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ee048b-eb94-4b44-a119-36745675eca5 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38cd57ad-b971-4384-95e5-d0349e67821f · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd2ac95-c35b-44e8-aa51-4b9e854e0c77 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Generative pretraining in multi- modality
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9523df4c-8de3-4e26-9389-7f5d77b4c31f · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization LLaMA: Open and Efficient Foundation Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127c0691-132b-4380-88be-fa6457fa7560 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Multimodal few-shot learning with frozen language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd4fe5d6-f837-4844-b485-01fa2be2b4a9 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5774657-632d-44e9-aa56-5ac20279ac83 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Neural discrete representation learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d637f0b4-aa73-49d5-83f1-6d1f5827fde4 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Attention is all you need
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5233e0-c6f6-4f7c-95a8-29db850161bc · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Phenaki: Variable length video generation from open domain textual descriptions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 40f906bf-7435-48b5-99db-ea375161fe36 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a0d732-ed91-4c61-9932-0e8cb2f64b06 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Omnivid: A generative framework for universal video understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 43331cb1-d65b-48e0-b710-798d2cb80e4d · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Omnitokenizer: A joint image- video tokenizer for visual generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 085300f3-bda5-4665-99aa-1b86c3deefc9 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Internvideo2: Scaling video foundation models for multimodal video understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 81439c27-75d2-4c21-8185-8d74e302f3e9 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization De-diffusion makes text a strong cross- modal interface
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd2a7da0-6e65-4afb-93e8-f4c3b646fef9 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization VideoGPT: Video Generation using VQ-VAE and Transformers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12804a7a-50fd-4d25-b397-e47ba164fd20 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Towards end-to-end generative model- ing of long videos with memory-efficient bidirectional trans- formers
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6086b9fa-b514-431c-ad34-c7ba61ba6bdc · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Vector-quantized image modeling with improved vqgan
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ea36884-0cf5-4716-b389-726e4ac4b9ae · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Scaling autoregres- sive models for content-rich text-to-image generation.ICLR,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a84aa0f-197d-439a-9d56-a797a50f47bd · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization MAGVIT: Masked generative video transformer
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2712b039-6360-480f-8be5-9e698c9bdf6e · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5ebfcc74-81b4-48a7-bafd-167b2d5b3256 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Language model beats diffusion–tokenizer is key to visual generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d55f7f7f-eec3-43ee-b530-bfd01eb1349e · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization An image is worth 32 tokens for reconstruction and generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3d8b81bb-0891-4f13-9d92-c23b56d3278b · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Codebook transfer with part-of-speech for vector-quantized image modeling
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b1f89d5c-a185-4574-b5af-94784b189269 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Few-shot action recog- nition with permutation-invariant attention
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 48e76d7f-93fc-47d4-8e9c-e396faaf397d · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Beyond text: Frozen large language models in visual signal comprehension
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5e2f6f3-0c49-42f8-80b7-7b51a6083305 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Model Implementation Details Visual Tokenizer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cf8fdd52-de64-46e5-bbf3-7775f13e1624 · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization More evaluation metrics We assess SweetTok using additional metrics: PSNR, SSIM, and LPIPS
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a6be956-8f12-4a23-a0fd-31d0cf93c06e · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e6d8dbe-d8ff-4ad1-becf-fe83f8bc503e · outbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization A Short Note about Kinetics-600
Reference 600
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdfdacf-f854-4699-b61b-7559fe8035f0 · inbound
TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.