Pith. sign in

Paper Citation Record · LEDGER

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

As of 23 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 10 inbound Pith citation observations for arXiv:2505.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07538 v3

Coverage vector

measured 100 of 136 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:21:30.601457Z

measured 110 of 110 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:34:11.795076Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:32.067048Z

Reference resolution

100 of 136 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 853fae54-055d-4765-9ddd-2781eec2210b · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.209938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.209938Z digest=sha256:e3d3525e5c9750506a78bb66a1a85a029bebeb632bf3dfdc4d9adf1ecb319b17

Observation eabafcd6-b845-4f1c-a6ac-6b2586ac66b6 · outbound

This paper cites Llama: Open and efficient foundation language models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Llama: Open and efficient foundation language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.215030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.215030Z digest=sha256:6e0123ddb3e6567931c5d5a862b941e1d9d79c0f6c47c2d0d5286c8f6c8d16d0

Observation 3917c776-d4a0-42a9-a912-8f28d7a203ab · outbound

This paper cites Cogview4: Next-generation image creation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Cogview4: Next-generation image creation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.218978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.218978Z digest=sha256:659491119aa7a6cd6c4ee92c953a384c42cb24cee4a7709d4dd740965ef50093

Observation 0757000c-989f-48a4-92ce-05ab452a7c3c · outbound

This paper cites FlexTok: Resampling Images into 1D Token Sequences of Flexible Length.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.222903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.222903Z digest=sha256:f5ebf9d53cb0259f9d9770e3000bfdbbbfebb00e73b857be71e144a599405811

Observation a7a9f234-dc8a-4e09-bfe3-d6380bdfb96d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.227428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.227428Z digest=sha256:c537444fb76bacaf14651f6eeb934bc8282780283b8dd19bda33a3214112928a

Observation 02dc6327-1738-45aa-8000-18c07da86843 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Instructpix2pix: Learning to follow image editing instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.233552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.233552Z digest=sha256:fc7295ea1cedcc256cf0f7d27760dd480924fcebe914b139c2484f79feeb063b

Observation d662ea7e-8b52-448a-b85b-7dce66cfc2ce · outbound

This paper cites Coyo-700m: Image-text pair dataset.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Coyo-700m: Image-text pair dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.238141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.238141Z digest=sha256:5e19985ee4eaf8466f2e521f12ccfc426405516542d7d7bfac8b92679e8020f0

Observation 3a6dd38f-3567-469d-a2a0-9842c5cbae80 · outbound

This paper cites Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.241982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.241982Z digest=sha256:11a2333b881bd73ebe66aa727f9d0cf0f2aa35bf7b2f52ef4122f2ae97cf5612

Observation db27d10d-1d20-424a-bbac-e9512bcdb96c · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.245880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.245880Z digest=sha256:9644862d46dbe8467eb91fc4b18563b9014f2832048bf15c45b7c85cbc9bee93

Observation e4e1cf9f-d723-45d4-9d4b-0ded1375c65e · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.250126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.250126Z digest=sha256:73d920f907f1b58ba9e2e229d4bf95ec5f301a6e865bc9f0ec2d5ffd756dd3f1

Observation ff94986d-04fa-41b5-8edc-3dc52c7149f4 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.254395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.254395Z digest=sha256:e9e8ac30cac700e20c81e1fadfffc229e2ee0b24787f04ee25c1142f77f32388

Observation 6e112dd2-b128-4dd6-8559-be9a6765b12e · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.258422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.258422Z digest=sha256:d696c3efdc6b2777b6a9cdd3348f09ab169d316d2bc27a71bdef743d55bc542a

Observation 0c411dfa-871c-416a-bd1c-e7ceac585b17 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.262430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.262430Z digest=sha256:3d34f540dfb980380de3bc0093352ac95a4edb106f1aac2507e98979c7a52c74

Observation d1b170c0-a401-4549-80ca-c1a75d75dc7f · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2024.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.266348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.266348Z digest=sha256:21049cbe399aa32ea6eccb5b1d7e6f98c3eefa569426f4deba07fb5d387935e6

Observation a2fc0355-1853-434a-9b94-3571188d362d · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.270022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.270022Z digest=sha256:98138aa179664d684caecd16ca22802e516aa83c4f8a4ee8a00f83aaacbfe1f3

Observation 991e4d41-d404-46ac-88af-909d179c1477 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.273912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.273912Z digest=sha256:b1f8aafb197e8598b6653a98e88b85b3b2b222a94ee026f8970a9bd6e134ca17

Observation 29020a1a-e7cb-47ec-9512-81f5351c1ebd · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.277996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.277996Z digest=sha256:329137f28e58f243a5f1e4ce16478b41f126289cc5b4cfc8817cad636926a5cb

Observation b07c87a6-8e2d-4c32-9d26-f3a4a23ca481 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.281682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.281682Z digest=sha256:7d037d97200ac3ccaca8736babd6109c9b17ee12cdde2142b59b286aac4ef2a7

Observation 9c650bda-b151-4f03-91cc-27e349ad988a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Imagenet: A large-scale hierarchical image database

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.285769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.285769Z digest=sha256:f9c49e904239c6468b39317f97f7c1926947bd2ec6833ab7c95ae0c9c4cc8ef9

Observation e33bdd08-ad5a-430d-9ce4-8e014621962d · outbound

This paper cites Diffusion models beat gans on image synthesis.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Diffusion models beat gans on image synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.289301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.289301Z digest=sha256:50f224bc190d5b3d7738e02f779d438a2653ee808d2d2e56f56be3cb28541adc

Observation a994d005-c83a-4597-b7e5-2cf6bbafc688 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.293327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.293327Z digest=sha256:1cfe0bba228011bf1a2c5e1be6948661b0e7d352d9ce7693a6f046acd3513614

Observation ccec577d-8cd8-4e0d-a320-db7caa8647fa · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.297617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.297617Z digest=sha256:18a7c79f5bffdff6c5a22adc087f57307dc7ed58856842c7910ab6b31a09d9d4

Observation 62a9d35e-df87-4935-9485-09363b4645a4 · outbound

This paper cites Adaptive length image tokenization via recurrent allocation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Adaptive length image tokenization via recurrent allocation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.301732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.301732Z digest=sha256:fd6764955d9ba46d007c8ed0f4d181d7bc14c6c8a333a5752c37ea68f368e8c1

Observation 80b3b58f-f02d-4e1b-9591-9e0af6ebb812 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Scaling rectified flow transformers for high-resolution image synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.305635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.305635Z digest=sha256:ebac88cfead7496bb5f94d01a6e10c36a185e75f71e32821d89ea0c1bbd4554d

Observation 76fee686-1e78-43b2-8562-e6b77d14173c · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Taming transformers for high-resolution image synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.309602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.309602Z digest=sha256:c9786d5a5382f31a7e8e008688bcff96ad1b3f207c1039b7712207ea0a6485a3

Observation 8427b3ff-1693-4cd0-9532-d975ee6f87e1 · outbound

This paper cites Guiding Instruction-based Image Editing via Multimodal Large Language Models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Guiding Instruction-based Image Editing via Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.313407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.313407Z digest=sha256:1180e76c5a69db570f28533060cfd4181f15352d71fe005f5e91a6b2fc236533

Observation 13386499-bd3b-45c4-a6cd-e137f993a7f3 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DataComp: In search of the next generation of multimodal datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.317518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.317518Z digest=sha256:29c69b8becf5983d81a998ab2affcf233c4bdaa98363f89713b7ee53535c07e3

Observation 18889c96-6f47-42c4-a1d7-b7a226c56420 · outbound

This paper cites Seedream 3.0 technical report, 2025.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Seedream 3.0 technical report, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.321605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.321605Z digest=sha256:3fb1f96e800218518a6586042ed133eca1f8dd5186a5afc079ad120cd60f2d06

Observation 6de8d79b-1cb1-42e6-9823-7387d02c7d23 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.325285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.325285Z digest=sha256:058660e827a3d9b90194f4b785bfcd0068e23a124e695f6d9e5ee18c694a367d

Observation 27d72071-fbb4-48e4-b8a8-924177433665 · outbound

This paper cites Instructdiffusion: A generalist modeling interface for vision tasks.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Instructdiffusion: A generalist modeling interface for vision tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.329926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.329926Z digest=sha256:6d5e9d6b75a4f13984986d6ad264edd265a50a4531f4e555de5ec6cd4ce31075

Observation 600b3329-c4b2-4fbc-ba4c-2ca452c5d1e7 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to- image alignment.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Geneval: An object-focused framework for evaluating text-to- image alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.334048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.334048Z digest=sha256:032d96bfb7baa59b495f6f1dd8658e18f9592058e8fee2274f1a9eee601a22bb

Observation 3e243bff-ef66-456c-bd6b-da1ab8b5c2ce · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Better & Faster Large Language Models via Multi-token Prediction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.337753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.337753Z digest=sha256:6052bad69d4ef9f87945d0026c86809caf1b8578afec3b915152670a423271e1

Observation 779507a9-4a73-46bd-8717-fdedb66aee14 · outbound

This paper cites Generative adversarial networks.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Generative adversarial networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.341665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.341665Z digest=sha256:3409cd7130fb8b2a37561711dcbc8d4661780a788a190d228278597cc6d50b2c

Observation e74b460b-b72f-4871-8416-8e9308c7f2c5 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.345353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.345353Z digest=sha256:253ec3707339dc1f2dc701933904f9d46057ef390502fdfe05839989e4b04df4

Observation 22104682-3486-4b52-be63-5681df19ef9c · outbound

This paper cites Multi-reward as condition for instruction- based image editing, 2024.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Multi-reward as condition for instruction- based image editing, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.349198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.349198Z digest=sha256:95b54ec8c44640d7db19624d4a5257bb692b43f3be1a7f64aec073982ef7ae3d

Observation d0d1cd56-4605-432d-be82-cf9ccd8d4fb6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.353022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.353022Z digest=sha256:ff67c400aa9cb430a927775249a12992e40d735c2e0218a76d0cc28f940ab295

Observation c375f887-d5d7-4c3d-aec4-040d63f3450c · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Vizwiz grand challenge: Answering visual questions from blind people

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.357040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.357040Z digest=sha256:d38e0bb2565d6e5796c1da034191a46d826db104736293629e28211a952c4c0c

Observation 9008339b-3c99-4b71-b876-bd54ceb307cc · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.360719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.360719Z digest=sha256:422bcf6b5433e9dd9639b7572da87b38cce2927e3624b4162ad4f6504b9c9726

Observation ea04f245-2ca7-43ce-86c9-de5e3eb8ccfd · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.364795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.364795Z digest=sha256:f028dcd57a86a789c232c4487cbe989debf56bbb4db65d575d22b3c50851aec4

Observation be4a4dcf-ca40-44ab-a1cc-4ae519122d56 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.368820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.368820Z digest=sha256:5618350f08d042e0997c953eab99726a5067ecc577c4eddab84affbabe1c1352

Observation f6dfcb11-85d6-4a44-ab43-ed323029bb19 · outbound

This paper cites Hidream-i1.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Hidream-i1

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.372389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.372389Z digest=sha256:827907ede2b2d3ffb5a9b5ca8a7f96c3ba1db4fe15f6cb5116d09b0a4ee03ca8

Observation f0981535-08e6-4647-857a-c8cc885f2a74 · outbound

This paper cites Towards a Definition of Disentangled Representations.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Towards a Definition of Disentangled Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.376368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.376368Z digest=sha256:5cde60a06d26d934b5368dfb1837aa7bf262cf06188b262fc795999de1fbb3e5

Observation 75fc73ab-7688-47e6-ae52-f7588a43c1de · outbound

This paper cites Classifier-Free Diffusion Guidance.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Classifier-Free Diffusion Guidance

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.380408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.380408Z digest=sha256:0d6414f544c09490badc2646afc8e4c5f2d94a04ab3c71a291a9288ee9f73c6a

Observation 52adf81e-6273-4639-95ac-4b5d69bccd39 · outbound

This paper cites Disentanglement via latent quantization.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Disentanglement via latent quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.384364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.384364Z digest=sha256:c4d335bcfd9c1fab912ca4c5abe717c84b8b4e75dc5af161d6a8d689f5ca6ddd

Observation d12a027a-dfe5-4133-9a00-33d89758b782 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.388080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.388080Z digest=sha256:fd2e373716721eb2bd142fe4ea1f15f1baf9ee9c55d07eb5c05efbb0c99287d7

Observation c963e1d4-ab50-4d81-9524-bb2df27de48c · outbound

This paper cites Understanding square loss in training overparametrized neural network classifiers.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Understanding square loss in training overparametrized neural network classifiers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.391981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.391981Z digest=sha256:176a162243fe14a9a3ffbc284882e3fe50e19ce026f9979057ef63096bf401dc

Observation 83e7ce75-2c5e-4f4f-9023-2dc9c52a0acb · outbound

This paper cites Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.395749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.395749Z digest=sha256:368f84a919a76091c809e59b1ed852bbb8d2e64620ab2486b6ce387f3d9e9fd9

Observation 6a26f0c4-bff6-4557-b824-6f5a2d25b334 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.399476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.399476Z digest=sha256:798b4a58a7bc5f98a6f91ce6d30beb5347390cef14c3c9a0b9667e2b003f8cc5

Observation 51c076d4-e036-45df-8d6e-4ffa10990a75 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.403344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.403344Z digest=sha256:d231499fd6e717618628e745e9f36e26f2072a69657923e7d31cedfc35c0525c

Observation 14de0ede-ca8c-47b0-be8d-1d1c7705329f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.407073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.407073Z digest=sha256:b81b33315285abc701de7e6ba5129901e3ef31505387353ebdb84c928770a5ce

Observation cc83c1e5-7927-4d0a-9849-b6b5e2cc9f5f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.410913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.410913Z digest=sha256:96f0becfdcb6466a511e382be523b43c156e624430c0f8407d30b19fee485ea0

Observation 31569132-050a-4f63-9b7a-11fb77f66b6e · outbound

This paper cites GPT-4o System Card.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning GPT-4o System Card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.414578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.414578Z digest=sha256:c46e4f49282ae08e67750b14c9df00a45054a6c765d118bd716bc1d35f34f871

Observation 4fa7caed-f4d6-4004-b5e1-ed066c383b44 · outbound

This paper cites Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.418517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.418517Z digest=sha256:23deb2642d8bfbaf0e40a3d2c89a756c402f00f450db6b5234468cce0a47e7d9

Observation bf054325-6558-464b-bb1f-89483c73c190 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning How Far is Video Generation from World Model: A Physical Law Perspective

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.422247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.422247Z digest=sha256:1ef6a136eee24ba0aba5d57c78307c3d97cdee53b440b094a778d43ee4b2c9c5

Observation ceab6c74-a369-4d70-a599-4527d22ebb47 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning A style-based generator architecture for generative adversarial networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.426541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.426541Z digest=sha256:f4aa0c7d483f76c98f644eeb0ede2765143cca0b66c7cdbe20d8e3e7cd950dae

Observation 4d0085df-d7ba-4c1c-8a1d-41f8b1f22403 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.430410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.430410Z digest=sha256:32fb7b258bc393bcc3a60f62e621c0e652e339d03210ae03704a5886a7b83918

Observation f702821b-c99b-41a7-b5a3-92b935fa8518 · outbound

This paper cites an unresolved cited work.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.434319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.434319Z digest=sha256:bc8abc2781dcbf0b8dcd3500ef1350699f054c0c0607bae9204c0f6270eb8092

Observation d8c11688-a2cf-40fa-8f0d-19d79cc1e8bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.437925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.437925Z digest=sha256:f5d7ca3f8bed14290f8d6585909926595791e2ac5f62132b795eba6d25466676

Observation 638fe5f2-d579-49c6-b51b-a6015cc8c741 · outbound

This paper cites mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.442907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.442907Z digest=sha256:23630baa0448ebd72df9cfed3ea8a27fb76d7d43e1fcd329bf5a332daec392fc

Observation 669ea675-efc2-4d0c-a58f-774f01abd0b6 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.446932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.446932Z digest=sha256:071799de0ad9c81921ba61bb313d4ff9e85bc0ddb3f894d9aafd81dbc91af785

Observation 4ff91796-5c1c-47e5-b54a-d5910d5e363a · outbound

This paper cites Treble Counterfactual VLMs: A Causal Approach to Hallucination.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Treble Counterfactual VLMs: A Causal Approach to Hallucination

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.451070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.451070Z digest=sha256:f162b113db99de2699d904fd6aa798e948245cb5a76d6e3aacdfa491917bf685

Observation 782a4c0a-4cfa-4dc9-a351-52f26267352f · outbound

This paper cites Autoregressive image generation without vector quantization.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Autoregressive image generation without vector quantization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.455088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.455088Z digest=sha256:404e6a4c94dfaca81063a0f9bf2f5175a04d044ca10e69a7365d342eeeb18a3d

Observation 2c8f42cf-5769-45ae-b5ad-a831411f96a8 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.459345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.459345Z digest=sha256:5a06c649acc5c8e0fd4c8b5be8fe0a1339d96bd633601bad58447a7a7ecf4f85

Observation f1d1ac06-1cbe-4fe1-b592-baa21ad21ed5 · outbound

This paper cites Dual Diffusion for Unified Image Generation and Understanding.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Dual Diffusion for Unified Image Generation and Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.463448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.463448Z digest=sha256:68287bf1028af3f841e003c17622cead5953926bab6a607946a04e9b60a6e586

Observation c3331cd3-151f-432c-8772-f813f68b1b87 · outbound

This paper cites Reasoning physical video generation with diffusion timestep tokens via reinforcement learning.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Reasoning physical video generation with diffusion timestep tokens via reinforcement learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.467625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.467625Z digest=sha256:dd479f4c4a239f5eca7acd406e1d190758839e064b3c4a3dd68dafc573b07437

Observation 650a9e5d-0bb0-4bc0-91cc-02f0d8a87581 · outbound

This paper cites Improved baselines with visual instruction tuning.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Improved baselines with visual instruction tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.471362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.471362Z digest=sha256:0dcc1ef77e97bc019ff046139fe2e852cf0f488eccd0088a8f67a4e1d3f16734

Observation 38e8b85e-d58d-45d3-974e-de7f82eeba7d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.475012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.475012Z digest=sha256:9778395961c99f5ee7b232dd1748c6902a219a8b4803706eda37f605361bb556

Observation d0a3c5b7-16bf-496b-9f39-89daf0ce0226 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.479803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.479803Z digest=sha256:f77c1024548e1daa19f5d452d8f7a44d757f0e61a64ef40ccac2b29c7f45eddc

Observation a8bb7a61-f1f1-4cb4-98b1-8dbc6d12f36e · outbound

This paper cites Challenging common assumptions in the unsupervised learning of disentangled representations.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Challenging common assumptions in the unsupervised learning of disentangled representations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.483451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.483451Z digest=sha256:304bcd57dfdf17c95479ab74a94037b04b1286f6a960bd9aafc838975c14a3f5

Observation 130c5350-07d0-4c92-861f-bccfcfe3b785 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.486994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.486994Z digest=sha256:5af5070ec330d44d1bab722ad832c39d36288ba315570213d7c0862e7f57b49f

Observation acedce9a-1579-48bb-9fe1-06eb4d750e88 · outbound

This paper cites Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.491285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.491285Z digest=sha256:b6047a0127061539264a8880f64280373b7f875f6b451f297073a3b7947045bc

Observation 4b433d4f-0686-4bbb-b350-b52ac039d73f · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.495041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.495041Z digest=sha256:dd0ddb2d040d32fded735c616e9a876d265664b531e76bfd2b9a5d9dd502ebce

Observation 8d68e5e7-dcc4-4377-9907-c7a83d748fe4 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.498950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.498950Z digest=sha256:fc8d2d4efbd4d6e26398f73528002de9d779e935c234d38934780d5ecdb6c272

Observation 3c96c7b6-7fbe-49f9-b0dd-369b6adf1f90 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.502722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.502722Z digest=sha256:ef9ff33b7b191e470cd74d635be4c3a7b34c73bda5c4b066f8141995b6afecd6

Observation a91bb8fd-1208-4e09-bbe6-003023b3a0f2 · outbound

This paper cites Harnessing Discrete Representations For Continual Reinforcement Learning.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Harnessing Discrete Representations For Continual Reinforcement Learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.505976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.505976Z digest=sha256:2a45b0d7c4180e9bcc6252e12f0079ff22bb2b74fadaa35a1e8d2fe23eb55b0d

Observation 295121ef-bf11-45ed-baad-4c4c7b2cd27b · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Ocr-vqa: Visual question answering by reading text in images

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.509513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.509513Z digest=sha256:0095ab0f2db2c4a36cbee698abf58f5f499dfeb3bae6e79dfd5acf1a959ea5bf

Observation 73b3651f-b484-44e8-95f0-3b96b46efae4 · outbound

This paper cites One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.513108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.513108Z digest=sha256:7802d3e4ea30792ed4a1d33e969dd51200261447b6b83651b541f15ebcb02ef3

Observation e4a2fde6-0611-4cbf-840f-c6c847e03f39 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Null-text inversion for editing real images using guided diffusion models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.517221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.517221Z digest=sha256:b126e282b881a1b955faaf1f053be8fe16fe8e77fab04aee08b6eb416acb0eea

Observation 14eed918-d081-43cb-9f97-983475a75d0f · outbound

This paper cites EditAR: Unified Conditional Generation with Autoregressive Models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning EditAR: Unified Conditional Generation with Autoregressive Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.520930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.520930Z digest=sha256:8fadfc923b41ad8dfbce0b8695326a4e487f32a5bc13fadf371c5ec8d7900217

Observation b9802df9-2822-45ff-bc41-d9f2cc0bf94d · outbound

This paper cites Introspective distillation for robust question answering.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Introspective distillation for robust question answering

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.524812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.524812Z digest=sha256:24e6179b96b6bbd5af4d35feb4cd2b2f12ddff770f74d165157374e7dd3660de

Observation 78feb04b-55de-456d-84a0-974af26d6921 · outbound

This paper cites Introducing 4o image generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Introducing 4o image generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.528336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.528336Z digest=sha256:c5183b5357146ce51e4712879427dec2075b2b5bcfb812d5af063b1b6c15ce62

Observation 38130b36-fe3e-456e-af8f-8028d8c4f217 · outbound

This paper cites Vr-sampling: Accelerating flow generative model training with variance reduction sampling.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Vr-sampling: Accelerating flow generative model training with variance reduction sampling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.535548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.535548Z digest=sha256:f1096cc7e8be5151034ff7e13e07b555d82059d0eb1f27428e56195c2c63a3e3

Observation 187a22fe-1b94-4757-8902-7af3cbf38d28 · outbound

This paper cites Generative mutimodal pretraining with discrete diffusion timestep tokens.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Generative mutimodal pretraining with discrete diffusion timestep tokens

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.539119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.539119Z digest=sha256:32e1bb01d42942457cef3d68af9c609a3e9002ddb1576d8afddd1f01ad5ef1ea

Observation 89ae4d76-8d74-4db0-9dfb-083496ba9fc5 · outbound

This paper cites Auto-encoding morph-tokens for multimodal llm.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Auto-encoding morph-tokens for multimodal llm

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.542780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.542780Z digest=sha256:689c8afaa9ae3e22cefea053b73beebbdcafa565521165cec591e9970dccd7e9

Observation 4e8b4275-9989-4de7-9c77-75cbc84b8e59 · outbound

This paper cites Zero-shot image-to-image translation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Zero-shot image-to-image translation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.546159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.546159Z digest=sha256:648b110d2f2eb10bad4712ac1a2b6bee6ab8134b93c22a168d1d3646f7158766

Observation 77b23f67-bd63-4b4d-a0ee-b9dcbd2e5b57 · outbound

This paper cites Causality: Models, Reasoning, and Inference.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Causality: Models, Reasoning, and Inference

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.549753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.549753Z digest=sha256:e5cc8af70ddc98a64cca1bb5e504cedee4e85c5362405707249f750484f7ffa7

Observation bb65c5c8-00e3-4078-a0d2-df32dae0e28e · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.553319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.553319Z digest=sha256:7cefc890902ac7eea034cc91cb9035a4e4a23fc55b000309a2068d408b042580

Observation 02149c5b-f559-4752-ac53-ef88093378a9 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.557129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.557129Z digest=sha256:e2583fa3908175a52f25b21c0f8199a18a0a2fd289d9dd7385711e87e7b860d9

Observation 2356c3ea-3ece-449c-b702-cc2e7d541be6 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.561032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.561032Z digest=sha256:5e0783b7052b8910ec8f5a33a6ab969c7f7f27710bdd8265151ddb37e85402e1

Observation 802ec8bc-4b74-460d-ae0c-f15b9c2b12b9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Learning transferable visual models from natural language supervision

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.564777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.564777Z digest=sha256:659ae02fd0a50e82d36fffa46d175e52d71dae8b08842cbe5b5e533f5cb97bfc

Observation 75424c96-1a2c-4633-8e32-dafe9d1d6321 · outbound

This paper cites Improving language understanding by generative pre-training.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Improving language understanding by generative pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.568462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.568462Z digest=sha256:d6614598e7357e8a922670fa05427e1bbd986d6dba91a9ef366ca4202aac177c

Observation f2497048-f8a6-4268-89ee-4bcd31d7fed8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Direct preference optimization: Your language model is secretly a reward model

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.572003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.572003Z digest=sha256:075276a07907392ff49cbbeb034232b3e1e0e6ec251503dff1ef3bc65a20ed33

Observation 83a59e16-9733-4722-905a-dd575b9d595b · outbound

This paper cites Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.575674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.575674Z digest=sha256:d339ba59e8ee45070114078904f669a9e1f9dee5d8028540356ba44db5cd58fd

Observation cb2ef13d-a3dd-4eca-83a2-c05ece23b561 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.579716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.579716Z digest=sha256:05dc5da770a30c8c3fb2f1f297f47bcb53aefa71eb639b6913507198d7d416b5

Observation e6e89895-4e7d-4fca-baac-551e9fb179e6 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning A-okvqa: A benchmark for visual question answering using world knowledge

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.583257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.583257Z digest=sha256:b9bd6c842f81085ec6710a03858d95cd8bed6e22c30a8f5910132613fda6efa4

Observation e41c0708-7ad9-43ed-ba2d-1d4e640e7c82 · outbound

This paper cites Neural machine translation of rare words with subword units.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Neural machine translation of rare words with subword units

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.586551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.586551Z digest=sha256:aa4ce2669b8e8d50a35b0a05f75d1f1212abae19128bb28557f975c97432a77f

Observation 0cbe20a1-c749-408a-b1f0-27a512be7887 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.590027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.590027Z digest=sha256:1ea88cd463e692f144329f361f8f45bc0d68943c773cc80922c6d409c514b364

Observation 2955a866-7621-422d-b9f1-339861e7bc32 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.593828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.593828Z digest=sha256:3d5cfae78e85d578b98b694398776e339e9b2d71dd94304d2bbceed7822eb497

Observation 907a28fd-2965-4100-888f-d76998e92674 · outbound

This paper cites SeedEdit: Align Image Re-Generation to Image Editing.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning SeedEdit: Align Image Re-Generation to Image Editing

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.597877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.597877Z digest=sha256:03dcacddb247bf40775d183a26dcb60532c6bc2880cce0397ca2b90e58279248

Observation 547786c1-c51d-4328-a325-36cbdf87e25f · outbound

This paper cites Improving Image Captioning with Better Use of Captions.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Improving Image Captioning with Better Use of Captions

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.601457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.601457Z digest=sha256:23d2809e91020e5dc6066f8a2ac4a2951e6950d05307a709e36b644ee9e54634

Pith citing papers

Observation 03ed1463-9d7e-49f5-bd31-d818ef5b29e3 · inbound

IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models cites this paper.

IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:11.795076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:11.795076Z digest=sha256:f82628d08036eb72b64a03bb1f64cf82c8efa1214653073374724ba5a7ead90e

Observation 5a97a6a8-26bc-4f56-9a80-f7f0ead67f4b · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.815176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.815176Z digest=sha256:7e144d7409e0c063dbe7e5775147d79eded9070e32130e2fddd236fc4554c219

Observation 37013941-5f88-47bf-bfa3-6ab407143bea · inbound

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal cites this paper.

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T21:53:31.402367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:53:31.402367Z digest=sha256:78c6446d2ca61548a0c7ad6d6ea3be5d2d919f13a92ad65de7f2c668e48f6d65

Observation 264d98ea-86c3-49f9-a435-05e7ea5f30e3 · inbound

(1D) Ordered Tokens Enable Efficient Test-Time Search cites this paper.

(1D) Ordered Tokens Enable Efficient Test-Time Search Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:10:08.999719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:07:42.245810Z digest=sha256:1ce190ee10fdc7927307cd251bd9eb74531f14209d6957de8680de3f5e87ea6c

Observation bea56d38-c42d-462f-8a86-8972860f0c02 · inbound

Autoregressive Visual Generation Needs a Prologue cites this paper.

Autoregressive Visual Generation Needs a Prologue Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:06.077496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T13:52:42.834440Z digest=sha256:e631481cd2062252b7dd1de554b050980a07a1e99c171626648aa128a3b75ba8

Observation a17bca25-1c3b-4e72-a6e5-dc9db4f3ed48 · inbound

Autoregressive Visual Generation Needs a Prologue cites this paper.

Autoregressive Visual Generation Needs a Prologue Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:15:46.140171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T23:36:51.148162Z digest=sha256:b88aba6975e241bc96c4a2c998e3e867c7c7be620960db63cc029d579f05aa28

Observation 4cb9248f-9406-4295-ad51-3e7438e6479f · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.068307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:1c8902149ee8920e05f868700416cacd6db16ce661fdf106425f407b65cfbe5a

Observation 0e76332c-b599-451e-b719-87cdbef12fb5 · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:14:26.769263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:07:44.343539Z digest=sha256:ce977518d3cfa43dab688662dcef7d33a17709f1f37cfede43139cbc8ebee0a2

Observation 7d076275-a792-43b0-b30f-23238cb1ce66 · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T09:37:23.456229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:37:23.456229Z digest=sha256:8726581e93434e001796bcec60ccb9d63748b24ec267f54b8002df965eff4c0b

Observation 2f0020f5-58c2-4cf9-9e0e-596ba65d3cd1 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:51.713110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:51.713110Z digest=sha256:6b5c02bfeffdb49d190eda5606bbe90579091d85114494a28cd1b78b45e3f8e7