Pith. sign in

Paper Citation Record · LEDGER

Making LLaMA SEE and Draw with SEED Tokenizer

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2310.01218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.01218 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.714439Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T20:16:29.588718Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 70162ff3-a8ce-43d7-86d6-46e45d12b224 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.453596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:85f0c8620d6aba2aa0b586e04569b0ac1c06a11cdeb0b6c3ec29c60cb56ba373

Observation 5892f0c8-53f8-4454-b1e4-d0210078372b · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.390334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:6ff9b2f2c504fbea33cfcc29eb49b6b32d8b02bf742ed11ad498e54f442246a5

Observation 239459fe-7563-48fd-a297-eca4d5de87d5 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.478355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:7307b34b66118329f9c79306ca2733b80fa628db60cf666930a4d61786b3a0b2

Observation 9ff8c8a9-dfd0-4fc9-850a-e8352882a5e2 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:09:16.881241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:ee741286583c6fbc16f7c5cc1bea2716a4d752952db93a3ec883f227ede45ec5

Observation 46622577-559c-4ea9-afda-31666510f829 · inbound

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models cites this paper.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.459326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:f32260456f38ea529594675deef7474e485631e365ab3caf0ab062b692667c16

Observation 6412ae67-e7cb-450c-b736-4f27b31d55a5 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.479085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:4976fdb761da9a179f0d4c1fbe3a0f8b5bdd51265071b9e33fbd42f0aec1ac0e

Observation 6a992442-55ca-4ece-8619-4e6d2a23cb7b · inbound

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era cites this paper.

Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era Making LLaMA SEE and Draw with SEED Tokenizer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:24.255889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:24.255889Z digest=sha256:6c48349c07514329eb647de46b03cdba3c68a77b9813e782b102eefd6ebb3f5b

Observation 050a9970-6069-4910-b918-cb01401a1834 · inbound

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding cites this paper.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Making LLaMA SEE and Draw with SEED Tokenizer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.366549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.366549Z digest=sha256:670fc63557b942d0794e98f3d5eb58c7d2bcea3651cf288a3832183bf17e7eb1

Observation f5df30e3-daed-4e0e-b44a-332eaf4051c5 · inbound

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation cites this paper.

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:10:54.624006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:10:54.624006Z digest=sha256:e1a4336f907a1d6a174bc4def02eba1d7674d279f9a429ffef5eaacf03af969b

Observation 2a352133-5948-4b17-8b7c-4bd3dd3493a4 · inbound

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models cites this paper.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.580658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.580658Z digest=sha256:5950c868d28ee284ee1526865968bc4e609eb00782082abd280cbf067790f232

Observation 340650dd-c441-4c0e-abe9-d8a31787b639 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.258608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.258608Z digest=sha256:63eae0834ced107ae2803ab0cf4179d20b8458aca0c3b40103c676664a481953

Observation 3179b8c2-6059-45e9-9987-8cb2b594a1a3 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Making LLaMA SEE and Draw with SEED Tokenizer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.136000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.136000Z digest=sha256:f56ef7392496f652cadd333f758551366d28c04feedee3d18790cedaa95fd994

Observation 78898fc8-66cf-4b12-b871-78e0c9f0cdf3 · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.789931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.789931Z digest=sha256:6daafa4c46ce530947fd5837dfcc685210a15ac70211b471564b2c8889e6d22b

Observation 06a20c3f-93ef-455b-9f24-5b6c4a46de20 · inbound

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance cites this paper.

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance Making LLaMA SEE and Draw with SEED Tokenizer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:29:52.601322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:29:52.601322Z digest=sha256:9a750e21541f3687bdb81fbdde0e6925581e7f4ab1e6692e0725ded10a7964a0

Observation 93142e31-20c6-43bc-96bf-f63d2547d81f · inbound

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer cites this paper.

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer Making LLaMA SEE and Draw with SEED Tokenizer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:13.153399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:13.153399Z digest=sha256:16ad02d8716c23dd30286ac3a7b53983073d0c8cbca3f235ccac8410b06a297f

Observation f155e2a1-a8a8-41b7-9cf0-7ce3355b4522 · inbound

IDEA-Bench: How Far are Generative Models from Professional Designing? cites this paper.

IDEA-Bench: How Far are Generative Models from Professional Designing? Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:39:28.390420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:39:28.390420Z digest=sha256:b9356a54ad89c1b6f9d4f94d1ffa43fc611a1be0bc5f038b7f4d0113505f8cbc

Observation c8b0ed51-e531-441f-914c-5fd9f51aba65 · inbound

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers cites this paper.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Making LLaMA SEE and Draw with SEED Tokenizer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.296562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.296562Z digest=sha256:773ed9670eab9abdf6e048e30403087e1696494b5931ec2b472a6c679a4300d2

Observation 50ed3788-2b69-4a11-939a-294476e8b714 · inbound

Next Patch Prediction for Autoregressive Visual Generation cites this paper.

Next Patch Prediction for Autoregressive Visual Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.952024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.952024Z digest=sha256:60dddbf56bea015ec398d3bc1e35cdc952a39bd3b54d8472eb6385dd82918953

Observation c35e9009-05ac-4cb3-b827-e2ef6d93259d · inbound

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models cites this paper.

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T06:06:07.038110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:06:07.038110Z digest=sha256:44e35ba3c0e3498142609a5e00e64909939d27a0b4dcc23583760b00cfb72cd3

Observation 643a5392-fc0c-4d17-9e0a-683554bf27d4 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Making LLaMA SEE and Draw with SEED Tokenizer

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.695742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.695742Z digest=sha256:77874c2c6a078f4a6ece09236c7ab2d4bd851e644773dd41a8cd83cb5ae10a33

Observation 200f4aed-69d5-4c0e-8afd-77519d92a790 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Making LLaMA SEE and Draw with SEED Tokenizer

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.798294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.798294Z digest=sha256:faa4b320bcfb2f975b8d1d0592409857b1a2ac87858ab2a9489017b545f2707e

Observation 4f9696d8-933d-4d91-8b84-0d05be0fc9a4 · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs Making LLaMA SEE and Draw with SEED Tokenizer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.899573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.899573Z digest=sha256:4534fc73d29dce93d2c26d5279f52a4cb2e2ae96f553379f63c8634c1246b68a

Observation 063107a9-8985-4b3a-ad3b-6b88fed99714 · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Making LLaMA SEE and Draw with SEED Tokenizer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.850621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:baaadabadf462b5d2e80a41e5cf19630bcf3258ba460bc55141ba7d9d172bb7d

Observation e25a5b84-4a3e-43ed-ba33-4c57a1d4e3e7 · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.714439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.714439Z digest=sha256:a01265b6b70610d00d12e6b0fce68c8da99968696975a5c61ca5cdd4bef36593

Observation 644111f9-966a-420c-b84c-94abc0b65c1e · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.826174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:1adcd2b449e72099bbb84155571209a1b5bad8e672b877a8369b09db8d023e14

Observation a2f59b40-9cad-4d1e-be7b-b3962e0fecb5 · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM Making LLaMA SEE and Draw with SEED Tokenizer

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:10:56.112562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:738b41464085f84fe4b23fd390bd09ff0d299069ebc30180cbf1ec50096a604d

Observation 11016a83-596a-4a6e-bee4-aad63d187c44 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Making LLaMA SEE and Draw with SEED Tokenizer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:54.366169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:54.366169Z digest=sha256:d1f6331b99ecee737a2c687e590005f52351b9c4331bb953ec45f2a2e0e5151f

Observation cda3c8cc-5278-4f38-8deb-67b5537f8dae · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Making LLaMA SEE and Draw with SEED Tokenizer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:55.995559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:551e33636b32c473c3c6fa155fec3c27b06200b0dfa95de0185aed1e0c43d0a5

Observation 24ac5404-754b-4718-b8f3-717d54631f7d · inbound

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models cites this paper.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.140712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.140712Z digest=sha256:c2b914c3fec3454c93cf331fb32260b9edd6cdaad7953e3e3c12d94a11c8bbd3

Observation d3487a38-dc62-4b86-bca4-cf93de2dbdb1 · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.502691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.502691Z digest=sha256:29f40716cab501a115ddfa1b2f299fe8695b4f6ad1157b79600872f49df24c63

Observation 87f0b4e0-468c-4b1e-8e6f-67262c9afe7f · inbound

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations cites this paper.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Making LLaMA SEE and Draw with SEED Tokenizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.655820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.655820Z digest=sha256:91be5fc61906000e3a76c768f8e7a8053fce270c08c8ca45fb04e7d2cf87c9f3

Observation 05fce673-ba5d-4366-9c62-e29eeb872097 · inbound

IGD: Instructional Graphic Design with Multimodal Layer Generation cites this paper.

IGD: Instructional Graphic Design with Multimodal Layer Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:41.979247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:41.979247Z digest=sha256:efacb768370c627160d3bceaf89b05f66a0da874a3bca81c162f0b3ad4b2d99a

Observation 0035c43f-fdc4-476a-b219-aceb791b2b21 · inbound

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models cites this paper.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.726451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.726451Z digest=sha256:0c5a79a1c7574a2738af47d851a5a13d142135cf993ae04450cfce119e45a2f0

Observation a97b6fa1-1915-4ead-9d96-76dff61ce866 · inbound

TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning cites this paper.

TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning Making LLaMA SEE and Draw with SEED Tokenizer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:41:11.885994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:41:11.885994Z digest=sha256:89e7e7c3f877be30d98de60055eca3d4c773600013f9223e1675208a7dcbe9a5

Observation 27c96fe5-b1de-46d0-b399-c6852e091228 · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.412176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.412176Z digest=sha256:74c832dabe8830bc96d0d3bcb62520db28e9b78e36b2bdf2697071c27fa94dc2

Observation 94b9f1ac-59f5-48b0-99f9-c221d39aee55 · inbound

UniECG: Understanding and Generating ECG in One Unified Model cites this paper.

UniECG: Understanding and Generating ECG in One Unified Model Making LLaMA SEE and Draw with SEED Tokenizer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T15:45:10.086435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:45:10.086435Z digest=sha256:c2a3ff82cb759c2f4c8154d9e6fd02f40d01a1a232d43929b80f9f8904c376ac

Observation 05cd402b-5186-48e7-9b3e-bded583af4b3 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:31.361767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:31.361767Z digest=sha256:ed6ab14e4b1dd7f6934d107abf221c701eaabc7723e48f4c3f134b12cc90dd07

Observation d778adb8-b3db-4394-b248-abdbd10180b3 · inbound

End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer cites this paper.

End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:05.964626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T19:41:03.302303Z digest=sha256:74c0f658aa8d749637a0b8fee999991c97060bde9a8cea69dcea6bb1e437234a

Observation f43ffe58-d574-4613-868b-02f1b8e1feb9 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Making LLaMA SEE and Draw with SEED Tokenizer

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.821835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:011536da93c32a145ebb38a1d5f2fc0fe3e79821c2e34e17e17d0b7266406cfe

Observation 186a5098-c54c-4649-8fac-6b9f4cbc284c · inbound

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing cites this paper.

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing Making LLaMA SEE and Draw with SEED Tokenizer

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:13.122612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:08:57.792229Z digest=sha256:4d9383805209138ac43a4caaad275f0ea5cba014d5d3ff7377fb480958438f5d

Observation 7805d954-f801-4294-a566-f745f9568dd5 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Making LLaMA SEE and Draw with SEED Tokenizer

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.919623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:074eb9e2e304fc49ca673c3264f5d5b73d86243e593e9424c4eaafa74a945396

Observation 8a32fd32-f805-420a-a075-5372a1a9d8d8 · inbound

InterleaveThinker: Reinforcing Agentic Interleaved Generation cites this paper.

InterleaveThinker: Reinforcing Agentic Interleaved Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.891323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:42:34.126336Z digest=sha256:9af3255f7d8fa34c4c8f4da30a19c32c4b3501a9a3945eec365d75f201685ff0

Observation f42d80b3-3afe-45bd-8138-7db669b41d8f · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Making LLaMA SEE and Draw with SEED Tokenizer

Reference 253

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.766332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:f99a13f5a305f5db9b2eff6fb6c70c3592def9855090e6ac056facfc7abaee74

Observation 1208ccc4-baff-48ac-b94c-2a2fb21b9c1a · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Making LLaMA SEE and Draw with SEED Tokenizer

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.718833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:3f9341c8308e1fce29440e7c3527aa3ea3f70bcb36d01943f0ca0b30393fe190

Observation a4de679e-c1e1-4667-ac3b-b459f967aebe · inbound

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation cites this paper.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:20.951281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T06:04:07.327934Z digest=sha256:914b1fd6eaeff86eb237333479a98a741190c1cd3be6784a57a81acdf6ec5089

Observation 7c8cef05-b7ad-465a-a62c-34c53062e132 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space Making LLaMA SEE and Draw with SEED Tokenizer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:3843d297fdcd8ce4a01f3a813a32c4cf28a37f56f6fc6072a47d3afa218f7aff

Observation a434d4f3-337d-46d7-978f-8e0dea4dad31 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Making LLaMA SEE and Draw with SEED Tokenizer

Reference 130

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:89d2c40922dba17017127d223d47125ca7affa8e4b970916e99d7c0f752c7097

Observation 59d22b61-4684-47c4-b9ba-0d776255867f · inbound

Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning cites this paper.

Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning Making LLaMA SEE and Draw with SEED Tokenizer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:16:29.589913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-09T19:57:27.683657Z digest=sha256:94aa694f556371a90a7b69b661beca923fcb6bd628db5c9b42966522d5b5c8e9

Observation 7ff494b4-408f-4432-80d2-6df5a068c518 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Making LLaMA SEE and Draw with SEED Tokenizer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:48.538845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:48.538845Z digest=sha256:6e50a6f62702e743068aa8df3ff2c1dff1d6b28033b2090c84c863185c92be0e