Pith. sign in

Paper Citation Record · LEDGER

Planting a SEED of Vision in Large Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2307.08041.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.08041 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:54.589720Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

10
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3fe3931-6c77-4fb0-ae5c-027a8e27dca6 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Planting a SEED of Vision in Large Language Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:59:50.525873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:5137da6a964fbe62a709be17c137a4bc321ecab3a46b017447c4878bf968b38a

Observation 26d03b63-a1bc-4499-97de-22bfc5ebf43f · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Planting a SEED of Vision in Large Language Model

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.025865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:86eb25e2fb820f974ee00f69a282cea4ca2ed5be040bb112d1426351a1d63732

Observation d8eff912-36a9-4665-8a80-3d3ee5cdaeb6 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Planting a SEED of Vision in Large Language Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:44:47.448796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:03d1d89936569d5b46c845466c17a82b04e17814fe6eece7b259f4538daa4147

Observation 8bc9297e-4349-43c7-891b-fde40426852c · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:48:36.350446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:81dffdf9c6ab3e51606890c7d2db40ad7ae1c3a4481dd7fb3387f9176795782c

Observation b9fe7687-ffd6-4a6c-b006-a12f68c18305 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Planting a SEED of Vision in Large Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:05:03.785210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:c6121c07efb1ac385ff2bb82776c4165c75740df7f8609b313ca85672d549aa9

Observation e61c6e61-1175-439f-9f01-4d34cc6d17e3 · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Planting a SEED of Vision in Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:26:21.394182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:57b6af4b4f2319486f38c6083403e7bebb850bfb2373d104234c921af5b36fcd

Observation 66db182b-a9ef-4077-b2f4-0460af504fe6 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Planting a SEED of Vision in Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.472580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:37cd6ee2065cb0848fcf1231360aa5a8889ea3e4106e0c2c73fe5b4de0dab3e2

Observation ef9d18dc-8bc5-4634-bde8-aa445cfd9309 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Planting a SEED of Vision in Large Language Model

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.261881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:8a217565ab50121b0dedf007ea7a09500eeba4322338b5233d5d7184bffb6b12

Observation cef15181-dc25-4773-93a6-0df94599faa0 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.400783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.400783Z digest=sha256:a4ac59f1076bc5935e4768c0b36fb4e0ac940b3a165178117259d4e9b6fe7889

Observation 7f3c46e5-a305-4c8f-8f23-2cdc7a32c81f · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries Planting a SEED of Vision in Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.162924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:f987b712430c46984370c6b5810a5675b64792bb4debadaf0d0b5572d9d9be4c

Observation 3d2fe74f-9855-4096-83cc-2da854771b37 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Planting a SEED of Vision in Large Language Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:24:04.821007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:5bf0a1643e698d01ead284e8010224916aec1ac46b8968bd37094cf2f1ac17c6

Observation 9121c834-8d28-4554-bafa-d53507f3dc9c · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Planting a SEED of Vision in Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:54.247781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:54.247781Z digest=sha256:72c590fd9e3880d89a14b86d89358859541d6f1a20100bf60dba2c9de04d9aae

Observation a314e9b2-206b-4574-adc3-c2b4999361bd · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Planting a SEED of Vision in Large Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.301701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:79c61705f887546cab5a8150efb16facd95c8813e5d8816d511ea3661bafe5ed

Observation 5244ed3a-2046-4a49-a4ee-77feb14e17cc · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation Planting a SEED of Vision in Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:02.997268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:02.997268Z digest=sha256:130c4e4cbaf0c0a647821d7fd681482f98b1015c4af3ebb0129d6b4cef937bb1

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:6ae07f6feee5e46bee11a1adb9495e0e0be93db297b79ac191d1ec28b7bdc330

Observation 81ff7b25-87d0-4a82-9adf-ca1f202ec84a · inbound

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos cites this paper.

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos Planting a SEED of Vision in Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T18:06:18.381075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:06:18.381075Z digest=sha256:3c1a2b9987077afc1a0f0419c22ec611ebd7581e781175a021df25b0ded844f6

Observation 4c277ecd-dbf2-4959-a401-7b8862c8793e · inbound

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving cites this paper.

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving Planting a SEED of Vision in Large Language Model

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:31:26.428162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T04:13:37.421188Z digest=sha256:bea569fd5a1a0bf5536db4bc299f36a38ac63e220affc5f80727f2182f852c7c

Observation adfc6a0b-763b-49c7-8317-ce7147a39521 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Planting a SEED of Vision in Large Language Model

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:22:56.374615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:098db92192b3234fce3618ab7e37742da91fafcdb7418622cb477016885bc18a

Observation 24696ebf-5e9e-48c5-9253-40d51874cfd6 · inbound

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing cites this paper.

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:13.132709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:08:57.792229Z digest=sha256:8c179f9d0ee96fb3d1d5f7e06ebf5567e370666bdd8387da9f2e9676a94cbb12

Observation 4df426a2-fa2a-467a-9f90-1df208e42540 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Planting a SEED of Vision in Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.864554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.864554Z digest=sha256:33d8d911a29068615a361a9ffc71df8de6e5ba2aaf7040a720d819e7519d8182

Observation c015a76d-6c53-43ef-b39f-fca19bfc22c3 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Planting a SEED of Vision in Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:54.589720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:54.589720Z digest=sha256:8895b6adaf83a764f33f576455a968b17e295bc20d8d2e8f4fe6f76887d32dad