Pith. sign in

Paper Citation Record · LEDGER

What Makes for Good Visual Tokenizers for Large Language Models?

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2305.12223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.12223 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:29:13.937112Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.837268Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7e37aa49-1f2b-4ede-b436-88c53bcdd7a2 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension What Makes for Good Visual Tokenizers for Large Language Models?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.579186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:08e08866da8e46882db75e3ed67e6a45e358b1f11737fc0a06aec27fd7c8c937

Observation e202018b-0e6c-4e9d-ae3a-b16738d7d2ee · inbound

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts cites this paper.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts What Makes for Good Visual Tokenizers for Large Language Models?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.685067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.685067Z digest=sha256:5f6f6dd26bc1ffcd357664c311cfb549d7820c74116d92f2b8098e3032e321eb

Observation 871e3fed-c550-46b7-bea4-775d98c30b7b · inbound

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models cites this paper.

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models What Makes for Good Visual Tokenizers for Large Language Models?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:25:25.620936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:25:25.620936Z digest=sha256:450cb14f5e3c05a6df7f98322da1d7f294f225cf14c66e3988ae3bcd1ee3eab1

Observation 5f8b5b5d-3e08-4d43-bdc5-0879d0161b8f · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor What Makes for Good Visual Tokenizers for Large Language Models?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.675146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.675146Z digest=sha256:b58c056666cee97463ef8dc64d8ab848abee521e132f9618bacf26f3937da981

Observation 8f933dd9-63bb-44cd-a0db-e6b8c29d1cb3 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models What Makes for Good Visual Tokenizers for Large Language Models?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.641935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.641935Z digest=sha256:1402c208e281da6f3bba6a0d9ca39aba4c4a43978d00fc7dc34456722c105cca

Observation 377aaa41-1de1-4f2e-afde-2aef9fdc42b4 · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs What Makes for Good Visual Tokenizers for Large Language Models?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.544321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.544321Z digest=sha256:1ea006fdf9546122049152111dc76c661c8e817850c283753ed2e755f0ae847f

Observation 95bf9e2a-1197-4922-86a7-0f49d37c7bea · inbound

AGI Is Coming... Right After AI Learns to Play Wordle cites this paper.

AGI Is Coming... Right After AI Learns to Play Wordle What Makes for Good Visual Tokenizers for Large Language Models?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:13.937112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:29:13.937112Z digest=sha256:54883e9d9cf7495346e8a253ba712bc8f84b2db21ebe30354fccdafc39640422

Observation b25a914d-23cd-437d-99d5-6b27d2e2e249 · inbound

MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models cites this paper.

MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models What Makes for Good Visual Tokenizers for Large Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:22:01.656664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:22:01.656664Z digest=sha256:331b1636c27d7394b327af5875467017a322183312613c81caca6b837a4c6bec

Observation ce548559-a3d4-4656-aa8f-3527fdad1767 · inbound

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation cites this paper.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation What Makes for Good Visual Tokenizers for Large Language Models?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.633448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.633448Z digest=sha256:8b38274082f125e5378cd85f82de3e661deee83abc6204b4ec5edab998f38168

Observation b199ac2d-0a86-4f1d-9bec-73bebedc31fd · inbound

ACTLLM: Action Consistency Tuned Large Language Model cites this paper.

ACTLLM: Action Consistency Tuned Large Language Model What Makes for Good Visual Tokenizers for Large Language Models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:09.441859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:09.441859Z digest=sha256:ffaafec28ac69922921e76efbf82f536596ac8837b77ee72edaddf337634e787

Observation 406c7cb7-1037-4bde-943f-acbb0c6c3085 · inbound

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis cites this paper.

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis What Makes for Good Visual Tokenizers for Large Language Models?

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:48.242753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T19:22:08.306946Z digest=sha256:92ecd949a7b3e2f7a62036ad532e7a2f9768306d76f56e188b4171b5c668fd87

Observation 748644a9-708b-4884-b0e9-6cceb8ce6bd6 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation What Makes for Good Visual Tokenizers for Large Language Models?

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.838650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:67f0c3e45253071cefa8bf87d95c39571ddb6dc304b8edd8c03f4baabbabb6f4

Observation ca9d1c46-48c3-4387-8f08-78f5a952be0f · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report What Makes for Good Visual Tokenizers for Large Language Models?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:4a7bba0aeb1c65e5308f9032f73d10d6064e63497e4e9effc42ffe3628c9000b