Pith. sign in

Paper Citation Record · LEDGER

Grounding Language Models to Images for Multimodal Inputs and Outputs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2301.13823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.13823 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:45:25.821030Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.458730Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ebd10290-b43b-4220-b057-70a2245ad056 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:32:22.938805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:77bada991180495923286307c1f53956d163ca73c99a269b02281712bac214e2

Observation 134c41ae-7380-4469-a8c6-554e05ccd16e · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:22:03.555113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:26b18e90757a23c414e8b77000efecf7f56b5fb741dc7f23646719ff501f5326

Observation e54222f9-02b1-4684-8c54-25eb49df8d14 · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:29:11.402781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:4b18a1065d9ac506b9f328f82cbaf9a3d9e0df6ae07bebfa2864d375b1783a6a

Observation 54c671f5-65ba-4484-b26e-c8180586a7f1 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:52:01.240496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:a8eb01200a25b6009f74fc87a3749552106f61a20828ce28bb06e05744653ab9

Observation de03229d-64ee-43b3-a381-5d181f1f934a · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.589991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:b4e625ef97210412532eed3f278d721157f710be0b08ef15eb53c3ed6dcbe86c

Observation d1ab45fa-6c92-4a6b-8bba-f7e69e8d0524 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.303752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:1a9feb4bf68b29903b5b898e220991906e8d2124b62b171931b845ecf9b404a7

Observation 966822c7-3917-4dac-86a2-54ec9ae3d0f6 · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:27:51.994241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:35cadb8ab6a6193746b45b44440dc5b624bfeb78ebf04d4b1b54762107488be8

Observation 023256c2-59cf-4509-9605-53896f22f191 · inbound

Visual Language Models as Zero-Shot Deepfake Detectors cites this paper.

Visual Language Models as Zero-Shot Deepfake Detectors Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:25.821030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:45:25.821030Z digest=sha256:c309a1dd87b715a8fba75bc005e2fe938e262887527ff6f8e9920bcd3e6dde08

Observation 92d288b8-2cbd-4977-b8c1-5676db307912 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.461606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:ab3de9a5879bac21d71951010e951374161c598fd66cfc2b587348611361039b