Pith. sign in

Paper Citation Record · LEDGER

Retrieval-Augmented Multimodal Language Modeling

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2211.12561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.12561 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:10:15.282937Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

29
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00b934be-6f7a-4749-956f-f7cd406540f9 · inbound

REPLUG: Retrieval-Augmented Black-Box Language Models cites this paper.

REPLUG: Retrieval-Augmented Black-Box Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:41:54.147344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:41:53.833754Z digest=sha256:4e05beafe193474509570747a7cab4a590d2fcfeeb7092bbb6ba3be004cc6f15

Observation 9add77e6-99b8-4d24-8f4c-05972f515c61 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:32:22.906269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:211ad4595871ef31488c24e0829e91685bfbdfc03e04300bb9905e42166b74d6

Observation 437f27d8-a5f0-49ff-8856-32d0ac7c6fee · inbound

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection cites this paper.

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection Retrieval-Augmented Multimodal Language Modeling

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T14:15:11.295731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T14:15:10.907921Z digest=sha256:0a50dc9929d11b5d78553e3d1db5a3eb34a44732ce1b521ebb69bdf996aed3da

Observation 00e2512f-96ed-4a1b-84a1-db95c39edd01 · inbound

Retrieval-Augmented Generation for Large Language Models: A Survey cites this paper.

Retrieval-Augmented Generation for Large Language Models: A Survey Retrieval-Augmented Multimodal Language Modeling

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:13:56.483980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T05:10:25.171044Z digest=sha256:fb4b42ff9432454c882cfba544efd68c01eb73680b3c255ebd385ae7299233cd

Observation a442456b-ee74-40b2-b22b-aa57346c2b3b · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.381184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:bca0c4aeff42b929e8a42819e3e1a9bafd933ba27d41f745319ee6b17e130450

Observation 5eebd766-39dd-4661-b414-7463e34fe5e1 · inbound

Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge cites this paper.

Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge Retrieval-Augmented Multimodal Language Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T13:10:15.282937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:10:15.282937Z digest=sha256:3b070b9882cc98ba9d083139246d683beefd34f9e32d43cc1478af1b63b1270f

Observation 93cc96aa-2d47-47f9-b54f-f2071ded5392 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Retrieval-Augmented Multimodal Language Modeling

Reference 212

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:52:16.195672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:a9e9bc62f4c61cdd695db8cd38fa9fa89e12aa4cc031359259e2d918af01ebbd

Observation 75c91901-3116-4ee2-adfc-63f1f8c7620b · inbound

Demystifying the Visual Quality Paradox in Multimodal Large Language Models cites this paper.

Demystifying the Visual Quality Paradox in Multimodal Large Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:56:38.297097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:56:38.297097Z digest=sha256:78a8ecea02156b76df0702432f6044cb92a27438af8c3eecd44e976db8353eb0

Observation 261c86f1-9a95-48dc-814c-d53994a0cc89 · inbound

Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering cites this paper.

Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering Retrieval-Augmented Multimodal Language Modeling

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.405681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.405681Z digest=sha256:6a651b6e25b39553f36ee5d210f905e30750eb46e44a5c9d9de18f1d2770ff80

Observation 9aac474c-5315-4279-94fa-81bbd9350349 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Retrieval-Augmented Multimodal Language Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.457752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.457752Z digest=sha256:0f5955692b10452c446bd78def63cab26715e0efb9ef05c0753ecbd94ef40eb0

Observation 5a0c5e72-11f4-45ec-996a-10d0e88c35af · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Retrieval-Augmented Multimodal Language Modeling

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:53.705938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:53.705938Z digest=sha256:89d763f8d4517dd8c6bdc29bff73556d707783d3dd8e7add8ae15ae2048d63da

Observation e5830b19-0053-4593-a880-68bf53e36a20 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval Retrieval-Augmented Multimodal Language Modeling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.322419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:5c50579c20366ec5024ffdc68fff77dac24beb117728180d769cc7ff8720e006

Observation 2811251e-3007-413d-949f-7a9332fa0d32 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Retrieval-Augmented Multimodal Language Modeling

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:22.630225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:22bbe17efc3f21eb1516a275d8a919515dbe3940b95298ae72934c1ef6f7ace3

Observation 0685a0b1-6834-4025-91b9-7be507767028 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Retrieval-Augmented Multimodal Language Modeling

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.598864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:21b3f390bce259a6cb3feafe9a2c9bba8ad55527c3e9ea6c4448f5b3b9e12a1b

Observation a1f8f8e3-7f02-4325-8f2a-2175f202321a · inbound

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning cites this paper.

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning Retrieval-Augmented Multimodal Language Modeling

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:08:24.781453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T15:06:53.377360Z digest=sha256:44cc2f12c1e9566cde638a63a9dcd717143fef8505e27594589aa6592a229dd7

Observation 432ce052-a6d2-4896-b08c-5b9c4fc652ee · inbound

Qiskit Code Migration with LLMs cites this paper.

Qiskit Code Migration with LLMs Retrieval-Augmented Multimodal Language Modeling

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-26T16:29:35.554353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T16:24:25.357338Z digest=sha256:9d467fac3fea350c51418b60764090b63d876eed9eacfb168a3c826c327673a4

Observation 448d343d-28bb-44cb-a1e1-046f9f495a9b · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity Retrieval-Augmented Multimodal Language Modeling

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:07.797148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:e59d26c51f872e2ff2d115b620cce26b47fb84276531f6bfa305f705440f577c

Observation e3bcc432-69a3-4971-ba94-5b99f908cc33 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity Retrieval-Augmented Multimodal Language Modeling

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.238082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:46b3720bc0bfce7e951e4967bba87430a507b9f346cd13d746951bfb96ed7600

Observation 792a6c93-babc-49f8-9a85-20b537027144 · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Retrieval-Augmented Multimodal Language Modeling

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:51.071227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:75ef984a9746aff17977f8f961030740911a5931e632dfeb165d5aab7ce1cea8

Observation 8e01de3e-ee56-47bc-96a1-57137851a3d4 · inbound

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models cites this paper.

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T15:14:25.526254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:14:25.526254Z digest=sha256:488578831c5e1e5ae8b1aeb72f52b042d473f21984cb574256c99cf9e376b18f

Observation 870d4085-1188-4380-8f83-5111b41c2d59 · inbound

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models cites this paper.

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T06:57:34.446788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:57:34.446788Z digest=sha256:8279a88fc8737d8f9ab1ff3bb271828dc98e7abcc4bb60ff5f4ea1b89bfb0548