Pith. sign in

Paper Citation Record · LEDGER

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes

As of 16 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.11396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11396 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:02:23.305959Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c44423f8-885c-44e4-9be8-222d7bb39dc7 · outbound

This paper cites In: Gold- berg, Y., Kozareva, Z., Zhang, Y.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Gold- berg, Y., Kozareva, Z., Zhang, Y

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.191882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.191882Z digest=sha256:fecf1c57902ff274517013fa528680dd12acbd3ad0932fdb0b25877209cd95e9

Observation 8df79f80-f676-46c1-8f88-4f1c50adc7f9 · outbound

This paper cites Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T15:02:23.425177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.197396Z digest=sha256:f9ce716a73f5cb5384339702a805d3691e60294da38c969ffb16619c241e08cc

Observation 79725cf8-9740-479a-9ea1-60e64f91eb80 · outbound

This paper cites In: Krause, A., Brunskill, E., Cho, K., Enge lhardt, B., Sabato, S., Scarlett, J.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Krause, A., Brunskill, E., Cho, K., Enge lhardt, B., Sabato, S., Scarlett, J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.753606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.202468Z digest=sha256:09234c784221d9153c2784c183147b9ce52a39a48a9501d539c2163525a8c692

Observation c8a35adb-04b1-4a50-a56b-dacada09c633 · outbound

This paper cites In: Proceedings of the 17th Conference of the Eu ropean Chapter of the Association for Computational Linguistics.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Proceedings of the 17th Conference of the Eu ropean Chapter of the Association for Computational Linguistics

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.733656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.207057Z digest=sha256:adbf286eeb2ac7f4db22db02ec5e1853ee146612af842ec923c96b55a1903bc7

Observation 1750bb6b-8846-4c6b-af9b-a4d987990dcc · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.211881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.211881Z digest=sha256:fa2894d370a1d8496d9daf6056191046bb6838317c440551f22bf7fd53904528

Observation 1ae961f3-1ee4-435c-b66a-1bb38611e1e2 · outbound

This paper cites In: Findings of the Association for Comput ational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11- 16, 2024.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Association for Comput ational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11- 16, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.716292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.217477Z digest=sha256:ad2a967a5d1750f714a5910fb7f2a53fb5a698d5214583e09858a12fec681923

Observation 525537a2-f56c-40d0-a742-1c6a16bfa4f5 · outbound

This paper cites In: Findings of the Association for Computational Ling uistics: EACL 2023.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Association for Computational Ling uistics: EACL 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.699918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.223857Z digest=sha256:7d201291d1355894aa8f1a5b8e6e5235b1612ca3f8ee1fecc68c684ad99e1ec3

Observation e0c1e7c9-a6ee-4500-843e-89878219bd55 · outbound

This paper cites In: Findings of the Associ ation for Computational Linguistics: ACL 2023.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Associ ation for Computational Linguistics: ACL 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.229370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.229370Z digest=sha256:0f730b765b9db24d46a61da69910b458460c7551938c2849d7c4cf9522889ac5

Observation 191b4070-d104-4d94-8f3c-489bb9b5f391 · outbound

This paper cites In : Proceedings of the AAAI Conference on Artificial Intelligence.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In : Proceedings of the AAAI Conference on Artificial Intelligence

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.234520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.234520Z digest=sha256:23fe0de7171c3f7795ce329573c56e817d3b3526e447cd85b3f1f8705877f65f

Observation 59d8c182-bd55-4414-8dea-36d1cc1d3839 · outbound

This paper cites In: International confer- ence on machine learning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: International confer- ence on machine learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.665352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.239294Z digest=sha256:828436b88045b431e535e00d99cb107551268acabc02ddfa1a190b829990cce2

Observation 9b64a958-ae5d-4437-b526-841b1381d5cc · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.245449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.245449Z digest=sha256:79f9c2ca1fa982768779673b29f3ca5b57937b1f527619111c524a99cbeedb6f

Observation 4996202a-9291-4e88-9038-0737510d766d · outbound

This paper cites In: Oh, A., Naumann, T., Glo berson, A., Saenko, K., Hardt, M., Levine, S.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Oh, A., Naumann, T., Glo berson, A., Saenko, K., Hardt, M., Levine, S

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.650212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.252442Z digest=sha256:6513d008198de581e2b3644f450fa8477e582eef878ad081e8d460414f26d861

Observation fc6484dd-47e6-41ce-ad46-f7132074b2b4 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.256879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.256879Z digest=sha256:1c016952dbefc4360a2d543884dd530c1b6f02a2afccb99de4f36ae39233b24f

Observation ccd8abb0-c467-4232-8c81-21890308095f · outbound

This paper cites In: Pro- ceedings of the 2021 Conference of the North American Chapte r of the Associa- tion for Computational Linguistics: Human Language Techno logies.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Pro- ceedings of the 2021 Conference of the North American Chapte r of the Associa- tion for Computational Linguistics: Human Language Techno logies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.262416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.262416Z digest=sha256:77b432248889e13a841ab4ec08dabc77d4229855741abfa38cc5f5d4a32db7f6

Observation 1eaf28fd-90cc-4403-befc-67767904cdd4 · outbound

This paper cites : Modeling event- pair relations in external knowledge graphs for script reas oning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes : Modeling event- pair relations in external knowledge graphs for script reas oning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.267112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.267112Z digest=sha256:f5de2ee5a18c10291be7ec31f11ce90cf8281add6f678d99048dfffaaffa129b

Observation 9f4c610b-c044-42b6-a754-0d0ffa837559 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.270947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.270947Z digest=sha256:367cbc7e4b6870105e54d4f98a87adbd262d569c8f26ff5b42c599e804862b5b

Observation ee4e5f92-be7c-4ccd-8e9f-f9154fdc5a6e · outbound

This paper cites In: Cai, J., Kankanhalli, M.S., Prabhakaran, B., Boll, S., Subramanian, R., Zheng, L., Singh, V.K., César , P., Xie, L., Xu, D.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Cai, J., Kankanhalli, M.S., Prabhakaran, B., Boll, S., Subramanian, R., Zheng, L., Singh, V.K., César , P., Xie, L., Xu, D

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.615837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.275985Z digest=sha256:a77c82b0e06ba6d3e1be8a5a28cd0e1ca981bbd252819eb4d4dcf4ea18b3a618

Observation 7a0b849a-5ef0-475b-85ba-f424723da061 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.285580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.285580Z digest=sha256:e62be8f3feb34935c9fb0ee3232e663c41d067774ef529a244a10fa0a9907fc8

Observation 14ef6d89-a516-4e00-8e59-0c5756570613 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.290415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.290415Z digest=sha256:845e3ce8ee7da5d2461a681ada7f99e56828e2695985980c7bcfafbbd1f325f3

Observation a44c310a-daa2-46a1-a8bd-10804d827395 · outbound

This paper cites Meta Knowledge for Retrieval Augmented Large Language Models.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Meta Knowledge for Retrieval Augmented Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.295039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.295039Z digest=sha256:4fcac882c1bebdd50a6f4c54671b6acdf927f9524fdc016f6f1da05950366c57

Observation e7083dd3-a474-4945-a193-6a637e572be6 · outbound

This paper cites In: Wooldridge, M.J., Dy, J.G., Natarajan, S.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Wooldridge, M.J., Dy, J.G., Natarajan, S

Reference 21

Resolution
verified exact
doi, observed 2026-08-11T15:02:23.595819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:02:23.305959Z digest=sha256:de93caa5d0f3bfa7a33d1be6a586e47b29f034bf6bf449fc27d01b0ebe134f13

Observation 357f6e0a-625a-4afa-b30c-f72a3bf62267 · outbound

This paper cites 8845–8854.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes 8845–8854

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.281390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.281390Z digest=sha256:d75886a11047082245253ebcfb139a0589178149f752f9858ada8fff098dfd5e

Pith citing papers

No inbound Pith citation observations are available.