Pith. sign in

Paper Citation Record · LEDGER

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.22074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22074 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:04:59.948955Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eed6f0b8-80c6-4573-b526-a83907fc9470 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.854687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.854687Z digest=sha256:621509fd340aacf2589e33dd9f1ba4929f92e99f1c7eb00470b200aacd3984e3

Observation a7e6307e-cf6c-4465-a6c8-7cd9e00a4038 · outbound

This paper cites Visual in-context learning for large vision-language models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual in-context learning for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.859067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.859067Z digest=sha256:ff6a445a388fd2115c166f1222ecfa3a8ef4741ef90600011df11f6863c61dc0

Observation daf19fc4-d5db-48f6-9db2-a38e434b1910 · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.862849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.862849Z digest=sha256:ec0d31b76d18a4e9c9e9e59c9b523424f412e8a3bba179b1b0102861464e217b

Observation e6de482d-ebe9-4236-9aca-1f6b4e645aba · outbound

This paper cites ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.867055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.867055Z digest=sha256:e468f64947e6db3d5cf28c422cca31b6c105531a44c60b81cedb18bd45f0de38

Observation 84520e5a-e62f-4d17-8039-a4c51d0e088a · outbound

This paper cites Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.870622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.870622Z digest=sha256:54af9ee0bfcb896ce2725b13c28eb01ff8d1cdbc3f0501a8252a5a7ace16ca90

Observation 76e8ea03-cdaf-457b-8e0d-efae0b1fc29e · outbound

This paper cites What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.343521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.874506Z digest=sha256:c76495923d1277fd7524cfb0914b75bd8703e0757668dd1582b93c780c3df621

Observation aece26ed-97b7-41bb-a991-f1d456d6b06c · outbound

This paper cites A review of multi-modal large language and vision models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A review of multi-modal large language and vision models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.331761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.878977Z digest=sha256:674a99175d6c79af4fb443cbe155663cafa69cd5983343d384fd1a92b29b3e4a

Observation 84fc1e5c-0a1e-4554-90ce-54a2f1285286 · outbound

This paper cites Visual instruction tuning towards general-purpose multimodal model: A survey,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual instruction tuning towards general-purpose multimodal model: A survey,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.319186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.883186Z digest=sha256:ea0f05703482913004257aafb66f3578d1b715918e1526098a8fbb223949160f

Observation 8ed50742-4e36-44c9-b86f-ad6ce0f8ac30 · outbound

This paper cites Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.306541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.887095Z digest=sha256:56e41776b1cd3965d77c27fb80df3f8ed26cfe51dd6052469fc6aa31829b6c9a

Observation 4520a800-44e8-4f19-8a52-ca216c3f6b64 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Sharegpt4v: Improving large multi-modal models with better captions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.294680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.890472Z digest=sha256:b5f6f2fdb491e80af50e178844531103e73a3b40d9f57fcd698aeebf7ac54689

Observation 1103b626-32c0-4f45-913b-80a277aa2a3c · outbound

This paper cites Multi-modal large language models are effective vision learners,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Multi-modal large language models are effective vision learners,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.894097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.894097Z digest=sha256:74622823d9dd83e39535e5a2eddde311b9dd67ddae127b74347b5a619ab73fa6

Observation c8b58121-f346-4249-8102-92961db53fe3 · outbound

This paper cites Modeling event-pair relations in external knowledge graphs for script reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Modeling event-pair relations in external knowledge graphs for script reasoning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.897499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.897499Z digest=sha256:40df68d5451a10c1da8b57ab87915f5f08a33fd01c14e8b6e561187c5410445c

Observation 06954d07-b834-4b41-a2bc-79ff76cfe250 · outbound

This paper cites Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.900566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.900566Z digest=sha256:7bcad8cccf24444b8af9e3a73b8cf7add642c9a76f10da8cc1c8d4aadacf8e30

Observation f9d596e8-1976-4f7d-b1a8-6ad16e1508e8 · outbound

This paper cites Eventbert: A pre- trained model for event correlation reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Eventbert: A pre- trained model for event correlation reasoning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.903623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.903623Z digest=sha256:435badee230b9c09d36e54e01780ea957b1f68ed97d439464b789ad4a8846883

Observation c9e8263d-7004-46bd-b587-ee7ad934f3ec · outbound

This paper cites UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.249670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.906812Z digest=sha256:1c20f3cb4c5663ba0c0840abacf9c13fb30bff6f5144e99087e0851def6fdef0

Observation aff49c02-5ff6-4a94-981a-7280134854c4 · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.910308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.910308Z digest=sha256:5a2a6170746dca18999d3c45a2c768393db983c7a5fe9ebe256a3f1c30728372

Observation fb191429-c155-4b43-a552-7a70ab5fff99 · outbound

This paper cites Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.238479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.913550Z digest=sha256:31221aa93f153c4784f872d4446b16b7a7e99a110f796c8b86cd86b1e6f90e04

Observation 99b3907a-3287-4daa-9672-4f4f4f0bcb4f · outbound

This paper cites Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.213765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.920890Z digest=sha256:1620ee5141babae39daac23242446a31810bbb61f6d9090bcbf82753087254d5

Observation a92279bc-8044-4e9f-8ccd-e3675463f366 · outbound

This paper cites Adareasoner: Adaptive reasoning enables more flexible thinking,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Adareasoner: Adaptive reasoning enables more flexible thinking,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.199911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.924161Z digest=sha256:f98bc7fc9c2464d27ea7058eaf1c769e0a796486782740f7b8ed05c887d73372

Observation ce37a8b1-cc62-422b-8209-403a76c8601c · outbound

This paper cites Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.187066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.927819Z digest=sha256:d3b32ed13a659ad3bd1fa803d49596f1260b1b3f7006ad2564d8ea2977899a47

Observation d3d7508e-0496-4981-9938-8fcff68440a5 · outbound

This paper cites Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.176327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.931589Z digest=sha256:9b8a3d75e16c2011da3d4efc30b3d277e8dea9bca76ea3cdc203abcf59f94c21

Observation a7b3811f-b8fb-4bc5-911d-e094928adcc5 · outbound

This paper cites Agent-r: Training language model agents to reflect via iterative self-training,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Agent-r: Training language model agents to reflect via iterative self-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.165443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.935334Z digest=sha256:9d35c3ab415996d4f9e9ccdd2f71197b207d9af91cc99f6ab342018b3798123c

Observation 2febc3bb-9713-491c-8298-99b41b4f4f2a · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.938273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.938273Z digest=sha256:1abce06207b67381fd502998fb820255eb0cbf3cb98570aec9f7039d0f5b839c

Observation 5392fb20-1382-4980-9e64-c6eacd653503 · outbound

This paper cites A theoretical under- standing of self-correction through in-context alignment,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A theoretical under- standing of self-correction through in-context alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.153345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.941550Z digest=sha256:8b8ca8fbc88133dad63e057f779dec0f2d593ef1ef10078c771fe26fa0c47167

Observation 45c89ec6-7aed-4a28-8da7-873179191f71 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.945292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.945292Z digest=sha256:ffc1bd443a1891457f29c1011678674cc6786e5415fb5362b90202c44a5ee130

Observation bcd54209-dbad-47dc-ba69-c710818bff93 · outbound

This paper cites Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.142180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.948955Z digest=sha256:28ee9158cbcd72a0786493f5bf50eaadf67c835ab36159dcd8932fded3089934

Observation 54312aed-0e1b-4ac1-a559-fe6981de1f74 · outbound

This paper cites 5014–5035.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs 5014–5035

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.225719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:04:59.917050Z digest=sha256:9a71f48382cd40d4450c384480355e81645bf751e575cb4626e1d9e4b915df6e

Pith citing papers

No inbound Pith citation observations are available.