Pith. sign in

Paper Citation Record · LEDGER

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

As of 14 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.22074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22074 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:04:59.948955Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eed6f0b8-80c6-4573-b526-a83907fc9470 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.854687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.854687Z digest=sha256:e6b93deb319309ae5a087b0070185a738655ef566558fd3b800bb1f035eb4525

Observation a7e6307e-cf6c-4465-a6c8-7cd9e00a4038 · outbound

This paper cites Visual in-context learning for large vision-language models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual in-context learning for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.859067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.859067Z digest=sha256:25edd9a0e844dc74b665d4f8eae265d333945327ab4cd5e5ca071ce370ff7e45

Observation daf19fc4-d5db-48f6-9db2-a38e434b1910 · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.862849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.862849Z digest=sha256:687f48192e9f8ac5a7fee4cef493608a2a8b2ef11be0e79bfcd33deadfeb5c61

Observation e6de482d-ebe9-4236-9aca-1f6b4e645aba · outbound

This paper cites ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.867055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.867055Z digest=sha256:f44b6167685a90632c138b708441848ed90da156919664e6d1144ae31e8aafc2

Observation 84520e5a-e62f-4d17-8039-a4c51d0e088a · outbound

This paper cites Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.870622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.870622Z digest=sha256:cdead6462a62fbc989d7ac6678ead7f487e348b9d152a6267fbc9dcf0523b60b

Observation 76e8ea03-cdaf-457b-8e0d-efae0b1fc29e · outbound

This paper cites What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs What makes for good visual instructions? synthesizing complex visual reasoning instructions for visual instruc- tion tuning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.343521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.874506Z digest=sha256:3e2cc184c339e17da11c3c7d002dd17c2a576f2766b7b52012d0fa01b875ed75

Observation aece26ed-97b7-41bb-a991-f1d456d6b06c · outbound

This paper cites A review of multi-modal large language and vision models,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A review of multi-modal large language and vision models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.331761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.878977Z digest=sha256:8d46163f1967faefb36e72df21e11c34509097d2e5a18e0b72845ad619cac7e4

Observation 84fc1e5c-0a1e-4554-90ce-54a2f1285286 · outbound

This paper cites Visual instruction tuning towards general-purpose multimodal model: A survey,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual instruction tuning towards general-purpose multimodal model: A survey,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.319186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.883186Z digest=sha256:fbd1efb6fac8352acab5d0b0c4dddb429c7e3b51487c54c64ecb5cf928d7ff21

Observation 8ed50742-4e36-44c9-b86f-ad6ce0f8ac30 · outbound

This paper cites Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Visual question answering instruction: Unlocking multimodal large language model to domain- specific visual multitasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.306541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.887095Z digest=sha256:c42fd2f9c258be51aef18ff95bbd11b6a0ee56ffcfc1e6f15797a690fc871b58

Observation 4520a800-44e8-4f19-8a52-ca216c3f6b64 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Sharegpt4v: Improving large multi-modal models with better captions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.294680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.890472Z digest=sha256:3adac3d3a92ca3add66769c4874add341a575d73a0e38f21f4a2f84f9890e4ce

Observation 1103b626-32c0-4f45-913b-80a277aa2a3c · outbound

This paper cites Multi-modal large language models are effective vision learners,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Multi-modal large language models are effective vision learners,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.894097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.894097Z digest=sha256:6af57d29f2319ea57c1eaf5ba73c8e3e1e27c9167a677dd6060545d9e67c896b

Observation c8b58121-f346-4249-8102-92961db53fe3 · outbound

This paper cites Modeling event-pair relations in external knowledge graphs for script reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Modeling event-pair relations in external knowledge graphs for script reasoning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.897499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.897499Z digest=sha256:d9cb870d59042a1c324d818c4657e7fd8b17bd4b540f6b44f2b76df927f235a9

Observation 06954d07-b834-4b41-a2bc-79ff76cfe250 · outbound

This paper cites Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Claret: Pre-training a correlation-aware context-to-event transformer for event-centric gener- ation and classification,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.900566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.900566Z digest=sha256:6026f4e84e6bd01038849d701ec7e4ad0b675b76c72ca55a65e1c0bb07ae96a2

Observation f9d596e8-1976-4f7d-b1a8-6ad16e1508e8 · outbound

This paper cites Eventbert: A pre- trained model for event correlation reasoning,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Eventbert: A pre- trained model for event correlation reasoning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.903623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.903623Z digest=sha256:e2106beb2fb1719436d7bc7721785fe7a775f8ce65de7107a9efaeac8375bbf9

Observation c9e8263d-7004-46bd-b587-ee7ad934f3ec · outbound

This paper cites UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs UNIFIED- IO: A unified model for vision, language, and multi-modal tasks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.249670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.906812Z digest=sha256:3c2d9f5f98c1b3863e5d7fa2e3b5f66437dc2dc26ef36f3d0b1239471061c04c

Observation aff49c02-5ff6-4a94-981a-7280134854c4 · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.910308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.910308Z digest=sha256:78bf5a37490c53d39415b2f8287ebdd10237f12de64dab3ac859fe8e36144fa2

Observation fb191429-c155-4b43-a552-7a70ab5fff99 · outbound

This paper cites Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.238479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.913550Z digest=sha256:5227321a2cc9c070763551c906aad41bdd521b811beb8e2ad5086b24f206dfd3

Observation 99b3907a-3287-4daa-9672-4f4f4f0bcb4f · outbound

This paper cites Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Au- tomatically correcting large language models: Surveying the landscape of diverse self-correction strategies,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.213765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.920890Z digest=sha256:4969c9aede3c299a286e479deca5f610a9d9b71635d615c1d88e418a1e195e46

Observation a92279bc-8044-4e9f-8ccd-e3675463f366 · outbound

This paper cites Adareasoner: Adaptive reasoning enables more flexible thinking,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Adareasoner: Adaptive reasoning enables more flexible thinking,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.199911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.924161Z digest=sha256:92801182aa1d3d5ea58760f07fd61be25034d3419acc564d1af2da6cb2019db2

Observation ce37a8b1-cc62-422b-8209-403a76c8601c · outbound

This paper cites Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Plan-on-graph: Self-correcting adaptive planning of large language model on knowledge graphs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.187066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.927819Z digest=sha256:6a5ca8b5c3ce45f513eee4762d8f0c48944de4abfcba941c61a361952660607f

Observation d3d7508e-0496-4981-9938-8fcff68440a5 · outbound

This paper cites Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.176327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.931589Z digest=sha256:186965cf05ee65a90ffa104192d2ba4641279c7fb171d88a95487ecb677f540f

Observation a7b3811f-b8fb-4bc5-911d-e094928adcc5 · outbound

This paper cites Agent-r: Training language model agents to reflect via iterative self-training,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Agent-r: Training language model agents to reflect via iterative self-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.165443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.935334Z digest=sha256:5322c1ad3fd09254fc050f2632ce973121cf7f75bab36867fa9cf5818d293b09

Observation 2febc3bb-9713-491c-8298-99b41b4f4f2a · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.938273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.938273Z digest=sha256:0dbe52877b70c61c5dc2b12bcfb81113da5e37b466e5a0d8af85abb272481d78

Observation 5392fb20-1382-4980-9e64-c6eacd653503 · outbound

This paper cites A theoretical under- standing of self-correction through in-context alignment,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs A theoretical under- standing of self-correction through in-context alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.153345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.941550Z digest=sha256:32692b044fe179b45415b491fc5a4b6a1dff5945ee1575b11625d6d45a84bfa9

Observation 45c89ec6-7aed-4a28-8da7-873179191f71 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:59.945292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:59.945292Z digest=sha256:575afb9747f1019d7e45327b278ca5551866b59b05819af603397709c2497f77

Observation bcd54209-dbad-47dc-ba69-c710818bff93 · outbound

This paper cites Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.142180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.948955Z digest=sha256:235adffd9ac339df6d85d9a80eaca291c1f0763ba05e57cfcac38d74e64ecdd6

Observation 54312aed-0e1b-4ac1-a559-fe6981de1f74 · outbound

This paper cites 5014–5035.

CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs 5014–5035

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:05:00.225719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:04:59.917050Z digest=sha256:e1ae14d5a2368652b7075b0fc4aafef16bfa1dc3d14e1a22ab81a018fbf613db

Pith citing papers

No inbound Pith citation observations are available.