Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:01.238326Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2501.02669.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:01.238326Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.191067Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:53:16.341543Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d9b929af-00a5-4a12-a07f-ffbbe87e0c48 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? shape type), the model first needs to correctly enumerate the attribute values (e.g
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4738033a-e60a-4a35-956d-6ed14f3af148 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 091f7cb5-6e71-487f-9f52-4ced86728735 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? backtracking
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e50b250-444e-4e1c-bf56-e893c2fd84f8 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c348bf4-4c2b-4212-a0e3-77dd9ac44850 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c049058d-af3d-4975-a6dd-142ba2e4f579 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b1f6d2a1-d1f7-456c-823d-73939131dfb8 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Are Large-Language Models Graph Algorithmic Reasoners?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3404e2f-cd44-4751-b728-0c2a890cb8ae · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978db19d-91f2-44da-93e8-ee725fb2899f · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? single-hop
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08ac75a3-4241-45ca-a8be-d40be35b0bad · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 06843895-1faa-4f65-bca3-c48f112ca72d · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1eb820e9-c6d1-4b40-9924-28cc28c105b4 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b419063-6839-4715-8156-031645c3f018 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7ae0808f-cb8f-4690-a4f4-9fc0568f3f2c · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff329e0e-2a4a-4a88-bc39-a405a7262fdb · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6309a005-e4a2-45e2-aa3e-7b52bb1c3aaa · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68bb3852-c4eb-48c8-ad6d-d2dfe5c2ff34 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 233 Figure 33.AHARDexample fromTable Readout
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 353eb138-7832-49d5-9b38-081d30de6fb4 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bb8675ab-8c9c-4b48-91ce-fa88e39e2ca1 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f517d880-22c6-4f55-8aad-13066805e3a0 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93957649-5c7b-4775-8d6c-c13322deba03 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 244b66e0-b917-4c17-9f2c-6a9a29e89968 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b3021570-a7b2-4bce-8e26-1ffd20f0dd7f · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff24848f-4171-4f07-85ec-70295e6d4e8f · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6d5e29a-2eca-46f5-a933-a7fba6130d87 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fd0a3a0-7f46-41c1-b50e-e97941bb5a47 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bde4c299-e96c-446b-bf43-ee2ffed36c6d · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a3da1c-c5fe-4d93-8c2b-3099b57c98ba · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 576
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc9d72c-a524-41a7-89f5-4345f31dbb77 · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Convert,
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a97e119-ad7e-4962-bab7-09c74feee0eb · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? doi: 10.18653/v1/2023.acl-short.43
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a01c31-c3bc-444b-8b05-59224c67ef9c · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? On Pre-training of Multimodal Language Models Customized for Chart Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db99015a-11da-4001-9936-ae37af6886cb · outbound
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1102cc23-3b0f-46de-8988-d3a531a0f46d · inbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f397b56-197c-44e0-9ad9-cc6c9852c6d3 · inbound
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.