Pith. sign in

Paper Citation Record · LEDGER

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2501.02669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02669 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:01.238326Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.191067Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:16.341543Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved16
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9b929af-00a5-4a12-a07f-ffbbe87e0c48 · outbound

This paper cites shape type), the model first needs to correctly enumerate the attribute values (e.g.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? shape type), the model first needs to correctly enumerate the attribute values (e.g

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.945341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.590633Z digest=sha256:66f29d5ca76a35ccfe5a9806d07ac148457aca7e91f4b2e5ee12b209cf3c9483

Observation 4738033a-e60a-4a35-956d-6ed14f3af148 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:04.316329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.491671Z digest=sha256:3f49fb8843ea4aa4904a0b0bc95f9a5e634013132d094a02cfa78f08aa6c0cd6

Observation 091f7cb5-6e71-487f-9f52-4ced86728735 · outbound

This paper cites backtracking.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? backtracking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.514749Z digest=sha256:f92c4de85126fa45a2ff6c7637651b195908788f2e1d142d0a94c5be4e411677

Observation 8e50b250-444e-4e1c-bf56-e893c2fd84f8 · outbound

This paper cites • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.583444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.715938Z digest=sha256:45a7ba381b60470f78eac658d4254b4210d15562163d452c5c884ddfef03bb23

Observation 6c348bf4-4c2b-4212-a0e3-77dd9ac44850 · outbound

This paper cites Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.534759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.344750Z digest=sha256:d6582a7ef58ef790ce1eb2a7612a25b7b1d9054a4dca97e13b83e02597367a8d

Observation c049058d-af3d-4975-a6dd-142ba2e4f579 · outbound

This paper cites Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.485464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.394167Z digest=sha256:1fdc38eab460e93d0897ea3fd5a2638743bbf13f8a7e18bb09cf38da1799ccc9

Observation b1f6d2a1-d1f7-456c-823d-73939131dfb8 · outbound

This paper cites Are Large-Language Models Graph Algorithmic Reasoners?.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Are Large-Language Models Graph Algorithmic Reasoners?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.399401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.399401Z digest=sha256:676290ff64bb99c6f84771708973c0c599084835d5f3eeadaaeeb2c5d8238f51

Observation e3404e2f-cd44-4751-b728-0c2a890cb8ae · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.424752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.424752Z digest=sha256:8b103a39a518a404e7658eed1193e8dad331a85bc21a3f65d14be0019f76be9b

Observation 978db19d-91f2-44da-93e8-ee725fb2899f · outbound

This paper cites single-hop.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? single-hop

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:13:04.444754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.454746Z digest=sha256:3f26ff36329a5a5a96f0999967d0decbf9222385d5b99756e68c4ea926f78cc6

Observation 08ac75a3-4241-45ca-a8be-d40be35b0bad · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.794769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.626799Z digest=sha256:922b439f221651c9f11b07b24f01d81133df9f70e868d16451537c3ff91ff9cd

Observation 06843895-1faa-4f65-bca3-c48f112ca72d · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.679614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.674277Z digest=sha256:bf3982c5db0acd4d2a59f9cd17e8ee1f0d7f52e32d0de90fe9792c5e0e816746

Observation 1eb820e9-c6d1-4b40-9924-28cc28c105b4 · outbound

This paper cites (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.474745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.754749Z digest=sha256:aa6533dd62f848e66bd7cc742b07a17843b608e1a470a6921b24f5aba71443e4

Observation 8b419063-6839-4715-8156-031645c3f018 · outbound

This paper cites 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.364749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.784808Z digest=sha256:6aeaa25ad7b6cb65f6b888c17a79a824c70a761eede103d9f9f137c2988c737a

Observation 7ae0808f-cb8f-4690-a4f4-9fc0568f3f2c · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.263094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.834746Z digest=sha256:22cb751fbe5f3ab0b35855a5b6fffd82ee8e589ecdabd3b24fdac4c8a79031bc

Observation ff329e0e-2a4a-4a88-bc39-a405a7262fdb · outbound

This paper cites Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.182008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.874880Z digest=sha256:9dcbd7e66507e77a9cbec3079933653ae6fd1f09833c416250c8ef25a2f2c497

Observation 6309a005-e4a2-45e2-aa3e-7b52bb1c3aaa · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.094820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.920852Z digest=sha256:20df18bd12174ffd75ca0a17baaab63104b1f6d18af86cb6264876f64a766a6b

Observation 68bb3852-c4eb-48c8-ad6d-d2dfe5c2ff34 · outbound

This paper cites Answer: 233 Figure 33.AHARDexample fromTable Readout.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 233 Figure 33.AHARDexample fromTable Readout

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.025587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.961279Z digest=sha256:285451057ae370a1ad91e5182eaaa919688001de6d543ff4558adbe8dcfcd7dc

Observation 353eb138-7832-49d5-9b38-081d30de6fb4 · outbound

This paper cites The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.904748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.994759Z digest=sha256:bc10e28d2af3abd9ce2b6e13147b9a3f4e199f52ea25a9dd4b4785ad3e549d91

Observation bb8675ab-8c9c-4b48-91ce-fa88e39e2ca1 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.788039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.044070Z digest=sha256:c5302658d0a47ce49b5f48e2b4b4c8dd393aec8112de343aed38cd873ba70162

Observation f517d880-22c6-4f55-8aad-13066805e3a0 · outbound

This paper cites The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.663678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.059180Z digest=sha256:9e1949a011b5338f08a3f60c243096256032c3a1b78ad1dab88e18e739c69bac

Observation 93957649-5c7b-4775-8d6c-c13322deba03 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.569937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.063216Z digest=sha256:8c01466b1e5e6a3957f57fadf65bf2c2d7a6981f952fee98f6bdf3a4345f8151

Observation 244b66e0-b917-4c17-9f2c-6a9a29e89968 · outbound

This paper cites 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.435490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.086083Z digest=sha256:58e7b7c050cec548e73b496792765cc642706c08e4b96b84c5dda1c7716dbc9f

Observation b3021570-a7b2-4bce-8e26-1ffd20f0dd7f · outbound

This paper cites … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.378205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.125099Z digest=sha256:7cf4cf5aa8f9aa913a97b48591ec3fc6bd1ddcde9a9e266a55b9cf8eef4f9913

Observation ff24848f-4171-4f07-85ec-70295e6d4e8f · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.234271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.146836Z digest=sha256:598c3514562046d1c9b518e5318a050cf8e83fcffdaa933c0b9b49d1727d013f

Observation c6d5e29a-2eca-46f5-a933-a7fba6130d87 · outbound

This paper cites color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.124845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.192600Z digest=sha256:8c96820fbfa7a933bbd0e0967470e1b114f33d81e0c6a98601f58b2298fa96d5

Observation 4fd0a3a0-7f46-41c1-b50e-e97941bb5a47 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.015722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.238326Z digest=sha256:aad9e63cbef236558dd94e381645e5ba24eba968a6a0ec782a9c3f21953addfd

Observation bde4c299-e96c-446b-bf43-ee2ffed36c6d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.234747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.234747Z digest=sha256:7afc2ed237a1e0f7b3cf98cc4c3e57b6e8071026c6bb7b0c0e08fe84151c3736

Observation a8a3da1c-c5fe-4d93-8c2b-3099b57c98ba · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 576

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.216677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.216677Z digest=sha256:fb948bbaf17ee7cb1008f322a7ff8498b47ea25f11eafaf5418a82b242b2f92c

Observation 5bc9d72c-a524-41a7-89f5-4345f31dbb77 · outbound

This paper cites Convert,.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Convert,

Reference 2015

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:13:04.074754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.544835Z digest=sha256:76233393023faba041661487a8cb50d7b018dfa27f1cdee0e0701ac77953247e

Observation 9a97e119-ad7e-4962-bab7-09c74feee0eb · outbound

This paper cites doi: 10.18653/v1/2023.acl-short.43.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? doi: 10.18653/v1/2023.acl-short.43

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.299514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.299514Z digest=sha256:fe0715b33aab4a5d5164513ba8a52c85596a7ee83f02c686380ab70daa4f6eca

Observation 42a01c31-c3bc-444b-8b05-59224c67ef9c · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.264753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.264753Z digest=sha256:c066c11e7b25136643d79f5e33e38080c33dd37ab196582bc57bf548b119cb3a

Observation db99015a-11da-4001-9936-ae37af6886cb · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.271469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.271469Z digest=sha256:9cdae9707a4988d3c801d140c041784345298424f6688b1201773485b8d63063

Pith citing papers

Observation 1102cc23-3b0f-46de-8988-d3a531a0f46d · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.191067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.191067Z digest=sha256:ce5b06d627ae029d7196582297e99f77ccbc4f1c38ceb45ee094317ded302b40

Observation 5f397b56-197c-44e0-9ad9-cc6c9852c6d3 · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.343122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:6bf09935e98871ffe3c7810ae4f617f8ca6bc637746cbd56ed3cd7c01ca7ba30