Pith. sign in

Paper Citation Record · LEDGER

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2501.05444.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05444 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:06.836442Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.233935Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59217165-9cd2-4838-ace1-1f645cccb2b2 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.840359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:9ada1fd3df8b09863ba0032e939ac299fd10fbdf117242dc004f4d4b07248fbd

Observation bea1f8ac-2ae8-4271-97b5-1c28b2ddafb2 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 241

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:40:41.471398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:16b44f558766e0ba47de6eccd9824dd29293eb3ce0c2768cb0d74f313be173f7

Observation fc20a570-a8c1-499f-a529-97c46d9fcd85 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.278400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:71980649cbef0985e1ce16e670f2bd4e7651f7a61e75a775b438b4bf3ad48f46

Observation a431e228-57cb-4df6-a848-fcc25fc8e8f5 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:59:03.400181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:5a030f36c1704ac70e387a49cea897459db34ee66d7389870e277f554ee9aa66

Observation 6de50937-63e5-4a52-8c9a-237f56fc1afa · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:17:02.892270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:91d38dbfcd7f6e98af8679e521ba6a5c1b5200e15bc0b34a39bedbf892b93d66

Observation ca4d68d7-8604-484e-b562-19be0e03b0cf · inbound

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems cites this paper.

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.836442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.836442Z digest=sha256:1e626f02991102728ce2c884870fcc77bc5157bb4bb3719df6956157c9e09f94

Observation 8cafd74e-8f58-4fde-a158-02ea21ec09c9 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.566161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.566161Z digest=sha256:9e334977820873dcc63b7e0a5afa3c84c0219ce6d5b6eb5c58cd3b40a53c9916

Observation 3269c397-9cf3-456c-a0c7-69b46502952b · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:56.253584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:56.253584Z digest=sha256:644aec3a1ca8af6916641eb3b4b077ae847911b4a954ff607f0ec222f28ba56b

Observation 13c6dd49-fa38-4685-86aa-7bee9f02578a · inbound

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow cites this paper.

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:16.478809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:16.478809Z digest=sha256:c8778f1a12a94ffa54e60382d581b346af1a9f5d4785782e724971e316add9be

Observation e892892b-c10d-45a6-8065-9432d3feab2f · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.468895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.468895Z digest=sha256:ace59aa4007a6c023b96d08e1e46ee01fb2c4f38ccb53e1ffb98a7b413302aac

Observation 82ab3d3d-13f9-4382-8cf0-d5cfdb58dcb1 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.650097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.650097Z digest=sha256:90de551a06c20562bfdecfe1cdb6c07b7643691396d64565460b9c2d96d37a0f

Observation 0af8138d-b50b-4120-96ae-298ddf751c9a · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.859705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.859705Z digest=sha256:735748749ee1e92295a2b4b3a73e725e1d1c500b48a903e955b856d89e4619f0

Observation 060b87ff-030b-457b-bbd4-e859554de971 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:57.253578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:57.253578Z digest=sha256:05e0bcab97ff7d838d98725a50bfb5b68f69af26825ec0f0ae876b6f5895f27c

Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · inbound

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs cites this paper.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:01.334554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:01.334554Z digest=sha256:ec1f9846121257ab52c3d4fb50f76f7a344f17954925c758884564546abcc0db

Observation a5964cce-ce42-4bca-9940-f8b194a075a7 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:21.244871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:21.244871Z digest=sha256:83619df43fc214d75631dd66ec19d6f29146b395beb860bb53979130de509439

Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.367946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.367946Z digest=sha256:f93c4211c576c74d47bc658ba52ef750c949cfbdfc76d9fe8582488af555a176

Observation c27708dd-cd6c-4f19-9f5c-898587f63c4b · inbound

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency cites this paper.

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:16.540534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:16.540534Z digest=sha256:4f25a659c964b1dcf8b64ebb282a51ea26eeec183756221e00b56bab98cd92f5

Observation 9024f7a6-09d5-44b8-a2c3-b8d1e33ac7af · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.211304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.211304Z digest=sha256:6e6c4276fcfe0aeb8810f8f7d68f29da9b2f06064ed66e16955ff34e38a2732c

Observation 55df135a-9ca6-4fe7-ad75-4bb19f440d8c · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:34.684819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:34.684819Z digest=sha256:62214fcd0001547a077553fc33f7fe890df9919fd0000e3df22e8dfc01012e06

Observation d1c9e39c-da3b-451e-b3c2-7c06538b096d · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.335779Z digest=sha256:1a76a1def302b1b1520acf4acf13fc74a4a0ae3e8a2c4691920ec035bdb6727f

Observation b7ae7c65-1c24-4eef-8281-58c46bfdc0d1 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.882839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.882839Z digest=sha256:0746299c4f34c5e18c2e05bee93008ccfdb3b3c919baed58f429adb6c11f7a60

Observation 9d5b0801-966a-45f1-9e27-1ecb18d0c745 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.934041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.934041Z digest=sha256:ec33a5eb3cec8635a1d19bc3c79672dd5e8da58414b0626b59bf6c40f23abb14

Observation f86f6a0d-31ae-4ba1-9bb2-869119a11ede · inbound

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts cites this paper.

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:23:52.822183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:23:52.822183Z digest=sha256:603ec653d655fa10b4ab3c7380af465dd78bf35bacd7470be5979ab86dc242b6

Observation aa570e30-2544-4fe1-8f67-24262a221426 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.653975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.653975Z digest=sha256:715ed983026c49736dc3461fffd684f47ec1952cefa8ea1c783b9c712646f1ea

Observation c66eeefd-99ec-410e-8435-a40e3634a2e5 · inbound

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe cites this paper.

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:07:27.398933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:07:27.277040Z digest=sha256:5ddb9ed6e46b2eb84d7cf01183a16db1a3918ba4053f9b8661cd3752c92d86e9

Observation b5e52865-c1e2-4a43-b5f1-2576227b5ab7 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:00:26.874973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:06405fa0e07cd1ce8fd86a0cf39d43cd25091ccec098a2fbe11fae95e8b8e167

Observation aa028a27-b5f1-47b4-99c1-6f0219a0a6b1 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:33.121778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:33.121778Z digest=sha256:9ac68bb2060b5bb898e3750e0dd51a25d8e1e3b230c3273ff08d55762717215e

Observation b8dcbd05-29aa-456e-a0fa-05bb7fc6ef92 · inbound

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy cites this paper.

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:04.457394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:04.457394Z digest=sha256:c7ca2a30f7eb31ba0d0b8ae906c860d012dcd2af08620c4a600ea9e9c8075d73

Observation 314fcb5f-e6d6-4dd8-9e5f-f82c27d3b6bf · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:45:14.287576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:f789b914afe596ef8155fefdbcec2b6fcd8c37f33122bcb8584526d58e045e80

Observation fabe2c58-eaca-4c41-9005-42d8daed444e · inbound

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving cites this paper.

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:04.311521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T00:37:55.147350Z digest=sha256:1084a5a2c24e563cfd1f197a139b00527906cd3eb59e7113484090f6e226d49f

Observation 3325c29f-6bc7-4f06-b15b-841b61ccdeb9 · inbound

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving cites this paper.

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:21:03.874703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T22:04:19.654714Z digest=sha256:14062cfb2e444a178fa14508a6d6d449318d672dce55a0cb6589744ba207535e

Observation 67764f88-0862-4b1c-9f4a-002e83632f6b · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:41:27.517349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:15981874b4f56cdc0eaf10fc3ec5fabad8635849ea05a849f671279fef2b66f9

Observation edaac7b9-5988-4235-82c7-2e1c8513f50e · inbound

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning cites this paper.

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:02.884007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T05:14:28.256032Z digest=sha256:6a0d6307aaa874db40849bb92e8b3bf2bea08313ee436ef0861ad238da7c505c

Observation eb529714-10e1-4e6b-88ff-28e8a71ad0a3 · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:48:53.490345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:3225318b806eadd1182082e1a9e07781de711c9e22bf1926fe1b2f8cec9818ab

Observation 7e1c70c0-7481-4bd7-af97-e107c43eae73 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T00:14:04.635099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:843a474a9d37dd0a04b2c5701d00f3496ff7615060bd39c1d7793765df50158f

Observation 41f675df-f331-4c6b-9d3f-3dd540b15384 · inbound

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs cites this paper.

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:57:06.715624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T23:15:38.962013Z digest=sha256:9239a2046b8cc1387ae75f34a7f405d2d311d7e4ee3009f3e60fc716825834f9

Observation 01469653-fd64-47f6-bc32-e6f8c0ffde20 · inbound

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP cites this paper.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:50.235779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:6a84434446150536c15d8d06f0b10ed33b6cb11b876866e8e41d68f0f263f2b6

Observation dcc576b0-35a1-4d70-8ba8-bdb99db859d1 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:17.410212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:473e665171a43379539d7bea1fde9419f16eb13a217d5ed1e2788398992e6b54

Observation 1ad6a5de-5851-4a93-8ca7-e9b1d05f83b4 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:46b048295546618bade111a13ec49e1ef3ad2ea32b0d76b5e2ce88678f373987

Observation e6992855-794e-4037-b7d3-00ab9c601dcb · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.384980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.384980Z digest=sha256:bd36193cc53395b79f2f92fe329e9245eb093f3c83b3a2bdc8f23b1bd0029af7