Pith. sign in

Paper Citation Record · LEDGER

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2501.05444.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05444 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:06.836442Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.233935Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59217165-9cd2-4838-ace1-1f645cccb2b2 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.840359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:1365fc2829db827d4738e510a047dec7f6271d1bf1595069719d3d11816edad5

Observation bea1f8ac-2ae8-4271-97b5-1c28b2ddafb2 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 241

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:40:41.471398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:ecee283b98185656cd8bbb63659504dcaa2b5e721d5d443e0d7eb3e400f0fc81

Observation fc20a570-a8c1-499f-a529-97c46d9fcd85 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.278400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:1fa2ca74d0ed5f965c8f262a782f7334b3a746262bc376328d9cd592b1b20875

Observation a431e228-57cb-4df6-a848-fcc25fc8e8f5 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:59:03.400181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:f0c2eed07484ebc3d89465bcbd2003925c9974957d18ffc38449bd5270d966ed

Observation 6de50937-63e5-4a52-8c9a-237f56fc1afa · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:17:02.892270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:db9ac62948f42e61c98265822395c187dc91c565df2c879c040997ee98429c41

Observation ca4d68d7-8604-484e-b562-19be0e03b0cf · inbound

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems cites this paper.

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.836442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.836442Z digest=sha256:1e626f02991102728ce2c884870fcc77bc5157bb4bb3719df6956157c9e09f94

Observation 8cafd74e-8f58-4fde-a158-02ea21ec09c9 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.566161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.566161Z digest=sha256:336b682285fd77191e0e30e6be3a8378cc88b943327d615f3f0f73b4180286fa

Observation 3269c397-9cf3-456c-a0c7-69b46502952b · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:56.253584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:56.253584Z digest=sha256:339b1c5adcebc14454f5ec03e8894aab851af67bdbe8f7ca803b30968eb7a861

Observation 13c6dd49-fa38-4685-86aa-7bee9f02578a · inbound

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow cites this paper.

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:16.478809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:16.478809Z digest=sha256:b6ef727f68e98931afad8dbd8687870ec843d2be52c36777bdcd88770e472b92

Observation e892892b-c10d-45a6-8065-9432d3feab2f · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.468895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.468895Z digest=sha256:c5ba142a1641d1f9f473404cb19170b76c2246bf3286e0fcb4fc6ab977bb10dc

Observation 82ab3d3d-13f9-4382-8cf0-d5cfdb58dcb1 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.650097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.650097Z digest=sha256:b3b266148576e7c93183449f2b365b6cb989a2ec2955dad624f7bb5773c5db67

Observation 0af8138d-b50b-4120-96ae-298ddf751c9a · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.859705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.859705Z digest=sha256:f1dadea898a0ab5d50de9f40318c92896e80868962bdfea7d0d0d1038cbbcd10

Observation 060b87ff-030b-457b-bbd4-e859554de971 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:57.253578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:57.253578Z digest=sha256:73a3d00d9e2ccc8710f363d155870e4da232e69928d26332b8a71a313c96241b

Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · inbound

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs cites this paper.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:01.334554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:01.334554Z digest=sha256:bb4c0de273d769994fb739fd5880a7ccc5713e0e97aeec720fc3d358ed5013fd

Observation a5964cce-ce42-4bca-9940-f8b194a075a7 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:21.244871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:21.244871Z digest=sha256:2703bf53a7d1f1f56e1f599a791b5b13ef49e893aeb905816eae457216d8df83

Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.367946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.367946Z digest=sha256:05724ca721422f0fa07c04261462ad249c0852fa4e8229bf06c5a824199e47c7

Observation c27708dd-cd6c-4f19-9f5c-898587f63c4b · inbound

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency cites this paper.

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:16.540534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:16.540534Z digest=sha256:4f25a659c964b1dcf8b64ebb282a51ea26eeec183756221e00b56bab98cd92f5

Observation 9024f7a6-09d5-44b8-a2c3-b8d1e33ac7af · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.211304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.211304Z digest=sha256:4e0cdb828db7682d539a7fc974ca28e4b2920ecff1b60ce5b384bda14be7aff0

Observation 55df135a-9ca6-4fe7-ad75-4bb19f440d8c · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:34.684819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:34.684819Z digest=sha256:60ce4f72d428eee4b7e4b5d41c86e3551b1b0bf5635c45b7f9c595b1b2f0eb80

Observation d1c9e39c-da3b-451e-b3c2-7c06538b096d · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.335779Z digest=sha256:1a76a1def302b1b1520acf4acf13fc74a4a0ae3e8a2c4691920ec035bdb6727f

Observation b7ae7c65-1c24-4eef-8281-58c46bfdc0d1 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.882839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.882839Z digest=sha256:9180fbd9f795e6c3689f3278eae7b12ee099deaedcf86fb22ecffeb1b4020153

Observation 9d5b0801-966a-45f1-9e27-1ecb18d0c745 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.934041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.934041Z digest=sha256:1c55f714695bcf0592bcdcd91290920a6c841af7fef425b796ebdea4561b4728

Observation f86f6a0d-31ae-4ba1-9bb2-869119a11ede · inbound

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts cites this paper.

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:23:52.822183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:23:52.822183Z digest=sha256:877e90ac34977acb68fc1fcbb00d4998ac369e67b578179908cdf9bd8dd1a81e

Observation aa570e30-2544-4fe1-8f67-24262a221426 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.653975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.653975Z digest=sha256:7ff5a998a15ffce46cba5b9d56cfb38733d3f03d073c4eb05e94950700376bbe

Observation c66eeefd-99ec-410e-8435-a40e3634a2e5 · inbound

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe cites this paper.

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:07:27.398933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:07:27.277040Z digest=sha256:3f00f8fbfb547b5e80eb8a97f0b1d1257f59d46a9514757a5c34eb8962edc9e5

Observation b5e52865-c1e2-4a43-b5f1-2576227b5ab7 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:00:26.874973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:2d3336bcb15b93cf2af13fa6e06c32288974b04b4f85ed4c4567fd44876018fd

Observation aa028a27-b5f1-47b4-99c1-6f0219a0a6b1 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:33.121778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:33.121778Z digest=sha256:fdf167d60c0b8446d0b1af8fc8aa49eaa1433a45b151917367ce730e5b6d2386

Observation b8dcbd05-29aa-456e-a0fa-05bb7fc6ef92 · inbound

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy cites this paper.

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:04.457394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:04.457394Z digest=sha256:c7ca2a30f7eb31ba0d0b8ae906c860d012dcd2af08620c4a600ea9e9c8075d73

Observation 314fcb5f-e6d6-4dd8-9e5f-f82c27d3b6bf · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:45:14.287576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:d07a6aaa2bba578765a407053e60fb390e0882423bb25e96aba6f7196a98d2a4

Observation fabe2c58-eaca-4c41-9005-42d8daed444e · inbound

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving cites this paper.

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:04.311521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T00:37:55.147350Z digest=sha256:e2f723b39ddec47e44471db28474d32a369b8e1a7d527ee605ca4875e6155e14

Observation 3325c29f-6bc7-4f06-b15b-841b61ccdeb9 · inbound

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving cites this paper.

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:21:03.874703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T22:04:19.654714Z digest=sha256:ded8da8dada855aa8f62ce97829c0493e0ed78a61b85342f3184243ddf9a32e2

Observation 67764f88-0862-4b1c-9f4a-002e83632f6b · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:41:27.517349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:020af5889e305fa1d44829c994c2a2755d834909bba124fd744a095d377887e8

Observation edaac7b9-5988-4235-82c7-2e1c8513f50e · inbound

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning cites this paper.

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:02.884007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:14:28.256032Z digest=sha256:c2e05ffd9aa81f3b54a2406caf389f27bb866e124730c235fd56cbf3e1341cb9

Observation eb529714-10e1-4e6b-88ff-28e8a71ad0a3 · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:48:53.490345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:17a42c1bdeca631db000b12d419a74f0652cb9a0d5ce2567d93e6cc408282bea

Observation 7e1c70c0-7481-4bd7-af97-e107c43eae73 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T00:14:04.635099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:d81da861789d938f86a01f35bec1d9ef67c045d446126a2c72295c297898471c

Observation 41f675df-f331-4c6b-9d3f-3dd540b15384 · inbound

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs cites this paper.

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:57:06.715624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T23:15:38.962013Z digest=sha256:9bffa4489d808d43d718f209c2ca2444307b96185d970ccfc68fafbb86b07391

Observation 01469653-fd64-47f6-bc32-e6f8c0ffde20 · inbound

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP cites this paper.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:50.235779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:2ebb7f7c03b4a76a994718a2fa68bbb94f2e555735694e5034e7596dfb8d0dfa

Observation dcc576b0-35a1-4d70-8ba8-bdb99db859d1 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:17.410212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:523e3cca5a5c8d78dceb74ec452249fa57162bebf9273b629577279aa53825c9

Observation 1ad6a5de-5851-4a93-8ca7-e9b1d05f83b4 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:46b048295546618bade111a13ec49e1ef3ad2ea32b0d76b5e2ce88678f373987

Observation e6992855-794e-4037-b7d3-00ab9c601dcb · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.384980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.384980Z digest=sha256:bd36193cc53395b79f2f92fe329e9245eb093f3c83b3a2bdc8f23b1bd0029af7