Pith. sign in

Paper Citation Record · LEDGER

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

As of 20 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 15 inbound Pith citation observations for arXiv:2412.11936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11936 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:29:05.241019Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:33.123081Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T04:32:32.800150Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dba59bef-4069-472a-b74c-f344e7c58290 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.076096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.107238Z digest=sha256:5219cbe46bedbe27c941c5f61fb362fa322274844dbfb24ca8f4fd53b533e655

Observation 8ecd0505-1d87-42a7-9340-5b828ef930b8 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.062169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.110683Z digest=sha256:bbb62bdc1a8caef1ae7216d7dd75bf06e347fdf08c2276ad884c1412162b69aa

Observation 97be51ba-aebc-406b-bc6c-3bed9092f080 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.048299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.113951Z digest=sha256:0e625d6f8e2fca4a5d27b537620d7966fd9c1f15d573decfe067b999e7e57411

Observation 39a73c8e-ddb0-44f6-bde8-197070c9fd57 · outbound

This paper cites RoMath: A Mathematical Reasoning Benchmark in Romanian.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges RoMath: A Mathematical Reasoning Benchmark in Romanian

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.004169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.004169Z digest=sha256:052fb89d53c4a005794f35ef059e410e3966dff018ab6f4752181c3f070cd541

Observation 99e823c4-d10a-44b7-8267-03c14e08053b · outbound

This paper cites Yes" or.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Yes" or

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:06.019128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.121229Z digest=sha256:e00c6aef58041870a6ce61b2d88ee34cceada401f53ee3c91668c99a91753953

Observation 8f3609ab-92a0-4350-b45f-acce94fb949a · outbound

This paper cites MathWriting: A Dataset For Handwritten Mathematical Expression Recognition.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MathWriting: A Dataset For Handwritten Mathematical Expression Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.013674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.013674Z digest=sha256:8d83878124a8a9db6589a4d6d68e60523c67a7783dd6322fc6d80d150c5a37f5

Observation 1a26adb0-27c2-4099-bdec-56177def482e · outbound

This paper cites OpenAI o1 System Card.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges OpenAI o1 System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.018816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.018816Z digest=sha256:b1e9ced0bb64ef30dcdc9a82533e0e9d697f20899328328dac4805fe053db8d7

Observation 2ba5b79c-1a85-47ad-8470-fae1c7eee7eb · outbound

This paper cites ATHENA: Mathematical Reasoning with Thought Expansion.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ATHENA: Mathematical Reasoning with Thought Expansion

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:29:05.575607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.029306Z digest=sha256:348f482d9285042b4150c7625ef9bc01ef9c50859da20cdd1cab810c2ba8bac9

Observation 3c05723f-389a-4926-80e6-a6452dd96748 · outbound

This paper cites Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.034781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.034781Z digest=sha256:f4eedc3e70b26c4c89fa72e2b93d2f063fc7aa847a8ae4d30403f88e9287598e

Observation 5e137c8f-3e85-4ebb-8143-dd8c7d61a794 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.061961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.061961Z digest=sha256:e38ac1a7d9b22214dd85797a80b7d6258742c73b0642084fed9b876c314abb2a

Observation 74039854-4833-48c0-b298-fe04cc271d58 · outbound

This paper cites Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.067239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.067239Z digest=sha256:fc0abb6426855adf2184300a93b7264ca22f8aa28fbff2b230160d37154cf5ce

Observation 0d3b8f3c-48c8-4bfe-af81-9f0bc5cdc11f · outbound

This paper cites Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.071773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.071773Z digest=sha256:b89294f8cedb0432025718f602cd5159a599a2539f2981405c5828eb160a8fe7

Observation ae4675dc-c168-4270-b17a-f8975df4f775 · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.077208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.077208Z digest=sha256:c3e6a0223819882df8b4d85b866e219f02c2ff04b19c092209ff55122953ead5

Observation 5205e4a4-ebd8-4836-9aad-4842212751ab · outbound

This paper cites Proving Olympiad Algebraic Inequalities without Human Demonstrations.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Proving Olympiad Algebraic Inequalities without Human Demonstrations

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:29:05.367816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.082766Z digest=sha256:d2aae82e91d1226edef6d34aff0567855f2490c2f337ef0812be2d5faa972c8a

Observation 6c1acb48-a2b3-44e6-80f5-9c98c69cdd80 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.092179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.092179Z digest=sha256:083b719aafd866309f265ca15c608889544e4d56e28c93349c42cda57c93d6fc

Observation e62a1cc0-3c9f-4ca8-8ad3-b97b642f8995 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.097298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.097298Z digest=sha256:a973efb6956c6bb9dc25d2841b96da23545c302c1f4721729ddf512f7a8d7b21

Observation cf14c607-42d8-4e3e-a264-24b39842d7f6 · outbound

This paper cites SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.102595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.102595Z digest=sha256:176b44c7b773ca4e39546041524c53416f0b15519d1c285907b35a8982e8bc34

Observation f126ae01-c912-468f-9899-8d718e7c2bc2 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.032432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.117743Z digest=sha256:8addfc103ded4e914ca0bc4b34fc0686135e8c4b524dad999c9012f66736a678

Observation d1ede7c3-80f4-49d0-a788-44747cc9c376 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.001215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.126110Z digest=sha256:178c903004b3bb85fc7aad1a4e890b2259d73ece24f6418db584c775bfa69895

Observation f516fd9a-df6e-45eb-853a-b9fc42bb8151 · outbound

This paper cites ❷ Bottleneck in Data Diversity:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❷ Bottleneck in Data Diversity:

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.987025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.130437Z digest=sha256:94fbf756081029a399b3ca88fe88e68c2e5e1a078d5324eea9dcbacbfb06e79a

Observation 7c0344c6-3a91-4975-9fb3-68349f7e9834 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.969714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.135877Z digest=sha256:c7c2b36be40e2413269681cf76d9626edec1e9ba911aa5348443895d11aeea45

Observation 7fbdf47a-d43f-4d72-8a29-8f07dccf4b64 · outbound

This paper cites ❸ Bottleneck in Data Scale:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❸ Bottleneck in Data Scale:

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.953781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.139316Z digest=sha256:3b7c29cce4c2433aa27a57bef9fc29b51c5ca62eac181cfa0f8cd679fda40194

Observation 42ae1ead-5673-4829-8ea2-84d0008536a4 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.938432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.143243Z digest=sha256:75f02cbee7cbd3ab9e4ba2932f3b22bd94166edd9e307263c92701fa8887eaa9

Observation f1fa7cb8-3a00-4e12-84cb-de91dca62478 · outbound

This paper cites ❹ Based on recent trends in the latest works, we further propose the following actionable sugges- tions to address these dataset bottlenecks:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❹ Based on recent trends in the latest works, we further propose the following actionable sugges- tions to address these dataset bottlenecks:

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.926069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.146955Z digest=sha256:1821c458fd35026b737081e575e80f008ca73b0367a3dbce10ea2d8ed31acb23

Observation 0509f27b-61dd-4520-a39b-4d72047310c3 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.913504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.150886Z digest=sha256:6dcdce0344a901179e7d611aefd3e23a5282e5ea3a21cdac8f2c14bc891477e9

Observation a6ea1e75-6233-4dea-8ead-24338f149242 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.901463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.154775Z digest=sha256:6aefb386c5aa27f354e6acc44b7723b91643718dc6262b80d8108c02deaec7cb

Observation 34f89647-9b89-4ecc-a502-643d4901d3ba · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.890428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.159774Z digest=sha256:573f83ad5559c8e4da0ffd4eb62ed464e6c22990bd624608f4d2100a2f32b983

Observation d2387a89-a605-4f65-8a6e-98ead558ba33 · outbound

This paper cites spatial reason- ing.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges spatial reason- ing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.879433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.165558Z digest=sha256:be92e7411bf06ab560db015d7610ac2f47b64061bab3976074c0a1abebc98d05

Observation 87b450de-ba20-4746-b819-e6677eb04a31 · outbound

This paper cites Use domain-specific few-shot examples (e.g., providing figure-text associations in geometry) to guide the model in switching reasoning modes.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Use domain-specific few-shot examples (e.g., providing figure-text associations in geometry) to guide the model in switching reasoning modes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.866455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.171547Z digest=sha256:9706a2fc12f361539882898299858025ea918ca165f73e1bf8a82ca9dcf29b53

Observation dc4caeb2-0385-4b9c-b915-3f51b631eedd · outbound

This paper cites triangle.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges triangle

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.851907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.176510Z digest=sha256:cfa57f21dba2bcab63a246d8ea12dfdb853d0981f3ce9e201d2245a85964bc05

Observation 578fe127-015b-4dfb-86e4-4ba4bc36f342 · outbound

This paper cites The limitation is that labeling error types is costly and it’s difficult to cover all long-tail errors.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges The limitation is that labeling error types is costly and it’s difficult to cover all long-tail errors

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.837944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.181253Z digest=sha256:e706154b2fd1bbde831e3cac956e2f59f343c276c9e7b900e9ab2bf3e5e3e638

Observation 72c164b1-2f2e-4abe-aa08-1a4a876ea596 · outbound

This paper cites computation-logic-conclusion.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges computation-logic-conclusion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.825041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.200899Z digest=sha256:1b59709bdd4c447137caaf661591d8e29f42f1ed95c0aa859641d06a766989f6

Observation 617e69fa-b12c-459d-8fc7-86a75370afc4 · outbound

This paper cites The limitation is that tool invoca- tion delays affect real-time performance, and some errors require manually defined detec- tion rules.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges The limitation is that tool invoca- tion delays affect real-time performance, and some errors require manually defined detec- tion rules

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.807600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.205542Z digest=sha256:6f1584b4031594333ad5e05d3c80a93cf5439da5409d44532b4b416bced787ad

Observation 17331e94-06c2-48d6-ba1c-c777f22b8354 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.794363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.209565Z digest=sha256:2d4cebc58be4a54299c04dc52d8eb922e7f093b2f8cea8bcca597840ccb4e1ed

Observation b3132461-8613-4f86-adea-aed2e29164c6 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.780786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.214385Z digest=sha256:58c134eeb009f1382063f862bf2d426a998bc92a71573773c5bdc109dfb77568

Observation 77173011-fb98-4b4d-95be-5e327b63df6b · outbound

This paper cites If the training data coverage is insufficient, test-time strategies may not be able to compensate (Ke et al., 2025; Chen et al., 2025c).

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges If the training data coverage is insufficient, test-time strategies may not be able to compensate (Ke et al., 2025; Chen et al., 2025c)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.768170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.218715Z digest=sha256:e6b1c0c140cc4a28e70e5ac297050d4470a180d1c9818dbe77e0c4065daf41f2

Observation 61ec9488-f54b-4770-b841-6691ff24c0d7 · outbound

This paper cites trian- gle.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges trian- gle

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.752711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.223162Z digest=sha256:9e5316fef14b5c16ff68a4449d6816cfcddeb63514e76dd6d0bd1121780c82d0

Observation 3433ae2a-0a2f-4284-8b6b-2ebc63c08367 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.738351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.228180Z digest=sha256:5df2f09503c10b66ad409d1c69bb5dc016f70d4425522f8e46e75c8a3737254b

Observation 385697d0-fe96-42a0-86e6-ea305380a1ea · outbound

This paper cites ❸ Error Feedback Limitations:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❸ Error Feedback Limitations:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.726506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.232068Z digest=sha256:c6c4ce92484b32f32e79bf90b504d4b62f29a8010b48d4a07316263565570cc2

Observation cf36e3f8-1cf0-47f0-99b0-e5df23e2bebd · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.714406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.236565Z digest=sha256:322fb512583406ed403b7044982b5a057fcfe3c483ae88c40afe261c48b7f9f7

Observation dc4f8666-a545-44b0-94fa-83dfa2de3e96 · outbound

This paper cites Long-tail errors such as rare symbol confusions may be overlooked.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Long-tail errors such as rare symbol confusions may be overlooked

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.699129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.241019Z digest=sha256:60d803e0de2a4ca1891b750509ce683aa783e6306796c164d8ea83111a66890b

Observation 2e35f305-c370-4cec-9ace-74211b0a6271 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:04.986025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:04.986025Z digest=sha256:959de1539fe75fd6defb4042beed0d67bb70486f2d217d829330e50434066821

Observation 2d6b954e-ddac-4748-a237-4c8a0eb053d6 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:04.991803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:04.991803Z digest=sha256:be1e2fa744d076679e474346b65c3fe215669f615b625a37a9ef3bad25e17d3c

Observation 100240ca-376d-4f4f-8463-db2c396533e6 · outbound

This paper cites MWP-BERT: Numeracy-Augmented Pre-training for Math Word Problem Solving.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MWP-BERT: Numeracy-Augmented Pre-training for Math Word Problem Solving

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.040249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.040249Z digest=sha256:fe276f7effe9a54648e63e038b12430b40cb068a6806ab35384a96d336750c9d

Observation b82d1d28-879f-4225-8e89-516e33005d76 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.023302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.023302Z digest=sha256:b9fe2d04e211ec3e09831bd0040559cd0027148a0d7f420fe489cb538b398418

Observation 17ade6d4-05a1-499d-bb98-db823689930b · outbound

This paper cites Advances in Neural Information Processing Systems, 36:5539–5568.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Advances in Neural Information Processing Systems, 36:5539–5568

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:06.088222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:05.009221Z digest=sha256:1b7d78b1503e7ee399ff6a631d2261eb96a42df5745f0a59e53f8e4064e803f9

Observation efd5c1c7-8f9c-4680-b841-34e91896f3d2 · outbound

This paper cites Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:29:05.652045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:29:04.996663Z digest=sha256:c71ec76ad4028adb68a0c95da044367c9693b11a11f43cf338a3def8ec9d5c41

Observation 9aea032e-f4f4-49d0-8fe4-e25da0fb9f47 · outbound

This paper cites arXiv preprint arXiv:2501.04686.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges arXiv preprint arXiv:2501.04686

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.057155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.057155Z digest=sha256:61653a3743c2a94c2b1f14912edf5eddbcc4d83528b0d21323573e188d660e72

Observation e2f9304d-de2f-4d5a-9e88-fc8542099eea · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges BloombergGPT: A Large Language Model for Finance

Reference 2256

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.087342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.087342Z digest=sha256:3ef322925e647b900a42d4c0036a63d570149be04befd2a1579b0f5d24c4cef2

Pith citing papers

Observation 82193361-4bc8-42a1-bc94-2b4db7320522 · inbound

Reasoning Language Models: A Blueprint cites this paper.

Reasoning Language Models: A Blueprint A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-10T18:36:55.267690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:36:55.267690Z digest=sha256:0006b288e60c2323465ce4153de20f6a12d4c5d6feba80be9f87354f86853f43

Observation 8fe2dede-2be1-4b50-9830-bc703c188f0f · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.803728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:47ff6ec2eab8028f59652f66b0fed0faf987b422b10067043a7ef8652f126a4a

Observation 71045df9-91b9-452a-a982-4d4a36911c30 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:11.970200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:11.970200Z digest=sha256:615a5336c85e8813a5053ce139ddbd7b09b7da93bb4cc2f003e7156b1239c946

Observation 6df75775-2aec-45fd-b1fb-a298f94f6a80 · inbound

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? cites this paper.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.553533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.553533Z digest=sha256:810e91d76154df31e02306a227b0ee1a37e0fa842b520b53450265976ec31de9

Observation 85ca2948-0a06-4aa2-a7c7-f3df5c4d17da · inbound

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring cites this paper.

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:37.737664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:37.737664Z digest=sha256:57df6602018d77c792101d4a060e3b7f13659cfa27c32289d387ee10ef300dee

Observation a5c85ee0-ca3d-4a0d-ba8d-3ebe4f30f42f · inbound

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method cites this paper.

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:42.587497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:42.587497Z digest=sha256:54dabe53a6c31ed3099e605eafa1e97795613c859538d678f604be57254cbe6f

Observation 5b5a7fe5-dd8f-4268-bb38-eaeb8f7e9bd6 · inbound

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey cites this paper.

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:00.698793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:00.698793Z digest=sha256:d4a5117bfdfe5c4fbeee698592390b25b8f443abac318f01a4aee80253d429aa

Observation 8aa1ca2e-288e-453a-b26f-315745af4d4e · inbound

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis cites this paper.

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:20.521456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:20.521456Z digest=sha256:1dd320afe16b78facb81f3f9d580383b4000e4eb8a3db7f8aed0b9cf2a98db50

Observation 0d5b4e13-f700-482d-b896-2f0c99511bb8 · inbound

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities cites this paper.

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:23.839050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:43:23.839050Z digest=sha256:02af55b835c3ba7b8ec446d8d915f5f659e8903e28f5a83c5c1ebd027268c609

Observation 1c2ee46f-164e-4739-9ac5-5f2c1d463460 · inbound

Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach cites this paper.

Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:33.123081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:24:33.123081Z digest=sha256:785c8e24bcd3455c5bd640b55bad51988860ddd122f8b9dd000c14d7a75a0f19

Observation 96b8d507-a2f3-417b-807c-1d4d8df6c155 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.511366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.511366Z digest=sha256:08f06af0c503acd7563d7fb667071247bcd750ee2e92637c07771e29ffb12cee

Observation eaa1e2b9-109e-4539-9e06-d948c0c055f4 · inbound

Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons cites this paper.

Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:50.011424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:50.011424Z digest=sha256:309069af82d6f409d94e07c7a60eb0d43e1a0ccbabc84dc44a88cf92a1960db9

Observation 37c2ba42-2db1-4040-a1bc-a13d2cc8e47c · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:53.783089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:53.783089Z digest=sha256:4ee3d84d4e09a74272328c7971d9f3f2b1beaa3bdade8dc9650920e1031a093c

Observation 73c582c6-f868-4444-b2fe-27fc3443fa46 · inbound

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines cites this paper.

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:41:56.049889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T00:38:43.897231Z digest=sha256:2be448c7660df74079224749750acfb60b27673fc2a7014ce8c37eee61dea2a7

Observation f34545ab-8bf3-4f76-a439-9cd855772702 · inbound

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning cites this paper.

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:29:33.508768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:29:33.508768Z digest=sha256:4784f533720cd0dc269cfb5cbf65cf307fe0e6389047c2f0a7fd3f9a4cdc3c3c