Pith. sign in

Paper Citation Record · LEDGER

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

As of 20 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 15 inbound Pith citation observations for arXiv:2412.11936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11936 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:29:05.241019Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:33.123081Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T04:32:32.800150Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dba59bef-4069-472a-b74c-f344e7c58290 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.076096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.107238Z digest=sha256:1e8809a8ccab1b9741f1329f48a4731cb5e5f4393f9c1208b514504dfcdf2d73

Observation 8ecd0505-1d87-42a7-9340-5b828ef930b8 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.062169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.110683Z digest=sha256:d762e03c93d42c543e096e2eca65da1da9937267d89beb09993b217e1dddc999

Observation 97be51ba-aebc-406b-bc6c-3bed9092f080 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.048299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.113951Z digest=sha256:3b405eadab683dbd62a6ef37ea376b0ba6bc461e465aeb32c6ff1115cdf650f6

Observation 39a73c8e-ddb0-44f6-bde8-197070c9fd57 · outbound

This paper cites RoMath: A Mathematical Reasoning Benchmark in Romanian.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges RoMath: A Mathematical Reasoning Benchmark in Romanian

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.004169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.004169Z digest=sha256:052fb89d53c4a005794f35ef059e410e3966dff018ab6f4752181c3f070cd541

Observation 99e823c4-d10a-44b7-8267-03c14e08053b · outbound

This paper cites Yes" or.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Yes" or

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:06.019128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.121229Z digest=sha256:62613065d0c17b7e1d20e3faeb961fd797830c1b21e606355fe7f137546ca581

Observation 8f3609ab-92a0-4350-b45f-acce94fb949a · outbound

This paper cites MathWriting: A Dataset For Handwritten Mathematical Expression Recognition.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MathWriting: A Dataset For Handwritten Mathematical Expression Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.013674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.013674Z digest=sha256:8d83878124a8a9db6589a4d6d68e60523c67a7783dd6322fc6d80d150c5a37f5

Observation 1a26adb0-27c2-4099-bdec-56177def482e · outbound

This paper cites OpenAI o1 System Card.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges OpenAI o1 System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.018816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.018816Z digest=sha256:b1e9ced0bb64ef30dcdc9a82533e0e9d697f20899328328dac4805fe053db8d7

Observation 2ba5b79c-1a85-47ad-8470-fae1c7eee7eb · outbound

This paper cites ATHENA: Mathematical Reasoning with Thought Expansion.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ATHENA: Mathematical Reasoning with Thought Expansion

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:29:05.575607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.029306Z digest=sha256:01fc4710e09eb0b78a26dbb5cf95c09277365f1eb533f7729e7f760a20db6f44

Observation 3c05723f-389a-4926-80e6-a6452dd96748 · outbound

This paper cites Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.034781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.034781Z digest=sha256:f4eedc3e70b26c4c89fa72e2b93d2f063fc7aa847a8ae4d30403f88e9287598e

Observation 5e137c8f-3e85-4ebb-8143-dd8c7d61a794 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.061961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.061961Z digest=sha256:e38ac1a7d9b22214dd85797a80b7d6258742c73b0642084fed9b876c314abb2a

Observation 74039854-4833-48c0-b298-fe04cc271d58 · outbound

This paper cites Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.067239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.067239Z digest=sha256:fc0abb6426855adf2184300a93b7264ca22f8aa28fbff2b230160d37154cf5ce

Observation 0d3b8f3c-48c8-4bfe-af81-9f0bc5cdc11f · outbound

This paper cites Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.071773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.071773Z digest=sha256:b89294f8cedb0432025718f602cd5159a599a2539f2981405c5828eb160a8fe7

Observation ae4675dc-c168-4270-b17a-f8975df4f775 · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.077208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.077208Z digest=sha256:c3e6a0223819882df8b4d85b866e219f02c2ff04b19c092209ff55122953ead5

Observation 5205e4a4-ebd8-4836-9aad-4842212751ab · outbound

This paper cites Proving Olympiad Algebraic Inequalities without Human Demonstrations.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Proving Olympiad Algebraic Inequalities without Human Demonstrations

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:29:05.367816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.082766Z digest=sha256:a765e9f208205ecb39631ef7cfc323ba7cf34747e6b495f1273258c9f659b77a

Observation 6c1acb48-a2b3-44e6-80f5-9c98c69cdd80 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.092179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.092179Z digest=sha256:083b719aafd866309f265ca15c608889544e4d56e28c93349c42cda57c93d6fc

Observation e62a1cc0-3c9f-4ca8-8ad3-b97b642f8995 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.097298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.097298Z digest=sha256:a973efb6956c6bb9dc25d2841b96da23545c302c1f4721729ddf512f7a8d7b21

Observation cf14c607-42d8-4e3e-a264-24b39842d7f6 · outbound

This paper cites SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.102595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.102595Z digest=sha256:176b44c7b773ca4e39546041524c53416f0b15519d1c285907b35a8982e8bc34

Observation f126ae01-c912-468f-9899-8d718e7c2bc2 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.032432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.117743Z digest=sha256:1b64fe26a6d9a3eecff11796e0047d498e985d6ddc6968a191c1f88dfe65f22c

Observation d1ede7c3-80f4-49d0-a788-44747cc9c376 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:06.001215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.126110Z digest=sha256:4f9df8ce1c3298c45197563b8531e5bdc0b153093d112b6da720bb53f1ab6026

Observation f516fd9a-df6e-45eb-853a-b9fc42bb8151 · outbound

This paper cites ❷ Bottleneck in Data Diversity:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❷ Bottleneck in Data Diversity:

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.987025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.130437Z digest=sha256:02fca834d39f6593b70f156628853304c91f4f8210718f1604f810e13a282283

Observation 7c0344c6-3a91-4975-9fb3-68349f7e9834 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.969714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.135877Z digest=sha256:473c4a5e470b1bee29d44ae1c8e05a4d4844c446aaa6f35dcc43c1e09b7150f7

Observation 7fbdf47a-d43f-4d72-8a29-8f07dccf4b64 · outbound

This paper cites ❸ Bottleneck in Data Scale:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❸ Bottleneck in Data Scale:

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.953781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.139316Z digest=sha256:4aa6eb8ed2866a340fba80b8b10dd1ea97019ca759a5a88b259b452b66119420

Observation 42ae1ead-5673-4829-8ea2-84d0008536a4 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.938432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.143243Z digest=sha256:54fca9dd1c70fe05173c6d321461368f78a133c33d1cb5298a47adb34c0512a9

Observation f1fa7cb8-3a00-4e12-84cb-de91dca62478 · outbound

This paper cites ❹ Based on recent trends in the latest works, we further propose the following actionable sugges- tions to address these dataset bottlenecks:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❹ Based on recent trends in the latest works, we further propose the following actionable sugges- tions to address these dataset bottlenecks:

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.926069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.146955Z digest=sha256:ab3f81d397b8cea5230fd0ed8f1e2af14575ec427a969f83f2edc1c6f58fa52f

Observation 0509f27b-61dd-4520-a39b-4d72047310c3 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.913504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.150886Z digest=sha256:4f99617ff14a848e0ea2a9b21a378bdd71b77e7b9290976837a2c1e760be4583

Observation a6ea1e75-6233-4dea-8ead-24338f149242 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.901463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.154775Z digest=sha256:d181bc54c410e7c9596936bdd2fa2cc83ae009341b69fedb86f1b5c226d32510

Observation 34f89647-9b89-4ecc-a502-643d4901d3ba · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.890428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.159774Z digest=sha256:9418b8f649bdd245fcb0c27da08db00e457e010771f7fcb2c9bf8b397b5b205a

Observation d2387a89-a605-4f65-8a6e-98ead558ba33 · outbound

This paper cites spatial reason- ing.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges spatial reason- ing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.879433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.165558Z digest=sha256:48986e9794e9ab36bfa8a794e3b9fa5bcd94d28222774122f5b165165775dda5

Observation 87b450de-ba20-4746-b819-e6677eb04a31 · outbound

This paper cites Use domain-specific few-shot examples (e.g., providing figure-text associations in geometry) to guide the model in switching reasoning modes.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Use domain-specific few-shot examples (e.g., providing figure-text associations in geometry) to guide the model in switching reasoning modes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.866455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.171547Z digest=sha256:af6d8629dfda3aea6d2f241e89433ca89f971cd2ca1722519166dab8cded58ab

Observation dc4caeb2-0385-4b9c-b915-3f51b631eedd · outbound

This paper cites triangle.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges triangle

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.851907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.176510Z digest=sha256:53035e0978dbb1e6c312a1b6bb7e300e421fe31c83195d842c283354ac3bb4e6

Observation 578fe127-015b-4dfb-86e4-4ba4bc36f342 · outbound

This paper cites The limitation is that labeling error types is costly and it’s difficult to cover all long-tail errors.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges The limitation is that labeling error types is costly and it’s difficult to cover all long-tail errors

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.837944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.181253Z digest=sha256:5f965800656e02d8d5ab5fbfeeb39d0301d03352540f83d939e1680f4a1f5e39

Observation 72c164b1-2f2e-4abe-aa08-1a4a876ea596 · outbound

This paper cites computation-logic-conclusion.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges computation-logic-conclusion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.825041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.200899Z digest=sha256:5d581c9cdb77d9e639cbc4e78be47961bd3cfb99cf0bb11718096f0a990f739d

Observation 617e69fa-b12c-459d-8fc7-86a75370afc4 · outbound

This paper cites The limitation is that tool invoca- tion delays affect real-time performance, and some errors require manually defined detec- tion rules.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges The limitation is that tool invoca- tion delays affect real-time performance, and some errors require manually defined detec- tion rules

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.807600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.205542Z digest=sha256:9b1ecf21575adb098862bf5f8cf8106389ccf0e28f322bd1e3d67f714ad3b85e

Observation 17331e94-06c2-48d6-ba1c-c777f22b8354 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.794363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.209565Z digest=sha256:17cb170a951fb645cc9fdb6940e24e0fc2c2fb729b56fd04313818f34f04db2f

Observation b3132461-8613-4f86-adea-aed2e29164c6 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.780786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.214385Z digest=sha256:21a252d7a736b5297abb04339beefc2dcf6637b772ced818e1ca76ddbeeaf881

Observation 77173011-fb98-4b4d-95be-5e327b63df6b · outbound

This paper cites If the training data coverage is insufficient, test-time strategies may not be able to compensate (Ke et al., 2025; Chen et al., 2025c).

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges If the training data coverage is insufficient, test-time strategies may not be able to compensate (Ke et al., 2025; Chen et al., 2025c)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.768170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.218715Z digest=sha256:3e148d4d897cb28202b97b4fcfd02260c5a5d63eefbd7d9088ad85016b4c4f7a

Observation 61ec9488-f54b-4770-b841-6691ff24c0d7 · outbound

This paper cites trian- gle.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges trian- gle

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.752711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.223162Z digest=sha256:ae4c9129e15f38724ddf2b0e05e5dcd33417c64381f22da5eba03ed80aa89aa0

Observation 3433ae2a-0a2f-4284-8b6b-2ebc63c08367 · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.738351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.228180Z digest=sha256:0741bef642655ee24269c1f4bb0ee57d65e6bf0d3c2668d2927ba9e5f6ea12ad

Observation 385697d0-fe96-42a0-86e6-ea305380a1ea · outbound

This paper cites ❸ Error Feedback Limitations:.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges ❸ Error Feedback Limitations:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.726506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.232068Z digest=sha256:d9074cef9015230036bb5cc312f1aadf633bc32e1f820afdcc28c9889888b663

Observation cf36e3f8-1cf0-47f0-99b0-e5df23e2bebd · outbound

This paper cites an unresolved cited work.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:05.714406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.236565Z digest=sha256:538b1fcd3e081258c165d4e8e313c2102e6a37d2e8966e8448f2fffecf5519a4

Observation dc4f8666-a545-44b0-94fa-83dfa2de3e96 · outbound

This paper cites Long-tail errors such as rare symbol confusions may be overlooked.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Long-tail errors such as rare symbol confusions may be overlooked

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:05.699129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.241019Z digest=sha256:213cbc41ed9a5f4b6d9398a957a4f7c947473f14ac589e773079cf203dca706b

Observation 2e35f305-c370-4cec-9ace-74211b0a6271 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:04.986025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:04.986025Z digest=sha256:959de1539fe75fd6defb4042beed0d67bb70486f2d217d829330e50434066821

Observation 2d6b954e-ddac-4748-a237-4c8a0eb053d6 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:04.991803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:04.991803Z digest=sha256:be1e2fa744d076679e474346b65c3fe215669f615b625a37a9ef3bad25e17d3c

Observation 100240ca-376d-4f4f-8463-db2c396533e6 · outbound

This paper cites MWP-BERT: Numeracy-Augmented Pre-training for Math Word Problem Solving.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MWP-BERT: Numeracy-Augmented Pre-training for Math Word Problem Solving

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.040249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.040249Z digest=sha256:fe276f7effe9a54648e63e038b12430b40cb068a6806ab35384a96d336750c9d

Observation b82d1d28-879f-4225-8e89-516e33005d76 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.023302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.023302Z digest=sha256:7c232026c848d6aef1a8bcccd5437c4e81bce81868cb612a44aeb5707cd9fc2f

Observation 17ade6d4-05a1-499d-bb98-db823689930b · outbound

This paper cites Advances in Neural Information Processing Systems, 36:5539–5568.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Advances in Neural Information Processing Systems, 36:5539–5568

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:06.088222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:05.009221Z digest=sha256:53edece5893f2149210975d9e9614ba5df5ef3670e377399d2827d6fc22fde85

Observation efd5c1c7-8f9c-4680-b841-34e91896f3d2 · outbound

This paper cites Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:29:05.652045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:29:04.996663Z digest=sha256:9e78578610b6f18f9772bc0d0f13ae53592f008fa61d90f5e26077103cdb2d23

Observation 9aea032e-f4f4-49d0-8fe4-e25da0fb9f47 · outbound

This paper cites arXiv preprint arXiv:2501.04686.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges arXiv preprint arXiv:2501.04686

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.057155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.057155Z digest=sha256:61653a3743c2a94c2b1f14912edf5eddbcc4d83528b0d21323573e188d660e72

Observation e2f9304d-de2f-4d5a-9e88-fc8542099eea · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges BloombergGPT: A Large Language Model for Finance

Reference 2256

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:05.087342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:29:05.087342Z digest=sha256:3ef322925e647b900a42d4c0036a63d570149be04befd2a1579b0f5d24c4cef2

Pith citing papers

Observation 82193361-4bc8-42a1-bc94-2b4db7320522 · inbound

Reasoning Language Models: A Blueprint cites this paper.

Reasoning Language Models: A Blueprint A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-10T18:36:55.267690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:36:55.267690Z digest=sha256:0006b288e60c2323465ce4153de20f6a12d4c5d6feba80be9f87354f86853f43

Observation 8fe2dede-2be1-4b50-9830-bc703c188f0f · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.803728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:532d70e6de6aca4cb92d76fe88907cebee41b56830cbd43c94030d1fdab5e483

Observation 71045df9-91b9-452a-a982-4d4a36911c30 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:11.970200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:11.970200Z digest=sha256:615a5336c85e8813a5053ce139ddbd7b09b7da93bb4cc2f003e7156b1239c946

Observation 6df75775-2aec-45fd-b1fb-a298f94f6a80 · inbound

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? cites this paper.

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:31.553533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:31.553533Z digest=sha256:810e91d76154df31e02306a227b0ee1a37e0fa842b520b53450265976ec31de9

Observation 85ca2948-0a06-4aa2-a7c7-f3df5c4d17da · inbound

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring cites this paper.

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:37.737664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:37.737664Z digest=sha256:57df6602018d77c792101d4a060e3b7f13659cfa27c32289d387ee10ef300dee

Observation a5c85ee0-ca3d-4a0d-ba8d-3ebe4f30f42f · inbound

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method cites this paper.

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:42.587497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:42.587497Z digest=sha256:54dabe53a6c31ed3099e605eafa1e97795613c859538d678f604be57254cbe6f

Observation 5b5a7fe5-dd8f-4268-bb38-eaeb8f7e9bd6 · inbound

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey cites this paper.

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:00.698793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:00.698793Z digest=sha256:d4a5117bfdfe5c4fbeee698592390b25b8f443abac318f01a4aee80253d429aa

Observation 8aa1ca2e-288e-453a-b26f-315745af4d4e · inbound

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis cites this paper.

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:20.521456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:20.521456Z digest=sha256:1dd320afe16b78facb81f3f9d580383b4000e4eb8a3db7f8aed0b9cf2a98db50

Observation 0d5b4e13-f700-482d-b896-2f0c99511bb8 · inbound

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities cites this paper.

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:23.839050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:43:23.839050Z digest=sha256:02af55b835c3ba7b8ec446d8d915f5f659e8903e28f5a83c5c1ebd027268c609

Observation 1c2ee46f-164e-4739-9ac5-5f2c1d463460 · inbound

Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach cites this paper.

Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:33.123081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:24:33.123081Z digest=sha256:785c8e24bcd3455c5bd640b55bad51988860ddd122f8b9dd000c14d7a75a0f19

Observation 96b8d507-a2f3-417b-807c-1d4d8df6c155 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.511366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.511366Z digest=sha256:08f06af0c503acd7563d7fb667071247bcd750ee2e92637c07771e29ffb12cee

Observation eaa1e2b9-109e-4539-9e06-d948c0c055f4 · inbound

Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons cites this paper.

Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:50.011424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:50.011424Z digest=sha256:309069af82d6f409d94e07c7a60eb0d43e1a0ccbabc84dc44a88cf92a1960db9

Observation 37c2ba42-2db1-4040-a1bc-a13d2cc8e47c · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:53.783089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:53.783089Z digest=sha256:4ee3d84d4e09a74272328c7971d9f3f2b1beaa3bdade8dc9650920e1031a093c

Observation 73c582c6-f868-4444-b2fe-27fc3443fa46 · inbound

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines cites this paper.

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:41:56.049889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T00:38:43.897231Z digest=sha256:77c76b8c9843a75849f61c621867acb814f3d1a15712dcf3c7584c1b38da9f17

Observation f34545ab-8bf3-4f76-a439-9cd855772702 · inbound

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning cites this paper.

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:29:33.508768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:29:33.508768Z digest=sha256:4784f533720cd0dc269cfb5cbf65cf307fe0e6389047c2f0a7fd3f9a4cdc3c3c