Pith. sign in

Paper Citation Record · LEDGER

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

As of 10 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 5 inbound Pith citation observations for arXiv:2507.03483.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03483 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:15:52.926857Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.598214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.092401Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved81
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0196b675-2020-4ad9-afd8-792cc86d3ac2 · outbound

This paper cites Qwen2.5-VL Technical Report.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.545119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.545119Z digest=sha256:aeb4f27c5ed4fda4d72a853c4598bb666bb26d950f1869cacdb4d4d9a08e63fd

Observation c001efd2-9060-47e5-8216-628980107bda · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.549797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.549797Z digest=sha256:c3c698718c01d1c19ce59f2b96381c8ec6f86336aeefc7f48580af74a2dda507

Observation 4e837ba1-77d1-4473-a9aa-2d5a7ee160d0 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.553874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.553874Z digest=sha256:280320415fc079c61437866783c95518572c4cff4b22ab1316aae33ae9b579d0

Observation 5e4c053a-6351-49d2-97dc-cef948232513 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.557584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.557584Z digest=sha256:996b4ddb77d21409ff23ef44ea0e4a71b36136c9e32cacefb9def9a9011e8d4b

Observation d0d01ff7-c1da-42f9-92ba-1fe13bdeb5ba · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reason- ing Benchmark for Expert AGI, November 2023.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MMMU: A Massive Multi-discipline Multimodal Understanding and Reason- ing Benchmark for Expert AGI, November 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.561578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.561578Z digest=sha256:ac0f7306743e68fb186ce9733f0546f266f59f6a7d8123e8dd7a692ae162654a

Observation 6dddc08f-5c55-4e0b-b062-4568f8cfaf57 · outbound

This paper cites Sci- enceqa: A novel resource for question answering on scholarly articles.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Sci- enceqa: A novel resource for question answering on scholarly articles

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.565003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.565003Z digest=sha256:85a3d6cf3e8ad72e2e4941db447a63bb446d2fa7aa9dee5a108f1b6e1794f621

Observation 461f4d28-97c3-4e46-83f2-50140d5ed506 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.568687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.568687Z digest=sha256:b269ce6cebc1f471967b68c054c1e7b050acbe1648472d8e0c8ad0c108d85e62

Observation eb3368ba-232b-4847-9056-24d3778465a7 · outbound

This paper cites GPT-4o System Card.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.572152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.572152Z digest=sha256:f813e2a8d8095ac2e9bd9f3c5493cbd7bce50e3e6d12fcec68ecccae8a625d33

Observation d73a7183-14f4-4193-b473-f8edbb384ad2 · outbound

This paper cites OpenAI o1 System Card.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.576901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.576901Z digest=sha256:edd9b0a456630242d8f040a71a90d7a89c199b3fadcd703f130479e3aee1951f

Observation 11ad014d-ed3b-4848-b195-3087ee8f5572 · outbound

This paper cites Introducing openai o3 and o4-mini.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Introducing openai o3 and o4-mini

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.581244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.581244Z digest=sha256:53fe6b9506a2ee6f15bce4d47adc4291ebe7f5591070492ac3ca56bb0a81f5fa

Observation 7a4441b2-d93b-4398-98a4-55d5c15d4457 · outbound

This paper cites Claude 3.7 sonnet.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Claude 3.7 sonnet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.585255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.585255Z digest=sha256:321334c0c26ab04216732bdcda0d4a67112fc3dc18ca2b4a65a0debea7db0430

Observation 7d4f9ae3-d4ac-478e-8621-9658363a762f · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.588780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.588780Z digest=sha256:5ad4bb25f080fe5a885f087779de70c38f11d10b2f9632ed00eeb7a0c582f262

Observation c54c5915-1fc4-4630-8e7d-3e80d278ea9f · outbound

This paper cites P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms, 2024.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.592208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.592208Z digest=sha256:589b628f6a72aa03386b21ee1c6a458c70738a2ce8c9fd989eeda9ce82121375

Observation a43d3464-a496-4d92-8aa2-7d56843564cb · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.596012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.596012Z digest=sha256:6f2158aefdfe139390d5d3110ea4c5ec47a6f08e4c59f6f9a226313bd75a1576

Observation 1e0f301f-09f3-41a4-9887-c4b61bbd5631 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Gpqa: A graduate-level google-proof q&a benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.600800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.600800Z digest=sha256:6cd0238d2a1849228b6b61d6ec1d9b9d62aa10bac09bd2d86388fe57cb3d4a9c

Observation 61b533d6-b753-402f-a077-0458583cd244 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.604524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.604524Z digest=sha256:81f3ee047e6f1b5555fbeed6b79bf821bf8586d26a42af4db611284783bdbf79

Observation 11a9397a-b159-4a05-a6f6-70fb31f82779 · outbound

This paper cites Cumulative Reasoning with Large Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Cumulative Reasoning with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.608615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.608615Z digest=sha256:3dc736f8b986f68d8e8ee5562a039ea67febbdd930fea03bf99a9f7e998368c8

Observation c048c313-80d1-4d7a-aaa5-70a70998b260 · outbound

This paper cites Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.617862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.617862Z digest=sha256:24260deea80183de9cfc4c9ae8dad92e7f05a69fd17da525a4b7e576ebf58f43

Observation 2771ed89-b22c-48f2-9a6b-6cae50a9822b · outbound

This paper cites International standard classification of education.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset International standard classification of education

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.622384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.622384Z digest=sha256:963d9d39e18fbf9a1abd05a9f9b0d4e5407ef6a4b0ca364558281af5c5a827d7

Observation 19dbde4b-59d3-44ae-a3a5-bae7614a5a98 · outbound

This paper cites Mind with Eyes: from Language Reasoning to Multimodal Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:15:53.550336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.625803Z digest=sha256:425cb0ee9e6b49ec7ec9f64a94bf3bcd162533569e0f761898296e66f751afbd

Observation b944415a-66da-46c3-85df-3faed605457d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.629640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.629640Z digest=sha256:db5d678562c5508e140b61ece65dd52c06a3146df268c9ea60d53fa280930d63

Observation 814105ab-5d6c-4ee3-92ab-bf4575a160b2 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.633699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.633699Z digest=sha256:7b49290badf9a4eee9c6536bc2adf7be9824918c73d294dbbcab591061184e71

Observation 31c21de8-0702-4dd8-a981-a1f2b467ffae · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.637922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.637922Z digest=sha256:72e285824956f52a561d99f48dd49a6e63d123b0c43bbad48da5bf796df62e30

Observation b76387e4-3dff-4ef0-8212-d05f5bca5e0f · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.641896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.641896Z digest=sha256:10739e1e820a33c158034c62c0618b764d7f9be31a4b7de7e478dd89c7f9891a

Observation e6bbca02-a527-4cb8-a7e2-05fd2c2b8596 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.646946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.646946Z digest=sha256:e83eef3d2f1887e1d75e4198ed2c145a3909e465395c17e5e3fbb4cc1a6017f8

Observation 14180049-22fc-4abd-be0a-90ce1b3dd6bb · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.651200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.651200Z digest=sha256:8fc6429f9b03d2e9fb00c78db1f52f0fe020c708b34c4d7e8a3cee77ec19ceb0

Observation f5a3d15d-6335-4ea6-8e7c-e02618a8b1c4 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.655037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.655037Z digest=sha256:1b3795eeb50d07271515a3fb569784935e1458f1f270706444cfe8bfa10f1d68

Observation 9c59b816-d211-4455-b145-30517c69511a · outbound

This paper cites HallE-Control: Controlling Object Hallucination in Large Multimodal Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.659099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.659099Z digest=sha256:d8abfe859cdc292f9a00b4248126ec89cdfe4d680f1353b18bf10a871eb13ae4

Observation 5fd9cdcc-245b-4092-bd2c-4ce3568ad1ef · outbound

This paper cites Hal-eval: A universal and fine-grained hallucination evaluation framework for large vision language models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Hal-eval: A universal and fine-grained hallucination evaluation framework for large vision language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.663171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.663171Z digest=sha256:b0103ef917286f1c958756031ac012a030ab729cfad6483b361129eb4e3baeac

Observation dff49deb-9081-4e57-a825-2d4fe27a794e · outbound

This paper cites The second half.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset The second half

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.666627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.666627Z digest=sha256:aec0681559713f87572c1586b1cf222457243c85a9f12ccf11f0c0a5e8d5dbc4

Observation 2fb3b996-4df1-46d4-aa02-6ed8afb6f6d9 · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.670814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.670814Z digest=sha256:2015cd105b82454c0ddd4ea83967ade4074537bb21a39e8b6c56fcd0506408d6

Observation 59a54d08-71cd-4773-a16a-f1ba0da0294a · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.674522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.674522Z digest=sha256:f7898afae32b35017dd8d46334b90d57ed61729a7f8efd1a389efbfc3cee2666

Observation 17c5c31d-b620-4493-b796-a8b6dd59cdd3 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.678109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.678109Z digest=sha256:34b455c75e68e7158ff0066885464182ee6dae0b6c55a2f8c9b2144b71a415b1

Observation eec4027b-0b94-4579-b9ae-994cd9e300f1 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.681595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.681595Z digest=sha256:c4ce7321107fffe21022a51a7ea558a7dded654a2f3f55936eb24d9cb9f52077

Observation d88e2028-b93b-4ba4-aead-19819e32635f · outbound

This paper cites Gemini 2.5.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Gemini 2.5

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.685339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.685339Z digest=sha256:c32a768c1495c57259feb4a86ea1e2b999199bd1fe6121ff9be0a1d87de50ffd

Observation a4ae9a29-feaa-4155-a224-c9d95132001b · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.689598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.689598Z digest=sha256:8c36e27753ca7c6c69b8209e70bd5edfd86dfe1637e8bd4927eaa2444df3cf03

Observation 57d9a831-24c3-49d7-b08c-97ae6260c925 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.693320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.693320Z digest=sha256:83a65326139a2ccbfda26aa95c994edb2c12442db97624761963e6c5d60d91a9

Observation 717384d5-dc58-4e87-b348-d36840ac1aa1 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.697483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.697483Z digest=sha256:ce955dfde3c4e08e3e45d87a2d24a0b9261ff331e5824995564a37c890ca6a39

Observation e509b6e9-ec44-41bd-865e-d1e0d170bfc1 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.701699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.701699Z digest=sha256:474fc943784672647d2efd6f375e4d892f85cde50fbea911f894e7d8bd139cff

Observation 59c29797-7406-4cc9-8f12-44a690f1425f · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.706978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.706978Z digest=sha256:039c19e32b25604f4a6536bc4013ba5f1eca09c610f6a68e6602cb5652272635

Observation 8bfe9a92-6741-43ec-9626-7cc730452232 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.710933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.710933Z digest=sha256:4ca07a7660ab87d960c286106c52e8906e7fb560f15a49e7b31e18b9183a5a03

Observation c983e567-740d-4f0e-b249-3590d7cef571 · outbound

This paper cites Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehension.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehension

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.934789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.714872Z digest=sha256:1e885633daf3387be1d7f18821f2d1a0c07600c6aee828bf686e88e4f770b46f

Observation 1a482397-f3b1-4315-9fa2-f3cf176d83b3 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset WorldSimBench: Towards Video Generation Models as World Simulators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.718735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.718735Z digest=sha256:ce2dbd6b458d23ebbdfa1242be1c0e8b1b6e702df33e7ddb2e971dbfb7fcb313

Observation 415e36ab-9b69-4de4-a2ac-106cb43d10e3 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.724401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.724401Z digest=sha256:c9d697ee9f6e0156cd881072fdcf5222a7fc44e9ad2daee69f0da5e7aa97b067

Observation c03f60c5-0ec2-4492-94af-b179664f1177 · outbound

This paper cites MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.729171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.729171Z digest=sha256:64b674689287d35c4480f3ddef11765e297b5a31684f466075c5967a603702bb

Observation c569df7c-8152-4629-8373-70c74568eca3 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.733114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.733114Z digest=sha256:8167c0603944e68df17fed09d26d0d8e3f4a784d01e35887d5ba560399587de6

Observation f5e5e55d-0c93-42bc-8c6d-ac9082eba907 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.737246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.737246Z digest=sha256:7835397dfe63ed54f45c68351d2d0d71d54d7fd6b9eeef71b89df1718546e25d

Observation 4d4fd369-42ff-48dd-aa16-8a3b2a588772 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.742706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.742706Z digest=sha256:ee9582347855ccbdec4575329d25a0de077ec93675c516d7c76f33898b4f7d62

Observation 8f0d0dbc-bb94-4e32-af2c-afdeb0ac4b09 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.747130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.747130Z digest=sha256:19bd5295383ad3fd68db73dc00a2d04a9c7bf37dd6fce7b0ccff5b210ecd8424

Observation ab0e998a-2ab1-49a5-a659-96c0d5b83aea · outbound

This paper cites Let’s verify step by step.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Let’s verify step by step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.752298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.752298Z digest=sha256:dfb06741f2b0a26767ad19d568e66f2e6cce68baa473ad57c9a3ed0a9c9227c8

Observation a8355300-10c5-4c64-bd8c-6442ff837d21 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Training Verifiers to Solve Math Word Problems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.756735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.756735Z digest=sha256:2c54271f5420a4f6d7151a6bc6ef699218b9b0080079bdffde03a135b1d9a3f3

Observation 849056ab-8732-4605-8bed-aa8161486bed · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.760921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.760921Z digest=sha256:c7dc7134d5cb4a8368193cc7baee9b04e795666f3607652b68709df846d4c5ad

Observation 7fae0959-5729-4815-b117-045361ba19bc · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLM Critics Help Catch LLM Bugs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.764833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.764833Z digest=sha256:f309c37833f78f7ae43eac24260b4ae809f56b09b995c60fb8303e8fb65d3915

Observation 081ce2bf-beb7-4cee-80f6-57e90bf31711 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.768877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.768877Z digest=sha256:7579f241539193d2e68ca19702a6c4b6b0065f15b058b423d5dc276ce609714a

Observation 29c5a000-9b78-4855-8090-ad9185c86536 · outbound

This paper cites Better Process Supervision with Bi-directional Rewarding Signals.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Better Process Supervision with Bi-directional Rewarding Signals

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.772660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.772660Z digest=sha256:426e682763948ac999cfacedb3dc231e7e3b4d1307d0bcc50f6e63314ee6e593

Observation 51cc0d4b-bd68-4089-99b2-3295328529ae · outbound

This paper cites Bandit based monte-carlo planning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Bandit based monte-carlo planning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.776986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.776986Z digest=sha256:077ca3fcfb3a67d2eaf044825b3459e88cd33bb5326495ce1c2be7635a6d5d76

Observation b15b2ea2-9412-4577-b323-58c9e367dc85 · outbound

This paper cites Efficient selectivity and backup operators in monte-carlo tree search.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Efficient selectivity and backup operators in monte-carlo tree search

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.899442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.780751Z digest=sha256:d5d563ab8eafcd167f21e9d233ae3e1a60603d401c4185695c0401f51a496deb

Observation d6a07ceb-aa06-48d0-9f29-7c11db37d6ad · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.784527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.784527Z digest=sha256:791013b1803fbc485efb66e92680a7322f9bee8b30dee4cb14a23b187e613343

Observation 8e28e85a-a52d-48c4-81a8-bcdde58098ea · outbound

This paper cites Process Reward Model with Q-Value Rankings.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Process Reward Model with Q-Value Rankings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.788301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.788301Z digest=sha256:e0f0b3e6edf13ec0223cf8d285174a8fe5c6e934cf80db332dd5a08ed9d6ff1b

Observation e9b66aa5-718f-4c68-a563-11dcb260dfc1 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.792261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.792261Z digest=sha256:93a8ed8b8b5bb621252908cfeb07b654d6c07312e0387b6e3c47de6ef33d11de

Observation e31829f5-60f1-4b74-87bc-5d1de299266b · outbound

This paper cites SALMON: Self-Alignment with Instructable Reward Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SALMON: Self-Alignment with Instructable Reward Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.796856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.796856Z digest=sha256:329dd057e95518d679cacaf8cf067939bde7ef575e31c0fe3affc708928f3812

Observation 1339cbfa-27d2-4f3c-b29e-0928a7dfcc23 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Constitutional AI: Harmlessness from AI Feedback

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.800618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.800618Z digest=sha256:9912e2003b04efbcd371ff281fca4aa496b05b32672f8a6f518051d5f08ed94d

Observation 97a9a0db-f0a4-4fdb-88dd-c2ff8fc7abf6 · outbound

This paper cites Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.803990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.803990Z digest=sha256:0a5d7a9c64af92e036bfb1f30789fa4e35808a9bb504772dab053f7f0897172b

Observation d92fb749-f394-4f1a-b187-4fc6ffdfd77d · outbound

This paper cites Examining false positives under inference scaling for mathematical reasoning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Examining false positives under inference scaling for mathematical reasoning

Reference 65

Resolution
verified exact
doi, observed 2026-08-06T20:15:53.023616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.808246Z digest=sha256:5fe43d1285796eed4f46b00e1b63aeb93c798119fc83c93ec69702cf5cf913e5

Observation e3725d2b-e79d-425a-8ce4-15a9d4f1ad1a · outbound

This paper cites LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.814532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.814532Z digest=sha256:ed3c2d1d1b74f1dd780874d84bec6be562edd9004d6d6e934a3d9b292d4d175f

Observation 820699b0-56dc-44d2-a1ca-65509caa0638 · outbound

This paper cites When benchmarks are targets: Revealing the sen- sitivity of large language model leaderboards.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset When benchmarks are targets: Revealing the sen- sitivity of large language model leaderboards

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.887422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.818291Z digest=sha256:4a1ae6aece6a2ed84767ff738e73ce7ca63a58cb195cee39d0acfc4dde5a44d5

Observation 881975a4-3c79-4864-95e0-367ca9a90bc4 · outbound

This paper cites LLMs may perform MCQA by selecting the least incorrect option.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLMs may perform MCQA by selecting the least incorrect option

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.874770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.827005Z digest=sha256:a7530af28b05e88a278f3be0f506eb988a12dd55e33608c8f7e6ede57f656b7b

Observation 29b595b6-4567-4e63-84df-20843a84b03d · outbound

This paper cites Llm-evaluation tropes: Perspectives on the validity of llm-evaluations, 2025.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Llm-evaluation tropes: Perspectives on the validity of llm-evaluations, 2025

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.830477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.830477Z digest=sha256:e1b7f65c11a1641c3fe7fdf6f844f244b5e3d6ba2c39480356455f47bf43ea13

Observation ac13ee45-6f5e-4d40-8333-138f75f524fc · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.834798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.834798Z digest=sha256:64d2b101075c9d9d97550d5261738586668c1a13c57eeec6267df332ff9614b1

Observation 2cbaac74-cec4-44ea-946f-42cf9c594d4a · outbound

This paper cites xfinder: Large language models as automated evaluators for reliable evaluation.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset xfinder: Large language models as automated evaluators for reliable evaluation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.863526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.838908Z digest=sha256:656ed696e6954bdd10c7dd3f05d804526e22a3a8de274f9ad2635a5dc4cdab7b

Observation 0f38035f-23bd-4724-9d29-014052b7e437 · outbound

This paper cites Gemini 2.5 flash.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Gemini 2.5 flash

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.851010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.842692Z digest=sha256:26b2f6d3d576bb6d3c2efd0f5593a240a6b125058fa45d1d653c8fedf599d478

Observation 7d9dcc68-1485-4200-93be-794bf78da20d · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.846633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.846633Z digest=sha256:13f58a384995c285971174d1a1560e8b0d3f62490ccea09abc72ec435990e135

Observation 855707f7-c8b0-46b7-b564-0ce9163f59c1 · outbound

This paper cites Internvl3: Advancing open-source multimodal models with native mul- timodal pretraining.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Internvl3: Advancing open-source multimodal models with native mul- timodal pretraining

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.839160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.850191Z digest=sha256:b1420651568f1e5866c6676f80a80f1a9dfb6ca00ec3c1855811e84e9e99004f

Observation 56920f09-1579-4f22-8f5a-91d6b6520a27 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Qvq: To see the world with wisdom, December 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.854871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.854871Z digest=sha256:afb6ce0747ba266cd85667ce53c2486dfa3b076e1b9ecb3cc41788c0e4ba9e38

Observation 1bca3adc-e7b8-4092-a8be-ad9def86939f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.858559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.858559Z digest=sha256:1a9a9a9902ca605bcdc90263a9ce3443358ef816ea90d8e68c275924547f9590

Observation c39b1df0-7116-454c-8268-c435a71761fb · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.863191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.863191Z digest=sha256:cc1fc9bc637f6e352af479ddd9b78358bda2659ed79b1c6d481c779395a0f4d3

Observation 333414fa-1cd1-44c1-88f2-de0b5dab8950 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LLaVA-OneVision: Easy Visual Task Transfer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.867152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.867152Z digest=sha256:4475208debdccde72ce7015b2f245f78ed432c7f6e08dee7d12acacf8c3582ad

Observation 86989673-95a9-4fe3-87f8-8f610e120782 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.871522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.871522Z digest=sha256:1efae5f10d50ea9f7f869b3223ad822a9743215a58ca7f7fe20b4cb3a10ccbd8

Observation c581f00f-528b-4a66-8d94-fa23c2970249 · outbound

This paper cites s1: Simple test-time scaling.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset s1: Simple test-time scaling

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.876328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.876328Z digest=sha256:46d4d783ab1f9c714f407e2041b5f1aa7408b3d74993ebe894817cfc317b679c

Observation 57807e59-0a43-4f93-b5c0-d06b67cc9349 · outbound

This paper cites Qwen3 Technical Report.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Qwen3 Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.880163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.880163Z digest=sha256:9150df68e8bbf10fb1f19124ba5d1e7fe9e154ee5aedb658812b08cbe7ad995c

Observation 6928dccb-5ae1-499a-a476-29a0ae3c78a7 · outbound

This paper cites DeepSeek-V3 Technical Report.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset DeepSeek-V3 Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.884044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.884044Z digest=sha256:9d34bc108f01a264558cfc78d42a18b479605b1cb34eb2cc4944721faefd0907

Observation 28faa975-416c-4b4a-abf9-c8e60afd8a80 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset ReFT: Reasoning with Reinforced Fine-Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.887684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.887684Z digest=sha256:22738306323cf1ce9692f725a4208224a391185b6d950f40ed939c5626a6349c

Observation 498d7beb-586b-413f-aca6-bce5709b0fc6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.891650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.891650Z digest=sha256:61265b6b856ded49ff2ca2aff1e899e6f71ea92ab7c45968cc80948c3ec3adeb

Observation ca9e36a5-7d26-4505-aa6e-fda642f5da38 · outbound

This paper cites Swift:a scal- able lightweight infrastructure for fine-tuning, 2024.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Swift:a scal- able lightweight infrastructure for fine-tuning, 2024

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.896058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.896058Z digest=sha256:cf2a4fd13421f86bb9bafa4ae8de5f0f22a994b09ba657c39f44f6182a8c1b7b

Observation 430102e6-d2be-48dc-87ed-cf6aa7da7169 · outbound

This paper cites fact verification.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset fact verification

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.812240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.900027Z digest=sha256:3c7afd1ae8f849a36c74401a743b92abe2a715d92a3ab77a6acbdf27830f6925

Observation ecfd68ab-fee9-431c-a6db-9fa838cf3f6b · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.800449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.904012Z digest=sha256:166a75f803d8f3b691ef1b33d9cd2cf30cf3baee51498ca1aadaaad3452d4ac6

Observation 554efa4c-683c-441d-bc93-e10c29d1600e · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.789297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.908029Z digest=sha256:a954a2d52ba7e3989ce01afeb41e17da2204a805e81d3408f2656be427fa8369

Observation 42294d97-2459-4b35-8f87-923266aed433 · outbound

This paper cites the southeastern part has rich forest resources.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset the southeastern part has rich forest resources

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.775845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.911579Z digest=sha256:859517ea263124f82b9dcb4012bdd1e2320f38a1135f736e2ee74c14152eedbb

Observation e3b84485-b1fb-4c67-8ed9-aa64db3d9c3d · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.763119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.915564Z digest=sha256:d11cdea43bc328e88eeace3d9c65f6a99ecebc5cd8d0c4f75bf8bb1b77f9377d

Observation 0912e1eb-79a9-48a8-a6d6-8a2c30e51ce7 · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.751893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.919130Z digest=sha256:69554547b9af72a81aa069f2fe893fbf96d146b4c87f2fcf2d0276531a6e0dc9

Observation 929a71a1-9ac5-4fdc-b706-9a5234b8099d · outbound

This paper cites an unresolved cited work.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:15:53.740663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.923169Z digest=sha256:b3a5a179cc653f0dbaca812742c4206ba23bcf8422278e1df9a0d5e1123e2d23

Observation 19e66c48-fdad-40ae-bc07-676440fc883d · outbound

This paper cites F Limitations and Broader Impact BMMR is a dataset that focus on multidisciplinary reasoning for multimodal models.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset F Limitations and Broader Impact BMMR is a dataset that focus on multidisciplinary reasoning for multimodal models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:15:53.728981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:15:52.926857Z digest=sha256:ec7f1b543cc1d8f7b9d400415d96efb1920011909064337b66eb3c5d61ec5a7b

Observation 85ad6546-f4a4-4d92-9676-795774282688 · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.744.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset doi: 10.18653/v1/2024.acl-long.744

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.822666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.822666Z digest=sha256:1cb59fa1b719d775c311eeb406990dc1228ccfb56ce8dfb29d221347bb8bc0f1

Pith citing papers

Observation 19eb5644-2edf-4276-bd43-6bfc7ba260a4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.884793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:b268da20ddde3a031a622c668192198cf77aa8be32d6151e8d812984ffb80662

Observation 043f1a4b-b753-4a0b-9e06-afe4d804ee16 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.067017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.067017Z digest=sha256:b5d209c5181eafe09c87ce5933a71f5cb2df2025fc62dc0f01d455ca0408ec3f

Observation ded2f978-9444-4286-af2a-5cfe152474db · inbound

Towards Characterizing Scientific Image Utility and Upgradability cites this paper.

Towards Characterizing Scientific Image Utility and Upgradability BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.854608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:10:37.254428Z digest=sha256:6f8b6756dfda67b08d32ca9b26161a153271a041343db037ea905e4421a8c4cb

Observation 6e2299a8-0fe7-4042-a0f1-785578fb7fe6 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.094168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:17e4a4b09a9f20eac3623c0985f15836a3666e022fc989f49c406fc1781458fb

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · inbound

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning cites this paper.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:703e9384b493d660d54cb82f7a303ad16599233464f1f9972841417e9046bddb