Pith. sign in

Paper Citation Record · LEDGER

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

As of 8 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2603.00546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.00546 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:56:35.246936Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T11:53:35.315405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:03:24.478700Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b95ece5a-b998-493a-afb4-1960e4c6ae80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.636015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.636015Z digest=sha256:43b91acda2d8f4ac846cba444bb28817f880ca18646e330a2f20b52fbbfb5f32

Observation ba55f520-b18f-42e9-a081-06f8bc823306 · outbound

This paper cites MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MLLM-as-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.702214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.702214Z digest=sha256:edd5c1bf544a91cf57b8c8db78738b260b6164a481b37ab729d1e5e54d6f7508

Observation 8fcd1560-87f6-4f87-900e-1b0d3c5986c8 · outbound

This paper cites Are we on the right way for eval- uating large vision-language models? InAdvances in Neural Information Processing Systems, pages 27056–27087.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Are we on the right way for eval- uating large vision-language models? InAdvances in Neural Information Processing Systems, pages 27056–27087

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.862650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.862650Z digest=sha256:72a159cd7201386385922c2e37d1f62028c64a204e0c2b79f09104afc1d34dc0

Observation ef5384c5-836c-40fe-8259-79616f60e36b · outbound

This paper cites M 3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation M 3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.933953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.933953Z digest=sha256:f34f8d0d5b226378824d9a72dccef567a01ba664008d9531708410949b6e1a61

Observation 84d50ffe-005a-44d1-9271-ab0a73a543e3 · outbound

This paper cites Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:27.990155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:27.990155Z digest=sha256:05cd95e4eb1dd647873820c9f6f46372be5cdae681eb79c916bb6e2c6ba82a6e

Observation 0941feac-e08e-45b6-a2d3-e47d321952a4 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.086458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.086458Z digest=sha256:eb9f49535c6d60c0ac8f075add97bb889d27602ed132ca5aa3a15b5ed328770d

Observation 0cd671ac-3da1-430b-83e5-ca62c916d4f5 · outbound

This paper cites Efficient selectivity and backup operators in monte-carlo tree search.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Efficient selectivity and backup operators in monte-carlo tree search

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.281455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.281455Z digest=sha256:315f21e843e1413ccdfca4658a346c9d41cb57038c005a4e62f5685ea498087d

Observation 2bb1915a-a348-4929-842a-7dd66ba6281e · outbound

This paper cites Mm-ifengine: Towards multimodal instruction following.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mm-ifengine: Towards multimodal instruction following

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.333590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.333590Z digest=sha256:cad5216713ee4106a6ab62fe14132456c32c0076a3e345aa9c439b63cc1dfb17

Observation 1ca2f837-1ad3-464c-977a-95820fb8ab4f · outbound

This paper cites MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.466278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.466278Z digest=sha256:392eb2251de34e63c3a3e7d3052efa4e0b5ba9a13f8959be902519f978ab5f57

Observation 335b71fd-9a91-4702-bd6d-6032358880bd · outbound

This paper cites Seed1.5-VL Technical Report.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Seed1.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.590628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.590628Z digest=sha256:68e244dbbf2e5f400ba6ccec7658464fd261b78cea3728066d8e63690710e75c

Observation 1d0c79e5-8689-4bae-998d-69272bda73bc · outbound

This paper cites GPT-4o System Card.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.712586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.712586Z digest=sha256:e69affedbcc29e2b9fd1e686b463bba7cbf074e19168925793f1dea25ead15e0

Observation 39b7f9c5-7f9a-4158-9cc7-c8e0108ca516 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Gonzalez, Hao Zhang, and Ion Stoica

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.870577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.870577Z digest=sha256:ba23bf09895809b942801c701908b6d872f0f9ce971a25f8212cee228e71b1f7

Observation 3aee8528-579a-4fed-8205-11dc1dea540c · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.025781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.025781Z digest=sha256:a5f55e24dcb65278d70c3170189f27564b2672aec151ab353d5efa51ee97ef2f

Observation 87760e01-565b-4123-b07d-57f1790e7454 · outbound

This paper cites From generation to judgment: Op- portunities and challenges of LLM-as-a-judge.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation From generation to judgment: Op- portunities and challenges of LLM-as-a-judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.234607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.234607Z digest=sha256:79d7c715846d63f63b63433663a951fe7ff81704310cac737f4565672d017224

Observation 71cc89be-fb92-444b-aeaf-a86bf525ddd8 · outbound

This paper cites Vl- rewardbench: A challenging benchmark for vision-language generative reward models.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Vl- rewardbench: A challenging benchmark for vision-language generative reward models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.459785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.459785Z digest=sha256:d67b6740ddba166b1d4385f7c57b91bc8b56e19c0d4d33a438eb3d6bd8e3f026

Observation 7ee98173-9fc8-4dbf-8b53-fa26734bad75 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.545585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.545585Z digest=sha256:5951ae5b9dc0d5c270370da9e2ec5ebed15f881eda801a3a918134b90556ee80

Observation 58616d15-a35e-4aaf-a966-0a5b2c48cabd · outbound

This paper cites Improved baselines with visual instruction tuning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Improved baselines with visual instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.603365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.603365Z digest=sha256:1541e245251748e3b4c3a78402b8ea6296d917813d82d16689f8f21a3b1f889f

Observation 6a18d149-1aea-4896-837c-e464ed7cdb4a · outbound

This paper cites MIA-DPO: Multi-image augmented di- rect preference optimization for large vision-language mod- els.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MIA-DPO: Multi-image augmented di- rect preference optimization for large vision-language mod- els

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.695853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.695853Z digest=sha256:85dbba62ba0d44b792c602f6a6ed8d1ad75e965c12abc766c3b9c0448f35148a

Observation 208fa9a9-3acd-4f28-8dec-c003cd5ceac1 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.797926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.797926Z digest=sha256:a4be03484ecdca0957a57ef47a4166352412402ac588d3d753d571300e83158b

Observation f61d1d53-b576-40ce-9acb-973822cea5b2 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.895022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.895022Z digest=sha256:40f5c66ae42ae951ac03a2dd009e68ffcd604fd45fe70fbb4ff304b0595ec7c1

Observation 2d4e6dec-6a8a-418e-8d02-bb3c703b7692 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.102690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.102690Z digest=sha256:e2af1c0b42658af0ca89d7ef09a1acacbfe948f3111f7e5dda7c9759a1c1e2f5

Observation c12b52a3-8990-4b27-9ced-1a516fc9929e · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.300126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.300126Z digest=sha256:0dab8fd612acd8a5d8f7944f25e8b33e39ca78bcb0ba957f215a679a9901f052

Observation 43515f9f-132e-49c3-b733-22885f3f06ae · outbound

This paper cites Qwen3 technical report, 2025.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Qwen3 technical report, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.487065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.487065Z digest=sha256:ba8c0f87fc27436d08de60ec2a442294d571aa6ea13a36737021ed59c1217604

Observation 9a820de4-a910-4a8d-a76c-6440165bd6fd · outbound

This paper cites MiMo-VL Technical Report.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.621170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.621170Z digest=sha256:0ad9b694b2fe69545dc260d798c5dbc44a508e3008e9f216dcd6d7e46f958389

Observation 0d27393f-7520-41ec-b445-bf42ce9cb706 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.762220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.762220Z digest=sha256:38ef87a691b2eb6063f70657f4ae40aeb1e16906747b9e84501d7cc1ab6f7bd5

Observation c956993e-8033-4bb3-b88b-75403199d96d · outbound

This paper cites Self-Taught Evaluators.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Self-Taught Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.964243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.964243Z digest=sha256:aa2db1e41cf98daa28b2c56c27ce8800a9806366920e8e1d7a53415020fe90a6

Observation a626effc-9642-4931-824b-5e833831e6c4 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.100419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.100419Z digest=sha256:a971b8c605020ddb3d46df8f11f1fe069e73566727efa3a79f4a588d48b110f4

Observation ce050607-bab2-4464-920e-7d6e9f9d0c97 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.186279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.186279Z digest=sha256:c8d101db80a22fcb2d1a087fb99166d6d742ab36a34d1694f3391454ec0c922f

Observation c2f28a64-5501-45db-9814-687086437bd8 · outbound

This paper cites SoTA with less: MCTS-guided sample selection for data-efficient visual reasoning self-improvement.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation SoTA with less: MCTS-guided sample selection for data-efficient visual reasoning self-improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.281051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.281051Z digest=sha256:e5f8d9720a3aa102debcb01b34fea342518bba1880b9ca3c890dd0eccd729fdb

Observation ad8e691b-00d2-4218-b3d7-c3be76064525 · outbound

This paper cites Unified multimodal chain-of-thought reward model through reinforcement fine- tuning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unified multimodal chain-of-thought reward model through reinforcement fine- tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.484261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.484261Z digest=sha256:c3d1d62b9a65566d7fa945fe9f8077117326462621caf61e096270bd11c71513

Observation 31c2c557-fc52-4cda-b980-d82f73e7fe6d · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unified Reward Model for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.618718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.618718Z digest=sha256:bb8224c9a73321fa72032b657ee2c2835b4e391d61b26eba54e3fa7fb4cfb84e

Observation b91f6f3c-613a-4879-8461-3f7d70e5a749 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.796208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.796208Z digest=sha256:e9d7a8fec82cb72158321cb7a77823fb40dfd13f2c961ff43bbbfbb570110826

Observation aa3fc127-7b93-4a47-adbb-472fafd0fd03 · outbound

This paper cites J1: Incentiviz- ing thinking in llm-as-a-judge via reinforcement learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation J1: Incentiviz- ing thinking in llm-as-a-judge via reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.858453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.858453Z digest=sha256:aeb3d1c1ac6e29b8e92d517fe1e892dfaa842a043d67eabfc16eb264eb53a89f

Observation 9bd999c1-d447-4212-867a-446d82d60593 · outbound

This paper cites Multimodal Preference Data Synthetic Alignment with Re- ward Model, 2024.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Multimodal Preference Data Synthetic Alignment with Re- ward Model, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:31.956457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:31.956457Z digest=sha256:0794cb4bcda13a96d500e445b71f52668a7788afd79ce4270dcdf9bd98eed0d9

Observation 0e781a24-5391-49b4-8b7b-e7aaef3e4f5f · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.055970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.055970Z digest=sha256:2c4ad3808e48babf94c56a97462a10d1ef9ffb19b438aae07da5f77eabf4f1d3

Observation 4663dc23-d0e0-44b4-9820-d6eee7906aa3 · outbound

This paper cites Monte carlo tree search boosts reasoning via iterative prefer- ence learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Monte carlo tree search boosts reasoning via iterative prefer- ence learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.196304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.196304Z digest=sha256:9e17da631a71342eca93df9a52adaac6e729ef8194ad2dfcfd9a4fad2ebc72a7

Observation b5119312-8e6d-43c8-b357-5e3047303c21 · outbound

This paper cites Llava- critic: Learning to evaluate multimodal models.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Llava- critic: Learning to evaluate multimodal models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.330762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.330762Z digest=sha256:a754afb421e05f65bcaf0483e4e6da8dbbb3b98df9ecd4eda677e13f0e53be0a

Observation 344db1ac-eb17-4a89-9d55-ceccdefb5236 · outbound

This paper cites Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, YuXin Song, Wenhao Wu, and Dacheng Tao.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, YuXin Song, Wenhao Wu, and Dacheng Tao

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.399862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.399862Z digest=sha256:6bf2f61c3f7fb3ea140e9234475064bf00a91137fe0a10abb1f106f0be986bdd

Observation fd9a0d36-c308-4e71-96c4-c1ac633e6e16 · outbound

This paper cites Mulberry: Em- powering MLLM with o1-like reasoning and reflection via collective monte carlo tree search.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mulberry: Em- powering MLLM with o1-like reasoning and reflection via collective monte carlo tree search

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.469544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.469544Z digest=sha256:4889d7d1b3f73e186c5504a5161f8c6bb3dae3e7d9e88d891b1f1e430d401e54

Observation 9f2c75f4-4926-474b-a789-d457d18ac3e6 · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.626573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.626573Z digest=sha256:ae86a4a49fa0c34412615d8f58e345d84d185fcc595b6742cff7ffab7b3d2b46

Observation 3cf4578e-7a22-4c43-abf0-b9ca285ac19e · outbound

This paper cites A survey on multimodal large language models.National Science Review, 11(12): nwae403, 2024.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation A survey on multimodal large language models.National Science Review, 11(12): nwae403, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.755599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.755599Z digest=sha256:5a8789b88fce40e98ef52742183e2ff7be763ad40e620ac69480e175ff593be1

Observation d11f5705-9a13-488c-a8c2-2f9dd3c1a2e3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.882052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.882052Z digest=sha256:ea0d2dffee9006b775e175f7a04c12cbb45ec8054fe95ad6b7e7cc7f477f2f42

Observation 9c4fd817-26f6-4cc6-b80e-2a2cb135ec94 · outbound

This paper cites Rlaif-v: Open-source ai feed- back leads to super gpt-4v trustworthiness.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Rlaif-v: Open-source ai feed- back leads to super gpt-4v trustworthiness

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.951359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.951359Z digest=sha256:a07550c31e5bdd757accb3a25f8197b31870c6f7d5bd6056576c9c751a7eeeb5

Observation 155dead4-f115-4b73-9b19-9b1e359ebcef · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.005768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.005768Z digest=sha256:f6c020570129281403602506eaec998805c11a54bad35c72ad3d6f5d8ca52be8

Observation 4a90b085-ec11-4f09-8888-9cca6fa2976b · outbound

This paper cites MMMU-pro: A more robust multi-discipline multi- modal understanding benchmark.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation MMMU-pro: A more robust multi-discipline multi- modal understanding benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.062176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.062176Z digest=sha256:ad1d40afebe867e2fb50173b85a72578ad755ef81d96002ac10f246f6f73ada0

Observation 03d601ef-cc7b-48f2-bd53-025463245bb5 · outbound

This paper cites InternLM-XComposer2.5-reward: A simple yet effective multi-modal reward model.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation InternLM-XComposer2.5-reward: A simple yet effective multi-modal reward model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.095853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.095853Z digest=sha256:c4638c19f9178f9046902e82f0b7935855f682a72927487561c379d59799f8b1

Observation b478a422-45dc-47ad-adaf-8b3ec038477c · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InComputer Vision – ECCV 2024, pages 169–186, Cham, 2025.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InComputer Vision – ECCV 2024, pages 169–186, Cham, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.160515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.160515Z digest=sha256:e050010ab160c4b0afa6da3334a743f6addc9791ab17e968a98b9aca79a17fc4

Observation e328df1f-2149-4c82-8992-e059970c3bd9 · outbound

This paper cites GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.229753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.229753Z digest=sha256:ceaa4786373e3191c356d147514dea148c777bf25daa50e430a6af090de665c5

Observation 0e55930e-bd20-4101-a527-91a1f656848d · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.297932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.297932Z digest=sha256:a142ced7177cf1a3653592d613481515efbe941f5bd130f83e374a2b0dedf770

Observation e46844e7-e775-4c67-9b1d-6f524c4336b3 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.389349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.389349Z digest=sha256:6eda3c63ed085dd6d43039df91b4cbeecdeec7cf033bb2f60d03e4c6c88691a5

Observation 0634c388-8415-4f9c-bcc1-e76a24a1cd03 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework, 2025.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Easyr1: An efficient, scalable, multi-modality rl training framework, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.492213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.492213Z digest=sha256:68cae198dad0bdf7d6fa9fb054371518982c75c6359c3cab3ec50effc425e9f1

Observation 30b42ec6-0547-43d4-b6b9-6efe512f3a97 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.552171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.552171Z digest=sha256:448ffe24c30d5299b0b45ff6fb0a42e27b7469a59f4f8a64ec89f2012b008d95

Observation e10bb0df-8bd6-4365-93ea-d163e706ac3b · outbound

This paper cites Capability-Oriented Evaluation Framework.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Capability-Oriented Evaluation Framework

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.610346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.610346Z digest=sha256:36841ac629ce7239599ddbe88185958f8128be2d0f318ebd5f85fe61a16cfce0

Observation b0bfeda5-4260-452d-9455-42e018e144f3 · outbound

This paper cites Open-Source Training Data Collection.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Open-Source Training Data Collection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.646776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.646776Z digest=sha256:449666c7d54d58b2640c4ba216b5a4fbddbba71d4f0ba6afd6bb119a660b0683

Observation 2569f6e2-ddd5-42d2-84e7-128362fddc86 · outbound

This paper cites Experimental Setup.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Experimental Setup

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.698935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.698935Z digest=sha256:71fec3fb0d208abfd0300a78d845c3dd8d2d2ce72fa034437be7baedfac46de5

Observation 4a291b2e-12d8-4054-81af-e02c462b0813 · outbound

This paper cites Benchmark Examples 3 A.1.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Benchmark Examples 3 A.1

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.759584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.759584Z digest=sha256:58393c526c5931c95df06f9c9e622eddd749e9f97f085295621004b94a1b41cd

Observation 07341223-79eb-4b64-9143-15eec125faf3 · outbound

This paper cites **Gross operating surplus**: Net operating surplus + Con- sumption of fixed capital = 240,000 + 110,000 = 350,000 Rm; 3.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation **Gross operating surplus**: Net operating surplus + Con- sumption of fixed capital = 240,000 + 110,000 = 350,000 Rm; 3

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.832523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.832523Z digest=sha256:4af24dfce6a9d4fd399047425abe45ad7a42d87d92edee22da3040443c80ee29

Observation 6284a8dd-0a6d-492e-a905-422ba40fa4c3 · outbound

This paper cites **Net operating surplus**: 240,000 Rm; 3.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation **Net operating surplus**: 240,000 Rm; 3

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.940890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.940890Z digest=sha256:021d8e4c5f212909117a1fe1a8453566ac9094f4a5acc9835ae8ef6dbbd7e2b9

Observation 362b74d5-ae8d-46c4-b058-ff5a35b0ea33 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.012350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.012350Z digest=sha256:7b5eda6d21f686b7817d84c6f9c456e1e03143566430ace65dd0e42b76042137

Observation 8396c4eb-63c1-4346-b453-3922c691afe6 · outbound

This paper cites **Sum of numbers in triangles: ** 4 + 11 + 18 = 33; **Sum of numbers in circles: ** 5 + 12 + 16 = 33.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation **Sum of numbers in triangles: ** 4 + 11 + 18 = 33; **Sum of numbers in circles: ** 5 + 12 + 16 = 33

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.045268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.045268Z digest=sha256:9d202180d9af287e56353b0c7d7f8722fbd0ec33738429c80ce2b48e1a24bf16

Observation 6367bd4e-400b-41d6-98a2-0d8c14f1fc27 · outbound

This paper cites We can apply this rule to the squares.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation We can apply this rule to the squares

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.111025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.111025Z digest=sha256:f3ef25ec363e1ed4a70d873d83ec46f94b1e9d35efe07872e3f65418cf07a389

Observation 1613519c-70a6-40cb-a12f-3594159d77c2 · outbound

This paper cites Key Observations.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Key Observations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.142051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.142051Z digest=sha256:ad2e48157e9e1d577483e03c051a2b8dd545dd776d15c2b35cbcb566a1cd6402

Observation 473dd069-4e65-4f6c-ba93-10ee49400532 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.206791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.206791Z digest=sha256:e1ed0b55e3eb1bca41ea16c3c74b55348a36d595c7420d5ded4134e70359ce49

Observation 0bb5a77d-dc22-40ed-9233-9c4746f41ebf · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.244048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.244048Z digest=sha256:77922d7972fe2f2a75bb6fdc8a499534d1922aa716f2e4cc0f23716a4085882c

Observation a9e52dd4-6919-48a0-b722-62c693ff2d28 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.310999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.310999Z digest=sha256:ded30064da15aeb893048cf9f50f6e6d0ee22023a2cf4092179f39e2baae7de2

Observation 281b89df-2018-40eb-a9d5-47ae0071c3b3 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.371808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.371808Z digest=sha256:2cd7024a5d061d753bc47218be9e9cd46855771fde0028059688a11647234ca7

Observation 9f1f8935-3028-4716-b6cd-a0407b82339d · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.464640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.464640Z digest=sha256:9c4ae7b51dba1c532f2c785bb51c2fd5b2a48079e757a16acfb001dfd9fe691e

Observation e5b0b1b4-e301-432c-b732-fe556ad5b261 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.535306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.535306Z digest=sha256:3d3e40cd260d3be25aecf18d5c7982a7858761057110bcaa064d9c376b74a37b

Observation 5745fa83-121b-4620-93c9-36b8e4490c53 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.568170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.568170Z digest=sha256:dc56c1d677e56f7d2caeac47bce306a6e10426f7fd31f3d6c75339712944d840

Observation 5fd8554f-5120-4673-a33d-06e2b89c8ab2 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.635972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.635972Z digest=sha256:ccde82e438302af58a0fa6ab2e5b3553ce7e92f83959d90c98eb85320ebe1303

Observation 1465f68a-7383-4e3c-819f-26928d009ee8 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.745434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.745434Z digest=sha256:96a7d918aa6a9d15a8845347c1643ad50f9a0ec0abbca21f762761c4fdb4d4c0

Observation 59d4d96f-b5a0-4ba1-bffc-3d8534be58a2 · outbound

This paper cites Modify the given correct solution text by introducing subtle, hard-to-detect errors.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Modify the given correct solution text by introducing subtle, hard-to-detect errors

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.814626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.814626Z digest=sha256:a4b1c8230c8be1affcedd326fcb5cb17c8539b4253961c2794206a76a0dc2b3e

Observation 9d0f57b1-70f9-4258-8890-eae572b8befa · outbound

This paper cites Prioritize factual correctness above all other factors.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Prioritize factual correctness above all other factors

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:34.951668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:34.951668Z digest=sha256:696b772aaf44c099793e3d0a1238fb87847d5aefac5bf87f577617e49b6aa4e2

Observation 12fabe4f-29c8-4cc2-b9ec-245d1b9af5bd · outbound

This paper cites Identify any reasoning fallacies, invalid inferences, or irrelevant logic chains that might affect reliability.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Identify any reasoning fallacies, invalid inferences, or irrelevant logic chains that might affect reliability

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.022760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.022760Z digest=sha256:86afeb43292047c3f6ef702acf7439aae856270c9c726e0cca013d3a9904441e

Observation 2432c871-2dd6-4796-ac40-3f56498da1ea · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.076387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.076387Z digest=sha256:c71e995dcc865296f52fac608d541d1eebfb350f2b9801e6e08651dd2ab649ef

Observation 06447396-6fab-46ae-8381-46821f8fe1bd · outbound

This paper cites The response should neither omit essential reasoning nor over-elaborate with redundant or misleading content.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation The response should neither omit essential reasoning nor over-elaborate with redundant or misleading content

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.154932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.154932Z digest=sha256:6f7bfc92a46f19ab8545d1175fe10b04fcac140b187a97ca8b4ea52c6bccc77b

Observation 386b1fba-a7de-4a2a-a35c-1eaf6c815f51 · outbound

This paper cites The decision should rely on reasoning quality, not stylistic fluency or writing pref- erence.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation The decision should rely on reasoning quality, not stylistic fluency or writing pref- erence

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:35.246936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:35.246936Z digest=sha256:86a0f8d263fa8aebcf82a88ad46532c16cebd490a03f5e0ef29d19a3cdcda499

Observation c91a1c5d-cc07-4537-81d6-5419be646f88 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.346396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.346396Z digest=sha256:02ecdd367f1237916dcb1ddb9f200e5c0a01411c1af52a16ee0e9e68d0f7f921

Pith citing papers

Observation ffe76620-05ea-44fc-a9dd-555a06543955 · inbound

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation cites this paper.

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:33.558625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T11:53:35.315405Z digest=sha256:f06f7f982569aa59cab237f190a7ab13d2c2f6abbb3df533292b760c58362ecb