Pith. sign in

Paper Citation Record · LEDGER

M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2405.16473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.16473 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:56:01.691015Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:31:24.804270Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa2db9a7-bd56-4ed2-8723-7abfc7ee438b · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.430760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:cf48ff5b34c85a961d9b60119a664989c995e006473fee9766f4995f4076332b

Observation c8baab00-adc6-4ee9-9422-05ba42110bff · inbound

Interleaved-Modal Chain-of-Thought cites this paper.

Interleaved-Modal Chain-of-Thought M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:38.175313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:38.175313Z digest=sha256:0061bb6496f56ad2e1a346f94bc5fe3c3ba2c6b59e4a99d07ad2778750f298fa

Observation df7ad01c-87f9-45a1-9afe-1dd4a27d19dc · inbound

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? cites this paper.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.567451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.567451Z digest=sha256:ecbe6093facfe014a0f8240cd2f3e1a82f8736eaa7d852a356c8f0ff752dee06

Observation 90671c1b-3e88-46b1-8aa8-74d8d62ec4d3 · inbound

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models cites this paper.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.986574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.986574Z digest=sha256:7ef5586db91bed8ff6df7f050b7838b5193655b581236331ee2aff3ff9f7df40

Observation c2262a3d-47e2-4439-bc02-52c12ec9e78b · inbound

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding cites this paper.

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:32:58.231052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:32:58.231052Z digest=sha256:05a6e28673cdd7c07ec3592e8860c7899594ad84696d51c341d46f926b2315ef

Observation 3b253375-e068-4098-b7f3-71a665c88a01 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.233339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:dfad2dbc0f8d554f378d6ee9d0edb7852c3e37c509a95323a8a915ea513fc2b1

Observation a768fc52-f4a4-45e4-8ead-41ade8a9bbb9 · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.955656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:24.955656Z digest=sha256:9e5db165c0edfa91d1bcda84ba7ff783e663c6484c43bf81d755bc67f0c6aa27

Observation 63d556ef-4e20-4813-b7d4-3824b8e545b1 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.740516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.740516Z digest=sha256:0be663a35173f39698935e9019d58d422134d35a17889c9739a48a0706fda222

Observation 5ef36ba9-0092-47a9-846d-f92be3f14d16 · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:32.456049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:32.456049Z digest=sha256:2ece2791dfaed40399b9d4349b047aee3cce338395ef702819c2a7ff7de71f47

Observation d6e6e4a9-3d92-4866-95c0-a526d30c03aa · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.142788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.142788Z digest=sha256:9433b703e923097ed68ca5302eba09fde95a82ffb2938c01f6e57c05d37ed74c

Observation c9ff8af9-fef5-462a-8545-4a856cbd90e8 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.014182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.014182Z digest=sha256:62d6ebc8f1162407ae7473c7d6014b55ee0bcc7b1977f5aa1c790733fb4d95c8

Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.193447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.193447Z digest=sha256:434b2839026f333611389adbcdcc35565d399d13386d503ff1b57d0795f36708

Observation 5cfa213e-4e85-4600-8e5a-d68347522cd7 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.344139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.344139Z digest=sha256:156959a0580bda81ebf420cedad6a68c06df167ec75a86bc7dc7ce4e871cf7ba

Observation 23ace000-230b-4d70-a2f6-a22c80d74a04 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:18.196208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:18.196208Z digest=sha256:1bbfa890ad3ddee7e3c57ebd0f83b49836c644a7639f262de4397e695d9af02b

Observation de2f71a6-6d3e-42c6-bc3f-951120f45647 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.820419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.820419Z digest=sha256:cae0aab1863542dc3f00e593971be6485d08b1f93fe41b811441fc1a76441668

Observation c5d95906-1dc0-47af-9f83-a6d76b8c75b1 · inbound

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning cites this paper.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.202517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.202517Z digest=sha256:7ba6f027dd572d71fc98ca55d901f4d9a9448423e3447fc441d9119ba1512637

Observation 499edf28-5844-4674-ab3d-bcb80ccd7474 · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.872966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:282005b42b40f3022215cd2b14f54c982e75ca78b41959e768b7112fab0d1022

Observation 0fb856ae-46bb-4b18-8f01-dd3c23c85fdf · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:06.028911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:06.028911Z digest=sha256:ac7165cdb0dfcdca66277a744999341db9e0889ae935a341a154480a1fb5bf13

Observation 93e0c80a-7cef-40ce-a675-08f108ee77f4 · inbound

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration cites this paper.

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T18:17:16.302417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:17:16.302417Z digest=sha256:bbd491c3f1d6dd31ec86e85d3d89f4097cdd174285c1255203217f342405cce1

Observation c59e1cb9-f502-49ef-a836-ce4960afb37d · inbound

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning cites this paper.

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:31:24.806625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T13:30:10.620448Z digest=sha256:57129ee10bbeef2cbca577b8fe053730550453a1fa967a9be8e7af42cc8fb31d

Observation c33ecb02-7b77-4bc5-b8eb-9a0400c2620f · inbound

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning cites this paper.

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T13:42:59.320438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:42:59.320438Z digest=sha256:7ad998f611bdb082683a9e224bb7abc7b9950f494ac91d9b4bff5a206a88affc

Observation bcdf53d7-3650-4256-8c03-e7f7d95c5f10 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:01.437413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:01.437413Z digest=sha256:ea4d32328adc74ebfe7b8ad30692273aead1f958cccf2492545d70f26f800c23

Observation 3098d1d8-c620-4482-bc2d-d54d5d21c7de · inbound

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models cites this paper.

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:50.458940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:06:45.417233Z digest=sha256:6721665b0d14040a46747a56c2972095d8374a3a2a7acf7089341090ade2ef49

Observation ce225f1b-afe4-473d-8aa3-b0b5418455d2 · inbound

Targeted Exploration via Unified Entropy Control for Reinforcement Learning cites this paper.

Targeted Exploration via Unified Entropy Control for Reinforcement Learning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:49:56.104547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T10:48:27.733823Z digest=sha256:c11efc93d4240bc2d03475787afe13cd72e4f8d92ad4f3ef9e52a5c9f03df493

Observation d1e68568-458f-4b72-a105-3d1bd065df69 · inbound

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning cites this paper.

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:26.437698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T08:28:07.531338Z digest=sha256:aab7a399e7a87a52b27cc54c57fb7c1e3b6fe524c150692ecec6db96ce26a4b3

Observation 73f1b7c1-46c2-4eed-bd86-6c79507e911a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 255

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:23d1a9816cffac0b0a513b3ac8ba7bcb4d6a13e017949810aac2fab0d62a943a

Observation ddfb857f-9aa0-4d81-9deb-23b1d1ad7ce0 · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:35.633314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:35.633314Z digest=sha256:b3e20e53900fcb75ef267aa7d0ffcf5e39a7b20ecb273841ea4ba09173f1f32f

Observation 8e4cad80-e85e-4f78-b76a-ecfb807eb8e1 · inbound

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus cites this paper.

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:56:01.691015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:56:01.691015Z digest=sha256:3ab421c018e9920746785b49eaf9edd9b5d30fb8a0e0e7724e7c83efafb8b428