Pith. sign in

Paper Citation Record · LEDGER

M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2405.16473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.16473 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:24.955656Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:31:24.804270Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa2db9a7-bd56-4ed2-8723-7abfc7ee438b · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.430760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:6c9dbbcc08c78437d8a1862c288cf106f2e2c7f3a191ddada4c499efa61758d7

Observation 3b253375-e068-4098-b7f3-71a665c88a01 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.233339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:0cd552729ea714dc7344310d693faac533d6467b00ea7e812f6d8d489399e448

Observation a768fc52-f4a4-45e4-8ead-41ade8a9bbb9 · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.955656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:24.955656Z digest=sha256:349381642d946ad2c557e1cc7892ae70d0c688f18f1ecdf2affdba7b4eaacc5a

Observation 63d556ef-4e20-4813-b7d4-3824b8e545b1 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.740516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.740516Z digest=sha256:6d50861a5a19ef72939e0c3b70d79590baf5425e7555c11eacc8bd530b8749cd

Observation 5ef36ba9-0092-47a9-846d-f92be3f14d16 · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:32.456049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:32.456049Z digest=sha256:33e3862bae3af954fbe8650832fdc0c20557356278a9a425b4a6910fb8205443

Observation d6e6e4a9-3d92-4866-95c0-a526d30c03aa · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.142788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.142788Z digest=sha256:9f34c219d6552b0dd23d84a163a7bf8569405960791727d90b85f6b8dae9a681

Observation c9ff8af9-fef5-462a-8545-4a856cbd90e8 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.014182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.014182Z digest=sha256:f4a81806008011222db9998dd44489a18a3393278637fb08b86cd81141e5f362

Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.193447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.193447Z digest=sha256:d5073cd014118a78dbc808b83cf2538856af097508301695ec151b05016719ff

Observation 5cfa213e-4e85-4600-8e5a-d68347522cd7 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.344139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.344139Z digest=sha256:69778680e8fa769d55ebdb5d45bf299aab1ebb70df3b311c33e15c696a364317

Observation 23ace000-230b-4d70-a2f6-a22c80d74a04 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:18.196208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:18.196208Z digest=sha256:8acbfc40951660e1fee2bbce86578b56647734788ab56710983e98688a297e68

Observation de2f71a6-6d3e-42c6-bc3f-951120f45647 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.820419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.820419Z digest=sha256:0ae2944abfd3a22cec62bfa78fa32c7e2ccfb16419292d9af8d12448db16e3fb

Observation c5d95906-1dc0-47af-9f83-a6d76b8c75b1 · inbound

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning cites this paper.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.202517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.202517Z digest=sha256:26e47c3a29fff510bd9ee6d86f04b9dc6a1d6635b27c2bdb93107a75e8cdaea4

Observation 499edf28-5844-4674-ab3d-bcb80ccd7474 · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.872966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:3655e610b4331d23ef5083898086f835822ef8ad87094b481eb5eae250b8311f

Observation 0fb856ae-46bb-4b18-8f01-dd3c23c85fdf · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:06.028911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:06.028911Z digest=sha256:e2c947dc7f18a7dfc0d48d72662d782df04f46ccd0f52d9a8d016f94176664ef

Observation 93e0c80a-7cef-40ce-a675-08f108ee77f4 · inbound

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration cites this paper.

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T18:17:16.302417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:17:16.302417Z digest=sha256:d62cb4044dddd1aa657f0896a1dbac27105e3b3940a8de9cd0360cfdb05f0f73

Observation c59e1cb9-f502-49ef-a836-ce4960afb37d · inbound

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning cites this paper.

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:31:24.806625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:30:10.620448Z digest=sha256:1d169e2970fb56cc377fc0765e358d38522592cade59947724eb90872b17500c

Observation c33ecb02-7b77-4bc5-b8eb-9a0400c2620f · inbound

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning cites this paper.

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T13:42:59.320438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:42:59.320438Z digest=sha256:af0b62a774d3defd4c288a11b1a6397101154d016593f42de4737bc37f5126e4

Observation bcdf53d7-3650-4256-8c03-e7f7d95c5f10 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:01.437413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:01.437413Z digest=sha256:c0c5c5da7a9a41d6c6a0aa1acd17be5e4737f97f0dbd7fbccb0b20e242fc4f63

Observation 3098d1d8-c620-4482-bc2d-d54d5d21c7de · inbound

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models cites this paper.

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:50.458940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:06:45.417233Z digest=sha256:3d6c27d00870bbf8f1031cd5aa1343a51331951a7177c40eba073fb791c4d50c

Observation ce225f1b-afe4-473d-8aa3-b0b5418455d2 · inbound

Targeted Exploration via Unified Entropy Control for Reinforcement Learning cites this paper.

Targeted Exploration via Unified Entropy Control for Reinforcement Learning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:49:56.104547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:48:27.733823Z digest=sha256:88b61d476d554e3b3c162f4a3b6ceb9e06d1fbe9e70dba4b2242090958604294

Observation d1e68568-458f-4b72-a105-3d1bd065df69 · inbound

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning cites this paper.

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:26.437698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:28:07.531338Z digest=sha256:9898a7a6ab4116fdcfe6cfb56c6c0528a94a468f25e18d458af01498f03222ad

Observation 73f1b7c1-46c2-4eed-bd86-6c79507e911a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 255

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:4b151900f286ead85f2c0c52306548eac37847663c161d0337551bc47ea36844

Observation ddfb857f-9aa0-4d81-9deb-23b1d1ad7ce0 · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:35.633314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:35.633314Z digest=sha256:16e6fe14dc789e0367a12936d1b2c6fa2d5c42fb53bfe03d1bdc23000e1f4b6c