Pith. sign in

Paper Citation Record · LEDGER

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2412.12932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12932 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:38:14.155401Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:59.500206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:15:46.368932Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8034891e-c37d-4063-98cf-2afe3b79d316 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.965366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.965366Z digest=sha256:d02d2a6555cd0a676ed566540742adbe4ad42271fa93d16143e8c4733cc81217

Observation fbd2b36c-273f-44ff-aa61-02a06070f876 · outbound

This paper cites write newline.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.969608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.969608Z digest=sha256:44f93d182cd34f9e67a6299689120f7aa4b8d8791b687ab866140f7c24016335

Observation 57a2ac40-063c-4295-8be9-01a130c220f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.973903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.973903Z digest=sha256:4b8761ab4757388a622560f7a21fae166762e2ad3ce57b63fa391f74710abccf

Observation dcd8b321-e3f2-449b-8783-839897f77078 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.669899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:13.978345Z digest=sha256:13a30cd003478b9d714b90b3d399478ec03e076becce84e46c5d2f4db8abb2b1

Observation 8abbe37d-f615-4282-abd2-c305ace76b5f · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.658652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:13.981961Z digest=sha256:e194ba7197f2bfd2591b5894a4c35e3f2c74cc952ca82f7cf94ec4616376f3b1

Observation 90671c1b-3e88-46b1-8aa8-74d8d62ec4d3 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.986574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.986574Z digest=sha256:84dd00cacf7a48ca01a2a81c5df8a3dbd79923386f6735dfedd1bbb3a5c23931

Observation dfd098ce-faa8-454d-bc96-4272b74bfba3 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.990476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.990476Z digest=sha256:6627c50d8ba78b9db7e21cf27bbb0fd9c9b556c3532e6bb9c700159761d66960

Observation 383e9cdc-f93f-4999-8613-e7e20e7c666c · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.994300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.994300Z digest=sha256:c83d22f654f5ece113a8c29370283a1e46dc8407b36e149c250424662d1c23c6

Observation 68a5ee27-1ea6-4fdb-87c3-c7093dfd2bf8 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.641535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:13.998232Z digest=sha256:0a689dd20bd9c663441ff492ce8390117c578dd02d52a9808df14a5836617456

Observation 69a46e92-6351-43c6-bd9d-ef0b65308156 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.630603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.001637Z digest=sha256:38ade46cf27910927c8f226f69c6a3c32a86e4f6f9239e726dd79b5743938152

Observation 273f5760-434e-4911-b0ce-05243ef18e73 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.620403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.005393Z digest=sha256:1f712e65f42de185e7af79e7e39d6d81549b810bc85834063b26ccebbc09b818

Observation 40d85930-82fa-496e-a1e3-df2c1e391c52 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.610455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.009044Z digest=sha256:f90d90958e786fb7ff87d1c6b0fb35db28b9da763d820fcc13d049610c9a0779

Observation c21bbc0c-d03a-441e-8f59-77fc46303491 · outbound

This paper cites P.; Poff, S.; Corredor, M.; Zettlemoyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models P.; Poff, S.; Corredor, M.; Zettlemoyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.599885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.012557Z digest=sha256:83ccb4976e7741b7f045c15cf51a7b40061009bfc4a116c2e5f44bdb36afa924

Observation 4f74b3fe-4c7a-4170-b8ba-e743f1333127 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.587672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.015834Z digest=sha256:df35b8fc040a95ed24ad5599646960f1a009d6349f812fb350ea299957f99e35

Observation ed6278f1-1757-48ac-ada2-5f1a4a9b2a37 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.019407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.019407Z digest=sha256:753a86d94effd1a6c9650dc5768a2fb01e344ca83cb07827f905a1c21f604323

Observation 19065e72-0b5d-4304-be0a-d6fc575e3cf3 · outbound

This paper cites Abstract Visual Reasoning with Tangram Shapes.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Abstract Visual Reasoning with Tangram Shapes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.022709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.022709Z digest=sha256:c8e79b46a3838cc77ce0297ca9de61ea210b042262672c34e173062f5d292146

Observation 1be8d884-cf09-49e4-a3bc-98a9b2254814 · outbound

This paper cites Y.; Fried, D.; and Salakhutdinov, R.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Y.; Fried, D.; and Salakhutdinov, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.571413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.026376Z digest=sha256:7137eb93c5b0c000952b973c58242b07c2c15a333f4fbfc62799f684bdb5e264

Observation c617cfbf-052b-43b6-a360-efd0e7859984 · outbound

This paper cites S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.559972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.029751Z digest=sha256:8ec5198535fa6b9596f2d8da2c8ac6ba4fdfd9d67cf86a1826ccd36a0ad24b41

Observation babae6f3-b5e2-4c33-8fee-c83d9131dda4 · outbound

This paper cites R.; and Koch, G.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models R.; and Koch, G

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.548776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.032918Z digest=sha256:929ee6e35f3f0bbce5dc54197bc9fbea2ceb2af2328b1872923b2b281846e1f2

Observation 2019cda5-9f63-43af-b405-3caa3415cba3 · outbound

This paper cites What matters when building vision-language models?.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models What matters when building vision-language models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.036134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.036134Z digest=sha256:5379b8f9b4d50f74322d87db05dbb11dc67084d2f38c4919da1c6d9d01b3aabd

Observation d9be3fae-36ec-4b96-906a-d84ef5d741da · outbound

This paper cites Multimodal Reasoning with Multimodal Knowledge Graph.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Multimodal Reasoning with Multimodal Knowledge Graph

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.039890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.039890Z digest=sha256:82ae29e1da8d48ab4a5aeb962c88834892d1a33980a683df83d8ce8a8298b8b4

Observation 6564c27f-cbaa-4ede-97b2-007aef125d7c · outbound

This paper cites D.; Strik, W.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models D.; Strik, W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.537450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.043346Z digest=sha256:e7a392f7f1158f3e8451164736d05d314e42aaf284448384d5894940430753a2

Observation 4544a36e-345b-454c-810b-ad01f34d769f · outbound

This paper cites Unified Demonstration Retriever for In-Context Learning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unified Demonstration Retriever for In-Context Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.046278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.046278Z digest=sha256:68aee6c272a82992589cf9682e527cfe58dff98e45520ae59293774bba3f2f8c

Observation 810e78d5-86a4-4ec6-8860-1ba54137ca5b · outbound

This paper cites A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.049627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.049627Z digest=sha256:5591f141be148140a67ce12b0ab916ddb6565da76bc592549d3d594f216e811a

Observation 4f169db1-9c86-4ed5-9749-71ed3d63aa53 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.053159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.053159Z digest=sha256:971514194322c2830e6b7aba09b0d09bf8911fa90ae2ee9af280c84a73c1d9ed

Observation a6208ccc-ec4d-430c-a07c-063dfc4d6b4c · outbound

This paper cites Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.056640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.056640Z digest=sha256:b368ccf8d72ff9852abe51d48c7fc73b2e69cb4b63e6fcf917b620cf58ab8283

Observation ec511551-9895-48b9-bd86-8286a527f007 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.525994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.060266Z digest=sha256:d9f97458d8d6d0e03d199b35cd71345e8d8159607fda17e9473de5ab0e36a87e

Observation 8a417068-b26c-4d4f-925d-5dbbcd1b084b · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.515788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.063317Z digest=sha256:e014f953e3c67973065730796d3b8f4589eac4624b4de32ea6705d901d97242c

Observation 2339f6bd-5f08-4f8e-85f2-f7fdbddac1a9 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.066249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.066249Z digest=sha256:27fae9e39a1b8ec21c62f743b2595a58d0059e17c2015a6585961f71cfd4ad93

Observation b2655f6e-31a8-403b-9d7c-e5ca4003f63c · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.504841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.069654Z digest=sha256:022630c28332a8c9f52f0ae07e736f3dd94d0479b9b2211f5b63731adb977d6d

Observation 85dd9dc4-7908-409d-9116-467825017e59 · outbound

This paper cites Chain of Images for Intuitively Reasoning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Chain of Images for Intuitively Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.072618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.072618Z digest=sha256:2cfa46e7a4f37061a0fba1c3f77d7f28ea8b8c12225ec333f0319754346f639a

Observation c7ca2f4b-34ea-4c76-aab7-aa5be1ba3c1c · outbound

This paper cites KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.076036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.076036Z digest=sha256:60ef8c0f37b3b0856d076b2281f79a1c2fcfba2366d044803885878e57f991ce

Observation 8807733c-7086-4704-831a-69d459326c88 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.079635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.079635Z digest=sha256:062028d51f58354ce894b858681198b8173aac00031b4fe5c8af7685e934132c

Observation e9324f2e-2037-44bc-8182-f6520df2b1f9 · outbound

This paper cites Large Language Models Meet NLP: A Survey.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Large Language Models Meet NLP: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.083304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.083304Z digest=sha256:6c87537313a10488b53ef4e813961dde690ea1711110ba77ee4564353b348ee5

Observation f5421ae8-7934-4145-b23e-4498daa03afa · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.494644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.087614Z digest=sha256:ee920f08cf8c182c7045b468f6de608c4215232313baf1a33dc8515d200619c6

Observation 713119a3-79da-4952-b398-a914aba96b10 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.090582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.090582Z digest=sha256:4c5d35ca6168fced7431cb3f4a9db717744ebbdd5188de34f07de72602233998

Observation 6d6b0eeb-931f-4be8-8f6f-59d562627b0d · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.477738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.093782Z digest=sha256:2c25b92f403b46018b8c0f80f4396e4627a55ca68f1509ba8d0215505091bf86

Observation 78611544-0b79-41f8-9451-2bac1a3eb059 · outbound

This paper cites A.; Yasarla, R.; and Patel, V.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models A.; Yasarla, R.; and Patel, V

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.466729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.096997Z digest=sha256:622beab472bf862fa20e370cc112a87851a46e479707add96987c2b9c260c285

Observation d45efeeb-9d63-49fd-8ff4-4b6bf220565a · outbound

This paper cites Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.100180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.100180Z digest=sha256:41ed68839a8705a8a50f4a40b430f812a86ff5f9942ba6a207ce79871699b59e

Observation 7911f6a6-a7a4-4fd5-97e7-50dfb40033fd · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.103515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.103515Z digest=sha256:2b4e0527ab660783f6e4d312ec2b9870f0d166effe06511fa21086f4675ee1bb

Observation ef7790aa-6b17-464f-84fc-46772035abfb · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.455530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.106809Z digest=sha256:c800c7402f85a97bfa405c62c4d56e0abac7833f4d1706920bd7d04dbeca92b5

Observation 164d0c2f-6277-4f85-9c05-44ffea2f4a6f · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.110031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.110031Z digest=sha256:48f77946b0eaf81b100a2f216529fb90e331c9de65e517b1435907e7632e3217

Observation 4cc2f2db-edce-425b-9e87-24c260e23e27 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.444870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.113646Z digest=sha256:0fe47a5d2cdbf7b1709add6cdf89d184223f2a047f60617d54766764952e10ad

Observation 3211c793-d67a-4f78-a89c-563cd12de61e · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.434445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.116780Z digest=sha256:4e67f98ae685853d29eebfcb45fe1a9aa365c4dd0bb9119c480a2773c82d5f05

Observation 142c6a96-f632-490f-b4e5-d58b7d8eaf93 · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.120001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.120001Z digest=sha256:a746bc682b96c3328afdee9e631a660cc46fcff7d949f2f74b1a93fb101ba56b

Observation 2a6769f1-860c-4fbf-a0f5-be13c7bc6c1d · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.123422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.123422Z digest=sha256:896c2bb046fd58dcf75b5e016925847bae6e4c553099fb9a01ebfdbe77585f7d

Observation 47751f60-6b9c-463a-ae0d-459d1c858fb6 · outbound

This paper cites Faithful Logical Reasoning via Symbolic Chain-of-Thought.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Faithful Logical Reasoning via Symbolic Chain-of-Thought

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.126778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.126778Z digest=sha256:3fff81f3d72173058cfb190befebcb570386df90771e12935cb8b528c3448c10

Observation 1ff5d7f2-ed6e-45dc-bfe8-47372802bdc2 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.130433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.130433Z digest=sha256:cd66a60b6d9c2a0797d575f330abf6470e254e9352ef16444bac57ea7d27e355

Observation 6cb29fa8-a492-4297-887d-5ef4b222f7be · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.424158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.134018Z digest=sha256:4d68fb9b3d5e1a7d039e2fa53648f387b9444e950ce23e1334f694b768543c99

Observation 85b6cef6-180c-469e-a7e4-e54085e4c354 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.137261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.137261Z digest=sha256:c69cff1ecd16c48799810fd920ada9584a5d390a9019fdba4fe827212abbe8f4

Observation 5a18b25a-5a27-454c-825a-af4e52c2f3e2 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.141106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.141106Z digest=sha256:a220966e5478101d159b4376396e2a7bf8310d44763c70e52cba349354478b5a

Observation 1cd17fb7-3032-4d16-945a-305b1cbdbe31 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.145181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.145181Z digest=sha256:e72d84b5a61b3b3a9e02e011f6b88b47426f4afc8308f435cdfa1ba9e69a66d8

Observation 29b5c9b8-f9f7-4945-8f4d-5cbf1548a5f1 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.148756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.148756Z digest=sha256:e3affeaa298397e6f2d3885416e99f9e50cd848a50bd1a5d09bb744adb18a945

Observation 0c7497cd-38ae-4c5e-934b-60964a4b1b1b · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.407242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.152223Z digest=sha256:b40e6bcca6892ec1164fb13c29552441827137b3e83b0a94b4a1131a9a49756a

Observation bd71a01f-2ac2-4f74-8caf-d03740639725 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.155401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.155401Z digest=sha256:9e344be51c4e0f217c411f968ddde8071ed979c5c7dcb72aef1c45631d7193d3

Pith citing papers

Observation ee5f873c-6dee-4731-8241-a898d9499189 · inbound

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark cites this paper.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.500206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.500206Z digest=sha256:c9ab0faa30187ca8b9d3e63f0b2828564826fffcfb23f8baf9a8431a2b443acc

Observation 257a7895-bdaa-4beb-a86b-a8c099750e61 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.371485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:466b50047e37d7b029b064271e7642e3fed39b67ea52e8b4fc554de8189b2342

Observation c92d0997-0214-4ac9-a7fc-d42ffd2e7a7c · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:42.000176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:dd452c0cd0c7e16debb8e9d5a98ef9e473a046bbcd10dd72eaa60a5cddf996c8

Observation da70c93b-f4ae-4780-8e09-a7717b5f7a75 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.293580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:6592e584c850e9a2e569f2ec2f3b57500028b67e34b6425fb7a270b4f885b4bf