Pith. sign in

Paper Citation Record · LEDGER

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 4 inbound Pith citation observations for arXiv:2507.09876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09876 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:34.963022Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.833017Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7cc29dba-3559-498c-aca7-61e5c70afd2c · outbound

This paper cites GPT-4 Technical Report.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.724807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.724807Z digest=sha256:a6dfd42853f79aa68d8ccbabe10e287a0a8f128e97d82750b46b7d8d8b3b9348

Observation 27dccd33-300a-42ae-bdef-f6fc5854b9cf · outbound

This paper cites Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:49:39.444946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.729419Z digest=sha256:32ee057c1e9feb4c225f067860dc176e74547bc2685d7da71152d52ad8f0bcc5

Observation 015d76b4-6404-43c0-8a9a-186d849caf1f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:39.276953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.734176Z digest=sha256:46f775d6510054e9a7e5a8ecea4dafedc8a18a7ff8ef0558cc2954191d2a9245

Observation 9f229cfb-9c15-459d-ba5f-197b93d0830e · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.738049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.738049Z digest=sha256:e497d6bf2953df930d4fd4833f08a23d1c35112a8786688b7db5a50098f4ec46

Observation 27225dd8-3d9e-4ee7-ba66-f364c93899cb · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:39.124749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.742594Z digest=sha256:1509b343c7b5fee184c75c967fd8cc35ac2500496d43e899e88ee2fef617280f

Observation 6b51a282-7bef-46aa-bf45-3d01fdc5038b · outbound

This paper cites AI4Research: A Survey of Artificial Intelligence for Scientific Research.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models AI4Research: A Survey of Artificial Intelligence for Scientific Research

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.746783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.746783Z digest=sha256:aa3c19bd500585ca308388593ceb4746bd2db320d91ddd12ef5004f90654a8ac

Observation 381c4fb1-79fd-43e4-874f-f9ace96b357c · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.751721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.751721Z digest=sha256:f2c4385397b032e2d091fecab7b1a319224e2fef72cf677ead8b374c60c971d5

Observation 7cbc93df-fee0-4c19-be81-59e670680f03 · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.756138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.756138Z digest=sha256:984dd0a4ce488c2d385d84d5ba2ddada1f4fcff5facd31ad132ad628c8a9c213

Observation 2d461dc0-9426-464a-83f7-fcb862011ed3 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.760649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.760649Z digest=sha256:91b60d3a62ab1e421692615c5572200726c7118b5a0ddb06d2f8b4ee2444875f

Observation 6a1f1955-fb2e-4578-932a-2a72e29473a0 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.973447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.764444Z digest=sha256:984159199bb1b0dfa8f77a3a656f22eebc761232281e6db066237c9e9f5686da

Observation cdb2164e-c249-4acf-8c87-f125a9f45085 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.768376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.768376Z digest=sha256:5348d6f4ed38c13c7bd3da5ba3d1ba87e616bf7fb564223e95bb31c27f3dbea0

Observation 1c9ba49d-6aa0-4872-a967-faf832a01226 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.817718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.771662Z digest=sha256:4ebcffdfb7f39b17240b7be5f888e7ba7794e14ec149338dd87ca70821f3118b

Observation e2612be1-665c-4f42-aef0-71123dcd7b7c · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.675109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.775143Z digest=sha256:f7c9774920b2102a00745cc74cb800b13c8d178255f2fd1e9aa5d4c440cd571a

Observation fc9b25a2-4771-49c2-9ba0-ada62172607b · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.456307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.778450Z digest=sha256:29ae9cc4bb8a09159b722c9bcee86898355557b9d1bd59329d258654e2d2b782

Observation 95d610e5-de64-4d0d-9cfd-0afd5792bf15 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.318985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.781657Z digest=sha256:1a572b819e81043aefbcf13e93d9c84ecb32f9c53bcf475fdf5247a2e96e73e4

Observation 47f20178-e7b6-4f49-b8e5-6375af45dfaf · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.784841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.784841Z digest=sha256:340c4d692390ff6f9109d7b11cd415c09c3feac392e9f4601bd8cfcbfa505685

Observation 40a9f754-3d02-443c-a189-4b4386002a73 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.144409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.788646Z digest=sha256:29a6ed5083be3f6671a00ced2118efd8f886600588d900ee75a30e0c64082516

Observation 34ee85ce-7b5b-4e31-ad35-2200692b9b20 · outbound

This paper cites Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.791982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.791982Z digest=sha256:e8f3aa63608ca172e1f6ebf663bb66a9277797e9f64e6e657f0b0034cf02fcf7

Observation 2608cca2-ea4e-443e-af2f-9c56d48093d5 · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.796119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.796119Z digest=sha256:f013e94327ec58f7c9c61e09ea2841af1c5ed7da1d4bac4fe13546a8468094f4

Observation c83bbcbc-7685-4a7b-a003-7427cf56afec · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.973782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.800952Z digest=sha256:ae5473d12d5eb3dab9707860d1cced5bf0d90d831a33c1fbcce1bacfcfdabc8a

Observation 8a6a185a-9cf7-4d98-8d98-077db591bbc2 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.797446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.804448Z digest=sha256:22b11ca9c53fa17238eefbee83bd9bf5c3470e95dec7f2d2a5e5ef7f66051820

Observation a2931641-5298-47e4-9ec9-3fe25df594a7 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.614690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.808242Z digest=sha256:ea6655450c72ba61c8b25ef7a8c76ad09c5b1067a7962578d5b4b8575a557d2a

Observation cf46c7f1-1d0f-44e2-aeb7-1e2d646f3f9a · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.812074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.812074Z digest=sha256:58c233020137a982be0289f503d33d7dc5e6279164a4c4a309e39f235ea5d4ee

Observation 705c3741-17e4-41b5-bba2-d05810d930e0 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.816549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.816549Z digest=sha256:56dff88fa0dd93d715615593e38200bb0f1ebd1fe9e547d410e2b5638bbc3c6a

Observation 99ae4b9e-a7b2-4628-8444-243c9f642b85 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.432290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.821127Z digest=sha256:c8834332dec70b4c9bcd1bfacfca99dbe9aefb624239b0ff6439dd6a7cb93e84

Observation 8f5a3dfb-fe81-4984-a699-8009d92c7ac0 · outbound

This paper cites Large Language Models Meet NLP: A Survey.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Large Language Models Meet NLP: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.825536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.825536Z digest=sha256:13894f5570fd64f7a30f7ced6d03331b0fd4e1a75506fd51beb229f41efe9c99

Observation 543ab5dc-eed6-4180-9e03-722476cbd432 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.243312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.829913Z digest=sha256:236a910dca6fd26d1ecb848564cb675a75e855453913531d1825a9e0c907d1d7

Observation 417c138a-f82d-4e25-8533-f5fecc693e0f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.833545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.833545Z digest=sha256:5da687e969614b222b708d7dd5b7363abed783f8dc339b9e4a5352ff78c8dd01

Observation aaf1e309-0123-40fd-9642-f6257d239dee · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.836951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.836951Z digest=sha256:ac584a1820629cdd1a805b0ea9069aeced04e4e793fa3494d3258ed1e607f4e7

Observation cdb59fb9-855e-4629-98b2-c2e8e7362bd0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.840557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.840557Z digest=sha256:12ea68b3c1b234f1239c50de6b7bcd5346da4233484d898be1f00143cfa97d35

Observation 0b118f7e-1b0c-483d-bf9e-cf69301697dd · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.066785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.844247Z digest=sha256:64a1e23be115addd13b048c10dc801813ad65c03dbae4c431135081b94763db5

Observation c8b75389-fff8-4870-8dad-a03a8435da70 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.848049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.848049Z digest=sha256:1d214e9c3e02b58004956156b02561acc807bf3ba331a5c85c55d9ddef5bc9c7

Observation d4367e60-f478-420a-ac14-fdb8ded6754f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.875213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.852024Z digest=sha256:7ce66d1202bc000280cf611b167885eb19fc2c47415842470715dd89b8721c2b

Observation e94e0bfa-8ec6-4cfa-aa30-3dd8f9994c8e · outbound

This paper cites ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.856382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.856382Z digest=sha256:5eaf9e9a3112aa100f770bd05a388abf80eac2c8170efb989e699ff9823a5677

Observation 96400931-ec18-4727-a50d-ab57e09dac6f · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.861248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.861248Z digest=sha256:cb921c02015fa1af745ff1d3d96cbb2d9f017322363752a2ce318696672e38c8

Observation 4039a2da-2b2a-458d-952c-ce10cad2ac97 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.865290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.865290Z digest=sha256:35c19bac7f016f688f772577c00bdcdefe233e3f4f10b4225e488f95a29f79f0

Observation a7b84ee6-057a-4995-bdb6-9daf7fb89eb0 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.762340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.870094Z digest=sha256:c8e980a154f66df5586befa370fccd99375c5868696208d49249771c23a6cf19

Observation 07208b62-5b89-4fe8-a7d5-d9468087a54e · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.621591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.874308Z digest=sha256:7203dfae77e6ee59bc8cd2734bcc1915e9ffedf02971bff6ef892fce64d38ffa

Observation 725d5f29-1d0b-4b20-ad2e-647f8f1ac83f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.877874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.877874Z digest=sha256:34bf8fd8c57a87a9dc7accfd9e2419690a9ae5c7dc2d7519eb7f892b70260844

Observation aa414a16-3ddf-4612-b5aa-f74256674037 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.340889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.885407Z digest=sha256:cbc1fdb571c6a900022f3944d19956786ac213dd960901e79da83e063f06e6a2

Observation 4021bd4b-b315-4079-99d1-370aadc0c1c5 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.190327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.889875Z digest=sha256:b0d702725ebfedd121ad94a026ae7da715285ddc00b12ef5aa25530a56eac400

Observation 75444443-fc2f-4161-a82e-a565c3d2ca3c · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.898886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.898886Z digest=sha256:4ce58f6602de32200afa9080b667f84ce74c900e5c70a202e33918c53f653c99

Observation 81bbbdab-83a1-4ccd-b08e-068c332d6e55 · outbound

This paper cites In Proceedings of the 32nd ACM International Conference on Multimedia.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models In Proceedings of the 32nd ACM International Conference on Multimedia

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:49:36.091285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.894306Z digest=sha256:e086a2534277635324f33cd97c135dcb1df233c14ba80c984302eb58aa7e9c8b

Observation f5c01e62-6511-4ff8-a8b5-38d91ddb0294 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:35.989539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.907424Z digest=sha256:c9f9859abe9cfbbbfa5ff36d0a6f7ba3fbc29545d9fb242f7c2f94948db5e4f0

Observation 738d0cee-20c3-4b20-9962-7b8c1bdad03c · outbound

This paper cites Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.903388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.903388Z digest=sha256:02aaebb71a6b52afaba3ce39c1ec6454ab2bb6db69c78cd2ce7163554ae848d9

Observation de0f54f3-7733-48c3-a7ca-14a2344bf4da · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.914835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.914835Z digest=sha256:d5a877ea07d2e42ca6d395335f025ea348fbce0295bc9c70925dfedebe16d7d5

Observation 2dce07eb-d539-47f4-b7e2-3ff2e92a9a68 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.911048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.911048Z digest=sha256:1f27233a01894c73b9281d3c6152dea240408db7c34514f9e2f43061eb6818f6

Observation da72bc3b-1d72-4d21-b00e-4fb3eae1daec · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.922692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.922692Z digest=sha256:641216b81e0c87ac40e48f61926b8febca367c6eededcc079effc7ac71397266

Observation fb1ad447-fb3d-417a-bbda-ba92963928ef · outbound

This paper cites Qwen2.5 Technical Report.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.919181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.919181Z digest=sha256:894f1ba02b67cac2b7573fbd285779d5cb9a60228af94928085950dbedd5184d

Observation c1269154-5f01-4e48-9464-a743702026bd · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:35.848311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.931465Z digest=sha256:c223af29317817023e1a30f5f3908ba16061815d878eb39b5fb7080d550074e4

Observation 5dfb38a4-4956-47d4-9875-0bea4f982ff9 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.927360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.927360Z digest=sha256:b81d0383474a966e5e26b4be0be73b203c0f6e3e6ec32a5a5657c904b8563b01

Observation 5e3c5ab6-bafd-4e4b-bf04-bdd6ff960e1a · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:35.695185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.939634Z digest=sha256:26eab301727c7a253689f06454ac4f0699fd225926f6b3fa6206a49304fefb4d

Observation 880af460-c720-488d-9016-1775980c506b · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.935033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.935033Z digest=sha256:dca7d02a9cbfe6a758007a5398cb5dd9d70b489a0723d8cd8d895bb114c0e6a3

Observation 77f6682e-d6f3-4804-97bc-14acf2efeb61 · outbound

This paper cites A Survey of Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models A Survey of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.948907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.948907Z digest=sha256:fc773e39e0a4cd6dfa3fd729cb175559e434dc2528a76cc60f64bd9cb1abdf6d

Observation e905d582-7990-4c7a-a245-8611b3d2a1ed · outbound

This paper cites CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:49:35.074557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.944193Z digest=sha256:9daf6d917428a5a45cbacf9734228eb0b2970f0173f96929900f9c2241ce9120

Observation dd4dff17-aeac-4d24-8f70-b2c93b50ac69 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.958286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.958286Z digest=sha256:03c1f984e600eb88dad301888d070ebf8ddaefd3cab8b2608dc807fe160aa97a

Observation 5048f173-f12d-46c6-b396-352c5b0f0fb0 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.953969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.953969Z digest=sha256:e444d3155e4c6f86e597b343f8204c8f57bd883bce1c7c225154df2078cac474

Observation 71857a80-a104-4735-aadb-a2a0ed40550d · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.963022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.963022Z digest=sha256:957d0dbe697dec5d9e4e88e95888d853a54576af1dcc026342be846a7861bab5

Observation 88264853-d586-461a-bbb6-a98f2bc7408b · outbound

This paper cites In European Conference on Computer Vision.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models In European Conference on Computer Vision

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:49:36.505134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:49:34.881390Z digest=sha256:7daf36f3c6f5169a4f2b8b6ec93c80b8fc4337f479c600f9d2c403d08d18aa8c

Pith citing papers

Observation 558a108b-cffe-40fe-a71d-db26eebee2c9 · inbound

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm cites this paper.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:55:35.070077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:d0cb562420f4ed9b53849b96cb45176c266d663f04d5b67c93eba0576efb3ced

Observation 51a2c363-4137-4637-9b3a-ed92ff0003fd · inbound

Act2See: Emergent Active Visual Perception for Video Reasoning cites this paper.

Act2See: Emergent Active Visual Perception for Video Reasoning ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.948194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:34:53.683729Z digest=sha256:32daf6cea086b79cd2f940be61e6457bb211ff9c7b9996e29aaa7491960dc0f3

Observation 05811938-693a-4663-9912-8d5666776b84 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.408448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:77943186ea82a9dfcbd5691001e96583fbcf4e6b859bb6145e87aa1a05343441

Observation aed15bde-fead-4786-81d4-3742e7779b81 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.834326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:7c66abf422ad1eeb218877fc4440a5fcba55f6fe847fafc0a6f077e24dc1ed9f