Pith. sign in

Paper Citation Record · LEDGER

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 4 inbound Pith citation observations for arXiv:2507.09876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09876 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:34.963022Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.833017Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7cc29dba-3559-498c-aca7-61e5c70afd2c · outbound

This paper cites GPT-4 Technical Report.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.724807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.724807Z digest=sha256:9461b1267cdeb8a43a787131e0cf6311dd44a02732bcb7cadc24912b2dd66035

Observation 27dccd33-300a-42ae-bdef-f6fc5854b9cf · outbound

This paper cites Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:49:39.444946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.729419Z digest=sha256:69107eb2109d23d866c229de8ad2daaceb8d9211577de70c0bf42615b5387afa

Observation 015d76b4-6404-43c0-8a9a-186d849caf1f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:39.276953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.734176Z digest=sha256:b91bec0c8f5fbc850e2777b7aacbfaf6971d34ae64e84f1394c9884054d4e12c

Observation 9f229cfb-9c15-459d-ba5f-197b93d0830e · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.738049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.738049Z digest=sha256:d645712d8c8d74b63331bde73ee3624dcd50e736af6e6e812941a141ae0fda38

Observation 27225dd8-3d9e-4ee7-ba66-f364c93899cb · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:39.124749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.742594Z digest=sha256:2a0b198fac42b11260f78ddc1c279a3ebf9873f315622b49bb7b8670eb7f3521

Observation 6b51a282-7bef-46aa-bf45-3d01fdc5038b · outbound

This paper cites AI4Research: A Survey of Artificial Intelligence for Scientific Research.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models AI4Research: A Survey of Artificial Intelligence for Scientific Research

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.746783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.746783Z digest=sha256:6f236fafbdb3dca0a194e17ac568d54d05f64bd82c924af5326e7de6054bb6f4

Observation 381c4fb1-79fd-43e4-874f-f9ace96b357c · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.751721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.751721Z digest=sha256:bda846cae2d4ff0d6ab27bcbec319f6aef3e6d1c4c4d1d80f8da9e94626492ec

Observation 7cbc93df-fee0-4c19-be81-59e670680f03 · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.756138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.756138Z digest=sha256:49a0951c6a78a088438c2a92521de872a5064ccdd5540c6bf99a660ff49d2460

Observation 2d461dc0-9426-464a-83f7-fcb862011ed3 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.760649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.760649Z digest=sha256:384715307247d2415f2adf36aae923293157bb46ace6b9ae0a64056abbb097c6

Observation 6a1f1955-fb2e-4578-932a-2a72e29473a0 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.973447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.764444Z digest=sha256:72e4548c081c60e807acc4a369ed0a011c57d90f5f4bc2861e5f4b9eb2f1c0cb

Observation cdb2164e-c249-4acf-8c87-f125a9f45085 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.768376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.768376Z digest=sha256:dd5e8608bcb6bdbcdfa6f98fa4f397657c4d13867346bb3450cbb0454e9defa2

Observation 1c9ba49d-6aa0-4872-a967-faf832a01226 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.817718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.771662Z digest=sha256:8caa8aa14cdf9858104703f9b1f89068326339dd867064a442999031117af37d

Observation e2612be1-665c-4f42-aef0-71123dcd7b7c · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.675109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.775143Z digest=sha256:32a951becca936fce35b6c095d5b044d47c544d9691884547408a63dd3a6dfff

Observation fc9b25a2-4771-49c2-9ba0-ada62172607b · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.456307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.778450Z digest=sha256:0b6d9f7962257c6a5f6d294b9987ba537011a8f7f7bffe58e77234430bc5fc55

Observation 95d610e5-de64-4d0d-9cfd-0afd5792bf15 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.318985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.781657Z digest=sha256:9a865443baf6f5f5a3b8e41779c38840eac4de133401391986dd644620bbe300

Observation 47f20178-e7b6-4f49-b8e5-6375af45dfaf · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.784841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.784841Z digest=sha256:80c02e99d6c167b76b25fa0bb1b595a93ed1dab850dec67f6b5466c4379d0088

Observation 40a9f754-3d02-443c-a189-4b4386002a73 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:38.144409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.788646Z digest=sha256:29c4af165142a7063cfc8e741e75af29c833596e1abde013ed5dc98c349182f1

Observation 34ee85ce-7b5b-4e31-ad35-2200692b9b20 · outbound

This paper cites Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.791982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.791982Z digest=sha256:61339132c951b15d2f575d56818992269a2909319d61e9a782f6e5bc4e74632a

Observation 2608cca2-ea4e-443e-af2f-9c56d48093d5 · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.796119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.796119Z digest=sha256:d79edbeb9f1bfaca46fd4bcc0e14c4fbde21a4ad77c544ef771721b9eb67921b

Observation c83bbcbc-7685-4a7b-a003-7427cf56afec · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.973782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.800952Z digest=sha256:e824d7937fc2e6897aed2abf6986809d754fe2a2a5948896cbd1fe3b681b2062

Observation 8a6a185a-9cf7-4d98-8d98-077db591bbc2 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.797446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.804448Z digest=sha256:78c7af55d7f381fb3c3443912ee2ca4f3a5eaf89a89a1250b797392d3b9d3572

Observation a2931641-5298-47e4-9ec9-3fe25df594a7 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.614690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.808242Z digest=sha256:a3d4edcd70c84947f99e9e417a3eb676d4b4f99345141ac9b6f218fd0a60cfdb

Observation cf46c7f1-1d0f-44e2-aeb7-1e2d646f3f9a · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.812074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.812074Z digest=sha256:34033753019e13abaee7c64ea0dbae591024f3063a76bc98a7e885d44c11663b

Observation 705c3741-17e4-41b5-bba2-d05810d930e0 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.816549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.816549Z digest=sha256:7e533740ecc7945eba82083297360ccefd130cc17ee2ed5900abca8eebf32285

Observation 99ae4b9e-a7b2-4628-8444-243c9f642b85 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.432290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.821127Z digest=sha256:5a2c0694f38a08f28930b07d17baeb737a439824ae1a0094ec522b1828cf8659

Observation 8f5a3dfb-fe81-4984-a699-8009d92c7ac0 · outbound

This paper cites Large Language Models Meet NLP: A Survey.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Large Language Models Meet NLP: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.825536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.825536Z digest=sha256:5b3afdf1d30347f12547c3eb771366b57825253a768c50d4b2249fec1f81dbd7

Observation 543ab5dc-eed6-4180-9e03-722476cbd432 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.243312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.829913Z digest=sha256:96077681032c8a729f6895eedc0fed6ba9bd613af4f0a8209d4554fd2d93f8bc

Observation 417c138a-f82d-4e25-8533-f5fecc693e0f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.833545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.833545Z digest=sha256:e42e2828a3e5b6ac34bc671903f2724e5f8fc212345eee1b115bbcc59b63a5e8

Observation aaf1e309-0123-40fd-9642-f6257d239dee · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.836951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.836951Z digest=sha256:6f1961b4e4aac4900492283370d0d45e9eb7c975b9a624486e032abac1d04752

Observation cdb59fb9-855e-4629-98b2-c2e8e7362bd0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.840557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.840557Z digest=sha256:1e4d5338f9fbc6246f377d36284d7d41fceb840c8f090f9533208f5d6ec0149b

Observation 0b118f7e-1b0c-483d-bf9e-cf69301697dd · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:37.066785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.844247Z digest=sha256:51a33fe4b15dbe2f4ff28ce7057a3b0f20bf5afc6e923b3300ce823d403d0f31

Observation c8b75389-fff8-4870-8dad-a03a8435da70 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.848049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.848049Z digest=sha256:c5b83f6e0256d8f9a1c944b51814fb76b49f72f6fee55f90db1651d81b41676b

Observation d4367e60-f478-420a-ac14-fdb8ded6754f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.875213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.852024Z digest=sha256:a858385fc61263c4c268d8aa5dd3693c11722de3777b47bbf5db2ebbcd993915

Observation e94e0bfa-8ec6-4cfa-aa30-3dd8f9994c8e · outbound

This paper cites ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.856382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.856382Z digest=sha256:522b6883c4357c62d35a0a4e0ed237a4ff3ea8eb6ec21a7ddf034b05d148265d

Observation 96400931-ec18-4727-a50d-ab57e09dac6f · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.861248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.861248Z digest=sha256:bfcfacb7470ae895f5d055505c60fc08e484fc33d395bf0ae96d3567df93d339

Observation 4039a2da-2b2a-458d-952c-ce10cad2ac97 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.865290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.865290Z digest=sha256:5d9cdc82267171da7cae0101871b47a8578538925c37e8c2d1feec4b7665cfad

Observation a7b84ee6-057a-4995-bdb6-9daf7fb89eb0 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.762340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.870094Z digest=sha256:64fa58f2f7ae9b8d1b9fb23d65c919375ece6a8999c042a982c9e201051c43ec

Observation 07208b62-5b89-4fe8-a7d5-d9468087a54e · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.621591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.874308Z digest=sha256:f425749b3fdd63ce974871509af89d68e41014190e59891a88324dc4398fb32e

Observation 725d5f29-1d0b-4b20-ad2e-647f8f1ac83f · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.877874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.877874Z digest=sha256:f5321ac70bb0b84dd6fdd6d489d503e24928a28808dda768a7b77d3bcd8ddf86

Observation aa414a16-3ddf-4612-b5aa-f74256674037 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.340889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.885407Z digest=sha256:ac130fc1b170a82230519f819318c1dd4662edc909ea37f983e5dd8f79706da9

Observation 4021bd4b-b315-4079-99d1-370aadc0c1c5 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:36.190327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.889875Z digest=sha256:bb5bc54dcbd1c72899123c729a31b6c6607e6167bc2e409d72be1aa4b7c1dd39

Observation 75444443-fc2f-4161-a82e-a565c3d2ca3c · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.898886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.898886Z digest=sha256:4b82bf2eed0f761670c5bba530199cb8199e152e99acc6a96a88778c32ff8dd2

Observation 81bbbdab-83a1-4ccd-b08e-068c332d6e55 · outbound

This paper cites In Proceedings of the 32nd ACM International Conference on Multimedia.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models In Proceedings of the 32nd ACM International Conference on Multimedia

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:49:36.091285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.894306Z digest=sha256:ce722aafc3b7883240db27c0553874e0676da55be86c03a6167afc2251b1e5b5

Observation f5c01e62-6511-4ff8-a8b5-38d91ddb0294 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:35.989539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.907424Z digest=sha256:65d6a29051697d4922b793465ef1cbb4cd5c9b0c09da2fbf908be87902b7c6c2

Observation 738d0cee-20c3-4b20-9962-7b8c1bdad03c · outbound

This paper cites Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.903388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.903388Z digest=sha256:f2acdb3fe162310cbd27d53e35a2ee454168cabf004409183ec101f1a80faecb

Observation de0f54f3-7733-48c3-a7ca-14a2344bf4da · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.914835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.914835Z digest=sha256:189b92a5575157dfc2747199f5920e607fa710559c18caa22b08f6e891357d17

Observation 2dce07eb-d539-47f4-b7e2-3ff2e92a9a68 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.911048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.911048Z digest=sha256:b0c8da72df10641fb196fec4bd4d55c0c2a4386036888b204255fdd612760677

Observation da72bc3b-1d72-4d21-b00e-4fb3eae1daec · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.922692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.922692Z digest=sha256:21b3be48a285260fff1dd91327c4e19b13483431c2a536a451b54af387dd3b0f

Observation fb1ad447-fb3d-417a-bbda-ba92963928ef · outbound

This paper cites Qwen2.5 Technical Report.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.919181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.919181Z digest=sha256:c7091fd12b3173c4828a3d32eefc92e49ddbd4c13ef6952de30e0dd70052e698

Observation c1269154-5f01-4e48-9464-a743702026bd · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:35.848311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.931465Z digest=sha256:2c8186f953223d0d7b345b91644a30e759cccfbc8671f77b1d26022f86b3416f

Observation 5dfb38a4-4956-47d4-9875-0bea4f982ff9 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.927360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.927360Z digest=sha256:643697023bd3860cfd553b7735a18d3f1bc2011fa2caf31647948a8b210d39c0

Observation 5e3c5ab6-bafd-4e4b-bf04-bdd6ff960e1a · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:49:35.695185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.939634Z digest=sha256:ddf9730cc20d4c6263204a21ca3e24d5f6a114286fe074c9e3e987f85ce2d1b2

Observation 880af460-c720-488d-9016-1775980c506b · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.935033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.935033Z digest=sha256:71babb8be27e548a4553710454a7643fa1385fd5fbac586f6d774a3d306f5416

Observation 77f6682e-d6f3-4804-97bc-14acf2efeb61 · outbound

This paper cites A Survey of Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models A Survey of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.948907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.948907Z digest=sha256:d1be2e0eba4987307ee80b9e4d5823321e1fa5d80575c9e003acdd3f8b7b0a65

Observation e905d582-7990-4c7a-a245-8611b3d2a1ed · outbound

This paper cites CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:49:35.074557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.944193Z digest=sha256:22c180e50d13cabeea66e7e989a4265d6049781c483db8fff7b338e19163cac7

Observation dd4dff17-aeac-4d24-8f70-b2c93b50ac69 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.958286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.958286Z digest=sha256:a40dbbcbd48a02438d416107bc8f272930d8381570d30bb83386bc995369cc7d

Observation 5048f173-f12d-46c6-b396-352c5b0f0fb0 · outbound

This paper cites an unresolved cited work.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.953969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.953969Z digest=sha256:b0d31ed31895f3b7bd1c78250153f842006b4967319dea61cfbbd104852ae8ce

Observation 71857a80-a104-4735-aadb-a2a0ed40550d · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.963022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.963022Z digest=sha256:5740116f817a7f16aa41422a3412ec007340060a65d97ebabc3fe12ee96f8de2

Observation 88264853-d586-461a-bbb6-a98f2bc7408b · outbound

This paper cites In European Conference on Computer Vision.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models In European Conference on Computer Vision

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:49:36.505134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:49:34.881390Z digest=sha256:9ce9fc00c6ca649e10ebe003c5209056a6716f3baf60f3186f41f7688961a556

Pith citing papers

Observation 558a108b-cffe-40fe-a71d-db26eebee2c9 · inbound

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm cites this paper.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:55:35.070077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:c7c617f723a07aaf060571ec0d2505f74bba39acdaecf19d9ddd260beeb4393b

Observation 51a2c363-4137-4637-9b3a-ed92ff0003fd · inbound

Act2See: Emergent Active Visual Perception for Video Reasoning cites this paper.

Act2See: Emergent Active Visual Perception for Video Reasoning ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.948194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T19:34:53.683729Z digest=sha256:898fccc53a56c836e3fd5cac7efd08fee6c37e4807fb11ff919db7eefbad2ab9

Observation 05811938-693a-4663-9912-8d5666776b84 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.408448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:1da7ed1ec1d817ec641c740069ed3deca1b80e20b9a4dec5190fc63385ea47a0

Observation aed15bde-fead-4786-81d4-3742e7779b81 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.834326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:b578764ac03f2d33b1155c756a273174e80564a6f2423a1bb8af3e4dbe67a393