Pith. sign in

Paper Citation Record · LEDGER

Vidi: Large Multimodal Models for Video Understanding and Editing

As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 8 inbound Pith citation observations for arXiv:2504.15681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15681 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:25:33.395157Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:32.341621Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:46:13.651443Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75545c5d-9008-4c3c-bb3d-62c7e814ba34 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Vidi: Large Multimodal Models for Video Understanding and Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.212021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.212021Z digest=sha256:50cd01c4eba4061a528147f883c3211bfe2ed528501a78bc7d2d6edb64ecd5a6

Observation c1b128ed-282b-4a18-be43-1223ecefd36d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vidi: Large Multimodal Models for Video Understanding and Editing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.217824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.217824Z digest=sha256:20f9ea225bed8e427d9c5f84f9e1e619bab2c6bed99e641790e157f940c1ba29

Observation 1dcb4d13-fd59-4cc3-98b2-9741a3fb604d · outbound

This paper cites Qwen2.5-VL Technical Report.

Vidi: Large Multimodal Models for Video Understanding and Editing Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.224325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.224325Z digest=sha256:666e11ca079e1ee318babf41e7b99d7786539d5d448185a7355a0bba03e5a016

Observation 908c11c5-134e-461f-9765-87df8cc8216b · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.230301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.230301Z digest=sha256:d209502cc5830dceea4dab1e37eba0e889d5f6f0cfd75313ec217559684acf95

Observation a55b34b9-08aa-4fbb-b82a-9c9a86b8f33a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Vidi: Large Multimodal Models for Video Understanding and Editing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.235593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.235593Z digest=sha256:26957df7e5e0cd3a9a28c488062ae1f901396a147a1a08030335a02a130042aa

Observation cfdf50ae-f463-4b04-ac97-3d186445b46a · outbound

This paper cites TALL: temporal activity localization via language query.

Vidi: Large Multimodal Models for Video Understanding and Editing TALL: temporal activity localization via language query

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:34.042280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.240648Z digest=sha256:29248695d388d99b5a9e1b65aa8fe503083745217943d837a3ed3841a6d80f4c

Observation 214ad4d7-1925-4596-82c5-de054e189fc1 · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

Vidi: Large Multimodal Models for Video Understanding and Editing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.246113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.246113Z digest=sha256:9bf9e5d96680f3f91a3420eb9295dfab4e737ed41308ec9a1312ac385b028635

Observation 2f40f9da-8876-4edc-b47d-0853101dd3e1 · outbound

This paper cites Fullstop: Multi- lingual deep models for punctuation prediction.

Vidi: Large Multimodal Models for Video Understanding and Editing Fullstop: Multi- lingual deep models for punctuation prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:34.026104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.251244Z digest=sha256:d4b70847b091f7ccfd097c621e411ce71b4db7b2612f8a25077d3614ea210d9a

Observation eca0cda8-904b-41e6-bd67-3ee360580d5f · outbound

This paper cites an unresolved cited work.

Vidi: Large Multimodal Models for Video Understanding and Editing Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:25:34.011321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.256100Z digest=sha256:7b26b2b7a48dacbb1261ae452774c3b0df81ed21cd5b98e7405c1e5d4d6fc4ea

Observation 2f08ae5d-be3a-4036-8e2e-b5338d041836 · outbound

This paper cites GPT-4o System Card.

Vidi: Large Multimodal Models for Video Understanding and Editing GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.260694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.260694Z digest=sha256:f3f905d354208bf8cbbe6f3ffe99347b0dfba6e19e10be1b4feaeef7453e0bb2

Observation 993668c6-c7f5-494d-b87e-68b41cfece4f · outbound

This paper cites Mistral 7B.

Vidi: Large Multimodal Models for Video Understanding and Editing Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.266005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.266005Z digest=sha256:50ce1273fc34dd2e24670f536cbb0231121bfd3ab203d9262120232f44548903

Observation 925c9294-5948-4462-937f-d5de7c897a1f · outbound

This paper cites Dense-captioning events in videos.

Vidi: Large Multimodal Models for Video Understanding and Editing Dense-captioning events in videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.996400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.270683Z digest=sha256:258fd46584643df8ae529e27a4823e1e0dcb18af553ed4c4ff00dd08046ae603

Observation 52eb6fbe-0c18-4296-879c-85f36cf4b945 · outbound

This paper cites D-Attn: Decomposed Attention for Large Vision-and-Language Models.

Vidi: Large Multimodal Models for Video Understanding and Editing D-Attn: Decomposed Attention for Large Vision-and-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.275272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.275272Z digest=sha256:407b001c8ba5da07e6c1545c68681c81a0f564dcd5eb6fb4ab8c8f37a860dbf8

Observation c4de6e23-b1fc-47d4-86dc-8339a6c0d366 · outbound

This paper cites Berg, and Mohit Bansal.

Vidi: Large Multimodal Models for Video Understanding and Editing Berg, and Mohit Bansal

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.980908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.280295Z digest=sha256:3d4bd34fb4f39301b5211debb22381aea9838b221dc8a3953f32d31e39dae1d9

Observation 921dc973-69cd-4219-bec8-e91e21bac565 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Vidi: Large Multimodal Models for Video Understanding and Editing LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.284877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.284877Z digest=sha256:73b2ea4ecd2dd89a9368c1378605cb6684b50a9a75917f761c4820f69f1bb1e3

Observation 79db2b3c-7825-41f7-b54c-cabe13a8bc0e · outbound

This paper cites Structured Context Transformer for Generic Event Boundary Detection.

Vidi: Large Multimodal Models for Video Understanding and Editing Structured Context Transformer for Generic Event Boundary Detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.289574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.289574Z digest=sha256:c678afcc250991c9943006f6b916071a4ae529aecfdcb4eb557d6f40b1841693

Observation ef0b583b-9242-4a78-a9a7-563b4c57b1f7 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Vidi: Large Multimodal Models for Video Understanding and Editing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.294219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.294219Z digest=sha256:c83779c424850001a563f2f394f708b610fadd3bd3fa76ebedb8eb8d9958d86b

Observation 5688aec5-1346-4d95-b5d5-46b36a739e0c · outbound

This paper cites an unresolved cited work.

Vidi: Large Multimodal Models for Video Understanding and Editing Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:25:33.965757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.299785Z digest=sha256:8389fd4406e1a29d9e5e60d25817415ed5c9a56375ceb20d854a12a1c6db9be1

Observation cc2472a3-33a7-4874-8443-8ee7660d6d0c · outbound

This paper cites Visual instruction tuning.

Vidi: Large Multimodal Models for Video Understanding and Editing Visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.950194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.305127Z digest=sha256:a54f2224c0665401247480b6bd47a0524d804fa55abfc141741568d37eb24e6a

Observation 4212afa9-c0d8-412e-92ff-3a793e0328f8 · outbound

This paper cites Decoupled weight decay regularization.

Vidi: Large Multimodal Models for Video Understanding and Editing Decoupled weight decay regularization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.933844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.309832Z digest=sha256:95d8ab1086633c08d5d5609d3b659a92aa487d43fc41561ff4b8158bb5e06511

Observation 7e6fb3ee-1f7a-4dc7-9853-8e70650fb3ec · outbound

This paper cites ZoomV: Temporal Zoom-in for Efficient Long Video Understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.315744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.315744Z digest=sha256:6c8544fdbf6cca6122f7511c0b5f4a2f99e86ef32256b0af99b57e335175a1a9

Observation aced7a25-51c5-4384-9cf3-48c9aa8dc898 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Vidi: Large Multimodal Models for Video Understanding and Editing Robust speech recognition via large-scale weak supervision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.913817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.320936Z digest=sha256:c1851d301784e784788babe4f947a8c72f85eda6b4c0de446d868b42ccfd5a08

Observation 03452171-a7be-4592-a52e-f95ebc2b79f9 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Vidi: Large Multimodal Models for Video Understanding and Editing CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.326194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.326194Z digest=sha256:e0615c770eba5140dfa1d8f14100e4ed9b1bcaf1f49e23a7c8d6d317d29c6090

Observation 175a2ce4-8450-4f99-8052-409ba80863a3 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.893528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.331326Z digest=sha256:7c79cd452fbfc25cc0becd510adcd0d9f61fab4b068b4acb0be3edf661652fdb

Observation 4feb0279-91c2-4237-ab43-02099d89dfe7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Vidi: Large Multimodal Models for Video Understanding and Editing Gemma 2: Improving Open Language Models at a Practical Size

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.336090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.336090Z digest=sha256:4a9f07142e13137d37b1c82bb0d13753db50e4dc2612cb5ec0b556d642962b28

Observation 550b9d51-acc6-4c81-ab00-9b987a757ac1 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing Moviechat: From dense token to sparse memory for long video understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.875255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.340709Z digest=sha256:2e8fad0cd7eb670d2352ba22f85b688e448174b7cc9c627cce6add404b0695a2

Observation dbe16d8b-ee4a-4e05-ac19-985e8cd3a226 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Vidi: Large Multimodal Models for Video Understanding and Editing SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.345897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.345897Z digest=sha256:7dc78fab789ef34fa54b4ae5b5169e6a9ad914d895dda5f76bf368022d53e9be

Observation 3a0225ed-6fde-4a55-91fd-4881953bc472 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Vidi: Large Multimodal Models for Video Understanding and Editing Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.858125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.351356Z digest=sha256:23d1feea640c67be45ceafb8a8c6c07dc1826f703419b6892df47379c8806283

Observation 5edaf4e0-9884-47d4-bbe6-5eda2d6cf1f4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vidi: Large Multimodal Models for Video Understanding and Editing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.356094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.356094Z digest=sha256:e756231dbe8670e1a2aa99a972b15ce8977e7fa95cf783d28da2ccaa2d5dc99a

Observation 654cb160-705f-458e-947d-6be61b1082fa · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Vidi: Large Multimodal Models for Video Understanding and Editing Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.360996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.360996Z digest=sha256:0c19159f5b8b3a6b358b2db74a4f6c2acfedbd7bde810ca79b3bd87a4ddf8245

Observation a6c0f6e6-1ee6-4bd0-85f4-c9cdd88e94e5 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Vidi: Large Multimodal Models for Video Understanding and Editing LVBench: An Extreme Long Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.365977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.365977Z digest=sha256:1bd77c8ae04b0b44c50dae8fe82c5e7137a99140c9614d11ed3616e99b0415b5

Observation c5303013-c728-4eb4-9848-818e6ea3fa1e · outbound

This paper cites Chi, Quoc V.

Vidi: Large Multimodal Models for Video Understanding and Editing Chi, Quoc V

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.841343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.370899Z digest=sha256:bf12b7dd37889f41ae207a47b9ebd50d9758602f984eaeeee331bc26d47c37c4

Observation e7b47912-3644-4c6b-96a0-0d3c51f5c5be · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.825454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.375805Z digest=sha256:c64e2e31bc19df9c6fc742377fc60275663772cd04fffe83329de106c725e824

Observation 3774260e-0ffa-4758-8876-f4d5e768af80 · outbound

This paper cites T*: Re-thinking Temporal Search for Long-Form Video Understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing T*: Re-thinking Temporal Search for Long-Form Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.380253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.380253Z digest=sha256:952213758e3804d7fac96116daaba7d2d60ec4e8cf4f3bf7b7ab598ea7c7e1b9

Observation 21125145-8afe-4382-a0dd-731c852c6066 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vidi: Large Multimodal Models for Video Understanding and Editing Sigmoid loss for language image pre-training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.809062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:25:33.385653Z digest=sha256:df788ac5ed3858a14f203def0faca1833d5ebe4ffcb6183a30d8bf89cc84f388

Observation cce8b1d4-b89a-4d8d-933d-e1a496a1b9a1 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.390194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.390194Z digest=sha256:74094846ff88c867cfb4084e2a464af1f73eb5d0429dc00210ef1560fe609ebd

Observation a27b6c80-bc29-42e8-9a38-94194eee2fb4 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.395157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.395157Z digest=sha256:026ba595582363868786c5ae7612053894b37997fcb1851096017105efae0285

Pith citing papers

Observation c784aa37-b1ec-47ec-b4c2-37215622e784 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:32.341621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:32.341621Z digest=sha256:cb6334ca17e04ddd28037fc5043b96bed114098dda0aaf69134c68b9b7ffe461

Observation e1dcda03-a4ba-4173-8a0b-7819012ba321 · inbound

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space cites this paper.

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T22:12:45.376792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:12:45.376792Z digest=sha256:bc7c72eae963feea85e70b89a43a14f95d333cbb10561045d1a0817c8d949aa7

Observation 60d8b032-c162-4b77-b4cf-53ec4b9aae5d · inbound

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation cites this paper.

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T13:21:10.692684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:21:10.692684Z digest=sha256:9e09b88b161db8d4ef5362b9bb6caa9bef82909f041b683f0830a7b8ee2d9963

Observation 3895ec79-b6ea-41d3-bc29-d9331292d0d8 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.126091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:db9a812d857854f2e1adc3bb521070bf8c1ae7f5d8a5ddbd8394a1a7443243bb

Observation 20247188-fc73-4b75-8c2c-6c629235a58c · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.660350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:89b43358e02bfce3052b43324c765fea97097510029e2728001207d3b765a543

Observation c6fcef02-2efa-4065-b867-a466e360b740 · inbound

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning cites this paper.

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:04:16.637378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:04:16.637378Z digest=sha256:4cae9e5be51913b90364f1dd833c7b9e84e26282019e50ed368a211e6db9fd05

Observation 8536bac6-8a77-49a4-be09-1afb22f38f08 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.844379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.844379Z digest=sha256:b84ce21e1e1f5b683e8952c278c9774695564f6885030ec34ba31c16c23dece1

Observation 48bf1810-df25-4868-a715-14b523634f5d · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.986807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.986807Z digest=sha256:2b02483ec02d853666a1d6d0eeef525a1eddd6acc36a05d30562024bfb474740