Pith. sign in

Paper Citation Record · LEDGER

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

As of 8 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2507.10302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10302 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:38:20.743420Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c201e185-eee5-49ee-a9b0-c2ca80d845f4 · outbound

This paper cites GPT-4 Technical Report.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.744309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.744309Z digest=sha256:596905278a5b05697e69887cf14ec3ba9abd4aa8df5c1609d70cd4603b2840ed

Observation 43efa36a-bb39-4e18-8419-7712a080ed51 · outbound

This paper cites In: NeurIPS (2022) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2022) 1, 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.800454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.800454Z digest=sha256:59f6024eb2264f534379afe65645442df05a8e1b1d15e6a6688feafd62c4486e

Observation 724bd4d6-c78a-45d0-a194-cfe9edf3479b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.887420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.887420Z digest=sha256:9202a6c313bec97a50cc68f6b037efcddbc3ce607a4c2db0478857b1124a0cdc

Observation 9f9a2c07-49b6-4862-a614-6beafbde6294 · outbound

This paper cites In: ICCV (2021) 2, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2021) 2, 6, 9

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.993701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.993701Z digest=sha256:ff600d19f199cab3c9d141b983884dcb6999bafc49a6ab19075e78aab69d9627

Observation 607be67e-a3c6-44a2-8956-93aab90f7d4d · outbound

This paper cites In: NeurIPS (2020) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.035353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.035353Z digest=sha256:b361b3989fc67d1302c5f65bfd417ed86047d04d5318a85c37278a95bd74e2df

Observation 03173b9b-2712-4504-8da1-8ca05b4a22fe · outbound

This paper cites In: ECCV (2020) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2020) 4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.105134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.105134Z digest=sha256:f9d924678fb27ac21db332d48f0afe6c1b01140281322b1b5be1a250b7bc6335

Observation 7836d8d4-84b7-47c8-a079-162823de13a0 · outbound

This paper cites In: CVPR (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.175728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.175728Z digest=sha256:b399b81b86f421685f43ae9bd23629a019f5f2b7038b4d82137a93d1fc083fec

Observation 4f3e1cfc-98a5-41bd-9d90-15aea3596228 · outbound

This paper cites Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.260667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.260667Z digest=sha256:b24ea9e0cd98d300078b0e95fac9a6315deb56d2386309a9ecf05183217bdcf4

Observation 345d4ff9-0653-4faf-8052-39fb18424663 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.332299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.332299Z digest=sha256:cd4e7c00304822b2da8c4950992f96901edffb4d8629dcff936ba5b599c1dcb8

Observation 66fe3c21-9130-4585-a951-609915ea4b3a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.390971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.390971Z digest=sha256:8eddfbbf0875f3413ba3b9ff2d5802dad8b6f39be209f3ab83597ab1c5ae8d36

Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.432529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.432529Z digest=sha256:6ac172d6acf364069558067ea9cbb9a810638363c1859e5b075502be638a0a78

Observation 86b558ce-f96b-4733-9f88-c59990bc0dd7 · outbound

This paper cites I see you, Batman!.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs I see you, Batman!

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.513288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.513288Z digest=sha256:4a7b597900c4e3779566b73e4ab09508c77afdf8835c9596f6acecac94016219

Observation 5c4a722a-da2d-42d3-bc83-d5940e9691a8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.586386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.586386Z digest=sha256:9b66c22b7889564a0d2e5c2e92e8650fa0af6b8b7a8842a9401ba80c46042b8e

Observation 5498de92-18fa-4680-97ac-b93c70780507 · outbound

This paper cites In: CVPR (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.671358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.671358Z digest=sha256:aef3b18b1609e5e14d953eb06b649d3f9d10f7016cee509ef5cb789885e13603

Observation b7ae5fcc-f264-4131-9fb5-268430b92ae1 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.763233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.763233Z digest=sha256:073a4f2607a7fbf2d0235984ac4abd81a7082b9a6ce792cb4dd870d775aa4c9d

Observation f0194408-f5ec-4a50-9cf6-5439a4022bcf · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.830921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.830921Z digest=sha256:69d32d3ed88e7bdce8cb44a3e3e8da80891c305f414ab64a87e418d235e1ea31

Observation 1eef3aac-44ed-4ba3-bf11-518e7d03ccfa · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.924271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.924271Z digest=sha256:c0142563e8ad40508f6500cfbc41c6b0226ede520612795e9eada698077f83a1

Observation eefbd9ad-4aaf-413d-9aaa-49f560c5227e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.022435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.022435Z digest=sha256:4aa9ef7f3aaf83debf41ebaa059d4b1622c71fca94979d1819ddd6bbcefc658b

Observation cb31d280-a509-4892-8e52-0cb174b0c121 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.079259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.079259Z digest=sha256:972268958d73a416f4da94ca3169ccf6bd9c1010817dc229d409852cb784d81c

Observation 41c84606-e81a-430b-9fc2-b1a174403ae8 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.184382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.184382Z digest=sha256:c2e751f1159c7f2581c1263cdbfd3e9a16e3590d5155bb11ab1d80a82ea16dd0

Observation 8cd5c9eb-8562-4209-a371-08b49ddf1272 · outbound

This paper cites The Llama 3 Herd of Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.276356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.276356Z digest=sha256:a92dda8363b0d8666b96149fc65e3388a485a3b8f6d6a1b9d15f458d0c0809f1

Observation b2823f80-b137-44c5-b1ac-0516804f4955 · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.378346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.378346Z digest=sha256:6eba113fe422d4dc05d579c9ec333638530aa44093c1669428ef73a8e2dc6db5

Observation e3776ed1-c023-4dcb-b04b-50b221e334e1 · outbound

This paper cites In: CVPR (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 6

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.426339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.426339Z digest=sha256:673504a56c8c95a7a582acddf748586cfacc0deff4cf2c223b453296fa3efb62

Observation 0ce230bb-953d-4a1e-8900-a855f7c56296 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.473346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.473346Z digest=sha256:2973faa51642d1fecd54ab26d7643c87bbab2a197608ce786e3379371c6c9c19

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:2adf5296dd22d2a79e15485f0e18ca5b37273b9a10b3e81f8b63f379d03dfcc9

Observation 02d3d1dd-4791-464b-b24d-bc252e7e3746 · outbound

This paper cites In: ICCV (2017) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2017) 6

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.614673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.614673Z digest=sha256:9415cf9b970732d8094007d5e26733378e39837a88833972956c8500ad43d5a8

Observation 1a15b683-b156-4d4c-9959-6d4ed0728984 · outbound

This paper cites In: ICLR (2022) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2022) 6

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.687789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.687789Z digest=sha256:b09864e78dd9b91438899b77327a936977d8bd4117e642ced447a3aeb566deff

Observation 607dafda-84ce-4743-9e65-46991d82fbb3 · outbound

This paper cites In: CVPR (2019) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2019) 2

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:31.066212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:15.781525Z digest=sha256:211a6cf0fcb427dc97365c98c55cafbe586de9335eab73fcccb47ee038e1c434

Observation e98d4342-9857-4619-b6b2-377dac9c3b79 · outbound

This paper cites Mistral 7B.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.849245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.849245Z digest=sha256:cc06b0edfbdff1ae48d6bfcbb4a3d22dfa752691ed0352fa5e6fcd85d44b6309

Observation 2c7f589a-3863-4465-a417-2f691152cc5e · outbound

This paper cites In: CVPR (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 6

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.848354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:15.963244Z digest=sha256:957ae57300d8975637278046ac5ead1014fe97fb5be03be3ee02600862ef98d1

Observation 6cdc7be4-f0c6-42f2-8e3d-3274e708f59b · outbound

This paper cites The Kinetics Human Action Video Dataset.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Kinetics Human Action Video Dataset

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.022354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.022354Z digest=sha256:3572a4b7a4dd10c57c71ab99fe73dfac6e9f240d7612e0ab01a17b9a4455053e

Observation 0959ef81-d36f-470e-82ac-a4902e94d8aa · outbound

This paper cites In: ECCV (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.087756Z digest=sha256:6087d36c1d01a830bd432d75d9edd144c256f0eaad5448a2f9d21492e1ab3c28

Observation 538bee68-24e6-4c16-ae51-3ecc023ed6df · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.116489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.116489Z digest=sha256:c4251ee52c87d9ea08ef51b000fe6d04a4a58487194bc1f12ffe489c352331a4

Observation b348445d-e214-4a77-a551-d348e98fc3e9 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.192030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.192030Z digest=sha256:f2c509a7dc8f8c6a9a2695e9f188e9544dc6815ec890c2c5ee7df9d62a289764

Observation d59a25e4-0fa2-4725-9687-d6d4acfc89ff · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.241699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.241699Z digest=sha256:fa6c194dc76baaa1f082f94e81e7ea19964f0059c2c5cc15c6212ba73d603523

Observation 4520737c-1b66-4709-bf54-2658a49e8343 · outbound

This paper cites In: ICML (2023) 1, 2, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 1, 2, 3

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.215814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.297011Z digest=sha256:1e1ec1ce00a015d54cb3f2d8a7963d21f1d7751d450d1613f65f409bbbc29665

Observation 7542d54d-f7fc-4283-b69d-2ada832e5bca · outbound

This paper cites In: ICML (2022) 4, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2022) 4, 6

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.902598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.371511Z digest=sha256:acfabc97f461c473b124eabae4454637abc775ac1c116188a1745176eab77111

Observation 7c86479a-1c60-4aa1-b402-216da35395b0 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.430652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.430652Z digest=sha256:a2ad8a31633ec799ff2806ed56b383255ec47d4d149b2f7739ea767cfdbbe70a

Observation d8289b59-9e87-4593-b8c3-81e9643c4161 · outbound

This paper cites In: CVPR (2024) 1, 2, 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2, 5, 6

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.657002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.489156Z digest=sha256:4a2ac145319bd3d60c46fbd881099ff78dfbebad8e9808262c8c2ac452ffff8f

Observation 442b2771-6305-4adf-85c8-671c77800522 · outbound

This paper cites TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:38:21.336547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.555603Z digest=sha256:5729244d769414c3a4fba6f1b48b8ee02a50cfafe052b5fe581f243b9ad329a1

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:ea95fb6fe080dcf812606ecdf07053613a06bf961460a76bd1acde580d3968dd

Observation 321d4b0d-5c05-495f-80a2-096a7362fa8d · outbound

This paper cites In: ECCV (2024) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.498383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.624560Z digest=sha256:08631e5cfecbb7f41e18dc4b35c32d5512a92e330d5e9698e596ab32ef75db66

Observation 32e494bc-9636-4bc8-9cc8-1b200e7e6b36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.675424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.675424Z digest=sha256:2c2520e6e6e061d5bc226983acc5081efa4eacbdb8c7c7742fdd9e3c745b1929

Observation 883d0dd4-0754-4599-8909-105cbb6e0d87 · outbound

This paper cites In: ECCV (2014) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2014) 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.304790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.723028Z digest=sha256:71bec7a86af9c3eb3b54f162328927c8d39f51151fa678e0da4659de95a89412

Observation e629ad1a-c779-4479-8d24-7deba78bd179 · outbound

This paper cites In: CVPR (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.062686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.765941Z digest=sha256:0ef4cad5afcba7b739371bf4a07b51c1bb7a7a6c0cfa17416df3e51bd3c1b9be

Observation 4a44d768-9562-4cc9-afbf-b67f35e7967f · outbound

This paper cites In: NeurIPS (2023) 1, 2, 3, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 1, 2, 3, 6

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.846803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.846803Z digest=sha256:ec10079c3d5a420fd8a2cd38c370028249e5b351425e720c1d2d572a7b6809df

Observation 81f964ad-84e6-4297-9683-bfdb99e9498c · outbound

This paper cites In: ECCV (2024) 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6, 9

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.889577Z digest=sha256:f1cfd819b92abe86bef61d5048611e5a5773f1670e94cc0db767943cc6536ca5

Observation 8461f797-6b66-4aca-a0ad-f46468ac7f1f · outbound

This paper cites In: NeurIPS (2020) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 3

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.539521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.916656Z digest=sha256:98a715e49f975c66c4c47895cb99e0b757a38556657afc3be1c990c0ae820863

Observation 7d6b7801-4119-411d-8ec9-75f68c7a6da9 · outbound

This paper cites In: ICML (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 2

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.318729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:16.966244Z digest=sha256:81cb6301fbbef6c7b5cb2610d87d046732fbfefaa2fcbe73e0a7652769dd0a1c

Observation 5bdbdf56-eabd-4041-96e7-58489e03b796 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.002664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.002664Z digest=sha256:942c3924882523369d358483b403beca6a67da1e145b4b250bbdc8662158d893

Observation 8e14157d-7c11-4d81-a775-cbeed9b6125d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.072193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.072193Z digest=sha256:3e9fe6a6f70da2600b716c06ded644bcaa05968bd426f408b8968c40e9403adb

Observation 67b073d9-1d6a-4165-9ec4-b8d6bcc8e4ac · outbound

This paper cites In: NeurIPS (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 6

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.175772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.130992Z digest=sha256:456f56bfb15220a049b7845565ddce008d388cbccf4c442e7438ba40c98965e1

Observation 386b79ff-461c-439c-8740-a77dfb464328 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.903975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.161209Z digest=sha256:10f3d198da167f7afcfbe0726c65378bb1474b198defc98653481c4c6fd37fb3

Observation 49a829dd-f3dd-4985-bef6-b4571903db11 · outbound

This paper cites In: NeurIPS (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 6

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.677008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.208737Z digest=sha256:34e72e9a53c35dd3daa21a436fda3622950be0b0f9635a15fd752df4cb6af9ee

Observation ee7759d0-0519-47ec-96a9-a2904642e3e4 · outbound

This paper cites In: ICCV (2015) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2015) 2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.487305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.278269Z digest=sha256:0d4c082fc4cb472b997eb9961546c39ce8b4c0c224187fd2af85f84134ade999

Observation 43c0b533-4ac5-4a04-9d05-714a5d021326 · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Streaming Long Video Understanding with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.320965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.320965Z digest=sha256:4ea86a1995a9ad1564c062623d2eaaccab0b60f5db7a77be8ee58d8584a4b2c4

Observation 7233cdfa-b97e-478c-92fb-9754b71359eb · outbound

This paper cites In: ICML (2021) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2021) 4

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.439568Z digest=sha256:bf731dd8414be0c0369fbe82497f177c5b5c382c4b9ac14fc2fa5d185d375fda

Observation b7ffccc3-a817-4768-aebb-b72c3e2920ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.535347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.535347Z digest=sha256:2a6876e83a43c79fd5de7ba413367026e3431425d0bb78971d6089c74d254c85

Observation 6a05af73-cc51-4fa8-a9fe-06b5697fa399 · outbound

This paper cites In: NeurIPS (2017) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2017) 3

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.011795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.677509Z digest=sha256:0c2eae68a6525d25444643a1a25924119efb13b9acc5f6f12698cc0a6557842f

Observation 8d2417f6-8f7b-4124-ad1b-50160aa41901 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.726972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.726972Z digest=sha256:b1a6b52e41d984c5be4a76a59ada695813a8cfc507c925ed0fd69c0f1cba60d1

Observation fbe2b09d-1871-496c-8ba4-195f40160540 · outbound

This paper cites In: NeurIPS (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 3

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.719355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.788180Z digest=sha256:f798b8d037f1ccf07a843d5c7d34ef2515855137999751948100b78d9c5b1a81

Observation d0e2a374-33f4-4108-a157-64655cfd577d · outbound

This paper cites arXiv:2409.02889 (2024) 5.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs arXiv:2409.02889 (2024) 5

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.876527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.876527Z digest=sha256:ebaadaf5caeddf6bde86a96a4b2b67713be559a679c916cab3b2cde12e4ab8f4

Observation a9e902cd-cb82-47cf-82d4-603c9f294408 · outbound

This paper cites In: ECCV (2024) 1, 2, 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2, 5, 6, 9

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.539128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:17.935120Z digest=sha256:2cebb7a717f3e615f0a2e9fd9d73af05fffa55ca45712a876db038fb7fcb1763

Observation 511b4426-8679-44f0-a05d-b86c191e34b7 · outbound

This paper cites In: NeurIPS (2021) 2, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2021) 2, 6

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.310005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.007029Z digest=sha256:f8a40d4dd0e41ecedb5c1713e831e51081cc82baef02564108b6a3ecf03eea5f

Observation 0de2c905-712d-4d08-b347-fd91b045fc66 · outbound

This paper cites In: CVPR (2021) 6 12.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2021) 6 12

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.983174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.031278Z digest=sha256:843a53a9538b0c7fee39427e74a5bbcc934641ca41f1e09c025129ef1419fe9c

Observation f7d431bf-cb92-4640-a746-f6f2056a90d8 · outbound

This paper cites Slot-VLM: SlowFast Slots for Video-Language Modeling.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Slot-VLM: SlowFast Slots for Video-Language Modeling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.120829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.120829Z digest=sha256:3b8e8e2b9f0d842bc1c6325e99b158c70b093f26d0a432ceb18bec4f16bd531f

Observation ebaffa45-cddf-4b24-9b6d-ca4799c54f16 · outbound

This paper cites In: CVPR (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2016) 2

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.628447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.244765Z digest=sha256:75e8e2f6591dd1917300bba93b427f1322fa9aebef8acbab58bf8deab8f2d12c

Observation 8da4eebb-4466-491b-b306-00581c9b7091 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.355450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.355450Z digest=sha256:77e90a4302d8826a6950f489cab454814769d8c05ef948dd00380afe4ebfac38

Observation 34460155-d1c6-48db-85f6-23598d400956 · outbound

This paper cites RAL (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs RAL (2024) 1

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.341643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.449200Z digest=sha256:1fc69e41d06ca717d2afaea6a73bd918a5a88774d2578c72870804f9a5adfd8e

Observation 7bf3ec7f-6b4a-4085-a110-40881a5299fd · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.564455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.564455Z digest=sha256:401c3779cb6b4b652b2bfe6005cbbca685f2ebe72fa712efa29bd30cd482e605

Observation e5b3b342-ae0a-44a5-b0b4-153c0da017fd · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.659223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.659223Z digest=sha256:205d498e169cd1c0b12336ef266e387c325a380225b8af350c9c767a9fc4c8ad

Observation 8e0faf98-bfa6-4278-8748-473a2e7926dc · outbound

This paper cites In: CVPR (2024) 1, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 3

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.052767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.725956Z digest=sha256:8b754556a7d97b14726c61b05064f0002bc48924c1f3d9b0940d6352cd3d0520

Observation 81106527-a8cd-4c8f-a833-43e854334668 · outbound

This paper cites In: ICLR (2020) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2020) 6

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.698975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.776700Z digest=sha256:1eadcf299e7a2f00501137e58685087a74fbb7379e5f4f94a1086cd3fd9833b8

Observation 0beed7bf-d1b4-4ce8-8fe0-58119cbd0be3 · outbound

This paper cites In: ICLR (2024) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2024) 2

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.434302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:18.945497Z digest=sha256:486d67546472da1b72ac4648bc84204d7b23658f2625704c4098db7f3ff12f23

Observation 0231cead-6e59-41bb-b365-ccd9cc2b50bf · outbound

This paper cites In: ECCV (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2016) 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.113166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:19.078605Z digest=sha256:ece67a98dc53340ff00b886e01342420fa3af7dff1736a8a3c281b5c9ba54eec

Observation bc52b499-29e2-4fb4-a2ff-eff8f96974b8 · outbound

This paper cites In: CVPR (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 2

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.824932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:19.188805Z digest=sha256:e1daa79060df2e505bc72a5c98ada700335ebe7afc681c2a7e999ede6401b9d1

Observation 3726dca5-5270-440f-946b-40c29d849afe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.313543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.313543Z digest=sha256:e9dbad5144cc8ed46d4b300083a34ad80eed63f800d77ff1d6315ced2abe0a6e

Observation 5c636545-9057-455b-86a3-c57df509d1d2 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.466169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.466169Z digest=sha256:05070fc32179e75991e8b7be3d4f1b120e9352dc16c87f3bc0313e9522340aac

Observation cf98a928-97fb-49f1-a8b2-0426a301e87f · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.548969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.548969Z digest=sha256:5f878fb73dfaf8e76293529b7fc26bac545859c95a1d1231af8787137b7a98ef

Observation bc289c9d-427d-4610-8baa-f295ed061a79 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.619551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.619551Z digest=sha256:ede2026f6b433b6db79845d93c80f4a4957e5d733fdaf8ad3e88a30966bba877

Observation e819ecc3-4df0-474b-b1df-8a031c12164b · outbound

This paper cites In: ICLR (2025) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2025) 5, 6

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.450916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:19.736433Z digest=sha256:6db3bf4d498b53925276ad9c8e221fd0294f5390a56d0b0bcb21aaafcd873588

Observation 4bc26048-147c-4687-835c-a95a3e75083a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:23.105549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:19.814174Z digest=sha256:cbdad4c39d3cc1d366c57fcd96d60b7f2c01ec974a0025ac8ce2b2f99a3bd4e5

Observation 904869a4-664e-4648-b434-788a107d910f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.891685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.891685Z digest=sha256:ab5d070d2a93473e28a03a2d04a6c6681da6701b50c74a9ea0aa7bdf48772f1e

Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.968980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.968980Z digest=sha256:df3af28b597fc1ff59168d8b5bf42d482a0774d8dc272da04870ebf54c271542

Observation 1cc34e76-338a-432b-b6d5-791c97873d5a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:22.784460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:20.040411Z digest=sha256:e0534e4741e1d3072ab0f32b37b337dd2937c63c9a03dd1292db9b3a7d7095d2

Observation 291812b2-9a35-4724-8333-e582821e6ec1 · outbound

This paper cites In: NeurIPS (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 2

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.525218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:20.130340Z digest=sha256:1d0867a9be7bb8b6af1bcf09ffa4feb870eb1037baaa31f96b951b09a21bc387

Observation 3aa8ffb8-1d66-486b-9788-233ae4cc58c3 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.186809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:20.238682Z digest=sha256:db6024522ba3c741459a8cd3111b432d46017316501ebf0301db4a9c7743d129

Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.326601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.326601Z digest=sha256:08420de81ba7f8086bf1dde19df11e86c7e4b1d0ffee6389473dc825d5b21a2b

Observation e973c4df-534a-4661-8449-cf606e65c9f5 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.409072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.409072Z digest=sha256:7efa37a5609f333268e8c3177240484a62d9eef95454053ebbe73983b1b14bf4

Observation e8ea2821-4834-40c5-8833-fc92a11bd334 · outbound

This paper cites In: AAAI (2018) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: AAAI (2018) 2

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:21.736590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:38:20.527167Z digest=sha256:2b04ff07acc31e651cca753e620dfcdb4e67b6c8d59cef50b228fbfc2f43f6b3

Observation 4144ffb2-608f-478b-b583-7ffa57eae19f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.626312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.626312Z digest=sha256:15095407010016ce874a3cac9625a3c7500c93f0b4b2f4b422127d2779db91fa

Observation f7f5f831-62fd-475b-bc0d-3474d9e72514 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.743420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.743420Z digest=sha256:4e65c2040f76ada277221af8790c0e6359d91d0698057fbea517cb0b0359ce4b

Pith citing papers

No inbound Pith citation observations are available.