Pith. sign in

Paper Citation Record · LEDGER

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

As of 16 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2507.10302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10302 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:38:20.743420Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c201e185-eee5-49ee-a9b0-c2ca80d845f4 · outbound

This paper cites GPT-4 Technical Report.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.744309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.744309Z digest=sha256:fa88540ed8bd8860bfac7f92b6370554ebfd52f74099188cda2033b8031bba43

Observation 43efa36a-bb39-4e18-8419-7712a080ed51 · outbound

This paper cites In: NeurIPS (2022) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2022) 1, 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.800454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.800454Z digest=sha256:843b7799807202daf4cadd652a6db9b3a56e321df8c50d620464c897cd96c44c

Observation 724bd4d6-c78a-45d0-a194-cfe9edf3479b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.887420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.887420Z digest=sha256:03591920d1e7d58912eff420979eb8d33e899fc36189e7f777d100443a7579f5

Observation 9f9a2c07-49b6-4862-a614-6beafbde6294 · outbound

This paper cites In: ICCV (2021) 2, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2021) 2, 6, 9

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.993701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.993701Z digest=sha256:2078980ef4459b832acda9bd1e9ec1c232c58b8b7e39fee357c891d06ccb0696

Observation 607be67e-a3c6-44a2-8956-93aab90f7d4d · outbound

This paper cites In: NeurIPS (2020) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.035353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.035353Z digest=sha256:ee676c11d78d85a17559f0c5971265578848727300c398e5c23b7411110c5f62

Observation 03173b9b-2712-4504-8da1-8ca05b4a22fe · outbound

This paper cites In: ECCV (2020) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2020) 4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.105134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.105134Z digest=sha256:bb7b32ffbdb521e65dead53b66581a7125044d3f67d82b18d89443db5fdfe79e

Observation 7836d8d4-84b7-47c8-a079-162823de13a0 · outbound

This paper cites In: CVPR (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.175728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.175728Z digest=sha256:403ba18fab9a8e1c668bf44646fa644220983641908a3b0dab8fcdd01c9896f7

Observation 4f3e1cfc-98a5-41bd-9d90-15aea3596228 · outbound

This paper cites Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.260667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.260667Z digest=sha256:1917d12d1b732af2f5d08f69159f36923c3e783baaa0eef542609eac45d3223c

Observation 345d4ff9-0653-4faf-8052-39fb18424663 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.332299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.332299Z digest=sha256:44f7e9a7414d1d22aac34ddaccfa4b74d5d82ebbee14b4aee57eada373194573

Observation 66fe3c21-9130-4585-a951-609915ea4b3a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.390971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.390971Z digest=sha256:8219999d552ff0f8104ec7d13f2a66d0db1545a8dea0e3da45b29612e9138ecb

Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.432529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.432529Z digest=sha256:d862f52dd3fc0c10f8a52c00fb560a11b853d211becabccb0f745e8a5cdbf1f2

Observation 86b558ce-f96b-4733-9f88-c59990bc0dd7 · outbound

This paper cites I see you, Batman!.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs I see you, Batman!

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.513288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.513288Z digest=sha256:94b0a16d0471021f0e060ec4e119f1dc48815d5afbdb0048c9f20e9b67e7624c

Observation 5c4a722a-da2d-42d3-bc83-d5940e9691a8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.586386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.586386Z digest=sha256:a4debd400d510495f02dc846b903572fab3edd286498a4418891e5b30d08fd67

Observation 5498de92-18fa-4680-97ac-b93c70780507 · outbound

This paper cites In: CVPR (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.671358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.671358Z digest=sha256:a2ba8d45514bf57f4c5c92e1b8b5802ffff57df333dd36697b436ae2070a3dc5

Observation b7ae5fcc-f264-4131-9fb5-268430b92ae1 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.763233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.763233Z digest=sha256:c035f1e3df55ba871abc2e0f479c623b2508ce0f0526581ee25e70532f2d089e

Observation f0194408-f5ec-4a50-9cf6-5439a4022bcf · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.830921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.830921Z digest=sha256:0463638b115f74030727204bc9f2f9237f4940f5968f09e6cb25a21f31f6927c

Observation 1eef3aac-44ed-4ba3-bf11-518e7d03ccfa · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.924271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.924271Z digest=sha256:623bfd8cecaa620fcef5b4da1adfcc9bff5aaf1c175ca46ec787282f3d21c147

Observation eefbd9ad-4aaf-413d-9aaa-49f560c5227e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.022435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.022435Z digest=sha256:809b490a49e65d33dfdffe3d29d2fd6e00f176b9c26cb7c8d5728a943eee76c3

Observation cb31d280-a509-4892-8e52-0cb174b0c121 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.079259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.079259Z digest=sha256:dcea7e06937fdcc1fbd5cded63cf16c2e2dad41b0c12b5b0bbb4aa8441671fa5

Observation 41c84606-e81a-430b-9fc2-b1a174403ae8 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.184382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.184382Z digest=sha256:6e287b829a0c7a08b92d66bf441a7e449e23667fd46e8b89d989c62de6c4b22c

Observation 8cd5c9eb-8562-4209-a371-08b49ddf1272 · outbound

This paper cites The Llama 3 Herd of Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.276356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.276356Z digest=sha256:8a791ea26d3a7909050eca0bbb3bc48599c0ee40fee2f2c513667d48503d6387

Observation b2823f80-b137-44c5-b1ac-0516804f4955 · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.378346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.378346Z digest=sha256:f9aa1579de0b1ad381061051efaee0e8a836aaa7cf1c6e7e6aa60179df8ff10c

Observation e3776ed1-c023-4dcb-b04b-50b221e334e1 · outbound

This paper cites In: CVPR (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 6

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.426339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.426339Z digest=sha256:52f2ccab9e5ef593e0b810917d5ac463304a048729ea3cf137a02db8847f7fd3

Observation 0ce230bb-953d-4a1e-8900-a855f7c56296 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.473346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.473346Z digest=sha256:cf8a6c0e12e8186406e7f74bdf227d2fa48afd07748e940d352177ff72250c0f

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:913e6e4c51dc400ae62964c08381cfa7a203090a033fb493e1e852c3f4fb7659

Observation 02d3d1dd-4791-464b-b24d-bc252e7e3746 · outbound

This paper cites In: ICCV (2017) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2017) 6

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.614673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.614673Z digest=sha256:e0ee4defabe550c3b3522be1057af962d9668744c4936ea96f30da4e0d312387

Observation 1a15b683-b156-4d4c-9959-6d4ed0728984 · outbound

This paper cites In: ICLR (2022) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2022) 6

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.687789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.687789Z digest=sha256:407695d65e7ad8ac1779d103b3cdff5c609c84eb53006700b76a24c6b7a0e4a6

Observation 607dafda-84ce-4743-9e65-46991d82fbb3 · outbound

This paper cites In: CVPR (2019) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2019) 2

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:31.066212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:15.781525Z digest=sha256:14bc74f4774a6f35ff850694ff66faf1d83ae0c947732f4558756d0d4a21317e

Observation e98d4342-9857-4619-b6b2-377dac9c3b79 · outbound

This paper cites Mistral 7B.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.849245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.849245Z digest=sha256:c1b0b5aa55c3aebe6a40fbc00118cc2d96e17ca92f00eda6a68992a0d0116f59

Observation 2c7f589a-3863-4465-a417-2f691152cc5e · outbound

This paper cites In: CVPR (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 6

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.848354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:15.963244Z digest=sha256:92eff0b7159def87b1c83081e50aa9a5d8aeeea60c29a5af604ca38a835e30ac

Observation 6cdc7be4-f0c6-42f2-8e3d-3274e708f59b · outbound

This paper cites The Kinetics Human Action Video Dataset.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Kinetics Human Action Video Dataset

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.022354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.022354Z digest=sha256:03aaa85f4beacadbc543f2ff6dd8caa1017310e35d23311b3423ab0bb930a7c2

Observation 0959ef81-d36f-470e-82ac-a4902e94d8aa · outbound

This paper cites In: ECCV (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.087756Z digest=sha256:8afb0bb258952cd7b77b785c117b0d579e9ce0f8768b02e812599efaa972e5b7

Observation 538bee68-24e6-4c16-ae51-3ecc023ed6df · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.116489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.116489Z digest=sha256:9afaf1a499fccafc5d591e72073772023b24647c708ca488e31965fadacba9a5

Observation b348445d-e214-4a77-a551-d348e98fc3e9 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.192030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.192030Z digest=sha256:722356bc56e607c2aa56e6304804af6bc47d4c64aa280e99d5a55b27f7b45e80

Observation d59a25e4-0fa2-4725-9687-d6d4acfc89ff · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.241699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.241699Z digest=sha256:723025c99b527dadc657948bb82b2f6a3585cf8d880fdc6ad14df681f939a4d9

Observation 4520737c-1b66-4709-bf54-2658a49e8343 · outbound

This paper cites In: ICML (2023) 1, 2, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 1, 2, 3

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.215814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.297011Z digest=sha256:3f65d691a32da7236021d7d9d438b9060963a62c7866915d8d7cb2070f0c5b00

Observation 7542d54d-f7fc-4283-b69d-2ada832e5bca · outbound

This paper cites In: ICML (2022) 4, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2022) 4, 6

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.902598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.371511Z digest=sha256:41909d924caf65b618ea79d4cc1f8d239a26038ecea4b2e96f1fc583679d9a72

Observation 7c86479a-1c60-4aa1-b402-216da35395b0 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.430652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.430652Z digest=sha256:c0387ca375adb77f837e29184509a9ed962ccb1b4773c01ef0663a4d1264dd23

Observation d8289b59-9e87-4593-b8c3-81e9643c4161 · outbound

This paper cites In: CVPR (2024) 1, 2, 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2, 5, 6

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.657002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.489156Z digest=sha256:0c4ef5e5cd0b37641a911fde01c4e50eb6c0ce2fe39f7b15d9738f803bc75a61

Observation 442b2771-6305-4adf-85c8-671c77800522 · outbound

This paper cites TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:38:21.336547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.555603Z digest=sha256:343e072a9fdbccfbe69423e9907cd38924b284bd75e0a6673acff8d3ed2ad19a

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:1b51f25605832d2cbd2c1060fc83111c01525308c89312eb12804780fc08d8c5

Observation 321d4b0d-5c05-495f-80a2-096a7362fa8d · outbound

This paper cites In: ECCV (2024) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.498383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.624560Z digest=sha256:fe0870822b04562b77711417e1df3de360bce3f9af2bf82e325a458392ab50f4

Observation 32e494bc-9636-4bc8-9cc8-1b200e7e6b36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.675424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.675424Z digest=sha256:c13bdb1d0e706fbb18bb419f951e5d409702c29257a1de43c581bc8d0da6f609

Observation 883d0dd4-0754-4599-8909-105cbb6e0d87 · outbound

This paper cites In: ECCV (2014) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2014) 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.304790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.723028Z digest=sha256:34e7335c6f73fb83e19a644c508568857fcec56eff4fe97ec79ee8b7e80e4f93

Observation e629ad1a-c779-4479-8d24-7deba78bd179 · outbound

This paper cites In: CVPR (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.062686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.765941Z digest=sha256:6f321e2e54a61785eb2ec2c66cd17dd456505622851220a78dfc1438a3b4911d

Observation 4a44d768-9562-4cc9-afbf-b67f35e7967f · outbound

This paper cites In: NeurIPS (2023) 1, 2, 3, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 1, 2, 3, 6

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.846803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.846803Z digest=sha256:5d933b14b25eeaecacd0d24f9d3b2efaf925c03c612ff9890e5eca2208ab9b93

Observation 81f964ad-84e6-4297-9683-bfdb99e9498c · outbound

This paper cites In: ECCV (2024) 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6, 9

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.889577Z digest=sha256:278ed43d238ae53e897655aaf83a2095c6d9c089386f711919fdd2aba690aa0b

Observation 8461f797-6b66-4aca-a0ad-f46468ac7f1f · outbound

This paper cites In: NeurIPS (2020) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 3

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.539521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.916656Z digest=sha256:6918649f3c89f635eda42e792ecdbce28d000918ab5519780849084f86135df7

Observation 7d6b7801-4119-411d-8ec9-75f68c7a6da9 · outbound

This paper cites In: ICML (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 2

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.318729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:16.966244Z digest=sha256:ea386ffa7ff44cbf97795538a215edd9b7e1707ac8b67be553f15b327d271fff

Observation 5bdbdf56-eabd-4041-96e7-58489e03b796 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.002664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.002664Z digest=sha256:a04d71b8f025536c8f66fa6c6ec93cfd3b2648f101e5a9adbf56896da2fc9145

Observation 8e14157d-7c11-4d81-a775-cbeed9b6125d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.072193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.072193Z digest=sha256:a70d6b1f11cc6ce548ea83dfac0e8fd1df386e0f0c8a5ef9ab87689ab76bffa7

Observation 67b073d9-1d6a-4165-9ec4-b8d6bcc8e4ac · outbound

This paper cites In: NeurIPS (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 6

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.175772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.130992Z digest=sha256:115b3d3203dc15dc2848958cdfe4f86f55f794812a8c2e311e7bc73a887c1005

Observation 386b79ff-461c-439c-8740-a77dfb464328 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.903975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.161209Z digest=sha256:82798d1e956a1e41beefd3ed8d88d7cf7ee482901ce68150bf92cf6cc66b0ea1

Observation 49a829dd-f3dd-4985-bef6-b4571903db11 · outbound

This paper cites In: NeurIPS (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 6

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.677008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.208737Z digest=sha256:fe11fd642439349b778516da875574d150fd7c02d91ba6950e39dc070ddcb621

Observation ee7759d0-0519-47ec-96a9-a2904642e3e4 · outbound

This paper cites In: ICCV (2015) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2015) 2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.487305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.278269Z digest=sha256:6a1b310cf280b427ee292ae8fe1423aea480ddbbaa534b44712c862f1b9e57c2

Observation 43c0b533-4ac5-4a04-9d05-714a5d021326 · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Streaming Long Video Understanding with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.320965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.320965Z digest=sha256:9e0d87ae10959feda56a21872c836f814e8ffcb43f14e55387b5644f8468b6f5

Observation 7233cdfa-b97e-478c-92fb-9754b71359eb · outbound

This paper cites In: ICML (2021) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2021) 4

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.439568Z digest=sha256:db8cfcfef4295036128cd62ddde688dc53b0d0c7c65cd48b34727f6ec5969798

Observation b7ffccc3-a817-4768-aebb-b72c3e2920ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.535347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.535347Z digest=sha256:016673633cd087733dd3d489c14d40135d6b6f8130b1c40035806ba76a6c6486

Observation 6a05af73-cc51-4fa8-a9fe-06b5697fa399 · outbound

This paper cites In: NeurIPS (2017) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2017) 3

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.011795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.677509Z digest=sha256:fc248c0ff15fb168063415f693e98265b8fe3a2363337ecd79fd7381a5c61241

Observation 8d2417f6-8f7b-4124-ad1b-50160aa41901 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.726972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.726972Z digest=sha256:40aa7ed6c511901fef4cfefd9ad7a4ba364732c8c9e764212512cffae11ba0d9

Observation fbe2b09d-1871-496c-8ba4-195f40160540 · outbound

This paper cites In: NeurIPS (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 3

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.719355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.788180Z digest=sha256:f87fdba3746f224e8113d061a6a4e6939cae087225ab2f11c5c441ccd81835ef

Observation d0e2a374-33f4-4108-a157-64655cfd577d · outbound

This paper cites arXiv:2409.02889 (2024) 5.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs arXiv:2409.02889 (2024) 5

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.876527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.876527Z digest=sha256:c1c1a1dd8074c6cc912e931fbf3915569fbe005c4e517f42fcb6e1d2a641e656

Observation a9e902cd-cb82-47cf-82d4-603c9f294408 · outbound

This paper cites In: ECCV (2024) 1, 2, 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2, 5, 6, 9

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.539128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:17.935120Z digest=sha256:478e8a2a4e3a73dd2cc0d4968caa359141d1cff6d6b262585d799cf1fa386614

Observation 511b4426-8679-44f0-a05d-b86c191e34b7 · outbound

This paper cites In: NeurIPS (2021) 2, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2021) 2, 6

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.310005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.007029Z digest=sha256:fd68d887f3db4a2e3cb18b2efaa3ebcb3df2fef06fb00c9e891812f084b56713

Observation 0de2c905-712d-4d08-b347-fd91b045fc66 · outbound

This paper cites In: CVPR (2021) 6 12.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2021) 6 12

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.983174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.031278Z digest=sha256:0c1968f7bb075d71592e2d6fb4a4628a8e714abcb91ee254ddd492bce873f4ff

Observation f7d431bf-cb92-4640-a746-f6f2056a90d8 · outbound

This paper cites Slot-VLM: SlowFast Slots for Video-Language Modeling.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Slot-VLM: SlowFast Slots for Video-Language Modeling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.120829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.120829Z digest=sha256:299ad6c2999fabfff1ff6b5af9870781ffa2bff04c3de316aaabf035faefb081

Observation ebaffa45-cddf-4b24-9b6d-ca4799c54f16 · outbound

This paper cites In: CVPR (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2016) 2

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.628447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.244765Z digest=sha256:6b12b24ff909c4079eec253f9d9e9a849e667c9601be968041abb457b230fe14

Observation 8da4eebb-4466-491b-b306-00581c9b7091 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.355450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.355450Z digest=sha256:0200fe885ee4fe80eca836cafb857d512aa6a11f307ea1ce8db1239ae1544c38

Observation 34460155-d1c6-48db-85f6-23598d400956 · outbound

This paper cites RAL (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs RAL (2024) 1

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.341643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.449200Z digest=sha256:c328fc37fa90fd955aecc8d4aa3edfb83bb88f2f82a3ecccef50e008897fc3b4

Observation 7bf3ec7f-6b4a-4085-a110-40881a5299fd · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.564455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.564455Z digest=sha256:4c4ec2d59eb36370921b7ba4130ac44b0afd25e6ea923ef67d1145bd178949b1

Observation e5b3b342-ae0a-44a5-b0b4-153c0da017fd · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.659223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.659223Z digest=sha256:13401a05e32b4013e9f55d4d50015d7ec7d163c8a483b327dbcb7c436c5f4d63

Observation 8e0faf98-bfa6-4278-8748-473a2e7926dc · outbound

This paper cites In: CVPR (2024) 1, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 3

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.052767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.725956Z digest=sha256:3381ace0f20679a2c96494ccb0e93e14d6b48f9045a899d3402da793ac6cbc31

Observation 81106527-a8cd-4c8f-a833-43e854334668 · outbound

This paper cites In: ICLR (2020) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2020) 6

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.698975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.776700Z digest=sha256:4047a5c8cb0a1ca6cb2aded0d87078686204df97b8ceffbe9ce7138f4545e5c7

Observation 0beed7bf-d1b4-4ce8-8fe0-58119cbd0be3 · outbound

This paper cites In: ICLR (2024) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2024) 2

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.434302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:18.945497Z digest=sha256:d87db8ef5cea294367d3be8b19be0d9fcaedc573ea0c091f09fc4248c2fcde14

Observation 0231cead-6e59-41bb-b365-ccd9cc2b50bf · outbound

This paper cites In: ECCV (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2016) 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.113166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:19.078605Z digest=sha256:ff54d540c11b6b1736b07876d26557b26d59a68bcf0485fca4e84b8e0ef5b266

Observation bc52b499-29e2-4fb4-a2ff-eff8f96974b8 · outbound

This paper cites In: CVPR (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 2

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.824932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:19.188805Z digest=sha256:245111ee88bab18b96231fd503481de5132baad2c71805309b0c54f6d6286e39

Observation 3726dca5-5270-440f-946b-40c29d849afe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.313543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.313543Z digest=sha256:67cc5672b86e814abb845e51c2ba13db29a954bebb6158d22414bb8df7843a48

Observation 5c636545-9057-455b-86a3-c57df509d1d2 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.466169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.466169Z digest=sha256:7594a88cd4f4862ffc6564161747bc0de4c1648f5324da48b9ad10696ecb7587

Observation cf98a928-97fb-49f1-a8b2-0426a301e87f · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.548969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.548969Z digest=sha256:7956089095c43b892120da3726a8cc23cae77d10d8b6772c5859aa3a95bb4a5c

Observation bc289c9d-427d-4610-8baa-f295ed061a79 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.619551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.619551Z digest=sha256:e172b5bc394f6fcc97a04d5c27f210a481846426cf809573204e9d7f7adcb310

Observation e819ecc3-4df0-474b-b1df-8a031c12164b · outbound

This paper cites In: ICLR (2025) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2025) 5, 6

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.450916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:19.736433Z digest=sha256:b33b9dd926929e5aaba1a15a2d32d862c8c9c6fd01643936df825176197164a0

Observation 4bc26048-147c-4687-835c-a95a3e75083a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:23.105549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:19.814174Z digest=sha256:4acb8f1752dabd7ef818d070eabad61a74d26351f7a5893f2f422c293659e83b

Observation 904869a4-664e-4648-b434-788a107d910f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.891685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.891685Z digest=sha256:5f474f1f38a1808f7fa04fcc4e8dcf0cb8f3379805da0a57eafecddbef7a2a7b

Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.968980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.968980Z digest=sha256:932f57f367bdc44c2136c2b4a6c0411cc3c13e00f66e79acaef2dc54a3823c12

Observation 1cc34e76-338a-432b-b6d5-791c97873d5a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:22.784460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:20.040411Z digest=sha256:91ae0e067b43a6cee7ae0b7f03371b3b5c0cf74cb9a384453499f8f147a802b3

Observation 291812b2-9a35-4724-8333-e582821e6ec1 · outbound

This paper cites In: NeurIPS (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 2

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.525218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:20.130340Z digest=sha256:90139d1e5d86dd1678e17e7320f44e7c43e7b2483cb2093281a3db0e9141ac39

Observation 3aa8ffb8-1d66-486b-9788-233ae4cc58c3 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.186809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:20.238682Z digest=sha256:75d200e9ccd9132e0be3bbd08bdbac79580f65c4518225d8a96b48299c50a0dd

Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.326601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.326601Z digest=sha256:e09a80d800a43ed9ec8a4718486d011523a859c0f24b8e4f95fad3be48f7fb97

Observation e973c4df-534a-4661-8449-cf606e65c9f5 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.409072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.409072Z digest=sha256:f150b9c8583437e9184f6ee3c74c24fd1018e504cea4bad6279bfc6a8e3183ab

Observation e8ea2821-4834-40c5-8833-fc92a11bd334 · outbound

This paper cites In: AAAI (2018) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: AAAI (2018) 2

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:21.736590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T17:38:20.527167Z digest=sha256:1fa99aab55eb4a2b936558553e33df010d82795ea7aea335b6bdeeb662e1c183

Observation 4144ffb2-608f-478b-b583-7ffa57eae19f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.626312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.626312Z digest=sha256:8654e0214220ab82f84f166522e31efb6d2c7ee12779d3aaf92ce0ade0637854

Observation f7f5f831-62fd-475b-bc0d-3474d9e72514 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.743420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.743420Z digest=sha256:6f37ead56ef961d838435bb7f2287aa3bacacd86338ff8748bd70086bc1c5d9f

Pith citing papers

No inbound Pith citation observations are available.