Pith. sign in

Paper Citation Record · LEDGER

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

As of 4 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 3 inbound Pith citation observations for arXiv:2406.05615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05615 v4

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T00:22:35.635679Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:37:41.721882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-23T07:25:28.511257Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact10
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 976b9646-8b96-49e8-925e-a3f348526e55 · outbound

This paper cites A CLIP-Hitchhiker's Guide to Long Video Retrieval.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives A CLIP-Hitchhiker's Guide to Long Video Retrieval

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T00:23:39.603943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:c39b7474f6ed3834a6a39e7a39d2f4a5e7cac218436b862a664b21f927e80ab3

Observation 123c8b5e-ffc0-47de-aec6-9e0e9d898a3a · outbound

This paper cites A Short Note about Kinetics-600.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives A Short Note about Kinetics-600

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:23:39.597715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:9afcd2296ec64323f3b863bd1c9bb5b0bdd06c61000b5494d30c0f53759b63bb

Observation 0302d9cb-47c8-42b9-b052-6c0b97b4e89c · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:23:39.659765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:8eb95edd8a36798a8b442eedbcba734ab14aed20cbdb2ea0ef51bef0d816f4a2

Observation 1c8d84ce-ab50-46cd-a837-aef0d9ea6995 · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Multimodal Pretraining for Dense Video Captioning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.584385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:ad04816521b780843d63bb3a9857352faa1b7ed561a1d4a7e603847a05978501

Observation 9058bebf-c827-460b-af80-512bfe559eb3 · outbound

This paper cites Temporal Tessellation: A Unified Approach for Video Analysis.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Temporal Tessellation: A Unified Approach for Video Analysis

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:23:39.629257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:7911803b044b8cc04cfbaf7a2a49b78a89826e06bd9209eca6a7a7b944ffc003

Observation 1fde12c8-86c0-49aa-87e6-857c917fffef · outbound

This paper cites In Proceedings of the 2018 Con- ference on Empirical Methods in Natural Language Processing, pages 1369–1379.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the 2018 Con- ference on Empirical Methods in Natural Language Processing, pages 1369–1379

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.854609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:dd04135cf082076a3cbd6d10e422a0d9b1ac585a7fedf63887408bd07e9c2577

Observation bd69b2c4-4d4e-44ed-91b8-551ffb3f7747 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives VideoChat: Chat-Centric Video Understanding

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:23:39.617174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:b779805e88573da9f2b0eb5f2e4f633f7345e8814335ce2867c1beb6518618c7

Observation f61f62ae-f6e4-414c-9216-2a8e3757797d · outbound

This paper cites Video Swin Transformer.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Video Swin Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.590980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:35af72e6f8ca61253a980246151cb40706e1a974bbe90e3728d7bbaa7cfae12d

Observation babc48cd-1c10-4681-aff7-5fdd4ed290fd · outbound

This paper cites KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.635473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:702ccd637ee83211ae0a8f5af55acee337718c9e5867a2a6f1b60870df483ac0

Observation 8ed671ab-8510-49f7-84d3-584a2422b938 · outbound

This paper cites In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18983–18992.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18983–18992

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.859286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:b8e1a1dbc60e5a82b632b98c8a159f68bb43bcddce742dcade6b6472b785ab99

Observation a6a1c026-ba9f-49ab-a19a-b13543fa9066 · outbound

This paper cites In Proceedings of the 29th ACM International Conference on Multimedia, pages 2871– 2879.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the 29th ACM International Conference on Multimedia, pages 2871– 2879

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.880575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:cc651fe0c8e51cc9b8a5b3eb6d8e2214814740747a056084da429c0d17dc77d7

Observation bc340d84-edd0-475c-a159-daf7183c25e8 · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:23:39.623096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:b7fb30a5305b0102dd048ea800582dcdfaf5af01c64f36e9ba434ccea667f6af

Observation 8141bcfc-d6f8-429f-a758-d68c80b39b17 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.648338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:797b6606104b9a91430a40c5fd5d94e97869b96841aaf475d9872989415c3b15

Observation e7f0a07b-598f-4cc0-a715-39259373732c · outbound

This paper cites Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.641818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:93b98a7deab05113a36132c03053f10c1599befc05f00280fa1bc5f463bc2c7c

Observation 3a21679c-6baf-4c27-8972-71413f860338 · outbound

This paper cites In International Conference on Machine Learning, pages 3891–3900.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In International Conference on Machine Learning, pages 3891–3900

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.884387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:2188a8783283fba56ea1b19ebc4394f6a083d2a9cd374c8eef17db49ae7a0055

Observation d0697f62-7cba-45cb-8c4a-b15b5003c3ca · outbound

This paper cites VideoGLUE: Video General Understanding Evaluation of Foundation Models.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives VideoGLUE: Video General Understanding Evaluation of Foundation Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.609764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:9536fa08f515216ece3526b76e8c317db79aa4b6c9863f6fe926ce7e949c2d76

Observation 09739cbe-1798-4c2e-95e3-13223114382e · outbound

This paper cites In Proceedings of the IEEE/CVF international conference on com- puter vision, pages 6023–6032.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives In Proceedings of the IEEE/CVF international conference on com- puter vision, pages 6023–6032

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.868250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:22e8d89c95f5a13cd46594e37453bb6719775ad01c22502cdf9f3fadca306b34

Observation a23216ec-1ea4-4366-aa10-9b5c3ca7fbb5 · outbound

This paper cites Advances in Neural Information Processing Systems, 34:23634–23651.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Advances in Neural Information Processing Systems, 34:23634–23651

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.872743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:e238596db7fb6f6e5db756a9135c003aa63a743ed8fa2e664d6943ff1ca8764a

Observation a348eece-2361-4d7f-a2e4-20785c3b50c4 · outbound

This paper cites MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.654123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:48be0aa2bbcfcc97e30c997d19e0006734370aca2af4e7a786a9d50c9f0cbc0b

Observation 16710594-e48d-4b97-b214-637010526ebf · outbound

This paper cites set up the stand ✔ 2.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives set up the stand ✔ 2

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:23:40.863410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:6e9c6f39d42f9dc951198039dadfb00b4cc87b37083189a13f7428685ae0f03b

Observation ddf9681e-4bdf-4e58-86b2-ad5f53b17c5b · outbound

This paper cites Video moment retrieval 38s 48s 60s 64s Q: People in scuba gear are swimming around.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Video moment retrieval 38s 48s 60s 64s Q: People in scuba gear are swimming around

Reference 21

Resolution
malformed identifier
raw_fallback, observed 2026-05-24T00:23:40.876556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:89cc91f41d964cb5ac6f4c9b734b8d91fc2104c2fd5814fce964330b4ff444d7

Pith citing papers

Observation 76956ff7-933b-42b2-a40a-ee85b8d72a96 · inbound

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation cites this paper.

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T07:25:28.513235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T07:23:51.435139Z digest=sha256:b5dc0dd9e60830b511731a007f08a3cb82854bb89cfc100c96b66c610479f8ee

Observation acd93843-1245-4273-9ddf-fc8d3bb2bc92 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.721882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.721882Z digest=sha256:daa8d2dae737da5a21cfc4b14c3747acab43a5c9c7ad32c2f853ced45a310db7

Observation 0e4e6dec-ef78-4322-abe8-37b149c03e8d · inbound

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos cites this paper.

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T09:11:59.440912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:11:59.440912Z digest=sha256:4e45cee388c39d79d15af17cb657ac1e9e1f5855e0a9a5444dfe399b5b1efc79