Pith. sign in

Paper Citation Record · LEDGER

VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2111.12681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.12681 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:31:42.408928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:57:30.204945Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 066ab485-dc68-4c03-afbd-710371346d83 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.252979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:1eb7e0448f54aa598220f3ea11cc82d935bfb7dda34edb04d4a4018944240bb1

Observation f1b3dfd1-d7f2-4246-a5ad-9a97e7ede817 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.335861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:4d00f54a684ef1615e044b2a076f9820018d00eb81666bef67acc588d146c358

Observation d54d0cf7-6a98-4114-9471-058a461bb26f · inbound

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models cites this paper.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.171526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:dd14aab21ec3b8ec99dfd7025835d7a5baf564c65e80927280bb10f0c953ea4a

Observation ab787156-95d0-45d3-be18-1de421898834 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.538637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:e64871d2879c3625281f2c927a8ad7d003d6f85f1fbb262a2a7f906e1f463d42

Observation 62b232a8-918e-44f4-bd7a-45d429e693fb · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.623094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:8bbf3ea5d154a25313f5dd8ba54370d9c0e22e789c39583d3eca72b68240fa62

Observation cc4d57cd-db08-43ab-ae1b-e401303b6ac2 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.198100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:98c8a77f9353af37185aceeaf559cf14d1d22acd5051920270dc779400aba764

Observation 067cc1fb-e7f2-489f-a50a-ad1df38a4057 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.059332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:cac01f6a2a9bc1a33d81e8d7df5d141e7f012afb8bae413b4eb4e2b398829f34

Observation 99c967b1-335c-40f3-b66f-b18e8b68e8d1 · inbound

Large Language Models for Multi-Robot Systems: A Survey cites this paper.

Large Language Models for Multi-Robot Systems: A Survey VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.110074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:32:05.138744Z digest=sha256:31096c4a372294dd81b7a770ad404066d6323ebc80731b9e7b3fa4399d88b29d

Observation 33cddc44-a6fa-4754-a237-163b82c9746b · inbound

Character-Centered Dialogue Generation from Scene-Level Prompts cites this paper.

Character-Centered Dialogue Generation from Scene-Level Prompts VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.708168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:31:43.083678Z digest=sha256:c2dedba0acf1c0fb24c91c737fde57ea1ed9f4203f6e0b82ac9c5e79ee69610f

Observation 505417fb-eae6-4fa0-b2de-e1c0b6e9f9d2 · inbound

EgoM2P: Egocentric Multimodal Multitask Pretraining cites this paper.

EgoM2P: Egocentric Multimodal Multitask Pretraining VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.408928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.408928Z digest=sha256:16258719df461cbfa5971ea1558c5ff7602108aa4535852bc748cbdaa8f4b9b0

Observation 5093ead1-14f7-42ba-a2bc-e0a2ff8d02f9 · inbound

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos cites this paper.

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:42.222668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:42.222668Z digest=sha256:8e4953ab6c9d270d0b8bfd013ce9f6419f2f929086c8e15ee13d98a02bfaca1b

Observation c74a9922-3306-40c0-8180-d6a268075e48 · inbound

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding cites this paper.

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:00:06.142630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:00:06.142630Z digest=sha256:01bde57456f12e8f3f1f3468d8d1793b999498b052f06b16ac6c7016183d7eaa

Observation 85ada825-88a6-496d-b8f1-1a9fce6c9643 · inbound

SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling cites this paper.

SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:07:38.649205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:07:38.649205Z digest=sha256:51acb3a80fa481e443031691bb235853b941e1e284213161ae559524fe8fde4a

Observation 2a4ed516-d4be-4c5b-a2a9-b686192812be · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.192897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.192897Z digest=sha256:9b5a17673cc6816432bd1fc40545be696a95d1b74a6291249c491a2ad22a87cf

Observation 04a5bd31-04a1-4392-ad73-e48d61960717 · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.551714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:2da375c5e15e1a62717e177c9ef03af8725c92a3f489f340c3ce627ca1283a25

Observation 16f6576a-d254-457f-bec5-29e5253bfc26 · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.157510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:9636aac7e1c02557856267e3cae46695c973ca789fc4ec95e1e9d2d79bc2b471

Observation bbea9e94-ae46-4d34-92c7-29194f78bfbc · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.206555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:e86d98073da472a9b1f41a7666167c31c7746d11bdcc09ee618fcab57eea40f7