Pith. sign in

Paper Citation Record · LEDGER

UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2211.09552.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.09552 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:52.113307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:08:49.809225Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df6acf2d-ce08-4040-9031-ec1f5337f94f · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.339682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:5ba34bd735b2e2f6bd4fc25d49c7170e40deb29e3e361dcba00fa90fa647fa49

Observation a5b874da-8fef-43d5-8329-ae9d9f30a5a9 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.561864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:553bf2138b7d26f77bdd56d4a4bddc202ca46e7d7aa7d507b3de65c30726beec

Observation 82c9ad0b-52d7-4917-8e8f-d6e77fa45beb · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.561008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:08fd0c0001fc4d8766f13941059caf5f3f06891b005fa4ff3cd59189b897b65c

Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.091018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0596b01af6c26e81a28a8265752a1fc147cfad9e2f82204263e71932fe2ded4d

Observation 215434de-9888-454e-8479-a6ce6fd57807 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.735657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:554ad83eef1a109b39b1d8e96edd2c1fd5937c72d4e3f8f1e6d13a47c0bad889

Observation 79847b2b-b338-4975-b4fe-70ff0311c6c4 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.360703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:399364e1d6736503c854bbafb5c0c3e63fc9a0482ca675c7ee6a293f9e07b61e

Observation f966cc04-5513-4b7b-ab90-7dd8676b441e · inbound

CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge cites this paper.

CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:32.639342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:32.639342Z digest=sha256:d84795065385984fab7474122d7c405c603058fe620a869b2d269036066174c8

Observation 64e43334-38e5-4c57-8386-dc24cde323e4 · inbound

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery cites this paper.

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:59.602264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:59.602264Z digest=sha256:a23a1046b33decd98a49a5eb8c9cf2a7ba11ecc9f0b736c35594b18dff3cd4c9

Observation 0df613e5-a734-41fe-955c-6a86b907c704 · inbound

FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection cites this paper.

FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:52.113307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:52.113307Z digest=sha256:913d933a8f82e49d6e22f2e40caf454aed01b6e15e9dd624196ae322c5c598de

Observation 2a772835-48f9-4514-9aab-aaccfa86327c · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:36.711194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:36.711194Z digest=sha256:31c518efaf1919128bfba197058e22b97b720fa57dbc5f6b5372d83e3f997f46

Observation 55fee3db-b3f0-49d2-a136-629b6b46e647 · inbound

Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound cites this paper.

Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:13.171863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:13.171863Z digest=sha256:ee36e2121046c996b6ea8974e95b6c0793c02f592ca3b903ac973431b0ca86a5

Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.646948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.646948Z digest=sha256:daa81614e46edf46727e3a65dd3d3b8a6bc332a696cbd4dabf63ba21b78de573

Observation 93f59791-a2f0-4a2b-8256-9388ba18df40 · inbound

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines cites this paper.

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:02.125281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:56:32.221330Z digest=sha256:45e33646ccc42f96125c6991297b8580b3fd5dd3abf72b9583db9d6103f02928

Observation b8ab7530-db06-4175-9376-079e41b6119b · inbound

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos cites this paper.

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.439765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:00:21.173667Z digest=sha256:0371372f1222bba92d95ec0c524e24d379fd7cab8535c11f6d60c5b8525d96c8

Observation 0db39fe4-456c-482c-a3ac-7d217f11adb4 · inbound

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition cites this paper.

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.693323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:27:32.335117Z digest=sha256:80fb547e03cd92eb01e60c2043b095fb6141f40354f156fbccb018cc93a7eed7

Observation 06a5d465-ba89-40c1-9c3e-bc18cd7bca3f · inbound

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection cites this paper.

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:52:13.618993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T07:51:04.580617Z digest=sha256:b448f56a97c01bce0bfd269eaecd447488ea1921f9a7d5016b72ebf6de4901b2

Observation ca830749-d822-408a-975e-5b5946051a1d · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.074915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:741fc97a7e81fb8e850619bd649e52b6e905f55aa347e61ee170f3bcf129cf6a

Observation 6ceaa5fe-b749-42d1-a426-f45c7a9b4b9d · inbound

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos cites this paper.

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:08:49.810953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T02:11:37.740576Z digest=sha256:79b6dac6d6fee9069dfa6b6363da6ec0f4d538b9527d15df6b7251e6b755b745

Observation 09b3d675-5426-4733-88da-c9b72119243f · inbound

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection cites this paper.

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:16.561201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:16.561201Z digest=sha256:21f1edce307d16c2e3151521be6fcefc44b8e99d67f8be4122035fbe382dc742

Observation 1661bf28-b79b-4f43-941b-0ed9a937da40 · inbound

PhiZero: A World Model Built Around Physical Language cites this paper.

PhiZero: A World Model Built Around Physical Language UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T01:50:30.642683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:50:30.642683Z digest=sha256:5b487d48a5bba1e359889340c1c18d1a4e6280738ab8fc9114c72c2ec9fe2753

Observation 623012d6-b50d-49de-be5d-6ead28915273 · inbound

Decoding Children's Gait Behavior cites this paper.

Decoding Children's Gait Behavior UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T00:41:55.526214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:41:55.526214Z digest=sha256:fc6df8753521f58507da2c9b87be12bb26ca58078bb8108ab75d228838ce44c8