Pith. sign in

Paper Citation Record · LEDGER

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2112.01526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.01526 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:42.384108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.067076Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 994f06a4-e599-46e3-a9c0-1ec28f86da27 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:46:09.985333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:c53364602fb4052c353beb1f84d87a70cf08aafec6906b8bfa3812b97811e877

Observation 30f1d953-db89-4e5d-841f-44dda78cb4fe · inbound

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition cites this paper.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.384108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.384108Z digest=sha256:bd4ccfb2b0c71cd545c8baa9857b57176a5c0b7b876947dde9154fd60e5d9bc4

Observation 52f2d07f-f3c6-40f7-a569-0c51a668d616 · inbound

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos cites this paper.

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:17:36.356493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:17:36.356493Z digest=sha256:0631df36d60017bce10fbe7f7e60bdc28ff4d4cb800b5262268175d64aec93be

Observation 39f2950a-f402-4421-b909-5b173f8fea03 · inbound

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery cites this paper.

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:52.497930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:52.497930Z digest=sha256:3810245c40a14a313575856718ce9be512f3e1d423d16da98367e36dc39626a1

Observation 47667847-793a-41e5-a300-01a84560394c · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:01.549104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:cd4e156f23697e1d00ec2367f8b238145286a26372941ed7e19478c1315770b9

Observation d0144242-2be9-4ce4-be83-9dee0d1a0f39 · inbound

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition cites this paper.

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.947512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T19:23:07.258575Z digest=sha256:9cf78be5acf763de1fd3033e532081f7643e0cb8fa79e87a406566a1f06448d1

Observation e0894763-fb72-4958-9a10-3107a2386b37 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.068937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:afae7d8b2e1d136845ec9e4e7ada8b7c7ca31dc0cca2e571e19edf5cfcbc176e

Observation 6dfce47d-8805-492e-b407-0cba9d26e496 · inbound

Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective cites this paper.

Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T23:58:47.097757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T23:58:47.097757Z digest=sha256:28468f2798b03466d079a32426f314a27d2c6355bc2e5f5637b085cd92552b14

Observation 3800e9ee-fe5c-4b74-bba0-876295b59135 · inbound

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines cites this paper.

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:52.587351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:18:52.587351Z digest=sha256:b1576d16d32eddec40f4489db4a2233ceebf74227dc7d1794b75f38865c0263c