Pith. sign in

Paper Citation Record · LEDGER

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2112.01526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.01526 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:42.384108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.067076Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 994f06a4-e599-46e3-a9c0-1ec28f86da27 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:46:09.985333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:1108b1891758827cffc4fcae441f7753cd73752ef87be66ea33ed3c7ab29b0bd

Observation 30f1d953-db89-4e5d-841f-44dda78cb4fe · inbound

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition cites this paper.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.384108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.384108Z digest=sha256:bd4ccfb2b0c71cd545c8baa9857b57176a5c0b7b876947dde9154fd60e5d9bc4

Observation 52f2d07f-f3c6-40f7-a569-0c51a668d616 · inbound

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos cites this paper.

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:17:36.356493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:17:36.356493Z digest=sha256:0631df36d60017bce10fbe7f7e60bdc28ff4d4cb800b5262268175d64aec93be

Observation 39f2950a-f402-4421-b909-5b173f8fea03 · inbound

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery cites this paper.

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:52.497930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:52.497930Z digest=sha256:3810245c40a14a313575856718ce9be512f3e1d423d16da98367e36dc39626a1

Observation 47667847-793a-41e5-a300-01a84560394c · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:01.549104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:1736fed56b3d773b439eca4b7ff3ce145f8d12364ff7031a032f700e120e6b85

Observation d0144242-2be9-4ce4-be83-9dee0d1a0f39 · inbound

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition cites this paper.

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.947512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T19:23:07.258575Z digest=sha256:a38d92e4ad7f2e70e62ab535917617b79ca5baf5d1dc375dad1c12e74f7e8090

Observation e0894763-fb72-4958-9a10-3107a2386b37 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.068937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:71f47d53313b8ece0193b34d2bcf5bc09322d59ba4bd92ecc9f05f89e07e906b

Observation 6dfce47d-8805-492e-b407-0cba9d26e496 · inbound

Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective cites this paper.

Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T23:58:47.097757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T23:58:47.097757Z digest=sha256:28468f2798b03466d079a32426f314a27d2c6355bc2e5f5637b085cd92552b14

Observation 3800e9ee-fe5c-4b74-bba0-876295b59135 · inbound

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines cites this paper.

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:52.587351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:18:52.587351Z digest=sha256:b1576d16d32eddec40f4489db4a2233ceebf74227dc7d1794b75f38865c0263c