Pith. sign in

Paper Citation Record · LEDGER

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2102.12122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.12122 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:05.774092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:30.165241Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4666b128-5415-49b2-96c6-efe0748b575b · inbound

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows cites this paper.

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:27:56.955940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T19:27:56.785857Z digest=sha256:6ba8269661071a44251e78909cac4d72f7a02762df1909a8ed27108e4157aeb1

Observation 1ef51377-045d-451f-a7c3-03c82601c186 · inbound

Accuracy Improvement of Cell Image Segmentation Using Feedback Former cites this paper.

Accuracy Improvement of Cell Image Segmentation Using Feedback Former Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:53:29.777586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T21:51:08.873764Z digest=sha256:81023274b4e24782349b2a9ce82f5844b953320bac73152fca309f99c63038e4

Observation d5c75ea9-506d-457c-9b98-a2b0e7bc9fe1 · inbound

Towards Robust Multi-tab Website Fingerprinting cites this paper.

Towards Robust Multi-tab Website Fingerprinting Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:05.774092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:04:05.774092Z digest=sha256:6e9e5a49a69cb0a36e8a1567d5d02f1bf9be4402d1d3c5ae454e4fe265229a9b

Observation 18492ede-7a4a-4434-8b1b-5b920aa5ed84 · inbound

Advancing TDFN: Precise Fixation Point Generation Using Reconstruction Differences cites this paper.

Advancing TDFN: Precise Fixation Point Generation Using Reconstruction Differences Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:55.098506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:11:55.098506Z digest=sha256:19b9fac764e8b884f151827f79c5055ab5a74865a4213c0e1af0dd074a769a78

Observation b86fb065-4a9b-4abe-ade3-ece7beb83fd7 · inbound

Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection cites this paper.

Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:11:04.363586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:11:04.363586Z digest=sha256:5a5051997cfd003cd39a0cc08fa829740b2ed2f1000d686ac543a9d4d0daeb55

Observation c17b595a-7159-4bad-90c2-4d4b222fd8d6 · inbound

Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection cites this paper.

Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:45:58.468456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:30:53.849769Z digest=sha256:eb744b016e333745054e6d377dbed8b536c39fe460fd73ba30760f7190c37b77

Observation 2ed92193-d393-4e15-aae2-39eec272d2b6 · inbound

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations cites this paper.

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.823647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:04:59.177437Z digest=sha256:e8cc7b0bc76cd196906a205cd77a42b5fd539292b10244afeaaa454107a549dc

Observation 3a9f71c2-89e2-4664-8d93-2afae46b4fd6 · inbound

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living cites this paper.

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.167424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T18:22:13.147215Z digest=sha256:d86f75285019e833130d925a5f16e27d9823f29c4c8e226f93f045dde950a8ee