Pith. sign in

Paper Citation Record · LEDGER

Conditional Positional Encodings for Vision Transformers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2102.10882.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.10882 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:24:18.962558Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:56.509722Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2212ddb7-89fa-46a9-899e-e24b0331cab8 · inbound

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows cites this paper.

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows Conditional Positional Encodings for Vision Transformers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:27:56.944252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:27:56.785857Z digest=sha256:3f63be8c62258c8653ed143851e6a6296e72185f32f7bc77f17d042098671d4c

Observation 27655d9f-a601-4848-8062-06195cc422f0 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention Conditional Positional Encodings for Vision Transformers

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:07:42.452808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:fd24d05aa729f65bc0f5f61664d39a9b640dd5420084a82e5e623c1fb3af1f9a

Observation 20bf5ead-4004-41af-b2ca-25e534e3039a · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models Conditional Positional Encodings for Vision Transformers

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.836536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:29ee37b397ec0af23808a3008cbcd873edfab9e2ad8192a80b616a183ee6974a

Observation db8502d9-d9ca-4485-8058-95c3c5ec85a8 · inbound

Generalized SAM: Efficient Fine-Tuning of SAM for Variable Input Image Sizes cites this paper.

Generalized SAM: Efficient Fine-Tuning of SAM for Variable Input Image Sizes Conditional Positional Encodings for Vision Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:28:27.563406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T21:26:32.186947Z digest=sha256:8d18fad55298db75ad0d061dc7479bffd8087cfa1ac8b896f429358ddb99d24d

Observation f33ecc9f-ff1d-466d-a9ce-f4532aee6c24 · inbound

Compress image to patches for Vision Transformer cites this paper.

Compress image to patches for Vision Transformer Conditional Positional Encodings for Vision Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T19:24:18.962558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:24:18.962558Z digest=sha256:ddcba94080703d3646dd4da3a52403adcc82b18eb9609ec40f8b3d6236be2029

Observation 2513cc64-b3f4-40ec-a2c8-d7071577fe94 · inbound

LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers cites this paper.

LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers Conditional Positional Encodings for Vision Transformers

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:21:55.991100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T18:18:56.091429Z digest=sha256:a392e24fc622ac118ee9ad59a3164676da95981a8e94c37cc3f8f73f31311753

Observation 2849dda0-aaaa-4426-8844-c9eb5ed6b04a · inbound

Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning cites this paper.

Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning Conditional Positional Encodings for Vision Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:58.611323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:13:58.611323Z digest=sha256:a65bc908aba9c62d17eab51f5ba439cc897410977f79026737a3ee3803e74c8f

Observation dcc5ce5f-76eb-4247-8dea-4353601196d6 · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey Conditional Positional Encodings for Vision Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.606136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.606136Z digest=sha256:74e82ca4cacff76912c5d52c625e861cab947ffdad8dc00602eb68c57ab23a85

Observation d3a3f5fd-6340-46d1-af72-99fcbb171e40 · inbound

Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis cites this paper.

Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis Conditional Positional Encodings for Vision Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:09.704815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:09.704815Z digest=sha256:5a4598d9a288a67528d0563edfad0d317b9fe879eef42f6ac72ab6610d06f611

Observation 8987f68c-8c40-440b-bc33-52f6ac536e7a · inbound

ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference cites this paper.

ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference Conditional Positional Encodings for Vision Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:59.017983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:59.017983Z digest=sha256:fa9663b71a9b12ddce2b1025c5637ec7c47cca20cea60bebf91abcd6d474bd3a

Observation 38c2ff56-6dd6-4526-84fb-d3c22f7b6628 · inbound

CoPE: A Lightweight Complex Positional Encoding cites this paper.

CoPE: A Lightweight Complex Positional Encoding Conditional Positional Encodings for Vision Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:11.776775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:11.776775Z digest=sha256:c61bfca738965735ebd58531f380b3de2f0806dc47b5c069f84c2fb3a5f23883

Observation b24d652f-9c45-4d74-a47d-4d3de200c058 · inbound

SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification cites this paper.

SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification Conditional Positional Encodings for Vision Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:59.387017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:59.387017Z digest=sha256:a7c5be949d874564b69638906f7932c8a7329cfa38e0d23532c6aebbc3ace1b5

Observation ba805d2e-9a36-4721-9bb3-b04ec0ed478a · inbound

Hierarchical Mesh Transformers with Topology-Guided Pretraining for Morphometric Analysis of Brain Structures cites this paper.

Hierarchical Mesh Transformers with Topology-Guided Pretraining for Morphometric Analysis of Brain Structures Conditional Positional Encodings for Vision Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:54.006092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:08:21.736766Z digest=sha256:b7b21c9296ab5d9b2ac729e87ea6a64fd407ae5865471a6e8ab143830f7ebadd

Observation 8183770e-7660-4844-a811-e3bc34453479 · inbound

Masked-Token Prediction for Anomaly Detection at the Large Hadron Collider cites this paper.

Masked-Token Prediction for Anomaly Detection at the Large Hadron Collider Conditional Positional Encodings for Vision Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:11:04.330841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T23:24:59.497990Z digest=sha256:115efd7824b3088ab1b9918e012b719b67534d1bcfd952c2417bd2a9f36f96f2

Observation 7bee24b5-3fa0-4a53-b292-afdb8feb4d2d · inbound

USEMA: a Scalable Efficient Mamba Like Attention for Medical Image Segmentation cites this paper.

USEMA: a Scalable Efficient Mamba Like Attention for Medical Image Segmentation Conditional Positional Encodings for Vision Transformers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.569616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:13:37.134769Z digest=sha256:49502af25906ab025520c669d055b6c240777033687d91b48d5a6bd0cbdd8d56

Observation 52959f8e-ec8f-4b95-9285-f7f071eae726 · inbound

Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers cites this paper.

Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers Conditional Positional Encodings for Vision Transformers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.986789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:14:36.959998Z digest=sha256:69899cb995c4a3cffd11bc94fa611001109f1016c8bbc2ee344fb03545e8c35a

Observation 931c443e-be63-4654-9069-576441a40e0f · inbound

End-to-End Context Compression at Scale cites this paper.

End-to-End Context Compression at Scale Conditional Positional Encodings for Vision Transformers

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:30.487775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:36:54.699174Z digest=sha256:3d252fe6654ff6ac9976cb186a84e695541372ff78744753f51b5aba93272e36

Observation 13d8427a-7d78-401a-a191-d5817240c59b · inbound

VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection cites this paper.

VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection Conditional Positional Encodings for Vision Transformers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:56.511381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:59:43.745654Z digest=sha256:b8a2ec17e48c17dfcdff5cf9cd3ba89df130c6bc7ee44698bde601ceaa0c9e2a

Observation 9a359755-f395-4fbc-ac6d-50e024e99554 · inbound

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers cites this paper.

DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers Conditional Positional Encodings for Vision Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.815071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:39:21.642584Z digest=sha256:1c669e12db316029626bd6e9e11bd768184bb550a1e8300a9f77757d80e2ba98

Observation b6aec985-8936-44c3-9c7e-4bb20c02a4c6 · inbound

Device Passport: Enabling Spatio-Temporal Pretrained Models to Generalize Across Input Layouts cites this paper.

Device Passport: Enabling Spatio-Temporal Pretrained Models to Generalize Across Input Layouts Conditional Positional Encodings for Vision Transformers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:37:18.453593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T19:32:45.679147Z digest=sha256:4db8ca0f3dd5698a0ab3faea7484b5190509b8768b897e06c7d9b278d7fe338f

Observation e08238e2-1901-404b-bb4c-ab96f4a925c0 · inbound

ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts cites this paper.

ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts Conditional Positional Encodings for Vision Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T02:23:52.225787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:23:52.225787Z digest=sha256:5f55c0e92de59cc04492529069104a6dbc34800bf148b9cd1dd9dea88a998e43