Pith. sign in

Paper Citation Record · LEDGER

A 2D Semantic-Aware Position Encoding for Vision Transformers

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2505.09466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09466 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:36:24.024024Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:24:59.868006Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation cd1b5b9e-6341-4f20-8b02-1ee2886cdf57 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

A 2D Semantic-Aware Position Encoding for Vision Transformers Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.334336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:23.917875Z digest=sha256:1adfaaf0f8899fdf97b7208a410c10a7a29602bfe05714d1b622a310abb8ea42

Observation fc6e9586-20f6-4370-8d02-f3441e01ada8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A 2D Semantic-Aware Position Encoding for Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.922527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.922527Z digest=sha256:aa18157d2d7d3d5075471681ec3a5b893e770aff334cbb81c03df1fb592165cf

Observation 906fac3d-049f-44bf-a8fa-6c30ef0d9671 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

A 2D Semantic-Aware Position Encoding for Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.926701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.926701Z digest=sha256:c39d9b8b588cd257867a11afa191fefb9022ecb9d5e67fb34f28245e472d15e9

Observation 0ae4600c-3bde-4971-a901-a4350fc96269 · outbound

This paper cites An Introduction to Convolutional Neural Networks.

A 2D Semantic-Aware Position Encoding for Vision Transformers An Introduction to Convolutional Neural Networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.930333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.930333Z digest=sha256:68bb97f22b8a80f09faf277fcc7475f260022036f2880bb0d2bf70643f13b89a

Observation fb579dfd-3b1a-47aa-ac69-d54769edda78 · outbound

This paper cites A survey of convolutional neural networks: analysis, applications, and prospects.

A 2D Semantic-Aware Position Encoding for Vision Transformers A survey of convolutional neural networks: analysis, applications, and prospects

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.316213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:23.935257Z digest=sha256:f25cf7d3ee2575a1253445b18a58554739af73fa4dc98dad1c5e242b4124df7f

Observation b6c9f95f-04e6-42ad-9528-9913bee3b495 · outbound

This paper cites End-to-end object detection with transformers.

A 2D Semantic-Aware Position Encoding for Vision Transformers End-to-end object detection with transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.939169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.939169Z digest=sha256:e71bd62107ac644c2e98fbf52ed332d87c09d30fa5ac567dce958b58ed4408c4

Observation d4370796-ab00-4037-9b60-b5e05f50bd18 · outbound

This paper cites Convolutional sequence to sequence learning.

A 2D Semantic-Aware Position Encoding for Vision Transformers Convolutional sequence to sequence learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.297597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:23.943348Z digest=sha256:d2a912bf530886a1a62f0e619db282714f73b62af3419352df8843e732fc4b35

Observation 39c52fd3-6693-4ce0-a5f2-8caf25e3de15 · outbound

This paper cites Self-Attention with Relative Position Representations.

A 2D Semantic-Aware Position Encoding for Vision Transformers Self-Attention with Relative Position Representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.946837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.946837Z digest=sha256:ec5209aac19eca8b5c10b886a96c5b785cb7a76eecf6d40e7ded70b8a68211eb

Observation 58d84363-93b7-4351-9ac8-95d3d762a5b4 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

A 2D Semantic-Aware Position Encoding for Vision Transformers Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.950709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.950709Z digest=sha256:98b141650309af316acf9e207cde91ff05ecc1c5b86e6bfb48d0961b625060cb

Observation 85863d54-5b64-409d-ae17-b9000d32696f · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

A 2D Semantic-Aware Position Encoding for Vision Transformers Roformer: Enhanced transformer with rotary position embedding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.954243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.954243Z digest=sha256:ccaa6dc5e15efc0339b280b046de59b5d5117af648246a3bcff9ed0d349ecef5

Observation 456d29df-0870-4d50-8866-9e2f81bccdcf · outbound

This paper cites Contextual Position Encoding: Learning to Count What's Important.

A 2D Semantic-Aware Position Encoding for Vision Transformers Contextual Position Encoding: Learning to Count What's Important

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.957639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.957639Z digest=sha256:66dbf1c3e70da384d73774bf684ed2324394917ef6a1887d391cf88746c98f4f

Observation db5f0865-fb75-4f5f-8d37-b8a964fb3f7d · outbound

This paper cites Rotary position embedding for vision transformer.

A 2D Semantic-Aware Position Encoding for Vision Transformers Rotary position embedding for vision transformer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.270496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:23.961146Z digest=sha256:642934bed6255606b6b20960c87ae8b11904c396701f661d1d7c4d7892d7a43f

Observation 0ad9b095-7edb-4b29-98e7-a269f547f01a · outbound

This paper cites Tokens-to-token vit: Training vision transformers from scratch on imagenet.

A 2D Semantic-Aware Position Encoding for Vision Transformers Tokens-to-token vit: Training vision transformers from scratch on imagenet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.964760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.964760Z digest=sha256:f26a287b86038e6944c46681282aee384a1d1e66bf02d94adf12007ac659d517

Observation 34694718-ad46-4887-8389-89f3b8e1c3d0 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

A 2D Semantic-Aware Position Encoding for Vision Transformers Training data-efficient image transformers & distillation through attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.968035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.968035Z digest=sha256:433ea0a44a0a7d9416ad4363a183ecfbf8262344d2add411764a3556a72f9bcc

Observation 58b16f48-2b77-4894-b7af-dcb475bdc746 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

A 2D Semantic-Aware Position Encoding for Vision Transformers Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.971642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.971642Z digest=sha256:3507ed8dd8abe988271f6e4da397aa4c5944428542a03db71715dd88ab280bbf

Observation 7a3ed9bd-6661-466b-a91d-b28cbe50b83f · outbound

This paper cites Rethinking spatial dimensions of vision transformers.

A 2D Semantic-Aware Position Encoding for Vision Transformers Rethinking spatial dimensions of vision transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.974985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.974985Z digest=sha256:072bb3d983aafc19835717bb0bbd6367459b2f1edad094bd6d8f114f960b9fd1

Observation 783e7021-b934-4616-9483-6f54f9266815 · outbound

This paper cites Cvt: Introducing convolutions to vision transformers.

A 2D Semantic-Aware Position Encoding for Vision Transformers Cvt: Introducing convolutions to vision transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.978738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.978738Z digest=sha256:1714de03f3f60c4610b0963633a0c9f461251982e001ed78908745a63c85f0c3

Observation 18511900-5c0b-4433-b983-461ac0e4ec9c · outbound

This paper cites Incorporating convolution designs into visual transformers.

A 2D Semantic-Aware Position Encoding for Vision Transformers Incorporating convolution designs into visual transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.221237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:23.982173Z digest=sha256:049aebb784d1e15e935052484cfa9366484945baf58e141a12ef9761c5155fab

Observation a91028e5-189b-4ce9-a799-a62cc3ead9b4 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

A 2D Semantic-Aware Position Encoding for Vision Transformers Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.986100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.986100Z digest=sha256:fadd7d10c1ff409f567974883e88f254c200fa7693d14e874f147f7141183867

Observation 7ba2335a-d9bb-441c-af8c-39a28af4863a · outbound

This paper cites End-to-End Object Detection with Adaptive Clustering Transformer.

A 2D Semantic-Aware Position Encoding for Vision Transformers End-to-End Object Detection with Adaptive Clustering Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.989698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.989698Z digest=sha256:30532f7effb0a0d8df7da31e8cdc20f4be6a06d9f5bd6c2af85ba97d87d2dfad

Observation 00f1041f-4053-4993-9d9c-2e15af692402 · outbound

This paper cites Fast convergence of detr with spatially modulated co-attention.

A 2D Semantic-Aware Position Encoding for Vision Transformers Fast convergence of detr with spatially modulated co-attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.208573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:23.994190Z digest=sha256:524762d4f8fc4815a0fdf55d92c4b5fa2c72c9f6df76e49eb107ca09374c5fc5

Observation 35ef02bd-7e58-4ddf-a11b-0b0d195462af · outbound

This paper cites Exploring plain vision transformer backbones for object detection.

A 2D Semantic-Aware Position Encoding for Vision Transformers Exploring plain vision transformer backbones for object detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:23.998118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:23.998118Z digest=sha256:83a4d19938087469350c2c5db8c16a8a2784df346fd85dd73c5de60cd9a3d1fd

Observation d76d772f-3245-4030-b0a2-951a991b79cc · outbound

This paper cites Vision transformer adapter for dense predictions.

A 2D Semantic-Aware Position Encoding for Vision Transformers Vision transformer adapter for dense predictions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.187850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:24.001487Z digest=sha256:76051a1816724454a5e113829f65d4bc13c33c216294550f655167b0e956cd52

Observation 6526cece-c906-45b8-8a6d-d6f2228911d3 · outbound

This paper cites Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers.

A 2D Semantic-Aware Position Encoding for Vision Transformers Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.176774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:24.005091Z digest=sha256:b6dab9a28365781df937cdce61c588a57320eb85637bb0f277e632cade1c09de

Observation b6d8ed79-3a7e-4815-9f99-b270636912b5 · outbound

This paper cites Segmenter: Transformer for semantic segmentation.

A 2D Semantic-Aware Position Encoding for Vision Transformers Segmenter: Transformer for semantic segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.164419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:24.008567Z digest=sha256:7467227a39a35ab8d604da8b4aa4aa443e55b29f70803a636a058e8c7d905687

Observation 85528404-788c-463f-b260-c52dd60bff93 · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers.

A 2D Semantic-Aware Position Encoding for Vision Transformers Segformer: Simple and efficient design for semantic segmentation with transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:24.011971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:24.011971Z digest=sha256:6719a9cc05eef9cad548cfba2713253a167ea52cc1f2bdf6555584221b657492

Observation 48fe5ba8-7ced-43d6-bb40-f1525e6399d6 · outbound

This paper cites Rethinking and improving relative position encoding for vision transformer.

A 2D Semantic-Aware Position Encoding for Vision Transformers Rethinking and improving relative position encoding for vision transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:36:24.015044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:36:24.015044Z digest=sha256:06f734501f2f556fd8a1ccd636a3e4d58b022b261d0d5375b59b39ad4cd4e0a0

Observation b73f2579-21cd-4093-8a7b-a01cf6556784 · outbound

This paper cites Lape: Layer-adaptive position embedding for vision transformers with independent layer normalization.

A 2D Semantic-Aware Position Encoding for Vision Transformers Lape: Layer-adaptive position embedding for vision transformers with independent layer normalization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:36:24.137738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:24.019554Z digest=sha256:3ebaa2b715eb3c474a544f81c678f57b36c1b6b0e2f96d3174876aa5d252696e

Observation 0729aed1-4fbc-4e0b-99b5-8fd1a9e4d851 · outbound

This paper cites an unresolved cited work.

A 2D Semantic-Aware Position Encoding for Vision Transformers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:36:24.122943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:36:24.024024Z digest=sha256:b774d915a6727eaf2ed35c495102b1ab324f6d4c729e46bfc974805ac76023f4

Pith citing papers

Observation 1516d0cf-18a8-4eb9-a0c6-551a561e9bde · inbound

Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers cites this paper.

Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers A 2D Semantic-Aware Position Encoding for Vision Transformers

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:57:03.694358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T14:52:33.394383Z digest=sha256:0b992057080de7b79b95aa72a530bc8e3d21c93051df438245e9e6d9d4c2f384

Observation e1fd5f79-d3dc-4291-bb49-048c01f8de92 · inbound

Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix cites this paper.

Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix A 2D Semantic-Aware Position Encoding for Vision Transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T03:24:59.868006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:24:59.868006Z digest=sha256:0caf7dd3e0975af9516da28a3431cc4699a32181ba06f1aab66137e87d015115