Pith. sign in

Paper Citation Record · LEDGER

Transformer Language Models without Positional Encodings Still Learn Positional Information

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2203.16634.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.16634 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:25:02.732069Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8501bb37-4668-4bfc-9689-f0dd56366d5b · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.088404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f1f2dd1c08eee9c54f0a6023d5a6b495b1ea2b61b202b86a0a62da6ba0c68b9d

Observation 653efb2c-3ef8-4d44-b276-5311aa9a4380 · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.862054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:2153f4436320a79fb0604ec1dd8f6e6173ba0c2a97c8ad3de47b8e0cb1234524

Observation 173022f9-3ad6-4e59-ae7f-05546af99168 · inbound

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers cites this paper.

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:56:50.070029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:56:50.009149Z digest=sha256:7b9afe9b4bb8f3193214a158020a3e31db77c34d51ed9df9aab78239393d99b4

Observation badd3ee9-56b1-40ae-8851-d1e1c3c20f38 · inbound

Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering cites this paper.

Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:25:02.732069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:25:02.732069Z digest=sha256:8f94166715d1a667bee02625c159fb52942dd515707e1b3d3ed4a640d55e8dc4

Observation 93ad4f48-f28c-4dd8-8a82-6aa22e0e9cbe · inbound

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding cites this paper.

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:03.081142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:03.081142Z digest=sha256:f9337cd94bb74dd725fff3e82914856a8761d5a4f623a1e6dcc0a323b62234d4

Observation efc5d93e-758c-404d-b3f3-fc444701bbdd · inbound

Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding cites this paper.

Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:44.811424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:49:44.811424Z digest=sha256:db7ef568bcf703765771c7289203a5c77f501806804b13a919b149137be8747f

Observation a9d86a93-f68e-4d2b-ab93-21b09c2b5825 · inbound

Transformer Vibration Forecasting for Advancing Rail Safety and Maintenance 4.0 cites this paper.

Transformer Vibration Forecasting for Advancing Rail Safety and Maintenance 4.0 Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:52.931195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:52.931195Z digest=sha256:f93563e17b17488da394916e04da1eebb80e8ecb1b4a9965ba5d0efc2b82d157

Observation f39a9532-9004-455c-8e24-47cfb80b006a · inbound

Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training cites this paper.

Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:06:43.880176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:06:43.880176Z digest=sha256:be45a6733e77ad623a75fbbe91215b3be1819340f404937d915b250c0c0403e4

Observation 50ff3154-c93e-41b1-81fe-ac5201bccfb2 · inbound

RoToR: Towards More Reliable Responses for Order-Invariant Inputs cites this paper.

RoToR: Towards More Reliable Responses for Order-Invariant Inputs Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:43.141542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:43.141542Z digest=sha256:6f5432155be5efcdb1f4632e4d3edfde449cebb703fbb5c81626f8daad09024a

Observation 645af7c5-3042-4dfa-8376-beebab301aad · inbound

Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach cites this paper.

Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:32.870058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:32.870058Z digest=sha256:829b84bf64d6f084a9ebf11cb7c375014a95bfca8fc63df70b1cf7863451756d

Observation aaca7eb5-b6f8-4d09-a665-c5d4e9145c51 · inbound

SeqPE: Transformer with Sequential Position Encoding cites this paper.

SeqPE: Transformer with Sequential Position Encoding Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:02.420065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:02.420065Z digest=sha256:a6fb569e041017d3b11805d1a2e62e8cbf8b928d5605302b0bec249c5e54a41b

Observation 76545f0f-3c66-46ae-ae27-bcac18cdf1ad · inbound

Multispin Physics of AI Tipping Points and Hallucinations cites this paper.

Multispin Physics of AI Tipping Points and Hallucinations Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:33.919802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:33.919802Z digest=sha256:2a3711e9d72c90d34171a9adbdac61e0a838c183d0054cb05a4fdfc47d59be48

Observation 2cf766f5-df18-47aa-b11b-5fea394327a3 · inbound

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting cites this paper.

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:29:01.272274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:29:01.272274Z digest=sha256:f449d71ccd974c4a75de69f96aac9683c4ac7aa31a62fc5afd4f8a930fe2bcd2

Observation 024e6ef0-6e81-406a-b1d3-a2608f3f7300 · inbound

CEHR-XGPT: A Scalable Multi-Task Foundation Model for Electronic Health Records cites this paper.

CEHR-XGPT: A Scalable Multi-Task Foundation Model for Electronic Health Records Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:59.188637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:59.188637Z digest=sha256:fcb22ce4b208679c432642367b2ca0413faee0e824ecc26422390804f894d0b4

Observation 7f806f3b-66bd-40bb-9b99-457706289e14 · inbound

HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models cites this paper.

HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T05:36:53.601520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:36:53.601520Z digest=sha256:a04b540871973566abf01f14e69ed145e268618ecc6948a338af3820614ba6ee

Observation d7436bdc-b19f-41b3-897b-1d6b760aa2be · inbound

Positional Encoding via Token-Aware Phase Attention cites this paper.

Positional Encoding via Token-Aware Phase Attention Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T16:21:36.609101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T16:19:48.318702Z digest=sha256:44c2000ed8ff32dc5a0d488aae9b6f298b34ff05c3baf053f94bf9eb2586379d

Observation 156be933-1f63-4f5f-9a6b-30a8988c3a98 · inbound

Group Representational Position Encoding cites this paper.

Group Representational Position Encoding Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.767612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:04:13.707931Z digest=sha256:74f212526c81fb498b9a3195109bee63418a69e57cef1c2137a299dc1f94a481

Observation a08c1a0f-9545-4905-8ec5-d1c1b43a7ee5 · inbound

A Foundation Model for Instruction-Conditioned In-Context Time Series Tasks cites this paper.

A Foundation Model for Instruction-Conditioned In-Context Time Series Tasks Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:13:21.732648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:11:43.037956Z digest=sha256:051e1a222585e0b68facdcd1f7fd0156a78d3bfb8b3da691acd4ba95829fd0b0

Observation dbebe238-8cd8-40ca-8c52-cc3e97245e66 · inbound

A Foundation Model for Instruction-Conditioned In-Context Time Series Tasks cites this paper.

A Foundation Model for Instruction-Conditioned In-Context Time Series Tasks Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:25:07.710016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T06:21:42.394251Z digest=sha256:78aa5d6d45ea7f4c0192120207dc39256797fe050b6b535e7974b585a706a166

Observation cc8cd4cc-44e2-4e39-9a9e-1e6b85392283 · inbound

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings cites this paper.

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:20:51.638374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:50:43.716813Z digest=sha256:965b06c09cf7924ee9d4354690a326757159f05aff2198351d492312295e99e9

Observation 5f0ca977-0fce-4456-aee5-387ace545b58 · inbound

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens cites this paper.

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:04.780449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T02:43:25.842048Z digest=sha256:942298a6086d71121b4b74343db3950c69e6fcad965e28851fc200dfa9f1c3f2

Observation 2ee66109-20c5-4638-8d77-35b795b06a43 · inbound

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices cites this paper.

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:43.848068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T20:00:27.987481Z digest=sha256:8f7aa81fe7af343c88f9349fb62ff7800ec48a43687d9e8df4a774edb6736224

Observation 406fdc27-22b7-4e10-b1f7-470c2cc7b495 · inbound

Graphical einops: bridging tensor networks and computation graphs cites this paper.

Graphical einops: bridging tensor networks and computation graphs Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.401271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T22:59:22.892338Z digest=sha256:cbb0ad7014d9ea63a4e58aa88cb99b48371d7a8f7b7be20944373cf49fe764c9

Observation 1c936d0b-22a4-4f5b-a87f-f560a3ab5e1c · inbound

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:09:59.040282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T23:52:52.985532Z digest=sha256:0eecdf1abec76fb57fe11120b7797b6c2015b34f9436c9c9a6b8eb4ac4c30f88

Observation 294b3462-0143-4891-9b6a-262638e79ea3 · inbound

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:44:37.728642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T09:38:43.250420Z digest=sha256:30692e05a69fa962d29d9a2c369ed5c826e9cdfdd683d056151509179191415c

Observation f60ed161-8f99-4c5a-b349-f1fd494fed80 · inbound

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally cites this paper.

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T12:17:16.288746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:17:16.288746Z digest=sha256:381073cb1028662e00c21ea93e50aa42bf70c62a9585a70fb11a9541a5e5eb73