Pith. sign in

Paper Citation Record · LEDGER

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

As of 23 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2509.22854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22854 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:50:23.244358Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:25:20.472516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:33:58.600956Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c25d9971-bf1e-4d93-a035-a836d7cc835e · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.033438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.033438Z digest=sha256:011999c9b73f2a422d5d24952749935f9a701c2b7f386bc36f1b9a77f1b8450a

Observation 9af70504-b30d-4139-af1c-afc375105dad · outbound

This paper cites More importantly, as the input length increases, the inference time of few-shot grows much faster than that of ICR.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention More importantly, as the input length increases, the inference time of few-shot grows much faster than that of ICR

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.239578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.239578Z digest=sha256:cd462d01f170db5c7a2da427e499311b0fbf48f715ff9522afe1a4521963054a

Observation a5af7271-d133-4aca-b2d9-32e2bec09075 · outbound

This paper cites cross-dataset.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention cross-dataset

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-04T14:50:23.244358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.244358Z digest=sha256:4fd5aa9dca74e4e2c34beba7e0189fa9610a0616309d8d322d79e87264792980

Observation 82aa3593-2484-48e3-888a-c38ca0010efa · outbound

This paper cites The Llama 3 Herd of Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.731921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.731921Z digest=sha256:5faab3847cbe8b60220e4f36b552982eb43863ef3cd6689fbf63484d6f765743

Observation 44a17e93-b972-45ac-9a3c-0beca0d9a932 · outbound

This paper cites In-Context Learning Creates Task Vectors.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention In-Context Learning Creates Task Vectors

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.808650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.808650Z digest=sha256:162cde6a8e58b03d878a5bd3aa1dd2dc0459144a0a6d736b73d06e5c1d63e106

Observation 0a9b24fd-0d51-4178-8f14-e357930dbe82 · outbound

This paper cites Language Models Implement Simple Word2Vec-style Vector Arithmetic.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Language Models Implement Simple Word2Vec-style Vector Arithmetic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.166106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.166106Z digest=sha256:f317c642856429522087183b69686af8bf4b02c2a9a21e989a7ffb77f71f83ff

Observation c3bce674-6557-48fd-b743-f82b92e071a9 · outbound

This paper cites MetaICL: Learning to Learn In Context.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention MetaICL: Learning to Learn In Context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.170294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.170294Z digest=sha256:2dd1d613264b897c39e10f70f13f5d8a2a534d52dd1b499c09d7e7cac3855ad7

Observation f400c71f-1e22-4677-8f4e-747d308f8f76 · outbound

This paper cites In-context Learning and Induction Heads.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention In-context Learning and Induction Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.174869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.174869Z digest=sha256:d269064a0c6c76a14b101e62d52383dc1185a246f6f1341679038c046b3b3a87

Observation 4a7fe9d9-a096-4e72-9fd2-9f96448617ae · outbound

This paper cites doi: 10.3115/1219840.1219855.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention doi: 10.3115/1219840.1219855

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.184071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.184071Z digest=sha256:f473c4886667b2d4cf9dcc31c37b73e850255c1404d82e18574f0c38b5b3b198

Observation 994a7e63-778f-4647-b355-323f51efd39b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.192290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.192290Z digest=sha256:0407c9d78165c155c6b80fc44600ca49159c8f7f58d96925245a87fda52b04c1

Observation 05319dbf-edd4-4a00-a543-cd34d99d4e15 · outbound

This paper cites ELICIT: LLM Augmentation via External In-Context Capability.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention ELICIT: LLM Augmentation via External In-Context Capability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.196943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.196943Z digest=sha256:51f71b060db76f3bff2840acf1fd3d3cb4ce13a7c0fde43379ccd26ddf5941ee

Observation bc499890-326d-4133-bc1a-28337bec954b · outbound

This paper cites Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.201430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.201430Z digest=sha256:193e320565503961a36a62816af248b0fec0f01ca83df7ebfa05c25a9811148b

Observation d09bab47-963e-40f8-9968-d96c6e3338be · outbound

This paper cites Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.206160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.206160Z digest=sha256:70f52bee7439e5b39d1820d90d47ad81ebe030a88f006f4462d29b34fbbaa880

Observation 4cd13016-bd0e-4f67-9b63-926595056edc · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.211091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.211091Z digest=sha256:6f8db82986c93be53d5a34c2a9226d1802e9b2ad9940a07d919bdb2533ea75c3

Observation 6d911264-1561-445f-a98e-36f6d29df409 · outbound

This paper cites Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.215388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.215388Z digest=sha256:66b8808cf38683367dd8d42b8d611d050b253a6916830b11eda2acc401252f5c

Observation 9d9ffb5d-df01-4b9c-b496-f5e54429b280 · outbound

This paper cites Qwen2.5 Technical Report.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.219645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.219645Z digest=sha256:2eabd62fccfb1dc2c6d93f23d9945f7e14fb2cb3a0e67f37336db683362f6b55

Observation d4b2d54c-1f05-4460-88aa-b56786235d86 · outbound

This paper cites an unresolved cited work.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.224062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.224062Z digest=sha256:b15b6f1ccd01f63c0f02626b97b2eaf4b46e20efb29477cc5920050f4fb87454

Observation b7157842-139f-4c28-8750-e890e6040551 · outbound

This paper cites An identical argument applies toU k.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention An identical argument applies toU k

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.229151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.229151Z digest=sha256:bdd405d7db6084c9db566ebc92af48be493820aed3df7126ca355d0bf377329d

Observation ca6c0cdf-c5b0-45e4-9c7a-f855583512b4 · outbound

This paper cites For training, we use the same number of few-shot examples as those contained in an ICL prompt during the construction of ICL bases, drawn from five in-domain datasets.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention For training, we use the same number of few-shot examples as those contained in an ICL prompt during the construction of ICL bases, drawn from five in-domain datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.234829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.234829Z digest=sha256:e65e9740eb1a8b7e8c2487fd61aaa66a46d046cab0596f21972fe7b7cadc4158

Observation 783df9d7-7fa4-4631-b747-30e336762a43 · outbound

This paper cites URLhttps://doi.org/10.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention URLhttps://doi.org/10

Reference 1970

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.552599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.552599Z digest=sha256:9940b47d2f5d95c600fbf5de767d0fcd39e712148f87a6440925b382a50592ed

Observation add3b2d0-56e7-4ae6-9388-8f9344101000 · outbound

This paper cites Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.031609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.031609Z digest=sha256:c6a585edd6e8b335d7724545f411a9c7d2e8644d0cf436490ec80b252ad07a56

Observation cba2eb2f-9aa9-493e-b8e0-c66fbeda1485 · outbound

This paper cites M2iv: Towards efficient and fine-grained multimodal in-context learning via representation engineering.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention M2iv: Towards efficient and fine-grained multimodal in-context learning via representation engineering

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.156936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.156936Z digest=sha256:f0c4ca96d9e4301e29a9cd067a07aab719c9689e0f8b35a89f6b52a6e6db525b

Observation 32e6b99f-c986-4658-a64f-414c0e48b27f · outbound

This paper cites A Survey on In-context Learning.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention A Survey on In-context Learning

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.627647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.627647Z digest=sha256:2e2ed5d7d980f282086264edb0d1c092fe2fabc25656a4cf39f8341b7a33e946

Observation ba35a10a-481a-42e2-9508-fdf5e60b80a3 · outbound

This paper cites Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.436932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.436932Z digest=sha256:9acb84d8867a0dd82d8e0d080bc0ce837b9640873d088034f1059ee2085734bc

Observation 16181b88-35f6-415c-9a9e-ef0583fae94f · outbound

This paper cites Function Vectors in Large Language Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Function Vectors in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.187923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.187923Z digest=sha256:817a861c57f7bd7d72e28484f0cfc4c589d74a995fa62a0da691642e9c9a9ad3

Observation 3b1a1d2a-a493-4ba9-9c45-de2bea18e80b · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:21.969963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:21.969963Z digest=sha256:c82d7358c8105760a7700becd2038aaf4be4ccfea8dda587ca4fde3bd08c6d7d

Observation 52ee78c3-2997-409b-9924-78f6d5982f6c · outbound

This paper cites In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.161687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.161687Z digest=sha256:0388b4afa1be3894ed9e940bd5efd784a1f6764d79cb38d1a631a8568d33a7af

Observation 092222b2-7b0b-421f-86fc-33828e0c044d · outbound

This paper cites CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.179507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.179507Z digest=sha256:b56f0bf515d0f0a19c75738faa98f8993f3083ab8d210b074084ff88fef9c764

Observation a5a7667d-935e-4e89-9372-e419cd2f311a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention LoRA: Low-Rank Adaptation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.913368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.913368Z digest=sha256:e8a350ba6a9bc8b5fadf111113ce6e9fe6ec6053fec00696c95ca3fca1641c09

Observation c7acb8e5-97fb-4cb9-93c2-2060b522dffa · outbound

This paper cites On the Relation between Sensitivity and Accuracy in In-context Learning.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention On the Relation between Sensitivity and Accuracy in In-context Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.163974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.163974Z digest=sha256:6bcadb9710989319c8141057cf5504fe82a8670bd12a884c7ac3edbd667c9267

Observation 096fbc21-f5ac-4cd6-bbd0-3a79b1fb0c85 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.311948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.311948Z digest=sha256:f009adb768117809b9f5f69cc430200fa15d8c66511e69d8017a46129a2dd2df

Pith citing papers

Observation 966e88eb-be64-4ced-9e10-ee70b5369dd1 · inbound

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning cites this paper.

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-03T13:05:38.627651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:32:01.059706Z digest=sha256:1eb7179f88958c1c5be4504a544babba503146ef67f3c6bade9e2e56181784e7

Observation ea781fbe-d942-4f7f-9b39-d617e33c5649 · inbound

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL cites this paper.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.472516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.472516Z digest=sha256:9c0340871cf78e0a1f0748ac3f67f061b088c4a2a3145333c0be1cca03adef40