Pith. sign in

Paper Citation Record · LEDGER

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention

As of 21 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2501.00823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00823 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:48:10.830682Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 68d2ad62-9815-4416-b550-c63db1adb201 · outbound

This paper cites Neural module networks.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Neural module networks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.737781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.737781Z digest=sha256:f4799ce341eba9148fcffed0e7787bea8304d5590a5c19fe0aab8e8db195494b

Observation 8bef3166-2b04-4471-9ce3-ff97781e5a4a · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Neural Machine Translation by Jointly Learning to Align and Translate

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.743325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.743325Z digest=sha256:a3b28c7cc04384ca170f4904c86931aef3bbabf9e615d289670e6420389fc036

Observation e9c29b99-5ed4-40e1-8696-cddc63fd653c · outbound

This paper cites Language models are few-shot learners.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.155154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.748756Z digest=sha256:2e9a0ae1f22b5640405df035e2b8c71ed36e93ba5d6214c33af932d88c502ea1

Observation 6212ed3e-ec18-4ca3-bd25-96f7476358ed · outbound

This paper cites Decouple knowledge from parameters for plug-and-play language modeling.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Decouple knowledge from parameters for plug-and-play language modeling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.138139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.753812Z digest=sha256:814dadc18515b44f64a6740b6a218e3e1fec28e94cef2cb5196e82b76b1ae024

Observation 2703d6f4-c895-4980-b9cf-592da8a953ee · outbound

This paper cites What does bert look at? an analysis of bert’s attention.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention What does bert look at? an analysis of bert’s attention

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.120482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.758933Z digest=sha256:9cf4b3e6d2a416e016f8344470c6161cae9b8f74d4b6e6b591ad43e45bbd8d1e

Observation 3a468c94-b02d-4473-8ab7-b63a145a299a · outbound

This paper cites What you can cram into a single vector: Probing sentence embeddings for linguistic properties.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention What you can cram into a single vector: Probing sentence embeddings for linguistic properties

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.102738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.764561Z digest=sha256:2a8fd063c4c48e5fb8774b25735cf332ccb6c2e914e15309bc554e2e82334ab1

Observation 61789194-cd27-45b1-a610-4383bef020ae · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.770397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.770397Z digest=sha256:0a3786799fff6fcd46e678538ebf74742aa9ec61ae1ddc7fa3bad0c3060fb7ea

Observation 0bcb0aef-e63b-426b-aaf7-c8d9283b198b · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Transformer Feed-Forward Layers Are Key-Value Memories

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.775587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.775587Z digest=sha256:5f8358a49be50f7db1f960a3cd15350ffa75880e55fc6c79966cfee794cf95bc

Observation ffc3439e-06eb-429d-a5c2-aa575890b0ae · outbound

This paper cites Looking into Black Box Code Language Models.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Looking into Black Box Code Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.781303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.781303Z digest=sha256:9ceaa180d32a69ab88c1fc162058fccfb66cfd785ef34bd9f7657061f64e4ac8

Observation 56832787-e3fd-49fe-bb12-b9386f8e838e · outbound

This paper cites A joint many- task model: Growing a neural network for multiple nlp tasks.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention A joint many- task model: Growing a neural network for multiple nlp tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.084345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.786538Z digest=sha256:a95b9c42808d3b76285e3e11c4dc71539eecec9ca1435d0c951f4b6cd0d97dd2

Observation 07c18f2e-5bea-4ecc-8ebe-e247a2b9b0ad · outbound

This paper cites Adaptive mixtures of local experts.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Adaptive mixtures of local experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.791462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.791462Z digest=sha256:1076e619ef4c5c0f504e9134dbc4d9d46eebf3c5903320e12734c92d1041e4e4

Observation 13eb4147-f5dc-4c5e-81c5-2f1091be96fb · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.056826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.796514Z digest=sha256:5dffea5e52e90b56ebaa886914d27f24591a87c764fa78b9449d7d762fe02754

Observation 569b32db-d4d7-4a47-ad52-2d462ce4d619 · outbound

This paper cites TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.801282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.801282Z digest=sha256:dae4a87d89794a01087ebac70fcd13efa4cf22438c8c440d04b01f5f60227c0d

Observation 1caf7237-ec8a-4cad-81d7-b3bd631f8da0 · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Universal Language Model Fine-tuning for Text Classification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.806872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.806872Z digest=sha256:375792daa3ed3ec14c35b05ecc63190217d1e5a1bbbca1f6b0965a11fd03b8ae

Observation 6a63faa7-031d-4473-b52f-fb43caae246f · outbound

This paper cites Self-attention with relative position rep- resentations.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Self-attention with relative position rep- resentations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.040321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.811887Z digest=sha256:2c472a47d105e44221a0524db0d3939f673a85259d1d5739d23bc4de60676ba3

Observation d757b1a6-0996-4525-9a98-14f71ea9ebf5 · outbound

This paper cites Deep inside convolutional networks: Visualising image classification models and saliency maps.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Deep inside convolutional networks: Visualising image classification models and saliency maps

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:11.023520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.816576Z digest=sha256:21b054a3786efcf6418d292f61df2bc87796e638e214246429b1420bccc929f7

Observation b6c41dcf-89f7-4344-8aac-b996332aa4af · outbound

This paper cites Attention is all you need.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Attention is all you need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:48:10.821185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:48:10.821185Z digest=sha256:2540b635a0c9057b72389305525c8acc36fe4c8dfb2335659536edc104b17c0b

Observation a2550a77-d33e-4388-8886-b27efd886529 · outbound

This paper cites Multiscale visualization of attention in the transformer model.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Multiscale visualization of attention in the transformer model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:10.995067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.825959Z digest=sha256:9cba58ac7185e096d33d6215715b671180e3228bd0e4fa33a365cf6aab3e111d

Observation 4ddf0089-0d06-4c2b-b199-8d3ef2db154f · outbound

This paper cites Parameterized transformers.

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention Parameterized transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:48:10.975700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:48:10.830682Z digest=sha256:1ecfc92ce0a3e12150ee00aeec3ea538279332beef53c3c82e3889680a0cc1f8

Pith citing papers

No inbound Pith citation observations are available.