Pith. sign in

Paper Citation Record · LEDGER

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2412.17019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17019 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T06:00:38.833910Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 620b57af-f5b6-4810-aa43-6f8e60cce290 · outbound

This paper cites URL: " 'urlintro :=.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.619270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.619270Z digest=sha256:bfdec8a0acc242709088e0d871d342c1bb9f1817ee9d4d86dc0e88735a544e23

Observation 69f100e7-0477-4de7-8800-2fd1de30ad1a · outbound

This paper cites write newline.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.624968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.624968Z digest=sha256:4d7ebb83fffa4904fb92cd03c3dc91a9dd97f4608ed87623be9c3e535f69641c

Observation 706ef6a7-62ff-416f-a3b2-ba663aba29e7 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.545007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.630305Z digest=sha256:2340afde6bb3ab56a09151e51efe571a7c7788ba6072d3cf9821c7ec4244e8da

Observation 6a5bf91c-6455-4d42-8180-174510c64450 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.531007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.635324Z digest=sha256:218c83fc85220f491d549ca6f3075ef02cb4226a15991757eff64581e3e75db8

Observation 2bac1fdc-ff85-4d0a-ae15-a49c47a9d475 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.640036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.640036Z digest=sha256:5f9c920f6e1e63b0bce9f16b85e37eae39b96ec08afcb698f9a823af5823839a

Observation 5df77d16-ad49-43b7-8489-d1637188bfd3 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.507491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.644787Z digest=sha256:beb99088d745b0e3dec2fe362bc0b5414517a06d6e6e41ad7926363e42c4bbec

Observation ec07b667-a7a2-4274-aaae-d1501f94de2b · outbound

This paper cites A Survey on In-context Learning.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT A Survey on In-context Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.649207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.649207Z digest=sha256:4f72f2384e432530b7858531261045d59cdfc2982fdb2b0167e464633b7cacce

Observation e657096d-f071-4693-b8a6-cd68c7627f2b · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.492220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.654160Z digest=sha256:6f290884b495333a66f350a8a0df1673eddc4a9e7147de6f59b7518bae1db15b

Observation cde2cdb8-d0f7-4262-95bd-efc7b8ab90cf · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.476823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.659109Z digest=sha256:e7b30693981a7924641dc8f7fe75103b87fd7e29fe5553a613f00078b95dd171

Observation b1699512-76dc-4c96-b206-acc3bb802b79 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.663356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.663356Z digest=sha256:332acc4db4dc795bfd11fd12e94ed50091b483e42e2e363e1e52154309eac17f

Observation d62eb6af-27ca-4680-9989-548c72e5e1cb · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.462566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.667694Z digest=sha256:ebc636fffb1970c961a4dc49fa0f7e7cac4548e4398a9c19b0a5383b47ce9001

Observation bd9352eb-adce-41bf-99fe-49c43e24eb7f · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.672348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.672348Z digest=sha256:8f4c83cda6701479db7014bc263d130f611bd3535a0ed933f9e85114ced9280e

Observation d70bf17f-ef97-4cd6-9465-d07d00584b7f · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.438786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.676668Z digest=sha256:4cdb0357226cfd84e5fa889c06e6d8c4f809353bb65e558803fa9d1fdbe35043

Observation 1bc18136-2b5a-4175-8ee6-5e89869569a6 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.424252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.681056Z digest=sha256:af45159582b978428d5e1181ff3e9003665b1b915d257abe7aebcfe8b1520408

Observation 0f3a803d-6693-44a6-acff-4d3a4bcd4d2a · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.685491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.685491Z digest=sha256:a17ae3c9061289b680b59c2865ee7004fb8eb5f7bec04189c2d7aec6fbc78fe2

Observation 2a505e70-2d60-4edd-a747-39212e57353f · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.690684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.690684Z digest=sha256:538f6ce3ec513bbcba346b7b17d81925fe32804941e1f748b4da6774bb3111a2

Observation 5c9b3b55-af39-42e0-8332-111d9fbe967e · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.695131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.695131Z digest=sha256:39424a396caab0f4ea799b4d7a075fb09d3edfac19e0de7417ba80854b74dc1c

Observation ef196a78-ab84-4a6a-9114-6b3a11bd9ee7 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.399933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.699371Z digest=sha256:29bfa9e9f29fa68d33311112cc3963807cabe81de4566fc60d62b73862945b44

Observation 7f9ab124-1f46-435c-8ee2-7b700317832a · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.704605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.704605Z digest=sha256:74ed109446cc9b83d8389e1bba32185b323a7e7758095ecb0673048f2d7e4cd2

Observation 3ce826c7-2d05-4e76-a979-72ae659567a7 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.376106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.709122Z digest=sha256:c75816a5622b83bd4521c5dc7a44a0482326753feabd25c1d15d77bb28053dc5

Observation 6217acc9-50e9-4d68-a09d-c66b0571c45d · outbound

This paper cites One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.713955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.713955Z digest=sha256:0a957b79266d4517cd47f262b5b92b4abcfb0fb308af52f5eae3e1874eef85bc

Observation 38352ba5-a066-4f92-968a-aa098501394e · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.361428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.718836Z digest=sha256:3a95b2c2d28d13f38e40c2d822d0fb41637967cf7ad640185a718f0055ba37ac

Observation 20e010a7-8ef6-4893-aa81-ef908345985b · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.346516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.722961Z digest=sha256:e80415b987892f5a6c3875935f4233a5ad6f0de4ed18146c030a83f7f1b2d0cf

Observation 8ee49ffb-ea23-49a2-b39e-474658e5fad4 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.727336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.727336Z digest=sha256:6d9374754b92b04c5efccba04d52d6d2c688bce110c184216fea8bc066cf386b

Observation 20436d33-a522-497c-ab4d-a76e17cd6d34 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.330583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.731732Z digest=sha256:370a09f823b1836268fb574c2b156fb84aa3457e05c39ed551859698bb40c7c8

Observation 73a5f332-a0d0-42a8-a400-7146b79760e1 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.735956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.735956Z digest=sha256:e12655d7f4614a89b84dcaa3139a842b3d656bb3dd56ce48a9736bfc10f1e5a8

Observation 9dd1fcd9-8c64-4087-af76-6a50821613f9 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.740134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.740134Z digest=sha256:cd41b9981bf0911655891125f89e2a56f3eafc1c5d635585ff2781a6e7ffa446

Observation 431ffcf6-bbb8-4bbe-b5a8-78212ad7b3c6 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.296442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.744794Z digest=sha256:50f5c85091efff36b0ed54d38d0c69bb102bd57805d8bb9e23b63b2e00bbcf62

Observation e457b792-71e3-4dc9-ab5c-4c2d390b1608 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.282013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.749185Z digest=sha256:180ff01164766e71f7bfad8f5dd4254d83c200967274960a5dc4af956a7e9ef2

Observation 791e9a40-2bd2-4b62-a36d-cddef6ec498e · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.267247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.753314Z digest=sha256:650b9b7fbc199475eaa62e17362f5f842a104142114652e9b2cd272f732d6d7f

Observation 6e953ab1-6409-4d05-bcc6-8b46d3c06f15 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.757583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.757583Z digest=sha256:918271d551ab0be6bd2314775a839be8d8465ac5030214c9cfdcddc8d856bcb2

Observation b951175a-0437-4df9-9c9d-a1006038d219 · outbound

This paper cites Transformers as Support Vector Machines.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Transformers as Support Vector Machines

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.761946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.761946Z digest=sha256:4ef9fc4a862b59ab3053fbb92d73f30fd8a30870cc82304555aef38be4878247

Observation 219b7a93-c3e2-4cd3-af51-c5ad16f840bf · outbound

This paper cites Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.766687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.766687Z digest=sha256:63ecc936a42ef4a1ba23b7f1bed63364a5a15c0abc6aee62c973f57e37e19073

Observation 1ced0d0c-f46c-4abd-8ae9-f5cf4a022ccd · outbound

This paper cites Function Vectors in Large Language Models.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Function Vectors in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.771236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.771236Z digest=sha256:a988556b8ac6a9b9355337ebe0abcc770888f18e7af881c70e70b16a9f10f7fc

Observation 77d51de4-20dd-4b9a-90bb-15d267e8f743 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.775927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.775927Z digest=sha256:9c3402b8107489d35bc6c784b28c98ba7dacac351edff269ba4b33adb25a7b08

Observation ca54b518-3220-4538-9c33-f4ff44602fe4 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.780568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.780568Z digest=sha256:7180ae72b85c18fab0421ccffdc92564d3a96d910491b40d08a971d3fe1dd0a6

Observation 58b4fcd1-b840-4e43-af45-84140e90eecc · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.233941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.785023Z digest=sha256:cb475139ea1f9d5ee4ffee3d3d3847410c7c595c3185b480503cdfea96e0f41c

Observation 94f9f378-430c-4e99-9658-2ffbe3eae93a · outbound

This paper cites Neurons in Large Language Models: Dead, N-gram, Positional.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Neurons in Large Language Models: Dead, N-gram, Positional

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.789496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.789496Z digest=sha256:81bff7c6e868c405a1561c3d10e59998057eb412271fd76f27f223f8b528b9c0

Observation 2cdca2ea-bb22-4248-9fb5-bd7ef5924ebf · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.793805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.793805Z digest=sha256:f7d326b28a920c5723410d3dde0539fb5c33316987f496f081be3897fbe7b253

Observation a81bf314-ec7d-45e2-ba48-6122f9ca61da · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.209156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.797942Z digest=sha256:cd60f53f7215ab198494b686a2594889c9ad19ab1513f33bfd4aa3e7175d96c8

Observation 8d090988-c255-4db0-94c6-f7a478228d2d · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.801931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.801931Z digest=sha256:3b268c2d846a0251bde31387202dc591d8aa49b93768744676458184d707d77f

Observation 40e42c3c-a4f5-4e5a-b169-babf83398177 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.806577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.806577Z digest=sha256:beca08a837825c4a4fb1fb6295fa9decf66a315e49299a6d49d5da3580fb77a7

Observation ca04a931-f967-496b-81eb-1a7c39cc6507 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.810700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.810700Z digest=sha256:14d7f9b662b65a2fd26b7da508dadb3c3649bcf9ca2ba9e40c44d4bb18d5c0b6

Observation edc2be9a-07e6-40e8-8323-bedf69ba8f90 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.815310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.815310Z digest=sha256:5b7df4c9b936e52200745da42f555a9be0a7db3d13a64fa800c46796119a9d95

Observation 007f243f-6f75-4793-908e-fb603c3caab8 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT OPT: Open Pre-trained Transformer Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.819686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.819686Z digest=sha256:22b19ef74bf5530705528099b05e0b2e0b14a4783c8c0d2a879b7f374a56a9ef

Observation 74857f5e-631d-4ebf-a964-3d0285a53349 · outbound

This paper cites @esa (Ref.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.824200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.824200Z digest=sha256:32c3f9101edef05714a38598dcb4c9ce5495e20a3023dd76c8b33bf3da708768

Observation 8c92a21c-8368-4e8b-81f5-2cb082a091b9 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.828910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.828910Z digest=sha256:4af9a777929178137e9885a6f637f9ec1b7e8aec35e41f77c0901f195f261949

Observation f341b33a-4a20-4ce9-8c1a-b46f7fcd57cb · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.833910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.833910Z digest=sha256:867a3a7084b5677fb002be47a11c530fc2804a3ba2d6154f11b22b29f9d83fcf

Pith citing papers

No inbound Pith citation observations are available.