Pith. sign in

Paper Citation Record · LEDGER

Foundations of Large Language Models

As of 12 August 2026, this Paper Citation Record lists 100 of 296 outbound references and 11 inbound Pith citation observations for arXiv:2501.09223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09223 v2

Coverage vector

measured 100 of 296 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:58.840590Z

measured 111 of 111 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:03.495496Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 296 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 907ff1c8-d1cd-4d71-a238-d064cd8e4500 · outbound

This paper cites write newline shortlabel.

Foundations of Large Language Models write newline shortlabel

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.509528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.509528Z digest=sha256:69d29804d5afa282b8e0aae67acf52eccc010327a829ef570c561e9c97c049a1

Observation 142d3c9b-39d8-49f9-a26a-991876b728b5 · outbound

This paper cites Etc: Encoding long and structured inputs in transformers.

Foundations of Large Language Models Etc: Encoding long and structured inputs in transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.514100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.514100Z digest=sha256:b1f334028d9dccc050e2ccd3b9e7ed83fcdd35cf5c9dcff9fdcac1f0e3411233

Observation 55f0c83b-20ad-49e1-a0f5-94551d180652 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Foundations of Large Language Models Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.517691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.517691Z digest=sha256:79b362e709d3af2125262e495c4767af6f881267a58fd828d3a5c72f2e938b32

Observation cb20103e-7338-4d06-9cfb-6b0b5aa8db99 · outbound

This paper cites u rek et al., 2023] 0.2em Ekin Aky \.

Foundations of Large Language Models u rek et al., 2023] 0.2em Ekin Aky \

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.521105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.521105Z digest=sha256:af34803ce326f200531a774ed2b795e07806b19a4e6ae204bd7c18b742036fe0

Observation 2a294053-18a6-4f59-99b4-dbaa26554a82 · outbound

This paper cites Revisiting neural scaling laws in language and vision.

Foundations of Large Language Models Revisiting neural scaling laws in language and vision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.524471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.524471Z digest=sha256:b282162005ab886a03544e5e1833bac820b028b4812522077471c9f7314022d1

Observation ccab39a4-88a4-408a-a311-93e92046d061 · outbound

This paper cites cosmopedia: how to create large-scale synthetic data for pre-training.

Foundations of Large Language Models cosmopedia: how to create large-scale synthetic data for pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.528037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.528037Z digest=sha256:77365b9a31d12e574e2db6ae2a5b255bba8b4f2891771a71d22815ad2e74ea52

Observation 1d64327c-867e-4a24-9eda-038ea4ffc4dc · outbound

This paper cites The Falcon Series of Open Language Models.

Foundations of Large Language Models The Falcon Series of Open Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.531546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.531546Z digest=sha256:5265c0d2b11dee1e3fb16acd443d3929ec1c600f1634115196224de6038cb70a

Observation 253bf201-a35d-4c8a-af92-2d970cb9a4c3 · outbound

This paper cites Neural module networks.

Foundations of Large Language Models Neural module networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.535537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.535537Z digest=sha256:ec64f4c8ae12e56eff045402e0406129f32a2c942093bddbbdbdb0904dfaf125

Observation 0cc30455-4c97-496f-83f3-946402b56882 · outbound

This paper cites Unitary evolution recurrent neural networks.

Foundations of Large Language Models Unitary evolution recurrent neural networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.538803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.538803Z digest=sha256:be3305c0e06f3f415491b49661b2d559d96bdaae48e46fb74909fce9fd38428a

Observation 45b24d81-263e-410c-a08f-fad3ede4a195 · outbound

This paper cites Situational awareness: The decade ahead, 2024.

Foundations of Large Language Models Situational awareness: The decade ahead, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.542168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.542168Z digest=sha256:b1487248a903eaa0e72630c66e2bf29df38f7c56a92bb0dd060b8612efd80a31

Observation 54be9f87-f993-4bbc-87a5-73f5f4173f7a · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Foundations of Large Language Models A General Language Assistant as a Laboratory for Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.545590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.545590Z digest=sha256:040a6bc5538cf70c6379e55a53103fc38eb3aa678ada8f4caece5ecc697b3f64

Observation 00bfb96d-f8d5-4aa7-9740-f31dd3a5670d · outbound

This paper cites Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V.

Foundations of Large Language Models Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.549094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.549094Z digest=sha256:8be17ba4fd887cc1a5dc29b591e51fe2705faa4928f89bd72243e32f410711fd

Observation 865bb5a4-bdd2-4858-8738-bb26eae8c007 · outbound

This paper cites A neural probabilistic language model.

Foundations of Large Language Models A neural probabilistic language model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.552216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.552216Z digest=sha256:a0a51a4a291b42a3ceccb040dccf6eec748aeab0d386e20b139ec0199db673ec

Observation 821ec499-ca58-4585-9b0d-169e4933ee6b · outbound

This paper cites Greedy layer-wise training of deep networks.

Foundations of Large Language Models Greedy layer-wise training of deep networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.555316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.555316Z digest=sha256:3b88949364d12b6f410dea9b2c749ffdf21c64240c6dd554de4175e69d399f84

Observation 5e8af4db-414a-4317-a738-8139b120335e · outbound

This paper cites Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atilim Gunes Baydin, Sheila A.

Foundations of Large Language Models Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atilim Gunes Baydin, Sheila A

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.558467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.558467Z digest=sha256:ae5386aab10171cef3a5ae31361b099afe51fe8a35b8ce78512f18db5a7dff77

Observation 27660fad-9372-4ca7-80e9-8cae5bc87a18 · outbound

This paper cites Pascal recognizing textual entailment challenge (rte-7) at tac 2011.

Foundations of Large Language Models Pascal recognizing textual entailment challenge (rte-7) at tac 2011

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.561595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.561595Z digest=sha256:dd8c9fcd117b35257cf0ccd413a3c1b6bff50eca53187fd823b922d33a09bb64

Observation 36098033-76b0-4a5c-ab8c-1b030d53e0c5 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Foundations of Large Language Models Graph of thoughts: Solving elaborate problems with large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.564599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.564599Z digest=sha256:4da4c413d44c50759c3d1a0ef6c025dee44b76ae62ecb7315542d012c6ababc6

Observation f8f5499b-05cc-4018-a046-10176a03ee08 · outbound

This paper cites Rotary embeddings: A relative revolution.

Foundations of Large Language Models Rotary embeddings: A relative revolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.567857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.567857Z digest=sha256:84a884712df80e3673450d6ee82a8854ae18befd5b6ce979eeac613a24f0e170

Observation 129b6346-9ee3-41f8-9dc1-1c870a533ffa · outbound

This paper cites an unresolved cited work.

Foundations of Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.571150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.571150Z digest=sha256:100a9ab9ef134f2af7a418b17499be719d8d365229b74639b48711a163fd2f44

Observation e61aa866-9aff-424a-bf0e-210baa8d1522 · outbound

This paper cites Combining labeled and unlabeled data with co-training.

Foundations of Large Language Models Combining labeled and unlabeled data with co-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.574729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.574729Z digest=sha256:8ac4b7b1cf6be54508c00ca55be2feb00effad7c84521cacbc8dfef04621a2b8

Observation 750ce1cb-3c4d-4e64-b591-9ba0d57e0006 · outbound

This paper cites an unresolved cited work.

Foundations of Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.578332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.578332Z digest=sha256:2ec0d840a1193305c4e94752afbb473d60cca478ae4667495c88ec7c6d6ace32

Observation c0982b95-36e1-4cb1-8f07-7655903b600f · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Foundations of Large Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.581590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.581590Z digest=sha256:8c26f55263afb6d1ff0781e92a1ea695f9cf2fb41486700ce9bdaaf013a1b880

Observation 80938da5-eb06-4999-bdbf-36fd6c42c68c · outbound

This paper cites A simple rule-based part of speech tagger.

Foundations of Large Language Models A simple rule-based part of speech tagger

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.584928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.584928Z digest=sha256:d018fc9b2070e59fa589639cc23093f593e68520f1c3147b3c63a53d224b64d3

Observation a71d4908-33cd-4ab8-bfa5-4609b318eeaa · outbound

This paper cites Brown, Stephen A.

Foundations of Large Language Models Brown, Stephen A

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.588223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.588223Z digest=sha256:6b4b5f0722ced7b0865c80a770881fdaa185bb1245f716904269b4b33b63ba52

Observation ac96e79f-ef89-4bb3-9f60-365eb4a452ac · outbound

This paper cites Language models are few-shot learners.

Foundations of Large Language Models Language models are few-shot learners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.591218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.591218Z digest=sha256:7f7ed276e94b26c33fbb0014220dfeb8ed1603e74980350b448a6ea6aa502329

Observation b7bbc6f2-0c3e-4ce7-8aa8-1a6d3ec6ba3e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Foundations of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.594432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.594432Z digest=sha256:e893a61a264a95d2dbdaf3059490f022c922960d8353227c70211eb1b85539aa

Observation 738ff9c8-3fe2-47c4-b087-6ce3a605fae9 · outbound

This paper cites Recurrent memory transformer.

Foundations of Large Language Models Recurrent memory transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.598126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.598126Z digest=sha256:978d8be06560cac89a8c009fc31849e999fe5a91b88a7297121907a14002c13f

Observation 0e4aacaa-e5f4-45cc-ad22-8509a7d7085e · outbound

This paper cites Learning to rank using gradient descent.

Foundations of Large Language Models Learning to rank using gradient descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.601401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.601401Z digest=sha256:5938fd17d9280db2ec2b5bd99c236439ef47d07d7819ccd0d43381ea562e640b

Observation 949605c5-2368-455e-b47d-eeaad574894e · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Foundations of Large Language Models Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.604615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.604615Z digest=sha256:91b4e097b32dcfcf34552c855b1cdd56bb6cd7ab7c11a5a0ec50d8c89894094a

Observation 8de533b3-eb4d-4bef-b5ee-28ad54472a6d · outbound

This paper cites Weak-to-strong generalization, 2023 b.

Foundations of Large Language Models Weak-to-strong generalization, 2023 b

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.608322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.608322Z digest=sha256:a63f5caac80fb92c0f6806018ecf13eef0e98e377211afe22c950c52507b3a91

Observation 5d3fd461-0ae9-461e-b8ea-0576d9bbd1b4 · outbound

This paper cites Broken neural scaling laws.

Foundations of Large Language Models Broken neural scaling laws

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.611345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.611345Z digest=sha256:e1f53c7d530b3ba2c5f2bb0cdf758b88de8bc865f8b4415138fdf299e47bc83c

Observation 9cf8fc65-05e0-4c0b-8279-c964e3f087f7 · outbound

This paper cites Learning to rank: from pairwise approach to listwise approach.

Foundations of Large Language Models Learning to rank: from pairwise approach to listwise approach

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.614506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.614506Z digest=sha256:fb4178851c6abbb42b234592619bfdcdfc40418996cf4f563b9794930e30d70e

Observation 62b1c3a9-b2da-4963-9c58-738f6d1cc788 · outbound

This paper cites Efficient Prompting Methods for Large Language Models: A Survey.

Foundations of Large Language Models Efficient Prompting Methods for Large Language Models: A Survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.617622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.617622Z digest=sha256:a325b0f0b6a87fc344a73b878f2b3379784270b30a9efa6240798f847fb9df8d

Observation 0b7f0568-7f32-443e-9f80-3c72107e2093 · outbound

This paper cites Statistical parsing with a context-free grammar and word statistics.

Foundations of Large Language Models Statistical parsing with a context-free grammar and word statistics

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.620989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.620989Z digest=sha256:1778a57ed7a156b8514e83340431f31a7d4e00a02e78c3e60046a6f2d5d8a156

Observation 667b4a7a-4c4f-402b-a827-b8e04d545740 · outbound

This paper cites Unleashing the potential of prompt engineering for large language models.

Foundations of Large Language Models Unleashing the potential of prompt engineering for large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.624230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.624230Z digest=sha256:fb0fa65122d6dbd118a3f7495d79af57ed984f140b5673e48b4aaae2dfc9389b

Observation 4090a95b-8a5e-4017-a7ae-475dfc51a018 · outbound

This paper cites AlpaGasus: Training A Better Alpaca with Fewer Data.

Foundations of Large Language Models AlpaGasus: Training A Better Alpaca with Fewer Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.627597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.627597Z digest=sha256:92c3b9b61eb27abaff7ff664c7bcc7230c67fbb1a1ead149bf0c90d0ec2b9453

Observation 1e9c5dac-7740-444a-ad52-699f2dd81e02 · outbound

This paper cites Alpagasus: Training a better alpaca with fewer data.

Foundations of Large Language Models Alpagasus: Training a better alpaca with fewer data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.630840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.630840Z digest=sha256:4dcee8018bc2f1dbef0e08001eedf27fe64003862176092e5687ae39bcc900fd

Observation 629829ee-48a3-457b-9439-9b00fdda6921 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Foundations of Large Language Models Extending Context Window of Large Language Models via Positional Interpolation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.634213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.634213Z digest=sha256:708d09c11bc6b6f4e474f489e707e52689fcbc8a3117243fcc93fc02d191b061

Observation 438fba3f-6165-490c-975a-cb6f8a3f4d98 · outbound

This paper cites The lottery ticket hypothesis for pre-trained bert networks.

Foundations of Large Language Models The lottery ticket hypothesis for pre-trained bert networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.637606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.637606Z digest=sha256:c22e6d9abae004cffa0424b4059f313dd0257584ac8f45501bda337418e631eb

Observation 90a69a1f-b00e-4a9e-a7d2-f8c0163d0333 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Foundations of Large Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.640625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.640625Z digest=sha256:89589bae717041bdcc98cf904a33be0ee13a65a1ecd99ca1087a6d044f8ba028

Observation ee536048-0697-4e59-ae3d-09cac467f832 · outbound

This paper cites Adapting language models to compress contexts.

Foundations of Large Language Models Adapting language models to compress contexts

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.643906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.643906Z digest=sha256:b8b2efba1012c3cf752e36bd5325d24c780428fc324a7d97c6c93445c1b752c6

Observation df910252-8385-451f-ba12-67ad10abe1c0 · outbound

This paper cites Kerple: Kernelized relative positional embedding for length extrapolation.

Foundations of Large Language Models Kerple: Kernelized relative positional embedding for length extrapolation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.646733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.646733Z digest=sha256:0b7298c51bfa420c5d4a868062c571ec5f0bce399dd6b6dd1e58d67ee4f36722

Observation 023e652a-ec66-4e87-99f3-e6492feefe4e · outbound

This paper cites Dissecting transformer length extrapolation via the lens of receptive field analysis.

Foundations of Large Language Models Dissecting transformer length extrapolation via the lens of receptive field analysis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.649729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.649729Z digest=sha256:07845e26cb49e4124737d5834270773ecaba31b7827c4425d0b5f00387e3dd9f

Observation a36f8d35-0e95-476a-8c58-181574d88f5a · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Foundations of Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.652630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.652630Z digest=sha256:1003a21174d52572045514cb6abb3bdf2efd56e7bbae85d98ee4b5d7eff5b934

Observation ad89e3e9-ead4-4c35-bccf-f36cbf181ea4 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Foundations of Large Language Models PaLM: Scaling Language Modeling with Pathways

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.655761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.655761Z digest=sha256:befb7fa7b2a7d2bdac4a03d04a5f7fe928d5894eba60414e2bbb0b50864cd441

Observation 005e5d47-f0aa-4f62-945a-bf2f58f55b82 · outbound

This paper cites Deep reinforcement learning from human preferences.

Foundations of Large Language Models Deep reinforcement learning from human preferences

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.659011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.659011Z digest=sha256:364e1f0350b894f1fcb0693f5aa6d2e4b1d97abebae1ec28e4a770c770adc095

Observation 04ff6043-2658-4bbb-a4ee-7b83e72148a7 · outbound

This paper cites Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future.

Foundations of Large Language Models Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.662121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.662121Z digest=sha256:198df927b2e8ea65bef037973f6d350dd8ebe08e113d1b54d46b339ca5159e45

Observation c0befe19-fcf7-4526-ae91-7268df30c0ba · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Foundations of Large Language Models Scaling Instruction-Finetuned Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.665528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.665528Z digest=sha256:c2f31ff65d3b922944f55258d4c65bdc918240c00dd2b349e553af37899e4ba8

Observation ea0ada8a-adeb-42de-b945-27d322a31a8e · outbound

This paper cites Electra: Pre-training text encoders as discriminators rather than generators.

Foundations of Large Language Models Electra: Pre-training text encoders as discriminators rather than generators

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.669466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.669466Z digest=sha256:0cf3dab4b29e9364bd6c5b49f2edd6ab09cfaa147f2cebf021720e8a466c6635

Observation bc92b21a-6ef4-4e34-9462-f12fbee7b9a3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Foundations of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.672489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.672489Z digest=sha256:fcf6dccd29849239947fac3c566dea6f109d9c26c11e1aecb9af4ae957965fd3

Observation ae0f367e-9693-41a0-871b-c5fec4dd387d · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

Foundations of Large Language Models Unsupervised cross-lingual representation learning at scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.675390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.675390Z digest=sha256:cc82fb9e93e539f033b7e3bdf26712822d1db28cc6bc08c644a649c057459c8d

Observation 19c5724f-26bd-438d-a0e3-32b75092cdaa · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Foundations of Large Language Models Reward model ensembles help mitigate overoptimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.679337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.679337Z digest=sha256:a76671f2e75f161776c97d7789b419cc2d5c86010976d0a875b1fe486db627c5

Observation bc04e841-fff6-422f-9399-851a6a699c90 · outbound

This paper cites ULTRAFEEDBACK : Boosting language models with scaled AI feedback.

Foundations of Large Language Models ULTRAFEEDBACK : Boosting language models with scaled AI feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.683371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.683371Z digest=sha256:46b52ece3ade6dae1ea8cea161bb03a34308c6895836d5dc8e149446d8e687a3

Observation 86e532b5-7031-4fb7-a3c4-fa58d534f1eb · outbound

This paper cites Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers.

Foundations of Large Language Models Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.686580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.686580Z digest=sha256:883c87a695cf9dcbaac5dc64c26dc7a21a75cab19b7de202c1ea044246fdafa5

Observation f103b72d-bab2-4b71-b063-0365be8237d9 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

Foundations of Large Language Models Transformer-xl: Attentive language models beyond a fixed-length context

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.689926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.689926Z digest=sha256:b75f4668994ed65a009cac4dbb6f80cf0725989db0ec43786fb9b1a7cd00e4a6

Observation a3b68eeb-7f42-4ed0-a92a-8747eb2180fe · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Foundations of Large Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.693294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.693294Z digest=sha256:bb538aafbaed6bce2a7e84b272d89276fad88ad802c19483ab3cca7ced6f02c2

Observation e3ad06cc-76d6-44a9-84c7-6a3856c33148 · outbound

This paper cites Universal Transformers.

Foundations of Large Language Models Universal Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.696834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.696834Z digest=sha256:cb60e3fd81b4cc7a8d82ebf4ca70b2d774df6071b23ea13fec033ed7bb10a03e

Observation 9b987ccf-238c-49fa-a245-c9f4e28a157a · outbound

This paper cites Language modeling is compression.

Foundations of Large Language Models Language modeling is compression

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.700711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.700711Z digest=sha256:0dcbffd7ba93c79b9975b971361f689cc73383c089f403a3f66539612ff5e910

Observation 0ae091fd-29d2-4c99-920d-ce72d21c4af9 · outbound

This paper cites Rlprompt: Optimizing discrete text prompts with reinforcement learning.

Foundations of Large Language Models Rlprompt: Optimizing discrete text prompts with reinforcement learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.703974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.703974Z digest=sha256:64b267546970556f49530ff300669695f9e2604f28b46e02f8e879c5ef3b1d89

Observation 34acb0c5-45e7-4e8b-af2c-06f045f83a6b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Foundations of Large Language Models Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.707089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.707089Z digest=sha256:e86ad66d4da6f6ca3ab80edc19ad949e24f2ddb142c592f64e354c131294cbce

Observation 2301d6e3-b5a1-4d97-9d7d-c74b63e1e25b · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

Foundations of Large Language Models LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.710382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.710382Z digest=sha256:3feb1c5b6c03b2721e5de54a0790fd8e38f69ffec1a97dc705b82d5a83c95999

Observation c000de96-a7b2-4279-9ee4-33af94d89411 · outbound

This paper cites Automatically constructing a corpus of sentential paraphrases.

Foundations of Large Language Models Automatically constructing a corpus of sentential paraphrases

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.714041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.714041Z digest=sha256:a36e829ef5b5e235d2f7fdec6be8f58c797f9fe3dd0e6bf2faaf72574ea5e726

Observation cf5e2e89-f322-439c-84cd-9827c5004f4d · outbound

This paper cites Unified language model pre-training for natural language understanding and generation.

Foundations of Large Language Models Unified language model pre-training for natural language understanding and generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.717134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.717134Z digest=sha256:7ca77171e5ba258ef5f503121f4fe810d9fe12bb64d017efabef73e85f033ebc

Observation 1e83ceb7-0974-4c07-b220-b9221b4e69cd · outbound

This paper cites A Survey on In-context Learning.

Foundations of Large Language Models A Survey on In-context Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.720433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.720433Z digest=sha256:39861cd653fba71db224287ed2d424f233304996693908ab2370c4945af7bde3

Observation fe5f2a13-26ba-4b6b-a146-d9ec38f48d8a · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Foundations of Large Language Models Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.723835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.723835Z digest=sha256:0256cb42e8e0c5e646d27d98a74475b85e69377815456c263bcdf4a28751b691

Observation a2375f91-3258-430a-91b1-d5938a51f401 · outbound

This paper cites a rli, Ekin Aky \.

Foundations of Large Language Models a rli, Ekin Aky \

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.726892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.726892Z digest=sha256:eef93d2680ef9766e5fe325da862b412602892c4ae4003378c763d417bfe33e8

Observation 4a307b94-658a-4f28-b6f9-6a0e1c2bdbce · outbound

This paper cites Successive prompting for decomposing complex questions.

Foundations of Large Language Models Successive prompting for decomposing complex questions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.730055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.730055Z digest=sha256:55b5e46ba4e2054198e39609e2966b8341deb343361b60290780285ea4e16c30

Observation 266b4cf9-742a-4d3f-b5aa-d053ef50bbc1 · outbound

This paper cites The Llama 3 Herd of Models.

Foundations of Large Language Models The Llama 3 Herd of Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.732989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.732989Z digest=sha256:bb024bd59391ca4f66964df65c1d5efb5f5e7786ce1baaa7c86b3eb4d5017789

Observation 4f1c0419-2d6a-427f-a8ce-2faf7c7feeb5 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Foundations of Large Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.736097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.736097Z digest=sha256:9b0068a9bd59d04755172f53bf39a3c61dae869c927f9b20f091cf9faccd6bc8

Observation ccff36fb-7f04-4eb1-af1c-744f6ccaa5a7 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Foundations of Large Language Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.739081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.739081Z digest=sha256:b5045304a8c4fb09fe2e82b578c756b07e1e53f94d90d85905f9cd1655017420

Observation a3c4c360-73df-4d8d-a8cb-d7f6844bbff4 · outbound

This paper cites Neural architecture search: A survey.

Foundations of Large Language Models Neural architecture search: A survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.742373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.742373Z digest=sha256:fa9c99be7d76503238956b13faa6e165b210a1c1b0e0ca80fa48d26e97f2bf8e

Observation 23b2c85d-29b8-4ca5-b733-f12fe18595f5 · outbound

This paper cites Why does unsupervised pre-training help deep learning? In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 201--208, 2010.

Foundations of Large Language Models Why does unsupervised pre-training help deep learning? In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 201--208, 2010

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.745408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.745408Z digest=sha256:3f34ab6abfbe8e94f229b18b6ab1c6559afd308e7df50f9ddbf5a0e8a11ea872

Observation 5b641803-11d0-47fc-9043-a63c8525efb0 · outbound

This paper cites Reducing transformer depth on demand with structured dropout.

Foundations of Large Language Models Reducing transformer depth on demand with structured dropout

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.748751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.748751Z digest=sha256:6a94892c09375c8e8a1eb2650c42b84405f9993757d91252706bee8965cb64a4

Observation 8c98ac07-7c1b-4b3f-a7ec-2f49877d57c1 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Foundations of Large Language Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.751736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.751736Z digest=sha256:3fd0db44aebec15584ac30ac24caa5520198a99212d15d204ed7a45ae916d978

Observation 7a0eaa03-4892-45e9-9a0d-b8229a5c82c6 · outbound

This paper cites an unresolved cited work.

Foundations of Large Language Models Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.754931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.754931Z digest=sha256:361eabca4355b9b6637e7e17baedf408403fa10af58d435a9b5539d8619ff2d7

Observation d110e0b8-02f9-43f1-b4c7-9cc41e14c074 · outbound

This paper cites Is it an agent, or just a program?: A taxonomy for autonomous agents.

Foundations of Large Language Models Is it an agent, or just a program?: A taxonomy for autonomous agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.758007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.758007Z digest=sha256:9e9039b76d94f94138e7a6414ccacbd638645a160c766eaa1de8f35f8cab00c3

Observation caa0fc4a-61d6-417b-bfc5-e1ae24c92f86 · outbound

This paper cites Complex problem solving: The European perspective.

Foundations of Large Language Models Complex problem solving: The European perspective

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.761564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.761564Z digest=sha256:3bb243714d0a485e94368b7951d08a6100281b823a5e691c1af8f95d76d5fd1b

Observation fafb59b2-8aec-422c-b692-88e2459022ff · outbound

This paper cites The State of Sparsity in Deep Neural Networks.

Foundations of Large Language Models The State of Sparsity in Deep Neural Networks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.764877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.764877Z digest=sha256:4959a02bf3559c791f666c2397226148d4fea0b0cc5fae2e346eefb233b7ff3d

Observation 63bb554f-d214-4413-ba0e-1629fc133caf · outbound

This paper cites The Capacity for Moral Self-Correction in Large Language Models.

Foundations of Large Language Models The Capacity for Moral Self-Correction in Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.768223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.768223Z digest=sha256:63d530f61f0f706eedc09b2bce3954c7ef0236a579d55a9cc90adc8ba3d77662

Observation a44d381d-e2fc-4429-8aed-814c38b598c1 · outbound

This paper cites Scaling laws for reward model overoptimization.

Foundations of Large Language Models Scaling laws for reward model overoptimization

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.771768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.771768Z digest=sha256:a9831928553425117d87afaf9426aeca4eab2d1849189233dc65c0575aeaef20

Observation 3189b0c0-67f4-4f22-8e57-d281f55aee88 · outbound

This paper cites Pal: Program-aided language models.

Foundations of Large Language Models Pal: Program-aided language models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.774917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.774917Z digest=sha256:e0105e07119d852796dde5561a0757952ac73efb7271cd633d5015b9fcb38529

Observation 767e325f-3e3f-4773-886f-cdee6563dcc6 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Foundations of Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.777809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.777809Z digest=sha256:e7776308449a943f2c76e991cfe17c491d791fe3edfa81cf3aa7961fa71968ec

Observation c076198d-4d16-4fc8-9152-3eee37c48f0c · outbound

This paper cites What can transformers learn in-context? a case study of simple function classes.

Foundations of Large Language Models What can transformers learn in-context? a case study of simple function classes

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.781135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.781135Z digest=sha256:9593a89466fee42c117f5177ca37c1f81f0886de29b61b0980ef34b5f9f6a912

Observation 7dce86e0-93ab-474a-94f4-84a61d16b73a · outbound

This paper cites Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation.

Foundations of Large Language Models Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.784378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.784378Z digest=sha256:c7f8512a8e75d0e1f8806ba38b3dbdde12d1ec311ad6114f0196022619e105d5

Observation d3ab9f2d-f063-4784-8dde-8b76f40f6d81 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology , 2024.

Foundations of Large Language Models Gemma: Open Models Based on Gemini Research and Technology , 2024

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.787819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.787819Z digest=sha256:3b215405a14548f361bd10b3224b1d10a23d391ab37d33137e5dc30444c42f13

Observation 0072450b-6b23-4bbe-a9c8-e33ed0836c84 · outbound

This paper cites Problems of monetary management: the UK experience.

Foundations of Large Language Models Problems of monetary management: the UK experience

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.792737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.792737Z digest=sha256:8dda71a06adad93dc3d4b7789eb65466fc303a2ba815ccd2b53b7e678944137d

Observation 8788c57e-87fe-4650-86d1-4d15edad0278 · outbound

This paper cites Data and parameter scaling laws for neural machine translation.

Foundations of Large Language Models Data and parameter scaling laws for neural machine translation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.796107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.796107Z digest=sha256:4ad2b174dc623463e3dc7699ca8cfdca1ade681e191d501968b165eac2b39508

Observation e395fbc2-f11b-45cc-8760-89285713e0d4 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Foundations of Large Language Models Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.799149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.799149Z digest=sha256:3d82bc55965c3c2f6487471b5d0025e53828714100a2cfa4475a26694bf87ed6

Observation 276e75de-886b-40d0-a1b4-4053e8ed2885 · outbound

This paper cites Textbooks Are All You Need.

Foundations of Large Language Models Textbooks Are All You Need

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.803274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.803274Z digest=sha256:4a79ae30578481a76910a693e89ba357d72d5f8b8a3c3579c8b9fe7a8586fd64

Observation 9f7e4c73-e72e-46b6-bd92-e9a1ace8ace9 · outbound

This paper cites Connecting large language models with evolutionary algorithms yields powerful prompt optimizers.

Foundations of Large Language Models Connecting large language models with evolutionary algorithms yields powerful prompt optimizers

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.806744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.806744Z digest=sha256:9d35f4d77bfe403b8065c4e3212760465b2b9abf3af5408c733969be650aaec0

Observation dba02226-7432-4924-8448-4b827c5ce234 · outbound

This paper cites GMAT: Global Memory Augmentation for Transformers.

Foundations of Large Language Models GMAT: Global Memory Augmentation for Transformers

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.810306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.810306Z digest=sha256:88d034ade88ba5d95c4791d6829b5dff36d072edb344c651685117c89bf18a08

Observation 86653ccd-2c25-4c36-a238-a44643e5ad7b · outbound

This paper cites Memory-efficient transformers via top-k attention.

Foundations of Large Language Models Memory-efficient transformers via top-k attention

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.813749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.813749Z digest=sha256:1bdeba5cf8773955aa00e6c294390db89a49081b2139c2d77512c24d40f3bac7

Observation c136eb27-bcd4-4741-9c65-e159d6ec199f · outbound

This paper cites Pre-trained models: Past, present and future.

Foundations of Large Language Models Pre-trained models: Past, present and future

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.816732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.816732Z digest=sha256:6cf10f4fa824ecf598a138b7d7f2eed7ee40476a58e89b0f185c4759f8ffae38

Observation 062ddc89-cbe2-4d9d-9371-822248700a9a · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

Foundations of Large Language Models Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.820000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.820000Z digest=sha256:7cef9d129817d61699e8c89d438ddcb2514aea9b50d3329ff1286c5c33681a7b

Observation e8c95e2a-01a0-4983-be3a-0787005b17af · outbound

This paper cites PipeDream: Fast and Efficient Pipeline Parallel DNN Training.

Foundations of Large Language Models PipeDream: Fast and Efficient Pipeline Parallel DNN Training

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.823417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.823417Z digest=sha256:16be3a14edb940363053bd1d6934302eb00e6a24b908b5a2b2fd2521b77761e4

Observation e02a2649-d01b-445c-ad02-4e8ddb2df8de · outbound

This paper cites Rethinking imagenet pre-training.

Foundations of Large Language Models Rethinking imagenet pre-training

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.827329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.827329Z digest=sha256:6d343c57460691ca5c77699f7fc0800c76186217a6c1a665f762ed151e3f26bb

Observation 11134ece-f9fe-4619-8fa9-d5496ef41b78 · outbound

This paper cites Deberta: Decoding-enhanced bert with disentangled attention.

Foundations of Large Language Models Deberta: Decoding-enhanced bert with disentangled attention

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.830890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.830890Z digest=sha256:7ef6e670d7241015bd2b0df84ab6c04ba5911c6e0058c82ea01ea6944e1b6db3

Observation 8474a021-a658-4e27-8886-2f0f1a4ab04a · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Foundations of Large Language Models Gaussian Error Linear Units (GELUs)

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.834369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.834369Z digest=sha256:eaacc81bbd57003e72f9df6cd72d4a40e538a33dfefa08fa8a5334ddbc5a99f7

Observation 48facd71-75a3-4373-ac6a-3e4fb6f23a97 · outbound

This paper cites Pretrained transformers improve out-of-distribution robustness.

Foundations of Large Language Models Pretrained transformers improve out-of-distribution robustness

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.837669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.837669Z digest=sha256:fab8a7fe13b882a1192c172629fb6facc841d7560f9a5d1dd6f4a62890ad3904

Observation 4bbd3860-ee47-465a-8feb-ed4b0c0e6344 · outbound

This paper cites Measuring massive multitask language understanding.

Foundations of Large Language Models Measuring massive multitask language understanding

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.840590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.840590Z digest=sha256:8f7c54a7290bfc6cee63864b104744301dc39eb4104174f823feabd61a2458b4

Pith citing papers

Observation c70c23d6-df86-41e0-8da4-2b2b588b97e4 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Foundations of Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.495496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.495496Z digest=sha256:6cbf9848568eda4bd81f02c88def82aad2b51a88dd24bf57ba331838b0372ec9

Observation b52d45d2-daa9-4dc8-ac95-5d65a0b60939 · inbound

Generating Privacy Stories From Software Documentation cites this paper.

Generating Privacy Stories From Software Documentation Foundations of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:18.347548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:18.347548Z digest=sha256:904f5824b0495b4827a3863ed4edb7bb4166ba7799f683de550694d2a51520e9

Observation 47daccca-fc97-4a3f-a989-a7a8b31a624a · inbound

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models cites this paper.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models Foundations of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.949552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.949552Z digest=sha256:d607c5552d48af61367b49eb1f106c150a4d8b79d95900f0888c5d792367af68

Observation 01594dd7-3460-4f95-93ae-ea1c3cde774a · inbound

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents cites this paper.

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents Foundations of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:21:56.459147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T00:18:07.968897Z digest=sha256:d6d5bc51f95f1ad40aff1a8d1c248487efab37135e941eb3e9b1460105efd9e2

Observation 88c00a00-7b70-494f-a432-ab08a843f7e9 · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework Foundations of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.391626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.391626Z digest=sha256:3eee754d63f11e18d1830c42503736f0b897f4b14a73a2d4baaaa5f115ad0a0a

Observation be4c536d-c6ba-4cee-a93c-99685d87b9dd · inbound

Learning to Shop Like Humans: A Review-driven Retrieval-Augmented Recommendation Framework with LLMs cites this paper.

Learning to Shop Like Humans: A Review-driven Retrieval-Augmented Recommendation Framework with LLMs Foundations of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T13:27:24.371121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:27:24.371121Z digest=sha256:ca9c059c70e85a59678259b12532b99b8fdb7b7278e37392989f691d269c983e

Observation 8b4c78b9-1641-4f25-ba18-cf06f97543c8 · inbound

SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis cites this paper.

SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis Foundations of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:43.084821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:39:43.084821Z digest=sha256:d8292d41637cbb306fa0257e6b694e2d55bfebb64634bbccb35bba0be5a6e438

Observation 70c746fa-cb11-46f2-84c4-93380dbc4243 · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs Foundations of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:07.854305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:07.854305Z digest=sha256:4d7e638d7fb75fa493156f149f736a77f7c562b4d13e0c628833f326ae5b6573

Observation 30bcfcb8-f671-491d-9007-66561ebbf890 · inbound

LARAG: Link-Aware Retrieval Strategy for RAG Systems in Hyperlinked Technical Documentation cites this paper.

LARAG: Link-Aware Retrieval Strategy for RAG Systems in Hyperlinked Technical Documentation Foundations of Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:54.816671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T02:15:18.274343Z digest=sha256:a974379b41555df3e93d9605fbf569c1f3bbe6ced249c27fc22d05d5de16ebcf

Observation 8fb911df-5a39-40e5-96a6-af4bf824280d · inbound

Qiskit Code Migration with LLMs cites this paper.

Qiskit Code Migration with LLMs Foundations of Large Language Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-26T16:29:35.646049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T16:24:25.357338Z digest=sha256:359894922a20237ba70c5ce9df0cc0bff02539fb7013a4751386b9de170c7314

Observation a2a8bf91-4de4-42fd-8dbf-d4b73015e0aa · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Foundations of Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.682417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.682417Z digest=sha256:895bf99ae8237bb2a83d8314923063a894452aaba0ee059678bb4dd51e41e933