Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:02:35.839991Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2502.03358.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:02:35.839991Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:42.748430Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T14:45:42.844290Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 29a4f6e6-c178-48e6-b19b-24dfe3bd3fc8 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763a9878-9177-475e-82e9-ad4738716eb4 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models L -eval: Instituting standardized evaluation for long context language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6acbe83c-b5c8-4b93-896d-a336bb61334e · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Introducing the next generation of claude
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 08565b99-00ed-4a4c-b455-f9a95c7fec0a · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models L ong B ench: A bilingual, multitask benchmark for long context understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6310a15e-cfde-4e15-96c3-32b56d2e753f · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Titans: Learning to Memorize at Test Time
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937847f7-f07c-4cd1-b8cc-c8af2ce6897b · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90abd7c7-3bdd-4023-9011-ef3566e3a131 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70f072c-0d17-47f7-ba8a-fd23d8a31ca8 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6aa7b0-c19f-4618-a00b-8fa8bf02a145 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models and Parrish, J
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a22c946c-ec87-45d2-a08c-4164611990a3 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0750460e-70b0-4218-9397-f65bcee6e05b · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7224ee54-4e77-45f5-9ef3-1dc6cabf080e · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Measuring massive multitask language understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed9a781-5a8d-4a1f-9841-82711df62484 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Measuring mathematical problem solving with the math dataset
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fda45a11-ae4b-4651-a8b3-ad2602087d08 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cd2227-cbaf-4606-ba62-292a8c05db52 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models \'E tude comparative de la distribution florale dans une portion des alpes et des jura
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b62547bc-e52d-47e2-99ff-837bf8f17546 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Needle in a haystack - pressure testing llms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92b11aeb-d4b1-456a-bd01-8b0ef69f66d3 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models L., Roebuck-Spencer, T., Short, P., Kabat, M., and Wilken, J
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58bd7cd9-a8d3-43e2-8c16-a4a8a8c9ecdf · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Natural questions: a benchmark for question answering research
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b95b20b-a29c-406f-ba3b-a6af96dcd958 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Needlebench: Can llms do retrieval and reasoning in 1 million context window?, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c66f5e6f-780c-4726-aa8a-31d26277e4c8 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models ROUGE : A package for automatic evaluation of summaries
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dda23753-5f5d-4645-b940-0c2720df951f · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075ecba1-1990-45a6-95de-eda68cce2c49 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models P., Santorini, B., and Marcinkiewicz, M
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a90cca0d-d06d-45cf-920c-a8fa6bbfb01d · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models and Jaggi, M
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13cf777e-9cc4-456f-901c-e6f1749dcc6a · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models S., Phillips, N
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d77c31fd-0d56-48ac-8840-c9e267636b88 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3ee4f6-2317-4594-92e0-52fca4eafe53 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc5541e6-dfac-484e-bd56-d2c606e70a94 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models N., Cowan, N., Hitch, G
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94587577-eda0-49d7-a245-d4e5a3b5998a · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b73982e-464e-4b68-bf70-1544973db5c3 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb49798-00d0-4640-aa9b-56a036f11a63 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c83ac0-c8f2-4c4c-b025-86d7bc3dfe3d · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Effective Long-Context Scaling of Foundation Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e259a5de-34a1-4afc-a0b0-281e0397fc96 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models Inftybench: Extending long context evaluation beyond 100 K tokens
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747ff55f-2722-43ff-9a37-ce7e8d5c38f2 · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models E., and Stoica, I
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04cf8462-dd02-4483-b395-4a7ed7672cdc · outbound
Minerva: A Programmable Memory Test Benchmark for Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e16ef8-01f0-4903-8d7c-ca63f47db4b5 · inbound
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Minerva: A Programmable Memory Test Benchmark for Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.