Pith. sign in

Paper Citation Record · LEDGER

Minerva: A Programmable Memory Test Benchmark for Language Models

As of 10 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2502.03358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03358 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:02:35.839991Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:42.748430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:45:42.844290Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29a4f6e6-c178-48e6-b19b-24dfe3bd3fc8 · outbound

This paper cites write newline.

Minerva: A Programmable Memory Test Benchmark for Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.661718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.661718Z digest=sha256:92b290df4e8e1332a649c321a7707de0c6eceb3f83e8a344ee89ef9461994a1d

Observation 763a9878-9177-475e-82e9-ad4738716eb4 · outbound

This paper cites L -eval: Instituting standardized evaluation for long context language models.

Minerva: A Programmable Memory Test Benchmark for Language Models L -eval: Instituting standardized evaluation for long context language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.667851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.667851Z digest=sha256:a0cc0c3f967d994f3734b8644ec2a22012634d0f5299c63928389b6db1a9f5dc

Observation 6acbe83c-b5c8-4b93-896d-a336bb61334e · outbound

This paper cites Introducing the next generation of claude.

Minerva: A Programmable Memory Test Benchmark for Language Models Introducing the next generation of claude

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.562838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.674479Z digest=sha256:ed26c42647aca147ac691f8046f280901c7e890ba814ebf92adf3900fb342e29

Observation 08565b99-00ed-4a4c-b455-f9a95c7fec0a · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

Minerva: A Programmable Memory Test Benchmark for Language Models L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.680215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.680215Z digest=sha256:67c679c13b6080258164e4eace49ba8d72eb0b6459609a66d19b3380f2a519d8

Observation 6310a15e-cfde-4e15-96c3-32b56d2e753f · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Minerva: A Programmable Memory Test Benchmark for Language Models Titans: Learning to Memorize at Test Time

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.685716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.685716Z digest=sha256:cbd871355ec3cd42410bdd7791c5279da72a29225f7878fc0794b8854449baea

Observation 937847f7-f07c-4cd1-b8cc-c8af2ce6897b · outbound

This paper cites an unresolved cited work.

Minerva: A Programmable Memory Test Benchmark for Language Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.692412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.692412Z digest=sha256:da7163aceef40e6fe16a3747a052ab6b5d3934e07ba961126980c5a19d5f5468

Observation 90abd7c7-3bdd-4023-9011-ef3566e3a131 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Minerva: A Programmable Memory Test Benchmark for Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.697882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.697882Z digest=sha256:e1a680e254d442368a3bbe09dd708191c70168dac10aa6864d96535709de6349

Observation e70f072c-0d17-47f7-ba8a-fd23d8a31ca8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Minerva: A Programmable Memory Test Benchmark for Language Models Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.703972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.703972Z digest=sha256:6afeec412fdb6be5da0f97b389f6af60113715cc99129cc0a22fbfd26fd15ca7

Observation ed6aa7b0-c19f-4618-a00b-8fa8bf02a145 · outbound

This paper cites and Parrish, J.

Minerva: A Programmable Memory Test Benchmark for Language Models and Parrish, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.533451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.709350Z digest=sha256:7d84ccb03acd2f50dda5da04de394fbf846d7061cb3e8644c332764e8bd857ad

Observation a22c946c-ec87-45d2-a08c-4164611990a3 · outbound

This paper cites L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M.

Minerva: A Programmable Memory Test Benchmark for Language Models L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.510293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.715006Z digest=sha256:319f33de671357f331785180aa521f651695f01631c72a977bd509fd550ed76e

Observation 0750460e-70b0-4218-9397-f65bcee6e05b · outbound

This paper cites Samsum corpus: A human-annotated dialogue dataset for abstractive summarization.

Minerva: A Programmable Memory Test Benchmark for Language Models Samsum corpus: A human-annotated dialogue dataset for abstractive summarization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.720104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.720104Z digest=sha256:8d9cf8a71e15d1513271b41bf38cbc6b4f3406cc62c0a5bdc80f6cb0a9ced660

Observation 7224ee54-4e77-45f5-9ef3-1dc6cabf080e · outbound

This paper cites Measuring massive multitask language understanding.

Minerva: A Programmable Memory Test Benchmark for Language Models Measuring massive multitask language understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.725514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.725514Z digest=sha256:fef149bdb682bc33dfbd4d6a8e92eb1acf2e0197ac41292cacfa39e02db37fe9

Observation eed9a781-5a8d-4a1f-9841-82711df62484 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Minerva: A Programmable Memory Test Benchmark for Language Models Measuring mathematical problem solving with the math dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.478963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.730308Z digest=sha256:f2d3af30a43e024736b68b3af720ffc913ebd4f4846189920e0cd0bfabdabf7f

Observation fda45a11-ae4b-4651-a8b3-ad2602087d08 · outbound

This paper cites RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024.

Minerva: A Programmable Memory Test Benchmark for Language Models RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.735311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.735311Z digest=sha256:37eaeb56bd337304912b98c1ab30d1c3b0a0486695796c35cb8543f6c0c71623

Observation d3cd2227-cbaf-4606-ba62-292a8c05db52 · outbound

This paper cites \'E tude comparative de la distribution florale dans une portion des alpes et des jura.

Minerva: A Programmable Memory Test Benchmark for Language Models \'E tude comparative de la distribution florale dans une portion des alpes et des jura

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.448313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.740545Z digest=sha256:2312dcb5d16c3a3ecd44a66301ca37b504ea091e696a3778c6b4e59cac92f443

Observation b62547bc-e52d-47e2-99ff-837bf8f17546 · outbound

This paper cites Needle in a haystack - pressure testing llms.

Minerva: A Programmable Memory Test Benchmark for Language Models Needle in a haystack - pressure testing llms

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.427582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.745844Z digest=sha256:f19ea75bd9066f9a22e2962d82388e99eccd8fbf1dd5c7e147cd7a87772d640b

Observation 92b11aeb-d4b1-456a-bd01-8b0ef69f66d3 · outbound

This paper cites L., Roebuck-Spencer, T., Short, P., Kabat, M., and Wilken, J.

Minerva: A Programmable Memory Test Benchmark for Language Models L., Roebuck-Spencer, T., Short, P., Kabat, M., and Wilken, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.408571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.750629Z digest=sha256:558c04aef62fd4c818b5e238988c357011bb242773a2e8d1c07fcd686fe7e9d7

Observation 58bd7cd9-a8d3-43e2-8c16-a4a8a8c9ecdf · outbound

This paper cites Natural questions: a benchmark for question answering research.

Minerva: A Programmable Memory Test Benchmark for Language Models Natural questions: a benchmark for question answering research

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.755897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.755897Z digest=sha256:f0342ca98f79da186800bd8ef34af1582ac97b3390cdd73da35b0715ee78a815

Observation 3b95b20b-a29c-406f-ba3b-a6af96dcd958 · outbound

This paper cites Needlebench: Can llms do retrieval and reasoning in 1 million context window?, 2024.

Minerva: A Programmable Memory Test Benchmark for Language Models Needlebench: Can llms do retrieval and reasoning in 1 million context window?, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.760913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.760913Z digest=sha256:1c41a576246408ba18a82083d9a1ebf5604d036b82457b9e4b40d1a41024cfd6

Observation c66f5e6f-780c-4726-aa8a-31d26277e4c8 · outbound

This paper cites ROUGE : A package for automatic evaluation of summaries.

Minerva: A Programmable Memory Test Benchmark for Language Models ROUGE : A package for automatic evaluation of summaries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.766110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.766110Z digest=sha256:d3a53bf3de50b6bfe9f6ac7ffd98dacdc1030ca0431a48b96a602886e75906aa

Observation dda23753-5f5d-4645-b940-0c2720df951f · outbound

This paper cites F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P.

Minerva: A Programmable Memory Test Benchmark for Language Models F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.771083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.771083Z digest=sha256:5ca75f74a0de440f827dc5b69219dba16839e034af954a3dea7b4d59f934c09b

Observation 075ecba1-1990-45a6-95de-eda68cce2c49 · outbound

This paper cites P., Santorini, B., and Marcinkiewicz, M.

Minerva: A Programmable Memory Test Benchmark for Language Models P., Santorini, B., and Marcinkiewicz, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.368502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.776578Z digest=sha256:9346d757cb9a3500f41d4355ceedcf9976a8ed92a2f409fca4dfa86fc3e5ae1c

Observation a90cca0d-d06d-45cf-920c-a8fa6bbfb01d · outbound

This paper cites and Jaggi, M.

Minerva: A Programmable Memory Test Benchmark for Language Models and Jaggi, M

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.781705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.781705Z digest=sha256:9be462562b8f300c139218fd3a7cafb149fc04f1fac160b2770de1b1648c3f3d

Observation 13cf777e-9cc4-456f-901c-e6f1749dcc6a · outbound

This paper cites S., Phillips, N.

Minerva: A Programmable Memory Test Benchmark for Language Models S., Phillips, N

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.340785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.786894Z digest=sha256:0f3d39073f7f8ea8562963a54ec9e853a9077a0809b48b09c5b4c4a2016441a9

Observation d77c31fd-0d56-48ac-8840-c9e267636b88 · outbound

This paper cites Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models.

Minerva: A Programmable Memory Test Benchmark for Language Models Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.791965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.791965Z digest=sha256:3a36b5881783b6f0e01c994b7e5675cceb93094be4cea5c83d1a64f2c60ed006

Observation 5c3ee4f6-2317-4594-92e0-52fca4eafe53 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Minerva: A Programmable Memory Test Benchmark for Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.797402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.797402Z digest=sha256:981e13d3c5dfaa7382187330081d992029660d8b7354994b2c16c48b3a5d8c7e

Observation dc5541e6-dfac-484e-bd56-d2c606e70a94 · outbound

This paper cites N., Cowan, N., Hitch, G.

Minerva: A Programmable Memory Test Benchmark for Language Models N., Cowan, N., Hitch, G

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:02:36.322784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T05:02:35.803868Z digest=sha256:73d02cb8e0f3ab082cd82a07d3cbb246b320ea6161f1838e91a71b7f59394572

Observation 94587577-eda0-49d7-a245-d4e5a3b5998a · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Minerva: A Programmable Memory Test Benchmark for Language Models Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.809700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.809700Z digest=sha256:0856dc2e5c77a1d9dd92a418ec2d994d5adbb794142c5f1064f3f0b9bbc5250a

Observation 3b73982e-464e-4b68-bf70-1544973db5c3 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Minerva: A Programmable Memory Test Benchmark for Language Models AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.814615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.814615Z digest=sha256:4911f7029d7fd2c6c72814c3c802d866467a7d4239a5be831115c830b60ed07c

Observation fdb49798-00d0-4640-aa9b-56a036f11a63 · outbound

This paper cites LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs.

Minerva: A Programmable Memory Test Benchmark for Language Models LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.819407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.819407Z digest=sha256:638de1560a60c19f96211ddde9b7250a9d5f17efc69c084d3d3e3b7c89b1824c

Observation 39c83ac0-c8f2-4c4c-b025-86d7bc3dfe3d · outbound

This paper cites Effective Long-Context Scaling of Foundation Models.

Minerva: A Programmable Memory Test Benchmark for Language Models Effective Long-Context Scaling of Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.824173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.824173Z digest=sha256:4a4de6c45fd83fc1f417426cd8a786dc7ba5ce8b8a556515800ed98cb76c6c06

Observation e259a5de-34a1-4afc-a0b0-281e0397fc96 · outbound

This paper cites Inftybench: Extending long context evaluation beyond 100 K tokens.

Minerva: A Programmable Memory Test Benchmark for Language Models Inftybench: Extending long context evaluation beyond 100 K tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.829428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.829428Z digest=sha256:e3a3e474f9e70380eb5f59acecc7e83cf5db5cffca1dab031f7c309738094f16

Observation 747ff55f-2722-43ff-9a37-ce7e8d5c38f2 · outbound

This paper cites E., and Stoica, I.

Minerva: A Programmable Memory Test Benchmark for Language Models E., and Stoica, I

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.834640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.834640Z digest=sha256:5b2bfbb3abe66cd5f9cc28ae936f19ccd2f53b07354b29121d3cd5c86ad9c3d1

Observation 04cf8462-dd02-4483-b395-4a7ed7672cdc · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Minerva: A Programmable Memory Test Benchmark for Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.839991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.839991Z digest=sha256:34390cee0d92e12922c9444a857e43ae0e3214e7e976711be5cc2abf94023687

Pith citing papers

Observation 91e16ef8-01f0-4903-8d7c-ca63f47db4b5 · inbound

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models cites this paper.

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Minerva: A Programmable Memory Test Benchmark for Language Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:45:42.847863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:45:42.748430Z digest=sha256:00a86c7436431f29a6e5129e009502fc2eec9f0525bdcb97a9b20cb0ead48a42