Pith. sign in

Paper Citation Record · LEDGER

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

As of 15 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 6 inbound Pith citation observations for arXiv:2506.17335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17335 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:19.859577Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T12:32:39.706055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.838826Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57159559-538c-43d7-8ad9-a324dc09a8fe · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.836512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:19.396084Z digest=sha256:973851dd1fdd78e668fd67413c09948408af0d2a8fe9f137ba8b727646a85715

Observation 87d77231-adf4-41bb-96d3-b30f3d6139cd · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.673889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:19.511451Z digest=sha256:15185d2182828a3103be3c1cdded6508770e1f445c3579319aea21d8c26ce8ef

Observation f33282d9-4110-4476-a0ca-49dfe6ab5815 · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.252786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:19.739258Z digest=sha256:a9351efcc540053e1de89abf2f883e43fb65c86caf873f47568d81c0bc72dc39

Observation 883b6106-d5c8-4d33-95d9-44f1c38e3cb9 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.113797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.113797Z digest=sha256:b92fdd075442914ddc35a555869cc60d4aa409ef1776c2d3494e658b9e20c7f7

Observation f9a25c14-4a9a-4ea1-a8db-50e3ec9c6a9b · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.213148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.213148Z digest=sha256:9bb11bc907b67a69a4c5e19f06ba4800d7cc47317a6954ee19d0936ba49734f3

Observation af802a7a-dda6-45ba-8d27-98008dd0d2fc · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.981951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:19.322363Z digest=sha256:9f6bcb5b12a3179efacffd194d2b38c6ae3bb17c0a02fc7fb5ee48bf4a0fd760

Observation baee7219-cc49-4df9-ba1a-11a586e082d7 · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.547627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:19.592820Z digest=sha256:6059de26e37651ebf97871cb6a0f0ef51f52821888a3e7d5c1843ef5f4155229

Observation 09dd6fd4-f00d-4c73-939f-63316f0f3191 · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.400595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:19.670576Z digest=sha256:7daf9d9d0426438a20c19ec16164bfaea6c2ac9979ac678caed5a0ac36f3d70e

Observation 8b116405-54ef-45cd-bdcf-025233dcdeb7 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.859577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.859577Z digest=sha256:d3b4bab032b2416ec7792dfa463ad1d102abf1c40132cadc57617d0153eecfc2

Observation bc2862dc-3e49-4de9-b15c-44655210cd94 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:18.855974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:18.855974Z digest=sha256:020b91e9552d3bc1ac9636fd1cbec7e65b3078b6cf13cd33bbf2edb6893cc08e

Observation 14d38a06-45a6-4507-820f-6cbd5426ef1b · outbound

This paper cites Qwen2.5-Coder Technical Report.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Qwen2.5-Coder Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.035337Z digest=sha256:82cd71cb81799e17ddd851421a98e6627aeb3b23f682425fdf5e1f1d154bc1a4

Observation be67592c-3bc9-444c-85d6-8b9f3e0caeba · outbound

This paper cites Scholar Inbox: Personalized Paper Recommendations for Scientists.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Scholar Inbox: Personalized Paper Recommendations for Scientists

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:51:20.085665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:51:18.942033Z digest=sha256:26367c6168a1491448a452aa9d9da01e8870b5973bc65ee220fbdc8bf833047f

Pith citing papers

Observation e6d145cd-0a8e-4942-83c0-02e112f8beab · inbound

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences cites this paper.

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:06.915158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T02:01:07.555124Z digest=sha256:be70d35cc18404b2910fd35e34eca3b636379be1406bba2d4bd8f73ba928b93f

Observation 60fbff89-35b2-4fa3-bcb4-5cb540e0de55 · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.510020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:e6a2679b3abe0124ce2d7810ef359840f3fd2ebc48841829e0004d1a2d66dee3

Observation 3ced7e92-6834-4b28-94c4-b79fb9cdeeea · inbound

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents cites this paper.

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:57:53.347049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T23:54:15.987953Z digest=sha256:7126944a3125ca019985b00c79e74b99e0f9c058c6e044423f2e1f740a472e69

Observation 0d1807e1-3a13-4bf4-93c1-7555d34ff371 · inbound

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation cites this paper.

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.748637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:39:00.460842Z digest=sha256:ad12ce43ffef541963596232464a34d14c6e6e5f9361e22343d165d82553f929

Observation d85ef8c4-0c1f-483c-b4cc-b16d2bd675b5 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.840731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T00:13:14.940915Z digest=sha256:cda3bff2a83b131120b7e6e2073d91031ba245b90f3dc72dfa95ad32170d3843

Observation c805db18-cecd-469b-8f1a-542ec7d94869 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-12T12:32:39.706055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:32:39.706055Z digest=sha256:021c700fef795d446e01fcee1bb7a515765b976bb98296d9d25c4c1a21ff9f36