Pith. sign in

Paper Citation Record · LEDGER

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

As of 19 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 6 inbound Pith citation observations for arXiv:2506.17335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17335 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:19.859577Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T12:32:39.706055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.838826Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57159559-538c-43d7-8ad9-a324dc09a8fe · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.836512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:19.396084Z digest=sha256:7a02a628f7148498e451c3a623d405fd3bb33356d5dff1df75fc66e871a773d6

Observation 87d77231-adf4-41bb-96d3-b30f3d6139cd · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.673889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:19.511451Z digest=sha256:d4b911f96c21eec45ea6f21324ffb47e8a133c1deb5a0c8dc4927a9ab82da6e6

Observation f33282d9-4110-4476-a0ca-49dfe6ab5815 · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.252786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:19.739258Z digest=sha256:ed8cf464638a720d08cd279e175ab400693a3dc10adda7406c1727fd28dba5bb

Observation 883b6106-d5c8-4d33-95d9-44f1c38e3cb9 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.113797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.113797Z digest=sha256:2220e24062b09cba8246682a3ef680fa17ece4690f09f92a5ac6bb48ec436972

Observation f9a25c14-4a9a-4ea1-a8db-50e3ec9c6a9b · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.213148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.213148Z digest=sha256:9bb11bc907b67a69a4c5e19f06ba4800d7cc47317a6954ee19d0936ba49734f3

Observation af802a7a-dda6-45ba-8d27-98008dd0d2fc · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.981951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:19.322363Z digest=sha256:a2afe5494830a7d09b632f11c5c673c44d737a8816ba965617ec8cd7da1381c4

Observation baee7219-cc49-4df9-ba1a-11a586e082d7 · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.547627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:19.592820Z digest=sha256:4f7369e7b57f3518bdfe01a0d17c445b295295eca33556b43e0e59d0154f60b6

Observation 09dd6fd4-f00d-4c73-939f-63316f0f3191 · outbound

This paper cites an unresolved cited work.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:51:20.400595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:19.670576Z digest=sha256:17b9bcfe076d020b46d6eeb47edea817e155afdaad549b709d98eb62b19c3480

Observation 8b116405-54ef-45cd-bdcf-025233dcdeb7 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.859577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.859577Z digest=sha256:d3b4bab032b2416ec7792dfa463ad1d102abf1c40132cadc57617d0153eecfc2

Observation bc2862dc-3e49-4de9-b15c-44655210cd94 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:18.855974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:18.855974Z digest=sha256:020b91e9552d3bc1ac9636fd1cbec7e65b3078b6cf13cd33bbf2edb6893cc08e

Observation 14d38a06-45a6-4507-820f-6cbd5426ef1b · outbound

This paper cites Qwen2.5-Coder Technical Report.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Qwen2.5-Coder Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.035337Z digest=sha256:82cd71cb81799e17ddd851421a98e6627aeb3b23f682425fdf5e1f1d154bc1a4

Observation be67592c-3bc9-444c-85d6-8b9f3e0caeba · outbound

This paper cites Scholar Inbox: Personalized Paper Recommendations for Scientists.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research Scholar Inbox: Personalized Paper Recommendations for Scientists

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:51:20.085665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:51:18.942033Z digest=sha256:23671a940fb5a76163d586a875e8f8fcaafe09a6fb48d5b80979012a60b5d323

Pith citing papers

Observation e6d145cd-0a8e-4942-83c0-02e112f8beab · inbound

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences cites this paper.

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:06.915158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T02:01:07.555124Z digest=sha256:542cb1838249828c47da3559719a9977e79b584c1d342aa10dff1ae2c0065265

Observation 60fbff89-35b2-4fa3-bcb4-5cb540e0de55 · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.510020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:daea77ea1db2cd8ab6e791bdf9556789872e868c3b12fa21e8b00ccb8c0eea7b

Observation 3ced7e92-6834-4b28-94c4-b79fb9cdeeea · inbound

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents cites this paper.

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:57:53.347049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T23:54:15.987953Z digest=sha256:4ef412a31f35fe11f55744925b924846182d6f4b67dc078bb09961a8a6db3815

Observation 0d1807e1-3a13-4bf4-93c1-7555d34ff371 · inbound

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation cites this paper.

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.748637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T12:39:00.460842Z digest=sha256:b1cd25cad40981e9cbad34890e370dfddf66d7e356baebc661dccead24deeda2

Observation d85ef8c4-0c1f-483c-b4cc-b16d2bd675b5 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.840731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T00:13:14.940915Z digest=sha256:afea830c06beb00d5bba128be502a80b586792009eb180d26cb638ef340074d0

Observation c805db18-cecd-469b-8f1a-542ec7d94869 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-12T12:32:39.706055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:32:39.706055Z digest=sha256:021c700fef795d446e01fcee1bb7a515765b976bb98296d9d25c4c1a21ff9f36