Pith. sign in

Paper Citation Record · LEDGER

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

As of 10 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.06301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06301 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:25.819108Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c5f7cd5-2b14-40f1-a771-e60121272912 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.840406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.616441Z digest=sha256:d92bb000abe3a3eb643478c6bf544cde70fb22ad89e943383ebffbe2022f447c

Observation 84c6ef51-45d5-495f-9a85-66d6b116523e · outbound

This paper cites Program Synthesis with Large Language Models.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.621540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.621540Z digest=sha256:0ec6d1cf066522f806ca063d2ceee2a5f99e812b2ad39afbe288bf0b8cd8269c

Observation 6b862b9b-a108-44ff-b7c2-bfc126a7f381 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.626627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.626627Z digest=sha256:6178cdce66a287307ecaa4868fb59b725d812d78006e871c55081afe79bb9b17

Observation 5fecd972-8608-4139-bedb-a94f5cec91d0 · outbound

This paper cites HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:02:26.542719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.639914Z digest=sha256:4cc05c61ed1cf9fc88a3f3efc7f57db17c667344f9eba62f1673be9c9a1658a2

Observation 06ae2d74-049d-4f54-a113-92f478ceaf97 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.823414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.644768Z digest=sha256:599a3fb31ffc344a9b4a55849dc3ce5fe90ae0cae387cf77467dfcb5a60872ef

Observation 7e24c6bc-e978-4394-a654-b4a3c2cfc91a · outbound

This paper cites Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.649833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.649833Z digest=sha256:18fd89137431371657637f61756e3e0b30c556d7c65fdb4a7463aa5a088c2f19

Observation 1ad6944b-e18e-41d3-b5e1-095ed755382e · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.655395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.655395Z digest=sha256:7c3542616b3fb11395e49af129e1ae1d1ab63c0c4a1a7bf36823aeea80fbbacc

Observation 4e70b33f-ef19-4b64-b8e2-73e9f5469edd · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.792632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.659935Z digest=sha256:e9024b26ed02680654d596df27406f8410e6989500db6f7c2258966d51d6e5ed

Observation 95aa86d1-88d4-415c-97c5-7a24d9379dfd · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.668797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.668797Z digest=sha256:214ac009feeb147d8e8cec412c7872270a3f807590b4b072b5c4e681dbe089ae

Observation 72497daf-5131-42b4-be49-78a98649a566 · outbound

This paper cites ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.673512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.673512Z digest=sha256:f6832eb0babf21e8d92f20ae9498dc01e2e8538fb043595fd172448662f1f1a7

Observation b114c3e1-3999-451f-90ca-dd70996b620e · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.678912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.678912Z digest=sha256:527c4ec067f6a1d64afdebab591025fb9a97055d24f0ef8d915cc56b52b18884

Observation 45050dee-36a5-4044-bb03-177129addcc5 · outbound

This paper cites Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.684162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.684162Z digest=sha256:4c0be7d7af850e58ce951489374dcb6b121f5b6955827e945de1c972ece415b8

Observation 59ab5bed-bff7-4a79-bee9-45284fe28a2e · outbound

This paper cites RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.689165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.689165Z digest=sha256:b5e325fa6122478984679148a01d31461a333dab5a6ee098436bc68bcd137a62

Observation 7a5ea708-3a4e-48e5-a781-7d8f087d9d65 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.694406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.694406Z digest=sha256:0a1ea62d5d91b0cd5b3897897ea19c1abe03c2294098e33fd55498d98187a7ca

Observation 832e847c-45be-4983-ae0d-e267fdca825c · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization GAIA: a benchmark for General AI Assistants

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.699110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.699110Z digest=sha256:5db90cfe13b8ef259a2ca520fed6dd2a7352110da74ce8d5632b3cb35b081d25

Observation ac423a22-3cfe-41b3-918c-a10db151494c · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.704618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.704618Z digest=sha256:611281a27be44fbfdd29a88ec38ea936c8ee591226374700991f5b646ec93d09

Observation 6d6365fd-311e-4770-9698-675129b7400d · outbound

This paper cites Opsahl-Ong, A.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Opsahl-Ong, A

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.710081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.710081Z digest=sha256:89acfbbd8d7e87dcd6cf7d8c180301f60b9fc497492f1ee41b9c5bfa503f6244

Observation ad972446-d52d-4349-a01d-73c312377a8b · outbound

This paper cites Ouyang, S.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Ouyang, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:26.773465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.714285Z digest=sha256:144cfe239c5cf7a40a3fd7e6ec7a25986d33d0df3dce28347b5c20909f8bcd71

Observation 6a3d847d-bc83-419c-ba9a-46bfdafc02e9 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.718442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.718442Z digest=sha256:b1596f4e94705ef57a39888baeb0867dd8b807f99e1e4a278be11c71e4af36bc

Observation f8c1677d-df60-4c64-b1d0-0381581b932d · outbound

This paper cites Romera-Paredes, M.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Romera-Paredes, M

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.731989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.731989Z digest=sha256:ccd8d02cdfd85b69e312538fb64cb92f476efb3f9c6c03f91d6c27da890c1218

Observation 72fcfa4e-52d1-4680-a932-1113a7b1aafa · outbound

This paper cites Ursekar, A.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Ursekar, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:26.755690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.736276Z digest=sha256:0ae4fb28a6fc0ae9cd1e2dbea2b67756c506beff282fafb37b6637b3d8190786

Observation 8c554198-4a8f-4996-add4-14b687b00309 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.737554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.740882Z digest=sha256:e95e2bd34c5f328ce7fccfbe4a9584132ceadcc623a4432a3d97075e6140418f

Observation 6b69b85f-28e1-4de3-95f4-1ebd8cd4499e · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.745383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.745383Z digest=sha256:b4483b891e3b1e5e121d4129e84104fe0947a2e789303e225ee699e7e727a077

Observation 5957b07f-23e7-492c-bfc7-8ae6bb7b80b7 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.716924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.749972Z digest=sha256:3cb31b239bfa648a0af82e6575246d55e26896612f73a9c9ba49c3132acd91f1

Observation 5cadbf08-4f74-4106-b420-6f947be9a055 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.692576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.754561Z digest=sha256:9b0ecd2afa93641b8975d25c3b36e90ced2cb1d6cef26f02beb2817dc4f0db32

Observation c3ee986a-611a-463a-941e-2e856cb77cad · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.759327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.759327Z digest=sha256:5931844edb3e705d22bc496af0d915910352fd16eaa522c885f027da11f5738b

Observation db2fe2c1-8763-4894-8e1e-9ca36870d67f · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:26.671001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.763834Z digest=sha256:713d15e8fa8e62fa82f5664e439b6bd3dc9ade23b1136417c3a8c9d492b0bcc1

Observation f0aa53e8-970e-4a68-b894-3261e7658649 · outbound

This paper cites Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.768245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.768245Z digest=sha256:d81d72ea26d701545d43c71c7533729004f20fa055def1f0c6c09e67531b7dad

Observation 07ccd406-2445-4169-8b89-6f761f2e1fb8 · outbound

This paper cites an unresolved cited work.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.772900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.772900Z digest=sha256:59f8936746b49d2c166408a642302c54bf999c55891d3ff7c2fe8cd21bb5940f

Observation 5853d2ef-5886-4c79-908a-c3d9a07ba557 · outbound

This paper cites G\"odel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization G\"odel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.787406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.787406Z digest=sha256:1e8a716181e71d046cd6cdb0c60c10c4df32b70029af6512bd3c1a377a22b748

Observation fd38340b-54e5-4194-bb30-b1108b5bc95d · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization TextGrad: Automatic "Differentiation" via Text

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.793445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.793445Z digest=sha256:ee67f1be7dc2709486e41b5579e9b710adfee89a92c24f4737bb4532a53aa35c

Observation 8b7b2ce4-092c-41fe-8d39-6e49194a1169 · outbound

This paper cites Zelikman, E.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Zelikman, E

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:26.647834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.799359Z digest=sha256:924b9dffdfd41119e08a3ab9e6c3f415249175e7f3a8378be93b8a74651999b4

Observation fc3bd319-50f2-4c22-b079-517dfd5d96b7 · outbound

This paper cites LLM-AutoDiff: Auto-Differentiate Any LLM Workflow.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization LLM-AutoDiff: Auto-Differentiate Any LLM Workflow

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.783053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.783053Z digest=sha256:addb36f60b42f601ee036cc1568de5f3ee70dc8d10e6e7d19b684b64d8d9388b

Observation e86ce37a-0105-4da9-9a23-55caca814dba · outbound

This paper cites Zhang, J.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Zhang, J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:26.629708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T06:02:25.809069Z digest=sha256:8a814565adbc200538a219a7193d39648fe0685585d3f042109ea01ba10c1652

Observation 0382457f-3707-41e7-b59a-3175ac6c7b7c · outbound

This paper cites Stop Comparing LLM Agents Without Disclosing the Harness.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Stop Comparing LLM Agents Without Disclosing the Harness

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-07T06:02:25.819108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.819108Z digest=sha256:f9ff361ddbaadc514417ef5a269e945564632b9fe91aca1387a908e22dd12d4c

Observation 73f9813b-71b4-42ab-adfc-fa76a52057a8 · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.804071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.804071Z digest=sha256:17954c5626dad41bb6cbdc17284e45d28400507a287f3cf8c6aadfedd7661ee3

Observation 9aa97db2-102f-483b-9dd4-c900c79cfe0e · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization AFlow: Automating Agentic Workflow Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.813777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.813777Z digest=sha256:9150da93fa78d5ca7348bfc7a5f47bf227ab32af3c1b7e4d5151e512188f9c1e

Observation f1e92752-39a5-4a6e-98a4-f093a899e2ea · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.635628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.635628Z digest=sha256:0afccb9a8332e67329817f4fe8e9016a81234cc9378b718e1b7f123ecda23751

Observation 234b9e0a-50f4-453b-8d80-4979d8b30ea7 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.664469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.664469Z digest=sha256:187866d724e92aac2aa6e9046720f1c7840ccd974373ce85471b0e137f38c232

Observation 5ae73031-9670-42b4-b8d3-9d6859d5658e · outbound

This paper cites A Self-Improving Coding Agent.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization A Self-Improving Coding Agent

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.727335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.727335Z digest=sha256:223ca8cc863bb977351052e728d405dc3ed6e29330b460e0f8fa1a8bb58eca3b

Pith citing papers

No inbound Pith citation observations are available.