Pith. sign in

Paper Citation Record · LEDGER

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop

As of 17 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.23002.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23002 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:57:40.804853Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c72079c3-cb88-4ff8-bbd8-11d56d9f99b7 · outbound

This paper cites HITS: High-coverage LLM-based Unit Test Generation via Method Slicing.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop HITS: High-coverage LLM-based Unit Test Generation via Method Slicing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.762460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.762460Z digest=sha256:ede643fba1cf77eca181717cfeaa9b600da9178e637f74cc15150ead57be8149

Observation 77756427-eb8a-4b8b-b500-35275fd6bb24 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop LLM Evaluators Recognize and Favor Their Own Generations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.766985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.766985Z digest=sha256:b856068362d0c466fc72ab023594de7ed28eceae26169f9f6464a24308d42e1f

Observation d1890c78-32a4-43c5-9ac1-cc910ac0fc50 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.771519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.771519Z digest=sha256:62c416d939aaed3e73a68344e52566d62b0abb22b89e79d4cd09b56d00910136

Observation 5edd23c8-e32f-443d-abcc-7c43b8f7eec3 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Great Models Think Alike and this Undermines AI Oversight

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.784183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.784183Z digest=sha256:40e5083b1ebf3cc9e0462950f96984f932f2c9e236e9a223afe815e482bd18a0

Observation b96d9aec-229d-49c0-8c0f-0037c26bb2c9 · outbound

This paper cites TestForge: Feedback-Driven, Agentic Test Suite Generation.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop TestForge: Feedback-Driven, Agentic Test Suite Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.788358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.788358Z digest=sha256:7ba6d342b598f30dc171975e7073eff851b64b5dbcc065e67fc69664de4af10b

Observation ec47ea65-66f6-4941-ac0c-7d60aa9fefde · outbound

This paper cites Registered Reports in Software Engineering.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Registered Reports in Software Engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.792701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.792701Z digest=sha256:403839996cec6ead98f807186f595ea796f893f15d2e8d76ce87cc1b306bd482

Observation 22056d44-b81a-4ae6-bab1-f2b8f129ee90 · outbound

This paper cites Fuzz4All: Universal Fuzzing with Large Language Models.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Fuzz4All: Universal Fuzzing with Large Language Models

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.804853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.804853Z digest=sha256:a7e85424e775ecd94393ec948f2b8bc6993c7320724ebf56b1b3745ad6938463

Observation 0b038097-1569-4f96-8f5a-a1e3dc6c6fdf · outbound

This paper cites Mutation Testing Advances: An Analysis and Survey.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Mutation Testing Advances: An Analysis and Survey

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.796638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.796638Z digest=sha256:6f494909d03dee47cb7de7f41bc11ade0dc02586d0e29580ee7810dba7ffd026

Observation 492a5aa9-ba2e-4100-ae62-220eb02b0a47 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.800920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.800920Z digest=sha256:a72d8d743ffae19a7ab4c72baea8ed39581b5775c2f0c9ac3dee626462c3ea74

Observation 9d0c53ab-c15b-4822-816e-caade29f8723 · outbound

This paper cites Pynguin: Automated Unit Test Generation for Python.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Pynguin: Automated Unit Test Generation for Python

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.758128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.758128Z digest=sha256:ae9bb3824c42df779cab96ae78a063ff91c6be6b458a2481479ff66864db8186

Observation 693eb81b-ba9b-48e2-aa95-db5bfc04361a · outbound

This paper cites An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.753104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.753104Z digest=sha256:146367bd9225b9095a843f323578066757bd2dca9819a37836053c9c42f6919f

Observation ba55738c-e6e6-46c6-b77c-84eb57d5b856 · outbound

This paper cites Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.775727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.775727Z digest=sha256:b983967967413a03db5004d577dd2ec2b25070a3c481450fa1a771fb04d75e21

Observation e272ef24-f536-4312-a1da-6ec30669cf8b · outbound

This paper cites Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.779886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.779886Z digest=sha256:913b6118b714731eac17b20d8d95b3599f0ad10dab702ce2271702e60d719185

Pith citing papers

No inbound Pith citation observations are available.