Pith. sign in

Paper Citation Record · LEDGER

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2509.03537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03537 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:17:13.821920Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a7a71da-eeb0-4562-b917-d3a509c04373 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.732880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.732880Z digest=sha256:f0858529ff9e6291c796d8ace4cad1154ef3cc8e777fc20f9154478f6b338f0c

Observation a8e3db9b-6914-4760-826b-fecc9e6486ec · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.738630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.738630Z digest=sha256:c3c5d9eb243f759b3bd0aedd3768a8359a82bd30e880fe8275bc530832c262f8

Observation b73d0ffa-457b-4b1d-85c5-b595229fe28f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.743935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.743935Z digest=sha256:21c055aa4b9ebb839c3d1c125e79c00301e16e4eb6fa00ca3df7cfe27c4d76af

Observation bbe70dc4-8e78-4d4e-88a8-b072a15b3a7e · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.749999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.749999Z digest=sha256:09cc0927776e345b4171fb13f1661d5a87093159888640d0bea44eec59d1fc06

Observation b8753fe9-fef7-4f22-9d23-1e4dc5afc420 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Competitive Programming with Large Reasoning Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.755616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.755616Z digest=sha256:98881c0388b0dbc4475cff7ba4054624a374cd2d81e02925161ce372f09593e1

Observation 397166e6-a4a4-44d1-8f37-0ffa783c02d6 · outbound

This paper cites an unresolved cited work.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:17:14.122287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T15:17:13.761076Z digest=sha256:dd1e2a170893c98270f277e353283773cf738414ad665a2f30cdb59794a262b2

Observation 16e4a717-7845-47c8-b931-7bb989f39dba · outbound

This paper cites Language Models Can Teach Themselves to Program Better.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Language Models Can Teach Themselves to Program Better

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.766596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.766596Z digest=sha256:981c7880ae492d3d189e0b1942aa3735acc2292bead17acf84e166a42dbd10f2

Observation f12c963a-9083-4579-8128-fc2233fb8adc · outbound

This paper cites Qwen2.5-Coder Technical Report.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Qwen2.5-Coder Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.771367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.771367Z digest=sha256:94132d68784e5eb4bd87073584b81ca197dcb82416773affd69e8b5c96bb7870

Observation 5e56cf67-6e4d-4b7c-92ac-30340b4a9d38 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.776157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.776157Z digest=sha256:371b86350348f5461e069e810f6fec11e7fe92828c7a641527db1fb774b52945

Observation 944c58a2-a88b-4d2a-9dbe-d3b2c24a94d1 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.781115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.781115Z digest=sha256:6f2f0b4811e5567e3f3ae12551c38a82e863bc73d1f9a46378be7837d0039d0b

Observation bf84de14-32ec-4fa4-bdbf-4020e10daa64 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.786271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.786271Z digest=sha256:60ea323fa886f99bd13ac337e28232488198eed2197865d8ffe737402149f68f

Observation c5759bce-d896-4232-b2df-88f3d6ad35f4 · outbound

This paper cites Qwen2.5 Technical Report.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Qwen2.5 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.791467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.791467Z digest=sha256:c2aff61c7c2d4276e8bfa5f055b30d4409170e07324c4a3d4861246f2184e85c

Observation 4cf570fe-e19a-413a-98c7-b137748981b8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.797081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.797081Z digest=sha256:af0d9322c87294a98a0d3fe6f52e59ecc099e8ed1ed158b6b63ef52e87154fad

Observation 826bf715-dc69-4193-92a9-815a966ded12 · outbound

This paper cites Execution-based Code Generation using Deep Reinforcement Learning.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Execution-based Code Generation using Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.801965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.801965Z digest=sha256:bf95f2792d8957cca22108c51681cde0dc4e380abe6dc4f0bc880e710356b311

Observation 44eb70b4-099c-4dfc-a3ce-bf381b2e2cd2 · outbound

This paper cites Sutton and Andrew G.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Sutton and Andrew G

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:17:14.107674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T15:17:13.807366Z digest=sha256:6e11055f605a4a3b9f983cf6ab7b9f761c615174caa54b936ed6b1fbd0e44ddb

Observation 4337c025-72a6-4375-9146-b13ac84b2cc5 · outbound

This paper cites Reinforcement Learning Enhanced LLMs: A Survey.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models Reinforcement Learning Enhanced LLMs: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.811953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.811953Z digest=sha256:01262c0bbfbedf020dccdcd12c92b9dfd3a64977c232e3acdafcc5d2adcf1e0d

Observation e33c6417-4a03-41ba-be60-dfa7156370dc · outbound

This paper cites the word is a subset of the puzzle’s characters).

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models the word is a subset of the puzzle’s characters)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:17:14.093333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T15:17:13.817245Z digest=sha256:713bb7e5497917345173addbcc2e3f593ddf0e19627b7759ec14f153fad56d40

Observation 3ece2b84-8724-4fa4-a6b2-6345d1d91dd0 · outbound

This paper cites includes the first character of the puzzle and contains all the characters from the puzzle.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models includes the first character of the puzzle and contains all the characters from the puzzle

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:17:14.077781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T15:17:13.821920Z digest=sha256:6ac7f034d52c0224ff59c9cbf7414f5c17d40f29e0161bfceca6c6fe4d8cacd9

Pith citing papers

No inbound Pith citation observations are available.