Pith. sign in

Paper Citation Record · LEDGER

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.19408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19408 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:15:19.895698Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1612773-4807-4786-bf15-1d21f04b9e95 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Dense reward for free in reinforcement learning from human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.155018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.155018Z digest=sha256:382b82e8ce036ed3058d7d7dbf2b829cf603d575e61878ec40185cee8615200a

Observation 40d49530-339c-4270-b019-946119d79488 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.226461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.226461Z digest=sha256:16ef4e41a91dcb1f517297e85602141604579f2a472ce1c2723f9c7875754792

Observation 074ee4f9-2342-4368-a7fd-4c482f5cae8a · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.305019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.305019Z digest=sha256:eebb7f39c3485c3b341cfc7bd38080a505410a5d9f4a03eee2be06fb384c76f3

Observation 6d1654b2-3948-476e-b468-3357d054e2fa · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.428563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.428563Z digest=sha256:872d3c26d145f294192221fa1c55b5eb48102cbc104fab5521f3b1ad9ab9e936

Observation 46f5c39f-8313-4aa0-9ac9-69f2cf5d1d19 · outbound

This paper cites To- ward semantics-based answer pinpointing.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning To- ward semantics-based answer pinpointing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.490631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.490631Z digest=sha256:3eef524dbb7bbbac156760ff23eadf2b630ffdc29d70cfb2f1bf99114ab97b7c

Observation bc50babb-2fbf-4087-9e68-88def37e6412 · outbound

This paper cites Learning question classifiers.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Learning question classifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.522302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.522302Z digest=sha256:aec0fea592c118a0d525b23e3ebd6935e95639fbac41d03acdb94f2aeafdcfc0

Observation a0dd0eec-1ae5-4e4b-9bbe-33b6c4f532e3 · outbound

This paper cites The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.589635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.589635Z digest=sha256:f81d3b1796e7d356a960b95cf948fac96c8e56e9c751cb0104c3a6e582210512

Observation a608cb4e-066c-46d2-a96b-0efe913ebcc4 · outbound

This paper cites Let’s verify step by step.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Let’s verify step by step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.810036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.810036Z digest=sha256:331f5afcbff5f16c9805e6211f766877dfa6ca1dc4ab6c9ecd4f37eae884d60c

Observation 2ebc1bba-b51e-4233-a9f0-d6d44130e546 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.844834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.844834Z digest=sha256:84d4bf0ac70d718ed61c0be9f1cff6864a4ab406341242511913d44f79970213

Observation 95156e19-0ec0-4a9b-8fed-b26b49ecb5f0 · outbound

This paper cites Lee, Danqi Chen, and Sanjeev Arora.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Lee, Danqi Chen, and Sanjeev Arora

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.849014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.849014Z digest=sha256:6e672afcc30f9b29a29d518810ea63cd4b295cc13ccaaa10dd8bf1da0053de3f

Observation 3b3dc49c-f664-4945-bf09-e5979f0b8983 · outbound

This paper cites Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.852682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.852682Z digest=sha256:bbf695d47ec48182ad0adb47ae3856cfbfe188c307bc5a119f24692b45daf1e4

Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · outbound

This paper cites Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.856120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.856120Z digest=sha256:a316d9451c3c8376fcaa1d6d6acfdd40fae71ae59973b1a8f737bd865896d0d2

Observation a1adf360-e185-4376-8a4b-a3b016651d6b · outbound

This paper cites Qwen2.5 Technical Report.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Qwen2.5 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.860627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.860627Z digest=sha256:76f8839ce64dc9d27be601d759aecb5d741c130aa0d7784729b8f964b77ea493

Observation cf62eb38-1832-4f7b-8297-a56ac2031945 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.864221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.864221Z digest=sha256:88507473975da00014cde0d400fe1ae3a0a88a12c481596fc536db21f3680cfa

Observation 48289a40-839e-4a98-92ad-a85a1f47eab6 · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Manning, Andrew Ng, and Christopher Potts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.867837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.867837Z digest=sha256:6681e3112ee6095a227b4c763275bda01e2828b5e26e8363cd745042ecd5a372

Observation 850ab672-14f4-424d-8163-5446b0ebc093 · outbound

This paper cites Black-box tun- ing for language-model-as-a-service.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Black-box tun- ing for language-model-as-a-service

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.871408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.871408Z digest=sha256:eaf8cc9b46330385f0eecd6452105e834edad6f07ccf70bae58c98494e22d139

Observation 4629c3be-53c4-4be3-892d-318b66d519d9 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.874508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.874508Z digest=sha256:4f8511a7492ba806d1a86a43a6e1abc06e83a27d79faedc53748fed649ed1fb0

Observation 2eb6ba90-93c3-447c-bb2f-ee3a8d42ce8d · outbound

This paper cites TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 18

Resolution
verified exact
doi, observed 2026-08-02T08:18:25.912187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-02T08:15:19.878269Z digest=sha256:cc772c56d9d3df076dd75c93947a099f386273abcd8c3e45c78e10d6c72f5209

Observation 0beb7bd1-65e7-4566-adbb-47e7684e278a · outbound

This paper cites Opt: Open pre-trained transformer language models, 2022.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Opt: Open pre-trained transformer language models, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.881783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.881783Z digest=sha256:4f883264acda26f8e998ad3ef96a3094addd7ac699acc1143cab5cddeb18c226

Observation 10f428bd-cdf0-4a13-b936-84da42b94da2 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.885366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.885366Z digest=sha256:05d3b2ed8b9a910e5c39a59712c9cafa9a1e0c5b1ae70ed6f1fe628f25cd0d2d

Observation 47647dd0-8559-4301-92c2-4581983fdba1 · outbound

This paper cites Total cost:2KB forward passes.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Total cost:2KB forward passes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.888832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.888832Z digest=sha256:6e525fa25dbfc437c967aa630c029560deb83c3502a0af8094049f5c8b44b464

Observation 8c6a4a03-9ce0-4561-b582-cdd0774c74d4 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.892116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.892116Z digest=sha256:e0d07ebbab5fefff74022632a75ba0f621e24483e8988c7b84fbfbc089f2521f

Observation 096c7ef5-696a-43fc-b7c3-5fe9d76ac254 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.895698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.895698Z digest=sha256:aabb6c80ff15154cc141008201f64d66458afb063388c56203e606a619b49b96

Observation 4a631a5b-30bc-42c8-bba2-f4afe7a0d02b · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.670505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.670505Z digest=sha256:30f902398b229fbdcba20644146bc1c95286f2cec20ad9b4fc07f1f0871ea5b5

Pith citing papers

No inbound Pith citation observations are available.