Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:15:19.895698Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.19408.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:15:19.895698Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c1612773-4807-4786-bf15-1d21f04b9e95 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Dense reward for free in reinforcement learning from human feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d49530-339c-4270-b019-946119d79488 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Training Verifiers to Solve Math Word Problems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074ee4f9-2342-4368-a7fd-4c482f5cae8a · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Process Reinforcement through Implicit Rewards
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d1654b2-3948-476e-b468-3357d054e2fa · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f5c39f-8313-4aa0-9ac9-69f2cf5d1d19 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning To- ward semantics-based answer pinpointing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc50babb-2fbf-4087-9e68-88def37e6412 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Learning question classifiers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0dd0eec-1ae5-4e4b-9bbe-33b6c4f532e3 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a608cb4e-066c-46d2-a96b-0efe913ebcc4 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Let’s verify step by step
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ebc1bba-b51e-4233-a9f0-d6d44130e546 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95156e19-0ec0-4a9b-8fed-b26b49ecb5f0 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Lee, Danqi Chen, and Sanjeev Arora
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b3dc49c-f664-4945-bf09-e5979f0b8983 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1adf360-e185-4376-8a4b-a3b016651d6b · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Qwen2.5 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf62eb38-1832-4f7b-8297-a56ac2031945 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48289a40-839e-4a98-92ad-a85a1f47eab6 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Manning, Andrew Ng, and Christopher Potts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850ab672-14f4-424d-8163-5446b0ebc093 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Black-box tun- ing for language-model-as-a-service
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4629c3be-53c4-4be3-892d-318b66d519d9 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb6ba90-93c3-447c-bb2f-ee3a8d42ce8d · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0beb7bd1-65e7-4566-adbb-47e7684e278a · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Opt: Open pre-trained transformer language models, 2022
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f428bd-cdf0-4a13-b936-84da42b94da2 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47647dd0-8559-4301-92c2-4581983fdba1 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Total cost:2KB forward passes
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c6a4a03-9ce0-4561-b582-cdd0774c74d4 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096c7ef5-696a-43fc-b7c3-5fe9d76ac254 · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a631a5b-30bc-42c8-bba2-f4afe7a0d02b · outbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.