Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:29.945895Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 8 inbound Pith citation observations for arXiv:2509.24372.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:29.945895Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:06.694233Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
83 of 83 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation ffed6561-48be-4931-a9be-ae1b5c170ceb · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2b85df-7f2a-42e6-a1c6-84f78c69bb68 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d92eb794-c4cf-461a-887c-578877622de4 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Llama 3 model card, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70bbafff-d02b-452a-989e-25dfb20a7668 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolutionary optimization of model merging recipes
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf56819-0c16-469d-a314-258c68a9f53b · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Introducing Claude 4, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a775b486-7838-497b-bbfe-946235b2934d · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation debbf38d-33ca-4e10-aaa2-9c6cfd9ecb91 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Understanding pre-training and fine-tuning from loss landscape perspectives
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e0d99b-048c-4be0-af24-9cac8f899fea · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning On the weaknesses of reinforcement learning for neural machine translation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1448cc4e-5477-4bc3-98e4-0d327e955730 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Back to basics: benchmarking canonical evolution strategies for playing atari
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccda3c47-12d6-4009-a746-794a5dffbac2 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1cadf89-8871-4d75-9cd2-ec44aed1fb11 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116f2248-7d3d-4cfb-b0cd-7c9ff0d9f382 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Knowledge fusion by evolving weights of language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b97881ce-5a75-415f-b8b7-f4646c3872b1 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Detecting hallucinations in large language models using semantic entropy
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97dabea9-1298-426d-8912-3e57462f0dc7 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Reward shaping to mitigate reward hacking in RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef754ffa-0924-4677-9409-758a7b7867d4 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective ST ars
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cbe0367-b8b7-4903-a123-a226527c892b · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Scaling laws for reward model overoptimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 289e40bd-8b2f-428e-afed-01bc37e65ae0 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deep learning, volume 1
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2178f773-8149-4329-991e-b487a03a513f · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d62c2a83-2353-42b8-a766-2c584ced21d5 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7036ecf-177d-481a-a876-e664d29aaf9b · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea0ae0a5-8475-4dfd-a79d-d2757b03ddaf · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dba326f-0240-4c76-9f03-f8b5a2b7d07f · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Completely derandomized self-adaptation in evolution strategies
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18b5cfba-bd1e-4db9-b35f-19b29e9e4410 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning When evolution strategy meets language models tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e173aea-9f86-4c5d-9ce1-329aff70b79b · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Neuroevolution for reinforcement learning using evolution strategies
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a764d1-4096-4b16-ac11-46cd21524b4a · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Do we need to verify step by step? rethinking process supervision from a theoretical perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec6ea8c0-8122-4c0d-ab22-977bbb2993b4 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Mixtral of Experts
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00a1de9-0509-45d4-888f-42b39fc7ed0e · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Derivative-free optimization for low-rank adaptation in large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bb53075-2a59-4fca-b33f-64d302e00a34 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Adam: A Method for Stochastic Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecbb5529-880d-4d57-98c6-d33f02e7a269 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning chatgpt for automatic scoring
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c35849-849e-4738-9715-b2b353ed2dc9 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 955ec27f-6a47-4a9f-bf3d-64b382b1f973 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee762dd-3f1a-4278-b0a8-5ac1b622d931 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeek-V3 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71fa5a7d-e6e4-49a7-8ff6-bf5646c24d07 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sparse me ZO : Less parameters for better performance in zeroth-order LLM fine-tuning, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa07711-73e6-4449-971f-3fdcc02cff0b · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5df2399-1981-44ba-968f-11fd4aadbd6c · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning language models with just forward passes
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1cfbdc1-a4a6-490a-95fc-77c111fe4117 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9788c601-c4b5-4d58-b29e-254452e3d465 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning What is artificial superintelligence?, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb4a197-ca86-4646-abab-89e9b6fe6c97 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c197115-645a-4f2e-acd0-9db7f5d71f97 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd170c12-df61-4bbd-a660-c882d507609f · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Tinyzero
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49155d3b-e903-42c3-a3cc-b920057a2491 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a2d506-5899-4534-a9bc-98b445aefd8f · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Semantic density: Uncertainty quantification for large language models through confidence measurement in semantic space
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7782ce-fd85-4227-b42a-3ca5eb9fa557 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Manning, and Chelsea Finn
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b669db8-8f19-4a24-b1d5-b8806868d090 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Rechenberg
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae9bac0-f017-4e12-b2f2-e81f9066e6fe · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce995643-c078-4e58-9949-c8c8c866b226 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Pawan Kumar, Emilien Dupont, Francisco J
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4167e780-7675-4cd8-8104-dc36299a2d36 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Code Llama: Open Foundation Models for Code
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd28f1ad-0bf8-4a59-bcbb-b5d4477a3d64 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning u ckstie , Martin Felder, and J \
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0563fd3f-7e40-414d-9928-66ac14c64292 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning u ckstie , Frank Sehnke, Tom Schaul, Daan Wierstra, Yi Sun, and J \
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2aa166e2-0e31-4c69-9a1c-63de63bc6595 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Improved techniques for training gans
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1127278-1185-41d4-b12a-b84bf56d8fc6 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c269cfe-a434-47ab-b6c9-81da6b750074 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning How well can a genetic algorithm fine-tune transformer encoders? a first approach
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56aaed94-1d56-4d14-9a96-75d295c6917c · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Approximating kl divergence, 2020
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a2957ed-851b-410d-9d65-94988b818858 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566bc3db-07ed-4778-b751-b2526b58273f · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Numerische Optimierung von Computermodellen mittels der Evo-lutionsstrategie, volume 26
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 021a19c9-2ada-4bc5-ad18-5b943bbf7621 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Parameter-exploring policy gradients
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5299c176-c970-41f7-96bf-b9b272b65ff1 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b23c43-5d60-420b-8886-818a98ec6671 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Towards Expert-Level Medical Question Answering with Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9c38cc8-5575-47ec-ae15-97c78348dc49 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning PRMB ench: A fine-grained and challenging benchmark for process-level reward models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 557e0319-cb47-450b-958b-5bc0fad95dd4 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b5aced-ca17-4bc6-8dc1-f4374deccddc · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning A Technical Survey of Reinforcement Learning Techniques for Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0991185e-9785-45f6-88f2-3b8b01088df3 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c6fe67-92a7-4f31-a208-e7efb0061d44 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning BBT v2: Towards a gradient-free future with large language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b60689ed-5695-4a7d-9fbc-0ae90570da75 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Black-box tuning for language-model-as-a-service
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250e12ca-d12d-4cac-8249-0cb6d5168f45 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sutton and Andrew G
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7e83d6-a8de-44e9-8cab-319f86980e50 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning mt5-based transformer via cma-es for sentiment analysis
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93e0ef8b-6fba-4725-ab17-0d57ec27dc60 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd0b78e-4c31-44bd-9bf2-edcd19b001ad · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Solving math word problems with process- and outcome-based feedback
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2139c937-4a60-4a80-89cd-74df5b9b7a4a · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Andrew Bagnell
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae3ebca-2647-47f9-8859-f6d4bec14fb3 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning When large language models meet evolutionary algorithms: Potential enhancements and challenges
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e3f036-81ac-425e-a968-a3352427da62 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Natural evolution strategies
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c6296a-1bc3-4eed-b3bb-27e7ec4859ad · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Natural evolution strategies
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa1b9614-df6d-42fb-a2b3-d94f7d0fad23 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning BloombergGPT: A Large Language Model for Finance
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d6254c-a23b-4b63-bddf-c1b91fc01fee · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolutionary computation in the era of large language model: Survey and roadmap
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d98d04-fe78-4afc-86f1-89a15c6fd4f5 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Qwen2.5-1M Technical Report
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e249551-1814-407a-93e2-139675d2d875 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c862fcb1-73f3-4966-b6bd-508ff12f6e4e · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e12e69a-8c2d-48c9-aea7-2f66a762bb1d · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Genetic prompt search via exploiting language model probabilities
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cea816d8-261a-442c-b0f6-df389bed24ac · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DPO meets PPO : Reinforced token optimization for RLHF
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d9c6a0-4abc-4204-8649-b723d2a9f70d · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning write newline
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3140f06a-fd27-456b-be01-aaf1d282cf6d · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning @esa (Ref
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b62c46-9c97-4f38-a72e-af0097894cdf · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c09985-6d31-4ba6-aefe-be9ccd556db6 · outbound
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08cfb055-de24-4a25-92bc-b230408cc6d1 · inbound
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d77328-35f2-43e4-a15a-13a73d071b8f · inbound
ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c079d1a8-4bcd-4cf5-b57b-8adc4393a860 · inbound
Goal-Conditioned Supervised Learning for LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4dbd30cf-5d48-4e6f-9b4f-3bcbc479737f · inbound
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 124a61b3-fdbf-48b9-b17a-ef78c1143029 · inbound
Mathematical perspective on genetic algorithms with optimization guided operators Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56bb7a7c-7b55-4bf2-ba20-fa69215bdf9f · inbound
Why can genetic algorithms work in high-dimensional search spaces? Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · inbound
Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9361b78-916d-42e4-ab8c-c08c8c56e984 · inbound
Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.