Pith. sign in

Paper Citation Record · LEDGER

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 8 inbound Pith citation observations for arXiv:2509.24372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.24372 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:29.945895Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:06.694233Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ffed6561-48be-4931-a9be-ae1b5c170ceb · outbound

This paper cites GPT-4 Technical Report.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:16.944824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:16.944824Z digest=sha256:b251321992ca8395041b2b154662c6e2079d8bc18426ff16ec3143b63be40e20

Observation 4a2b85df-7f2a-42e6-a1c6-84f78c69bb68 · outbound

This paper cites Intrinsic dimensionality explains the effectiveness of language model fine-tuning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Intrinsic dimensionality explains the effectiveness of language model fine-tuning

Reference 2

Resolution
verified exact
doi, observed 2026-08-04T14:43:27.589291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:17.104745Z digest=sha256:243ad0bd1975b763eacf8e6ea6eb73b940f4b242da8751911c1656123ebd352b

Observation d92eb794-c4cf-461a-887c-578877622de4 · outbound

This paper cites Llama 3 model card, 2024.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Llama 3 model card, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.244748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.244748Z digest=sha256:4fc9e05263a8ad290de1315876075ae3c897cce4a6305251de24d90552fa9644

Observation 70bbafff-d02b-452a-989e-25dfb20a7668 · outbound

This paper cites Evolutionary optimization of model merging recipes.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolutionary optimization of model merging recipes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.443331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.443331Z digest=sha256:0327258a32db909a784e40563cbd14bbfd9d4f44d1310f644a8ea91ff7a37aaf

Observation 1cf56819-0c16-469d-a314-258c68a9f53b · outbound

This paper cites Introducing Claude 4, 2025.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Introducing Claude 4, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.695291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.695291Z digest=sha256:f8d23f7d0779518e2319a6b031f515a46a9147dae0c4e09d50020ce6338fd6e8

Observation a775b486-7838-497b-bbfe-946235b2934d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.934749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.934749Z digest=sha256:4556e4c2a248fd32a80af534ab0044fe7fe5e39688632a8224ffd2ab7169ab8c

Observation debbf38d-33ca-4e10-aaa2-9c6cfd9ecb91 · outbound

This paper cites Understanding pre-training and fine-tuning from loss landscape perspectives.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Understanding pre-training and fine-tuning from loss landscape perspectives

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.060517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.060517Z digest=sha256:d9d938e7ee507d927308e655d54258f1b139159a318de6f6e6888c558453402e

Observation 40e0d99b-048c-4be0-af24-9cac8f899fea · outbound

This paper cites On the weaknesses of reinforcement learning for neural machine translation.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning On the weaknesses of reinforcement learning for neural machine translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.304745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.304745Z digest=sha256:1a634233c359755e452c35941d106d6bf2ec17cadf4d1d1b0ed062d55b39596e

Observation 1448cc4e-5477-4bc3-98e4-0d327e955730 · outbound

This paper cites Back to basics: benchmarking canonical evolution strategies for playing atari.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Back to basics: benchmarking canonical evolution strategies for playing atari

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.515288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.515288Z digest=sha256:6117e2f0cc47f1c961e694c7ce0850e16d7915adacb1729269d78e4bb0d6b9ec

Observation ccda3c47-12d6-4009-a746-794a5dffbac2 · outbound

This paper cites Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.684764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.684764Z digest=sha256:278c1c7899b51a80f620575df1ba9211c186da36a61a70440996917769feaa7b

Observation b1cadf89-8871-4d75-9cd2-ec44aed1fb11 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.815282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.815282Z digest=sha256:815f4a685137799e8499a5789520202f5a2fbf911520897ed9174e1123426d09

Observation 116f2248-7d3d-4cfb-b0cd-7c9ff0d9f382 · outbound

This paper cites Knowledge fusion by evolving weights of language models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Knowledge fusion by evolving weights of language models

Reference 12

Resolution
verified exact
doi, observed 2026-08-04T14:43:27.258406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:18.912503Z digest=sha256:151353833c63934daacc8eec9567c5af955b2a22caa53c497b1cc711b85a3b37

Observation b97881ce-5a75-415f-b8b7-f4646c3872b1 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Detecting hallucinations in large language models using semantic entropy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.144753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.144753Z digest=sha256:5004e1b26bee71cdc28db2088e0cc2e2236fa1a06fa9b1ad1693a8a41e580d50

Observation 97dabea9-1298-426d-8912-3e57462f0dc7 · outbound

This paper cites Reward shaping to mitigate reward hacking in RLHF.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Reward shaping to mitigate reward hacking in RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.404754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.404754Z digest=sha256:0ffaffe1d9a8365ae36b99003a73d788af252d254783580780d9721248388354

Observation ef754ffa-0924-4677-9409-758a7b7867d4 · outbound

This paper cites Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective ST ars.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective ST ars

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.604788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.604788Z digest=sha256:eeb80fc7e6078459c0657059134513a1a86c91559b2246f231d03892d7295830

Observation 1cbe0367-b8b7-4903-a123-a226527c892b · outbound

This paper cites Scaling laws for reward model overoptimization.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Scaling laws for reward model overoptimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.857855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.857855Z digest=sha256:dbdb3dd42fdb43e69981fa0c8dbfacb81887a416cb862df7470c80b6e6159da1

Observation 289e40bd-8b2f-428e-afed-01bc37e65ae0 · outbound

This paper cites Deep learning, volume 1.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deep learning, volume 1

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.014401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.014401Z digest=sha256:571461dddf5556cb849988e9e906b75e797d63098974b25d13fce7dfeb2896cb

Observation 2178f773-8149-4329-991e-b487a03a513f · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., 2025.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.204798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.204798Z digest=sha256:9aa6eddd3de07bbec4177f0e126525122062e1bc8229366f32e3fc7a883d57fd

Observation d62c2a83-2353-42b8-a766-2c584ced21d5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.390576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.390576Z digest=sha256:995782cdc9575075e116cd9302b3e20de18e7a8c48f6a94504f3edcbccebcd67

Observation a7036ecf-177d-481a-a876-e664d29aaf9b · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.764759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.764759Z digest=sha256:ac610cc33dd3bd53c4256ebf9bb2c2699ad25262ca1d4b16acc2512e6035534f

Observation ea0ae0a5-8475-4dfd-a79d-d2757b03ddaf · outbound

This paper cites Connecting large language models with evolutionary algorithms yields powerful prompt optimizers.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Connecting large language models with evolutionary algorithms yields powerful prompt optimizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.051034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.051034Z digest=sha256:c4ca4f69022c324d111550444472d3122f8c75ab109f5f723fc755c81d45c031

Observation 1dba326f-0240-4c76-9f03-f8b5a2b7d07f · outbound

This paper cites Completely derandomized self-adaptation in evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Completely derandomized self-adaptation in evolution strategies

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.314845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.314845Z digest=sha256:641438f1a4446b591b6154fbe2d63c2387c83e23cc3564a60ee4ec21babfd15c

Observation 18b5cfba-bd1e-4db9-b35f-19b29e9e4410 · outbound

This paper cites When evolution strategy meets language models tuning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning When evolution strategy meets language models tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.479943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.479943Z digest=sha256:371bc3a2c52e159adf6c0ff28e8cc4c861a984998b2683dfef74d4419f75f007

Observation 9e173aea-9f86-4c5d-9ce1-329aff70b79b · outbound

This paper cites Neuroevolution for reinforcement learning using evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Neuroevolution for reinforcement learning using evolution strategies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.755222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.755222Z digest=sha256:f159cca0eef9cdd18c9a81ff201c78e134015318c5889fce17129286fbdf30a2

Observation 66a764d1-4096-4b16-ac11-46cd21524b4a · outbound

This paper cites Do we need to verify step by step? rethinking process supervision from a theoretical perspective.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Do we need to verify step by step? rethinking process supervision from a theoretical perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.904752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.904752Z digest=sha256:a63430d8d111f3e44e9eb225ae05e795b8be64c5c33778839e9111c27df31b5b

Observation ec6ea8c0-8122-4c0d-ab22-977bbb2993b4 · outbound

This paper cites Mixtral of Experts.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.124747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.124747Z digest=sha256:aff6d5b6cbe664fd0e20061d4a8400eb354e3e9e39f71a878ba6fbd22d7c99b8

Observation b00a1de9-0509-45d4-888f-42b39fc7ed0e · outbound

This paper cites Derivative-free optimization for low-rank adaptation in large language models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Derivative-free optimization for low-rank adaptation in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.334749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.334749Z digest=sha256:ecd7949fe8c4441e2874c4e489c8ab0e56d5965c4a462c0c668fd12bf8612e8e

Observation 2bb53075-2a59-4fca-b33f-64d302e00a34 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.504761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.504761Z digest=sha256:7fbf932ddbbd809120fb8cbb9fc57b0a6e719703f8c3c4a4a2212da4b471fcd5

Observation ecbb5529-880d-4d57-98c6-d33f02e7a269 · outbound

This paper cites Fine-tuning chatgpt for automatic scoring.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning chatgpt for automatic scoring

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.674753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.674753Z digest=sha256:790c4a598d60710b63fbd9b2f9c8c20fc338e0a2033dfabd0e1f59ef53a6ab0f

Observation 26c35849-849e-4738-9715-b2b353ed2dc9 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.920603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.920603Z digest=sha256:eaeb31e7df1359136b14249e8262d95a9e8e66641993d32b87e676b63f975856

Observation 955ec27f-6a47-4a9f-bf3d-64b382b1f973 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.084747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.084747Z digest=sha256:06dcaefa736c956de2e8a6ed5ecf6384ae15e06a2f7d69acd1e5e7baff17c120

Observation 3ee762dd-3f1a-4278-b0a8-5ac1b622d931 · outbound

This paper cites DeepSeek-V3 Technical Report.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeek-V3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.187732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.187732Z digest=sha256:39734fce87a989047642cd8cbc2c15406cf02ee4c8f843585477c5911662b568

Observation 71fa5a7d-e6e4-49a7-8ff6-bf5646c24d07 · outbound

This paper cites Sparse me ZO : Less parameters for better performance in zeroth-order LLM fine-tuning, 2025.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sparse me ZO : Less parameters for better performance in zeroth-order LLM fine-tuning, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.344787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.344787Z digest=sha256:f1cc5308b0378e158c58cd7c1ac04edfac0102c2ff3f4ffc97effa86b03ee282

Observation efa07711-73e6-4449-971f-3fdcc02cff0b · outbound

This paper cites Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.484928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.484928Z digest=sha256:0a12b68205864823fe33537acfecbfb842ae63ebc707e9da55503cd4c8f341ad

Observation c5df2399-1981-44ba-968f-11fd4aadbd6c · outbound

This paper cites Fine-tuning language models with just forward passes.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning language models with just forward passes

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.574795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.574795Z digest=sha256:1d8faf444a6be02699448cc2a7038f0d95043070996ec6d6187333aecaf6dfef

Observation a1cfbdc1-a4a6-490a-95fc-77c111fe4117 · outbound

This paper cites Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.664754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.664754Z digest=sha256:40aecb5ae9f268a06eb40e8c1fb9fb930c03aa30a83493fb25a5859bb842c5de

Observation 9788c601-c4b5-4d58-b29e-254452e3d465 · outbound

This paper cites What is artificial superintelligence?, 2023.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning What is artificial superintelligence?, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.764751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.764751Z digest=sha256:b51fb18cb1071f35b3f9f104d47d8b1d7ca6575aac1def8b68b668eac69ff05c

Observation 3bb4a197-ca86-4646-abab-89e9b6fe6c97 · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.915163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.915163Z digest=sha256:284652553bfc8e622591e416483dba45b46e200764bff5130289bf821be520af

Observation 9c197115-645a-4f2e-acd0-9db7f5d71f97 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.294757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.294757Z digest=sha256:1ca8e319e98203109ff9b5e25901bfdf2b8761345f7101febbe064c3f6c29f59

Observation fd170c12-df61-4bbd-a660-c882d507609f · outbound

This paper cites Tinyzero.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Tinyzero

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.415276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.415276Z digest=sha256:46e7afe7e48e387d73bc1331440ac919d152010aa3809990066bbe71533121eb

Observation 49155d3b-e903-42c3-a3cc-b920057a2491 · outbound

This paper cites Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.624754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.624754Z digest=sha256:dfd91485003ba8f1ca435b3636df9ca3c6dc3a775cc5c7806e48df7dd11828f1

Observation 49a2d506-5899-4534-a9bc-98b445aefd8f · outbound

This paper cites Semantic density: Uncertainty quantification for large language models through confidence measurement in semantic space.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Semantic density: Uncertainty quantification for large language models through confidence measurement in semantic space

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.787960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.787960Z digest=sha256:b48434d6e9e052e768f872e8d13767070dcbbe2f65ac9cf9be3ae0380fb8216a

Observation 8e7782ce-fd85-4227-b42a-3ca5eb9fa557 · outbound

This paper cites Manning, and Chelsea Finn.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Manning, and Chelsea Finn

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.934759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.934759Z digest=sha256:19d7039b5a721810f91ad4ea84de7f2221861a82045419df88d6c167331576ed

Observation 9b669db8-8f19-4a24-b1d5-b8806868d090 · outbound

This paper cites Rechenberg.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Rechenberg

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.124738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.124738Z digest=sha256:dcfd4dcde11f598d590bf9871c4a540da5a9edcd4d42a656b142b9a31b235600

Observation 4ae9bac0-f017-4e12-b2f2-e81f9066e6fe · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.304763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.304763Z digest=sha256:7dbe12457fcfca5141a568e77001e05f908e327215f05ccc8fee33cd3a963949

Observation ce995643-c078-4e58-9949-c8c8c866b226 · outbound

This paper cites Pawan Kumar, Emilien Dupont, Francisco J.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Pawan Kumar, Emilien Dupont, Francisco J

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.494752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.494752Z digest=sha256:d358f34650677bc88269c02fdabac5e9ed75a94086cb8f49ae059178994fd25d

Observation 4167e780-7675-4cd8-8104-dc36299a2d36 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Code Llama: Open Foundation Models for Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.564745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.564745Z digest=sha256:6e055df4d4558b096ce0bdb0f7d126e2d17c335486a85562d4f1e00ddbcde20a

Observation cd28f1ad-0bf8-4a59-bcbb-b5d4477a3d64 · outbound

This paper cites u ckstie , Martin Felder, and J \.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning u ckstie , Martin Felder, and J \

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.646284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.646284Z digest=sha256:9162c42e7b667f7b3e2a7056e2c2f95c22cd9bd854c5327066560a4792374930

Observation 0563fd3f-7e40-414d-9928-66ac14c64292 · outbound

This paper cites u ckstie , Frank Sehnke, Tom Schaul, Daan Wierstra, Yi Sun, and J \.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning u ckstie , Frank Sehnke, Tom Schaul, Daan Wierstra, Yi Sun, and J \

Reference 49

Resolution
verified exact
doi, observed 2026-08-04T14:49:22.811963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:25.755143Z digest=sha256:d55cd8ca4a66972e05100c93a79f3d59d69ae796d94b0edc30ce221dc2d8ae4f

Observation 2aa166e2-0e31-4c69-9a1c-63de63bc6595 · outbound

This paper cites Improved techniques for training gans.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Improved techniques for training gans

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.834738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.834738Z digest=sha256:193a76a60389e85405836ef76c87254ba6703b7b1aefe7db5ddca5489311a56d

Observation a1127278-1185-41d4-b12a-b84bf56d8fc6 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.910796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.910796Z digest=sha256:e17bc64b8104f392e965dc6f8a5ed2d28b8c28c0eb6739b69522c7b9b0682ef5

Observation 1c269cfe-a434-47ab-b6c9-81da6b750074 · outbound

This paper cites How well can a genetic algorithm fine-tune transformer encoders? a first approach.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning How well can a genetic algorithm fine-tune transformer encoders? a first approach

Reference 52

Resolution
verified exact
doi, observed 2026-08-04T14:49:22.554438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:26.187680Z digest=sha256:979a0c654730038b71ee608d30e648836880d979f6c3785f5a949d56cc6e1944

Observation 56aaed94-1d56-4d14-9a96-75d295c6917c · outbound

This paper cites Approximating kl divergence, 2020.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Approximating kl divergence, 2020

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.317534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.317534Z digest=sha256:cc11e761eed620a6d3f4c0bc21e48a919302bf3a379e390ad7e5dfc31c39b004

Observation 1a2957ed-851b-410d-9d65-94988b818858 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.444749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.444749Z digest=sha256:197a0150618bd6bdb56305aa88c0ef04e83c2f7088fd7b48105d23f65894a5cf

Observation 566bc3db-07ed-4778-b751-b2526b58273f · outbound

This paper cites Numerische Optimierung von Computermodellen mittels der Evo-lutionsstrategie, volume 26.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Numerische Optimierung von Computermodellen mittels der Evo-lutionsstrategie, volume 26

Reference 55

Resolution
verified exact
doi, observed 2026-08-04T14:49:22.266846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:26.595147Z digest=sha256:357235b2ffd22cb53669442aaeb46c3d3de88c4cb9950ec795f977a015464768

Observation 021a19c9-2ada-4bc5-ad18-5b943bbf7621 · outbound

This paper cites Parameter-exploring policy gradients.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Parameter-exploring policy gradients

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.744734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.744734Z digest=sha256:7bd9f266d5a9a4b38b99bfa2e649e31279dfec1b03f45f6d66a04c3814c1abf3

Observation 5299c176-c970-41f7-96bf-b9b272b65ff1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.852975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.852975Z digest=sha256:9890bb62cc94a0343be568e2ce1036e06cfc12f0c57df99c97c954ceb76ad806

Observation 80b23c43-5d60-420b-8886-818a98ec6671 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Towards Expert-Level Medical Question Answering with Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.960792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.960792Z digest=sha256:ddc682f8aa3b85a09eda628a3f630d8690de09fe21472239e6c34e0b1c1445ad

Observation f9c38cc8-5575-47ec-ae15-97c78348dc49 · outbound

This paper cites PRMB ench: A fine-grained and challenging benchmark for process-level reward models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning PRMB ench: A fine-grained and challenging benchmark for process-level reward models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.025445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.025445Z digest=sha256:a2011e6a79eaa780202ded38553428c5e852bdd08bf23b41d5fa192cae63755d

Observation 557e0319-cb47-450b-958b-5bc0fad95dd4 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.126234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.126234Z digest=sha256:080cb8183cfe85eb2c769fcfb64e62ffe75ff9616ed3c7ac9d0bac72a3ee8115

Observation 73b5aced-ca17-4bc6-8dc1-f4374deccddc · outbound

This paper cites A Technical Survey of Reinforcement Learning Techniques for Large Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning A Technical Survey of Reinforcement Learning Techniques for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.174129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.174129Z digest=sha256:c6fbb0fc90bdea6cc093acf54ab4138873c342a8f559d7dba25fc09d72268a74

Observation 0991185e-9785-45f6-88f2-3b8b01088df3 · outbound

This paper cites Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.246511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.246511Z digest=sha256:98361d6bbe461ddeb289307949fea7985c5beb8f115668bb56980bb12be9b195

Observation e3c6fe67-92a7-4f31-a208-e7efb0061d44 · outbound

This paper cites BBT v2: Towards a gradient-free future with large language models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning BBT v2: Towards a gradient-free future with large language models

Reference 63

Resolution
verified exact
doi, observed 2026-08-04T14:49:21.997293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:27.374747Z digest=sha256:f179f414731c792b64ec028047cbb915c32265fb9442ac8a10d30be44512bb35

Observation b60689ed-5695-4a7d-9fbc-0ae90570da75 · outbound

This paper cites Black-box tuning for language-model-as-a-service.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Black-box tuning for language-model-as-a-service

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.448579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.448579Z digest=sha256:d0bc1cd293f9ac52f464f88d658c3fb0bbd61ff043e9f6d0fea703272372a9ea

Observation 250e12ca-d12d-4cac-8249-0cb6d5168f45 · outbound

This paper cites Sutton and Andrew G.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sutton and Andrew G

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.506331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.506331Z digest=sha256:f62b8ae92cde178856a320fa716e1523982c04c841d95cdc788f2916e03bc2c7

Observation 3e7e83d6-a8de-44e9-8cab-319f86980e50 · outbound

This paper cites Fine-tuning mt5-based transformer via cma-es for sentiment analysis.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning mt5-based transformer via cma-es for sentiment analysis

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.586741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.586741Z digest=sha256:ee661dc373ebccb8e7b0a67adf130e05a13480757a4e8da75883cfe5504b6071

Observation 93e0ef8b-6fba-4725-ab17-0d57ec27dc60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.686155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.686155Z digest=sha256:7e007747a0478df9bad7c109309821ded094959eb3013a0bae24b0280d9fe079

Observation 8fd0b78e-4c31-44bd-9bf2-edcd19b001ad · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Solving math word problems with process- and outcome-based feedback

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.886374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.886374Z digest=sha256:d843c5ca0da24d662846118f5cb2732194d4aa231299ea51774fa65d9af012a5

Observation 2139c937-4a60-4a80-89cd-74df5b9b7a4a · outbound

This paper cites Andrew Bagnell.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Andrew Bagnell

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.045066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.045066Z digest=sha256:e76abfc7987fcdfb0969fefb5b78758f024b2fd5de7e02487daffacc94596df1

Observation 4ae3ebca-2647-47f9-8859-f6d4bec14fb3 · outbound

This paper cites When large language models meet evolutionary algorithms: Potential enhancements and challenges.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning When large language models meet evolutionary algorithms: Potential enhancements and challenges

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.108064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.108064Z digest=sha256:2bc9428303fba062d5a92885530632d7c3b0c2adea3e090aa1a79ce08534ec9c

Observation 03e3f036-81ac-425e-a968-a3352427da62 · outbound

This paper cites Natural evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Natural evolution strategies

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.295794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.295794Z digest=sha256:3d313b91babfefcffad89f4bec42c1c58a3bacd1a0745bcf76a86462f23bd560

Observation 20c6296a-1bc3-4eed-b3bb-27e7ec4859ad · outbound

This paper cites Natural evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Natural evolution strategies

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.396359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.396359Z digest=sha256:f47bbaae5fa976cc33e583358bc85c95a82a4cb1a3029654b007f6f063be2c33

Observation fa1b9614-df6d-42fb-a2b3-d94f7d0fad23 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning BloombergGPT: A Large Language Model for Finance

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.534733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.534733Z digest=sha256:244f45aab84a0951d1d209022a845e85af19e8c73e2cb70585ae933ebfd73a5a

Observation d2d6254c-a23b-4b63-bddf-c1b91fc01fee · outbound

This paper cites Evolutionary computation in the era of large language model: Survey and roadmap.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolutionary computation in the era of large language model: Survey and roadmap

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.654757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.654757Z digest=sha256:244c78891f2421ffcdd27aa7685e42d311390f897f30608e15fc2e009f5617a4

Observation 09d98d04-fe78-4afc-86f1-89a15c6fd4f5 · outbound

This paper cites Qwen2.5-1M Technical Report.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Qwen2.5-1M Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.822086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.822086Z digest=sha256:34721abf494ecde61ce4f5383a9fcf3fc89775e8bac4049f6c79a7e2e8889a57

Observation 4e249551-1814-407a-93e2-139675d2d875 · outbound

This paper cites On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.925179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.925179Z digest=sha256:7c0949ac27fd3b5cd154b1fc633c5527ac78b22c984ce086f1eafbf174be1a76

Observation c862fcb1-73f3-4966-b6bd-508ff12f6e4e · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.015865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.015865Z digest=sha256:a4c1c8472eaad4496718cf13fa270553569673ec918ead16df723a04ea12c099

Observation 7e12e69a-8c2d-48c9-aea7-2f66a762bb1d · outbound

This paper cites Genetic prompt search via exploiting language model probabilities.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Genetic prompt search via exploiting language model probabilities

Reference 78

Resolution
verified exact
doi, observed 2026-08-04T14:49:21.859507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T14:43:29.112085Z digest=sha256:5b47a4f4c7b307c7bc69a482ac97df103de05a4f4f12087938baf5ad839cc87e

Observation cea816d8-261a-442c-b0f6-df389bed24ac · outbound

This paper cites DPO meets PPO : Reinforced token optimization for RLHF.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DPO meets PPO : Reinforced token optimization for RLHF

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.366209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.366209Z digest=sha256:901aae423b4bc3849781314f12b6776027271896d0352c2d7ac69147aa96ce36

Observation a6d9c6a0-4abc-4204-8649-b723d2a9f70d · outbound

This paper cites write newline.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning write newline

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.468661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.468661Z digest=sha256:47482587f6ea127a863b3b05ebe5c821dd4595feb635acf66541b83fb62c7660

Observation 3140f06a-fd27-456b-be01-aaf1d282cf6d · outbound

This paper cites @esa (Ref.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning @esa (Ref

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.624735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.624735Z digest=sha256:fae08472ecf5ee0cb4a7dfa18e9f19ed0ab95b0b24d347dd3b2625f3f16cd53f

Observation 96b62c46-9c97-4f38-a72e-af0097894cdf · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.761120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.761120Z digest=sha256:6b875b0b6dcffdba57837400cf3f00f29803bb19cb16087b4caf5fac3979e9d1

Observation b6c09985-6d31-4ba6-aefe-be9ccd556db6 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.945895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.945895Z digest=sha256:d2dc1976994717bcc98b731e2f657a3427eb4c94271ca233d35184ff712123c0

Pith citing papers

Observation 08cfb055-de24-4a25-92bc-b230408cc6d1 · inbound

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training cites this paper.

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:54:47.861802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:54:47.861802Z digest=sha256:ab14d462b1c8379df6649125bca79b7891ca1d09f70ad8b07f62d5bb51263b38

Observation d7d77328-35f2-43e4-a15a-13a73d071b8f · inbound

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning cites this paper.

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T08:24:04.873255Z digest=sha256:2997c77b08d10ec0bdbed42e85b7e945bc8aa141cc60af7d85a923151e94722a

Observation c079d1a8-4bcd-4cf5-b57b-8adc4393a860 · inbound

Goal-Conditioned Supervised Learning for LLM Fine-Tuning cites this paper.

Goal-Conditioned Supervised Learning for LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T22:37:46.345159Z digest=sha256:d6d6403b5abb343cec45a59c2232a46faa01b49ff35c44ccaf532e5d6f859fc8

Observation 4dbd30cf-5d48-4e6f-9b4f-3bcbc479737f · inbound

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play cites this paper.

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T21:37:56.570173Z digest=sha256:2e23c08fcd85b164747e41ed99c05b80eaa54d89cbe41a8a86cff4eb568c5431

Observation 124a61b3-fdbf-48b9-b17a-ef78c1143029 · inbound

Mathematical perspective on genetic algorithms with optimization guided operators cites this paper.

Mathematical perspective on genetic algorithms with optimization guided operators Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T07:34:35.422699Z digest=sha256:f8a3ee94df6614b31ce01cc30f57c808071bf3b518b91387a4fa4fb9546ce56a

Observation 56bb7a7c-7b55-4bf2-ba20-fa69215bdf9f · inbound

Why can genetic algorithms work in high-dimensional search spaces? cites this paper.

Why can genetic algorithms work in high-dimensional search spaces? Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T03:11:04.083187Z digest=sha256:fcdc823d47781e1cda7d87a8aeae25084140d38e784dc89c9400e9dcc0abcebc

Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · inbound

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning cites this paper.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.856120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.856120Z digest=sha256:1d31d862ec91c802abe38a37e44c1d9880cc16352f24a0c02d0eeb1c33a31bd3

Observation a9361b78-916d-42e4-ab8c-c08c8c56e984 · inbound

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging cites this paper.

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:06.694233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:15:06.694233Z digest=sha256:a99864b0cac21a5dc7bb38b0a2f1032451f29be1832bb336bc8c9c325d3758c0