Pith. sign in

Paper Citation Record · LEDGER

AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2305.14387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14387 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.258829Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

54
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c633b906-40bb-48c1-b2ad-7cb858c7dcb6 · inbound

Large Language Models are not Fair Evaluators cites this paper.

Large Language Models are not Fair Evaluators AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:10:42.405853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T12:10:42.248005Z digest=sha256:fc7589a1dec3d0760386e6b1028a73c81a5d28caeba55c99238634a05ca4345a

Observation 41c776e8-d9d2-4e35-a3a2-e329d193b2fa · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:52:59.178783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:628a65990f50be9380d625ff4343936290b5b92424343124d9e3f6e79e08f895

Observation d04362a3-2474-402b-83e2-bd6d5380fe24 · inbound

Textbooks Are All You Need cites this paper.

Textbooks Are All You Need AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:44:03.204657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T04:44:03.148223Z digest=sha256:0cf5984c0657577fdb9857420d0478c22e730d631e53cec6141eecf1ee836185

Observation 0404f66a-62af-479f-a6d8-65ffa17db7b1 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.557301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:36ef3c16c9c2f3ebb3aec665eeeea1f26447dea23057ec829058656e1800ba7f

Observation cd184d77-cc36-4485-9d2c-7d16d87aa969 · inbound

Baseline Defenses for Adversarial Attacks Against Aligned Language Models cites this paper.

Baseline Defenses for Adversarial Attacks Against Aligned Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:24:40.111516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T23:24:39.835347Z digest=sha256:da5aad85fe44a065144bccb2d99948a8b204017e947d4c56082d308dd80dd7c6

Observation fbf82ab5-4dae-4e5f-ae5a-b0b1bec77a67 · inbound

Chain-of-Verification Reduces Hallucination in Large Language Models cites this paper.

Chain-of-Verification Reduces Hallucination in Large Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:06:50.331746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T01:06:49.811982Z digest=sha256:f5b0874f5b2a1cf8d9d28fa1c4d585096fcde6ccc7a36c9572ec058b13a29fd9

Observation 59a8c904-4c23-47ce-8f37-ff774f2cd88e · inbound

Aligning Large Multimodal Models with Factually Augmented RLHF cites this paper.

Aligning Large Multimodal Models with Factually Augmented RLHF AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:58:17.843448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:58:17.699042Z digest=sha256:46a8b3e294b6187e775a9dcf551abac88c695adfe8ad2841c69756b67cac0fb4

Observation d214e8c3-495a-4794-b92a-6fdfe98f9429 · inbound

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection cites this paper.

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T14:15:11.113391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T14:15:10.907921Z digest=sha256:aaa79b7811ffc822f58efba414d3475c1e8eaa6036e58e6c0dca8abb113ef65d

Observation 6182808e-2263-4bc4-ae30-791c538dcd7c · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.083261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:14f1ae03a2c2a7065009a90860ea45f2928d654bc461933741ea5a8da8ceb742

Observation 650f1700-58a3-4bc1-8196-94d98ab8091e · inbound

Self-Rewarding Language Models cites this paper.

Self-Rewarding Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.411055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c737c33a5544a5247a323b4ebffba0b10912d9ea2e56395158d5b90fc09c65a2

Observation c0008a37-ac01-4e72-8eb4-452843efdb90 · inbound

Corrective Retrieval Augmented Generation cites this paper.

Corrective Retrieval Augmented Generation AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:19:17.301138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T11:19:17.120464Z digest=sha256:d5937583c2ed97c3e10a540e1e614599cc5e98024332b513b785569aa6591b22

Observation 98eb55fd-94b7-446f-b2ff-518e5b4fab1b · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.926678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:0f519568d40c5cf3a889619caa7f8a34f761621df51747323baa09a747539ce8

Observation a4007580-dd99-4d7b-99b3-570889290493 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.165372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:35ed2d2ae44125043e7bfa21fd1b4b0231a92277a99e0c33ccc348a20d6e13af

Observation d359e29c-cdbe-41d1-a64e-0fa9d0f514e9 · inbound

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking cites this paper.

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.258829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.258829Z digest=sha256:97035195dc1119996abb55b7b3e74f6798924fa494450322b67acb6e7ebc4e83

Observation 662f2b23-c9d8-4462-b276-496aa0e30648 · inbound

Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution cites this paper.

Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:22.616939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:22.616939Z digest=sha256:aa97035040ad126af6b0c9d5598af738f5b5780ca22f0558868fb6c95c12392d

Observation e8c042f1-7573-493f-aac2-c9a443c337f4 · inbound

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints cites this paper.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.639987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.639987Z digest=sha256:2b3307f38d28d6a9dba1d8199a496b83fa4b8525513b08ce668790d237456be4

Observation 7ec1677e-173c-4452-ab99-ce10bfeee767 · inbound

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models cites this paper.

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.887056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.887056Z digest=sha256:9bae0f1119c3116199c94c39a4b8715802c5b395641248fff05a91fa403e1bf6

Observation c0e8ab4e-2af1-4680-84be-b13782bead09 · inbound

RecoWorld: Building Simulated Environments for Agentic Recommender Systems cites this paper.

RecoWorld: Building Simulated Environments for Agentic Recommender Systems AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:56:38.763239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:56:38.763239Z digest=sha256:060849eca78ef7fe4d24ecf4ce63d150b4118fcbdbb5a9f1dab4d3a1050087d5

Observation 0954badc-258c-4482-983e-fd21e59143d4 · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:56:24.568395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:c031cf0819fc830ff75aa30319c7fdfb058bd96f8d2923e0c7d9446fa12582e7

Observation 603a3ffd-b380-48b2-b709-f295c5a3a440 · inbound

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems cites this paper.

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:47:53.751787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T12:47:28.248540Z digest=sha256:b44536754e5bce38fe04ab07ada9bb401cd13668ab0f2d6e7b2bd844b7cdf7cb

Observation 9fb779b0-903a-4a0c-898c-e1821bbcfcf5 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.112040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:ed62ff6d7f0948ee171e364cb3816fe78f1cf55927561e9029a02532fb8a1091

Observation 42f50f90-736f-45cc-a9cd-a95063fe62fc · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.198968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:27fe4da8fb5aa97222a906935fa6d58005ef4c31552bc6b2ef3dc0d64c2ee239

Observation 642a627a-3497-4036-8e18-857282d8e03d · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.800129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:bfa830ab014132ea61eb1ba1cd3941dce01457d28637217a3199625acd43b21e

Observation 24387372-a44b-410c-8017-87796548bf21 · inbound

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning cites this paper.

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:24.092587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T04:51:05.359636Z digest=sha256:2400b7e82ef289032c18e59da8609fb236546cd1120abc66ef0a4acc3e591585

Observation d819dc1b-c8ce-47a8-913b-7bdd4b79b5c3 · inbound

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise cites this paper.

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.797161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:51:06.448678Z digest=sha256:c2a61dd937561a3d8621da8f8572a8ebbae3856b6901cae2034a8a77b97c116c

Observation 87967254-7898-4ef0-9b2e-114ae8fcd293 · inbound

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation cites this paper.

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.378053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:24:30.660372Z digest=sha256:dc8469cc399b3807227339de1ff669ae030c24c8f2ac59f2328d7ffb07337120

Observation 1fa36e31-3765-47a6-896b-81e21a2f5453 · inbound

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation cites this paper.

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.021269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T23:10:03.733636Z digest=sha256:24c9bdc7abca4264d3839125f3c4f792fe18493c6b8ec4499ac6ab19124309c2

Observation 5b97ba62-2938-4c33-adcf-a8c4fc23d29d · inbound

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing cites this paper.

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:40.156723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T15:48:26.303462Z digest=sha256:418972a303bf1d1e356be07813cb746afd65a749677db0705916b766d62089d9

Observation c68d1aae-9a53-4eb8-8a41-dcb39cc5f058 · inbound

Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs cites this paper.

Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:31.548571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T18:09:17.414031Z digest=sha256:304f3ea3795b59e4c50aafeb49ddcb2ce07512ee680b144283a08b0e1ff165b2

Observation 004c6577-3963-4e08-ba4f-a124e17a89a7 · inbound

C3-Bench: A Context-Aware Change Captioning Benchmark cites this paper.

C3-Bench: A Context-Aware Change Captioning Benchmark AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:10.108934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T21:02:52.529391Z digest=sha256:17f60108743a5ea8d1154b5dc80d39f0c15f0871fa32b1738c005b714cda6015

Observation 4fbfc001-c736-41d4-ac46-390a0ee1ca24 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.664257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:024d4561d0cb7d3edb405537ffedf3dfb7bd35982f60efd4afcddb3a9497ae02

Observation 08d8ff0c-5ae4-41e2-b896-c58ddbfc6f99 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 268

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:5ca5ddcbc1792b0bddb5b20fa3befed0d85e90e66227eb752b865bf540e3df57

Observation 7cc20e6e-6e09-4de3-ace2-c50a812b577f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:03.696680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:03.696680Z digest=sha256:366aba32dd41cbc195e17fcbe559c6babf153af45ef96b9a5e8bd44ecb76b971