Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:51:18.719130Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 8 inbound Pith citation observations for arXiv:2412.11605.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:51:18.719130Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.104812Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T01:36:24.079882Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57de65ab-616b-4a8b-8d9e-2ccfd04ca63a · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a56eb26-e843-4d45-b3a8-09231215c565 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Language models are few-shot learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67baca8-8d1f-49f2-afdf-b572993add76 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Towards Scalable Automated Alignment of LLMs: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e84509-0d18-462d-9e33-f65eb5610008 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a27e19-125c-4ff9-b0e4-4f732b92f71b · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e1bd5c-833f-4365-8255-6f15070d0008 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b8148eb9-779b-4dab-8a70-0681e1a5905c · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Palm: Scaling language modeling with pathways
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689aa6b7-a959-4f68-8645-e72b84129c9c · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a477687a-35f4-4cfe-846d-b374b1036f91 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae51ea5d-be46-4f1f-9106-86da25c7f207 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 772a4b83-580c-490c-859a-a1a9ed70ee4a · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Reinforced Self-Training (ReST) for Language Modeling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 595c4fb4-5d71-4f7c-8004-2f4ce6bffdad · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Measuring Massive Multitask Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5c5a8a-18c7-4af0-9504-f385e2c6df51 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dece077-fdc3-4aee-807a-4a848688482b · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 716ad6ad-80f5-412e-a7af-39dba8696df4 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54c3f68-2f5b-4f59-b1f4-e6d3cb21bab6 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 414f43c0-a9f7-4bbe-addf-8cde6501204e · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-Alignment with Instruction Backtranslation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bada4aba-f69d-4041-baf9-fe357f83e9df · outbound
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18cf2028-cf49-4779-8b9b-8e4147233836 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Best Practices and Lessons Learned on Synthetic Data
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65a52709-18eb-4c80-acfd-56ec685d144a · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AlignBench: Benchmarking Chinese Alignment of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc9e422-2042-4afc-a1fd-d39c43b2f1ea · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AgentBench: Evaluating LLMs as Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b76b4e4-7708-42ec-9747-67b22d96984c · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec14df0c-c265-4b60-b059-9cd8186aa51e · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Large language model instruction following: A survey of progresses and challenges
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eefc9a6a-6c60-4004-a156-5cdad6aae022 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models SELF: Self-Evolution with Language Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1196511b-e4e2-4d97-8ba8-4b17a71decda · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Introducing meta llama 3: The most capable openly available llm to date, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b96ed1d6-e0c9-490a-be40-41886f7c2c03 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Training language models to follow instructions with human feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ff5b6e-9251-432f-a558-181e977afaa8 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Instruction Tuning with GPT-4
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1750171-cdc5-41f6-bd96-f1c6d3e70c9c · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a12d52-6e08-44f8-9934-e1e3d3174476 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc19121f-69ac-4dfe-b319-23e7e9a0d9ae · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ab0306-7eb0-4176-b81d-1747ad80aa4e · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Ai models collapse when trained on recursively generated data
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f3aafa-d995-423f-a1e0-320f5be12d38 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d93f471-f4fe-403f-915b-ea17b8aecbda · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4c7e1c-cbfb-4c83-9fdc-ac514986013f · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17207f7b-0d60-447a-abcb-15503639536b · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26dac730-f7a6-4c1b-9eb9-6d2f3d047b06 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-instruct: Aligning language models with self-generated instructions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 298a892c-9522-4bb3-b10f-685557af61ca · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Benchmarking Complex Instruction-Following with Multiple Constraints Composition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e28c132-3a02-448e-aa9b-d43b8ed73fba · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e06ae4f-ec94-4352-b590-66e552058bb6 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439bf495-eaa2-4b12-9582-53c0e242f58c · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac176cf6-fe98-4f0d-9ebf-3430cf731dfb · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-Rewarding Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5295b7bf-ca1a-4c53-9fc8-f3f5a76ed7e1 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdde6bfb-9072-4a5c-baab-4baf4322095f · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models GLM-130B: An Open Bilingual Pre-trained Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4101d8-ce7e-441d-88ab-e5be788eea72 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Evaluating Large Language Models at Evaluating Instruction Following
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321cbda7-ef41-4c1e-a222-db8c3246af63 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e495f3-aa8c-4527-8f89-a30ed693fa2d · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb81bb3-77e0-42f7-8b26-0a1c5e131395 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7695402a-4a38-4639-b6d3-00d37d93e8b7 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Lima: Less is more for alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48ead95-c7d3-40d5-b2d2-db50a437c5a4 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Instruction-Following Evaluation for Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22be3101-05c4-4780-8986-b71ad87d10e6 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models write newline
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12695c63-4b4d-44d4-b70b-9f731d3b1beb · outbound
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d0961ef-68ec-4c3d-beff-c20d67fa8e58 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7b2024-f3df-43ac-9225-66af1264f474 · outbound
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Such an ability is well-suited for and often optimized by preference learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789ba3c5-d7d7-4917-a534-0f9f7cb2afbb · inbound
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74deeb1-b09f-41a6-bdbe-7bb7827abdc6 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0e22977d-aa26-4dce-8dc4-f2f6e484023e · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98993b75-5176-470c-9c55-df619b6fea16 · inbound
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901a5319-8782-48ef-9be7-675dd91527fa · inbound
SEIF: Self-Evolving Reinforcement Learning for Instruction Following SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9694d02b-353d-4c9b-ae25-d9858c05f098 · inbound
Anchored Self-Play for Code Repair SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064ca75f-2f59-4c0e-88c8-753332ffd3d5 · inbound
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb77a871-757a-4ad2-842f-b800da01771a · inbound
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.