Pith. sign in

Paper Citation Record · LEDGER

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 8 inbound Pith citation observations for arXiv:2412.11605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11605 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:51:18.719130Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.104812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:36:24.079882Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57de65ab-616b-4a8b-8d9e-2ccfd04ca63a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.562507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.562507Z digest=sha256:61132a2b98471b86c6ab3f01fd8332fbdaa9523a685ca70e1fcdfefd300a0acf

Observation 2a56eb26-e843-4d45-b3a8-09231215c565 · outbound

This paper cites Language models are few-shot learners.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.566692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.566692Z digest=sha256:ae445dfe9311d81ad229ef48948d1ba0458e208fb6390a7e9dfc0b336f85105b

Observation c67baca8-8d1f-49f2-afdf-b572993add76 · outbound

This paper cites Towards Scalable Automated Alignment of LLMs: A Survey.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Towards Scalable Automated Alignment of LLMs: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.568912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.568912Z digest=sha256:ceb343bdb4bc56c4531605cffd3828307dc13ac1c505da2deebd8e91357ea4b5

Observation 92e84509-0d18-462d-9e33-f65eb5610008 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.571471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.571471Z digest=sha256:708c5a3fa87bc6b211d6b5fadc1dec62ae2f1290772d6dd18f06a1409f163903

Observation a1a27e19-125c-4ff9-b0e4-4f732b92f71b · outbound

This paper cites Black-Box Prompt Optimization: Aligning Large Language Models without Model Training.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Black-Box Prompt Optimization: Aligning Large Language Models without Model Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.574514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.574514Z digest=sha256:8deb43dbe236d09e33b3160c5a03e0d795834d6cd6c7c9c016097f470cfa135a

Observation 46e1bd5c-833f-4365-8255-6f15070d0008 · outbound

This paper cites AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:51:19.026520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T14:51:18.576982Z digest=sha256:f876be776c3ad690fcd90c7cc880eb702d114c8dbf300e3545e36c6bed080e42

Observation b8148eb9-779b-4dab-8a70-0681e1a5905c · outbound

This paper cites Palm: Scaling language modeling with pathways.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Palm: Scaling language modeling with pathways

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.579920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.579920Z digest=sha256:e1378ad0073b6373963ca92603216b952804bb236970b7122ffcce00cc663af7

Observation 689aa6b7-a959-4f68-8645-e72b84129c9c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.582282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.582282Z digest=sha256:28df91f6f48ae7096c23750e330d759ee2dea265fa1bc0955ab9f595f01f07c3

Observation a477687a-35f4-4cfe-846d-b374b1036f91 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.585672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.585672Z digest=sha256:a00cfc7192a8db679801f8ca2774d8e836ac2a10fc01d24d26a51aedf5929686

Observation ae51ea5d-be46-4f1f-9106-86da25c7f207 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:51:19.119819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T14:51:18.589200Z digest=sha256:375f9412b9e0ba2cf81ee51f45a785a336597c55d368c30c9dee23f290c4d3cf

Observation 772a4b83-580c-490c-859a-a1a9ed70ee4a · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Reinforced Self-Training (ReST) for Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.592288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.592288Z digest=sha256:826693c3bab180eb168458bd1b40a758040fc1715b1047fe9bad0fcbb840f654

Observation 595c4fb4-5d71-4f7c-8004-2f4ce6bffdad · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.595505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.595505Z digest=sha256:47a1f04e82dca7158b34180cd687ac04c4b16d6b6e340be01b6355c1404e9fca

Observation cc5c5a8a-18c7-4af0-9504-f385e2c6df51 · outbound

This paper cites ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.598349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.598349Z digest=sha256:791830a58c6645f7cd36c996151e35502fa5b779d11f2e04bf6bb0b5d4b0c04b

Observation 2dece077-fdc3-4aee-807a-4a848688482b · outbound

This paper cites Mistral 7B.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.601853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.601853Z digest=sha256:9e19571539717070f2e750cb3ffc24dc368cd4f872f4cfad89a6a0254534dcf7

Observation 716ad6ad-80f5-412e-a7af-39dba8696df4 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.605394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.605394Z digest=sha256:906c2acdfea59ecdd3a9f8ed91fee2d389a4f0b82784d9804e7d2f2c43cfc8b0

Observation e54c3f68-2f5b-4f59-b1f4-e6d3cb21bab6 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.608543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.608543Z digest=sha256:c546d4e8ac87feab2423642196c7c764cfd80b06684f5c7f269fa8bae5791a72

Observation 414f43c0-a9f7-4bbe-addf-8cde6501204e · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-Alignment with Instruction Backtranslation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.611036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.611036Z digest=sha256:bfd3fdc42ac66f3edc1c0bebc1d9c08c8c54eb89bc5f20b599ce828c09f17893

Observation bada4aba-f69d-4041-baf9-fe357f83e9df · outbound

This paper cites Hashimoto.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.613965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.613965Z digest=sha256:d573a582cfe204a08253c3c5d497f5d78d00c204ce253cd32c91f3c6e67ff991

Observation 18cf2028-cf49-4779-8b9b-8e4147233836 · outbound

This paper cites Best Practices and Lessons Learned on Synthetic Data.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Best Practices and Lessons Learned on Synthetic Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.617017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.617017Z digest=sha256:5328da1af37461c543821ffeb228b00a4a92eb43f248ec596ee27648dc0749c4

Observation 65a52709-18eb-4c80-acfd-56ec685d144a · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.619871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.619871Z digest=sha256:6eda531c21df51c68c812252554bc5735bd06dab9ad959e455d32598f6a2ed5e

Observation cdc9e422-2042-4afc-a1fd-d39c43b2f1ea · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AgentBench: Evaluating LLMs as Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.622897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.622897Z digest=sha256:fc40af7202bfb730d00c0c3707e921c25da7188cf6b36f3708ce674a12d8af08

Observation 6b76b4e4-7708-42ec-9747-67b22d96984c · outbound

This paper cites MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.626064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.626064Z digest=sha256:f155f3a16cdf95ebe4f03106f83a1a6f818c7284e96e0f679df274263e943aee

Observation ec14df0c-c265-4b60-b059-9cd8186aa51e · outbound

This paper cites Large language model instruction following: A survey of progresses and challenges.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Large language model instruction following: A survey of progresses and challenges

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:51:19.105492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T14:51:18.629482Z digest=sha256:28b5575a042e252905b60ad35050a2680e86ab021d8fc121d67b8faf613e267e

Observation eefc9a6a-6c60-4004-a156-5cdad6aae022 · outbound

This paper cites SELF: Self-Evolution with Language Feedback.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models SELF: Self-Evolution with Language Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.631807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.631807Z digest=sha256:85b7a98fa24095f5200f483aefd649b77b039ef6880710da2c7edea898db3134

Observation 1196511b-e4e2-4d97-8ba8-4b17a71decda · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, 2024.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Introducing meta llama 3: The most capable openly available llm to date, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:51:19.097794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T14:51:18.633886Z digest=sha256:8f5da4d6eaffbf288d73839fadb148debe2f3ca6cc17a8bbc1377ead005d79eb

Observation b96ed1d6-e0c9-490a-be40-41886f7c2c03 · outbound

This paper cites Training language models to follow instructions with human feedback.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.636876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.636876Z digest=sha256:fe95e0d373001454518ab37ea3a1c85ca2d1ce71ff2f0383f8cd6294afc9027e

Observation d8ff5b6e-9251-432f-a558-181e977afaa8 · outbound

This paper cites Instruction Tuning with GPT-4.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Instruction Tuning with GPT-4

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.639128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.639128Z digest=sha256:452e23870484e631f9969d75391c95d998807fdcdfff5d1534c26bb811c2f349

Observation d1750171-cdc5-41f6-bd96-f1c6d3e70c9c · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.641907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.641907Z digest=sha256:875d8d5626503e8bbd39e952a88a56d02eedb7d3c661726446d2e6f5395549ba

Observation a8a12d52-6e08-44f8-9934-e1e3d3174476 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.644813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.644813Z digest=sha256:9ec384752d81f391732a76f635cdcea37c46af22a179fd94cafb89e9ad4999f5

Observation fc19121f-69ac-4dfe-b319-23e7e9a0d9ae · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.648035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.648035Z digest=sha256:50666919ae95afc872db4ce6070b3e9db7db16ca259858cb4a6f1491cd3a72e1

Observation 74ab0306-7eb0-4176-b81d-1747ad80aa4e · outbound

This paper cites Ai models collapse when trained on recursively generated data.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Ai models collapse when trained on recursively generated data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.651675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.651675Z digest=sha256:cd4bf5469538bb68b747a3bc8cbdd3e85c14f1b9a7fcb2436c69a05b1dc72769

Observation d6f3aafa-d995-423f-a1e0-320f5be12d38 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.654555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.654555Z digest=sha256:044473a0037e5d4477ce92bc86adde074f53d2ad05433636ed9076e4f483a8de

Observation 2d93f471-f4fe-403f-915b-ea17b8aecbda · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.657854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.657854Z digest=sha256:22292c33fa6070b3425fb5eab6354f153a9fae76a3069a7c4c052567b3930fd7

Observation 3e4c7e1c-cbfb-4c83-9fdc-ac514986013f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.660593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.660593Z digest=sha256:f2506eb187e7bfdbb5e41f7b3943566f03625f5662b11e0a406de2d03767994a

Observation 17207f7b-0d60-447a-abcb-15503639536b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.663299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.663299Z digest=sha256:35bcbec82d90759959e35bb6c2375a252df5e19a0a173bcc1283dae46c8844d0

Observation 26dac730-f7a6-4c1b-9eb9-6d2f3d047b06 · outbound

This paper cites Self-instruct: Aligning language models with self-generated instructions.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-instruct: Aligning language models with self-generated instructions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:51:19.079237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T14:51:18.666177Z digest=sha256:f8f8ee08d33952388658b0443c44b747e824a857608abeb55f743404d3f3ec9f

Observation 298a892c-9522-4bb3-b10f-685557af61ca · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.668935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.668935Z digest=sha256:92990e6fe4cc3d7365ef05bbf7be312fbd53ce0f9642fc55aee97b1785034c9f

Observation 9e28c132-3a02-448e-aa9b-d43b8ed73fba · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.672843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.672843Z digest=sha256:600858035cbff4b288408122ef517367cb86ee5e15538c5f63750e1610005d16

Observation 7e06ae4f-ec94-4352-b590-66e552058bb6 · outbound

This paper cites FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.675698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.675698Z digest=sha256:5d9c1611e733253e394082c3df76559706b2ffc3fdb79714e3942d4764de5218

Observation 439bf495-eaa2-4b12-9582-53c0e242f58c · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.678510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.678510Z digest=sha256:eea774f62c092e71f2cfc0427f84913cca0a633fbd4c2083eac5d10c862e7094

Observation ac176cf6-fe98-4f0d-9ebf-3430cf731dfb · outbound

This paper cites Self-Rewarding Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Self-Rewarding Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.681880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.681880Z digest=sha256:6bef2cfb8fd06874d41ea1d2febf954fc8a4c48d0cdd4dc348c96935a7e5a213

Observation 5295b7bf-ca1a-4c53-9fc8-f3f5a76ed7e1 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.685681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.685681Z digest=sha256:2323611982b44ad83fe451dc7e8f0fe0917909b1d5d23957dc6c2b85b0b4f2d2

Observation fdde6bfb-9072-4a5c-baab-4baf4322095f · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models GLM-130B: An Open Bilingual Pre-trained Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.689154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.689154Z digest=sha256:b0b9eb4952e2ac993146ebd43fee744b103da6f4c9468ed2adae9af8adba8009

Observation ad4101d8-ce7e-441d-88ab-e5be788eea72 · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Evaluating Large Language Models at Evaluating Instruction Following

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.692054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.692054Z digest=sha256:168eec14f466a6cac7eeefa0362f528af430c78d2a279536901e0a0d51e09048

Observation 321cbda7-ef41-4c1e-a222-db8c3246af63 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.695790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.695790Z digest=sha256:18dfd9a99465511cf08ec558b4d4e8bbf3238907fc62ea4aaee5c8186e566074

Observation a6e495f3-aa8c-4527-8f89-a30ed693fa2d · outbound

This paper cites Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.698426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.698426Z digest=sha256:5f3bd4869579018e8c703275845d00a02d1d443a7dc267b036f525fa12b96cfc

Observation 2eb81bb3-77e0-42f7-8b26-0a1c5e131395 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.701267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.701267Z digest=sha256:4ab5f7ffc5f87d56b8469352df72761031eed8d43c0719dffc9c56fc5e9bd528

Observation 7695402a-4a38-4639-b6d3-00d37d93e8b7 · outbound

This paper cites Lima: Less is more for alignment.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Lima: Less is more for alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.704382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.704382Z digest=sha256:a252c9af91c2729cd57b9c0b705af6f5b54177e0ca040084e6778b60c9ff87f5

Observation c48ead95-c7d3-40d5-b2d2-db50a437c5a4 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.706443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.706443Z digest=sha256:c065974dbb9903261e826c7b8b3c12f5f1f7f16c2dd6f058559bf9c04e323965

Observation 22be3101-05c4-4780-8986-b71ad87d10e6 · outbound

This paper cites write newline.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models write newline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.709497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.709497Z digest=sha256:98c95fdf37b193f69aec1b2f9d9b42c9def8483166bc2c74cc345db14216e639

Observation 12695c63-4b4d-44d4-b70b-9f731d3b1beb · outbound

This paper cites @esa (Ref.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models @esa (Ref

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.712913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.712913Z digest=sha256:dc05dcd2ed582e77eda844d178ab3d87efea76db9b540dda77c8fe2bde3b522f

Observation 0d0961ef-68ec-4c3d-beff-c20d67fa8e58 · outbound

This paper cites an unresolved cited work.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.716694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.716694Z digest=sha256:3e936ee783d9cc13e95b60a28bfbadf9433373351e7f3ed50612a7bea3a2333f

Observation 5a7b2024-f3df-43ac-9225-66af1264f474 · outbound

This paper cites Such an ability is well-suited for and often optimized by preference learning.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models Such an ability is well-suited for and often optimized by preference learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.719130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.719130Z digest=sha256:9515cbd9f71eaeec3f37454dbc142241f8b0f7744fe17e5648514af47dade5f5

Pith citing papers

Observation 789ba3c5-d7d7-4917-a534-0f9f7cb2afbb · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:51.982894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:51.982894Z digest=sha256:565d313b9be30f543c33604e9eab4cab7feb3c61b9cdbbf9b54f0be97fe9c9ae

Observation c74deeb1-b09f-41a6-bdbe-7bb7827abdc6 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.082230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:61a78ef09e5bedc12c962312c06dee6a6c743dd5b750b5fc40c0b03cf3cfe409

Observation 0e22977d-aa26-4dce-8dc4-f2f6e484023e · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.964189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:fa5f07876a213c5663861c8b1d624d72811e1595d684c4766e05594543842c7b

Observation 98993b75-5176-470c-9c55-df619b6fea16 · inbound

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following cites this paper.

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:17.196428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:17.196428Z digest=sha256:b4fb245ee8ce4f3932e1578b15005303a6d5cca506f3dbf1d406056878b9aeca

Observation 901a5319-8782-48ef-9be7-675dd91527fa · inbound

SEIF: Self-Evolving Reinforcement Learning for Instruction Following cites this paper.

SEIF: Self-Evolving Reinforcement Learning for Instruction Following SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.978752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:51:09.514927Z digest=sha256:9c644ef6b6719d55f13943e1651050b20c79e8b0acc7b6590d05ff687fad6018

Observation 9694d02b-353d-4c9b-ae25-d9858c05f098 · inbound

Anchored Self-Play for Code Repair cites this paper.

Anchored Self-Play for Code Repair SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T01:53:44.674517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:53:44.674517Z digest=sha256:00fac68efa9a77d513464622bd5eebb1353d94b0b912e661cc416e2b72e90482

Observation 064ca75f-2f59-4c0e-88c8-753332ffd3d5 · inbound

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills cites this paper.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T04:28:45.255280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:28:45.255280Z digest=sha256:af2fb6f23a04a8b7f4368996ec3c6a704cf64e944b0b98272a207b9a3d3492aa

Observation bb77a871-757a-4ad2-842f-b800da01771a · inbound

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information cites this paper.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.104812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.104812Z digest=sha256:6be52f5a472bc4fab37af5ff070b6da77fc4c9880a84c326ff47ec7c1dd6a00f