Pith. sign in

Paper Citation Record · LEDGER

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.04412.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04412 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T19:23:10.328625Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86ca782d-3e33-40d3-9bff-62a237458a54 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Online difficulty filtering for reasoning oriented reinforcement learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:261ecb210063fa08631632cd7a0ac5d4b98b70cbe9fd5968e85be160502ea51f

Observation 2d87c94e-8486-482f-a55d-a04a067b6410 · outbound

This paper cites Curriculum learning.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Curriculum learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:b6d477fecab5ce6de944c79283ad8ac9060670c189f4e51f223b440accfce01e

Observation 67c3a10c-c837-4e45-ab52-cdb093e1daac · outbound

This paper cites Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:b92c147236046c6ae9c9a714b4c50d0d89ff95da828360e6e14d94adf48bbbfa

Observation 1f201116-9926-4dba-8371-ccc8f0416cd6 · outbound

This paper cites Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:790e58c4834044536f373a020dc31ecc0ef1ff58c489ce81aae8fe3875856b74

Observation c7525426-0c1d-4d56-9d41-89c4128777ad · outbound

This paper cites an unresolved cited work.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:1832c10c99bf432c4b8ee6510ed58e1f7299a582909f59b1cd93183637fae133

Observation bfe5363c-8cd6-43b6-8511-d624128d2c81 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:129a1a9488a114bcf4a840267626b72dc580257107284127843090f6cb80724c

Observation 4a067c88-8e7d-4ff9-b524-31d4db693b3e · outbound

This paper cites Advancedif: Rubric-based bench- marking and reinforcement learning for advancing llm instruction following.arXiv preprint arXiv:2511.10507, 2025.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Advancedif: Rubric-based bench- marking and reinforcement learning for advancing llm instruction following.arXiv preprint arXiv:2511.10507, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:aeb2961c9f204e3bb6d53b625e09197e419533475d4d544bb8ee0154da4a179c

Observation ec597410-04d9-4162-9c9e-1b79e06bf902 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Distilling the Knowledge in a Neural Network

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:524f83726922a11395d935f79f709c04d890e8710e95e96141885ca6ebc83e2e

Observation 70b3b8f8-8bce-4f17-8ae9-08b994026120 · outbound

This paper cites Large language models are reasoning teachers.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Large language models are reasoning teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:ab757888c652a5a655e775bcb001ecec7a90a8893f521ba402b46864d43d9bf7

Observation 10135b71-3735-4e28-ad15-dad414f60591 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:2675b1886c613e23b4bd88edb0117fb4590d83cab42a50540d1de658f547d0bb

Observation 71a1ee3f-722b-433a-b011-e4cf3f80c624 · outbound

This paper cites Vcrl: Variance-based curriculum reinforcement learning for large language models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Vcrl: Variance-based curriculum reinforcement learning for large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:383d83be7f8f5e0faf98f6c8137797f42d61f62a1b9ee86172236b2dac7baa35

Observation 3f2bc3ba-3117-405a-b674-af0cc9a9f8b8 · outbound

This paper cites Followbench: A multi-level fine-grained constraints following benchmark for large language models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Followbench: A multi-level fine-grained constraints following benchmark for large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:5f4dfe511823af6bf20a4687a6971894faf0bd10db6c1de5cb9a0f513ad08296

Observation 7f84151a-9e60-47a2-96a0-3e0a6fc49322 · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Prometheus: Inducing fine-grained evaluation capability in language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:939a81cc0f070fb86c1fbf7b650f0be8bc09ae7961b0d1e720f2515901b75d9e

Observation 8c0fc07a-24b2-428d-b102-4531fd40fac9 · outbound

This paper cites Language self-play for data-free training.arXiv preprint arXiv:2509.07414, 2025.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Language self-play for data-free training.arXiv preprint arXiv:2509.07414, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:fc2b0bc53151f12ee1f54b25d22326dfb228558d7ebb3a280c5b73d434bb15d2

Observation ab0e0874-f0df-400b-a20f-d5bfa6ccf4f9 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Gonzalez, Hao Zhang, and Ion Stoica

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:0b2ba9156255d0af9e816860cbb486342d6d7880bb1af610c60afbf4bd48c91a

Observation 7e86ef38-5da0-4936-ae8d-2890fc3262d7 · outbound

This paper cites SPICE: Self-play in corpus environments improves reasoning.arXiv preprint arXiv:2510.24684, 2025.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL SPICE: Self-play in corpus environments improves reasoning.arXiv preprint arXiv:2510.24684, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:c45fe62a25374136e343c8d060b829df4b2bdb4c90277a7e298e8b00989d547c

Observation 2b7f806a-675f-4eca-a3ba-3f1f26204c5d · outbound

This paper cites Openrubrics: Towards scalable synthetic rubric generation for reward modeling and llm alignment.arXiv preprint arXiv:2510.07743, 2025.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Openrubrics: Towards scalable synthetic rubric generation for reward modeling and llm alignment.arXiv preprint arXiv:2510.07743, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:dd25c11f8cf685d0af52fe8691e58bc7a264dc95aeee3897efb90af6f66d8d49

Observation e89443e7-8cb3-4988-a80b-b78964317481 · outbound

This paper cites G- Eval: NLG evaluation using GPT-4 with better human alignment.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL G- Eval: NLG evaluation using GPT-4 with better human alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:75eee459c87395833dff65c356eacf6ee5562a9de0691d2f9ace208c24550b74

Observation 5687cb72-c2ef-4600-b05e-73c575c5c4d6 · outbound

This paper cites Aligning with human judgement: The role of pairwise preference in large language model ev aluators.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Aligning with human judgement: The role of pairwise preference in large language model ev aluators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:b69dc32f3de6fdab996aea1153437abc5660ce3d7755ae2a24031d906e8e0ed5

Observation e55804b4-9c82-41dc-bb7c-b7f481b1ef26 · outbound

This paper cites LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:3985b63128944ddfc3c867b70b1d618599ad7089aa688337cdb8dbca33d0136b

Observation 983e9400-539e-405b-866e-a1491cdbc756 · outbound

This paper cites Teaching small language models to reason.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Teaching small language models to reason

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:f8c051e9f03c871476a0e1c386d56e4d099004169c0c1fb12ed6c8e22d014d11

Observation 1278c5ee-1690-430c-bdbc-3e3d2874e800 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:213072ede552a6c3d0ff412bed5fe3e56e1a7ec2ee279f6d8ca070f4ed5d9338

Observation f35c9b4c-b09f-4ce6-8843-6f0de04c0504 · outbound

This paper cites Infobench: Evaluating instruction following ability in large language models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Infobench: Evaluating instruction following ability in large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:710c35e1bcf0ef5eb15ba210ec85036405932c781d0df0af5c7b927732d6b0b9

Observation 2bce996b-8755-4747-b4dc-2fe50057cfac · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Direct preference optimization: Your language model is secretly a reward model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:3a15d1e9180a2a718b752d8fd30d1115ea614f0677c400533a93192fbb3c9a5c

Observation 2aa8ab28-6e57-4d51-8fd6-89f7137d66ab · outbound

This paper cites Proximal Policy Optimization Algorithms.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:ec9a736cc71523573c3e1d6364caa774e48f414df7cb78119ead0072754855b9

Observation 2f65330f-e8bd-4e57-8198-5616dd2a66d6 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:d58c7802dd1929ca10f90005a3e79e795d41bc4f9523807af3fbcb1b0b3540a1

Observation 1053f9bb-bfb5-4c7a-aa61-bd5c0dfc8d9d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:0a40d310e766bcdfc56c95657a03547f500487d63b4b63a081458073ddbd5649

Observation ca9b0ce8-4367-44b0-8725-92dfd438c0cc · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Hybridflow: A flexible and efficient rlhf framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:7b580a7f80dcf2155b9c9261db9fd2a6dd0b884f4aa162454e8f8c0c3f0dd2ae

Observation 38e24633-ec37-4cc6-b08f-13e96808deae · outbound

This paper cites Learning to summarize with human feedback.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Learning to summarize with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:9cdbdfd83c481ed2e3e9db88b9aa1572bbb9b2fe8a7d4cd54f7d9c943ba5253d

Observation 6c52ea9e-8729-4ada-85cb-f5d7487a4626 · outbound

This paper cites Sutton and Andrew G.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Sutton and Andrew G

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:9b109cfc87499d48012ea83f976313c538070a3dcac3affbba34cbc41a79493f

Observation b52b05f8-23fd-4f93-9e5b-f95633968524 · outbound

This paper cites Proximal curriculum for reinforcement learning agents.Transactions on Machine Learning Research, 2023.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Proximal curriculum for reinforcement learning agents.Transactions on Machine Learning Research, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:6d185288ae76921c0ae4d9d9ee4c208c3e95a79319b19436bd4453281eeb87f8

Observation 8fb5304f-16d3-423e-bbc4-f6616d4033a6 · outbound

This paper cites Checklists are better than reward models for aligning language models.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Checklists are better than reward models for aligning language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:be47a31310a87de9190b96b4cf417cc8406e854365d46b24a58491cc9200319f

Observation be063aa9-ba5f-49f8-9873-c5a00b58fc70 · outbound

This paper cites Williams.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Williams

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:1b90cec841d0a021122bc532698be39bcd75fc96f09810dae078344beb7419dc

Observation 337e28fb-23ad-4217-b57b-4dd1cc1dcc5f · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:da4048ad9b933a34cdd853efa800f3890b027dde9e55ec0c28a2d4630b013c1c

Observation 6a9f4e00-9704-4ff9-85f3-e8357cc351ab · outbound

This paper cites Qwen3 Technical Report.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Qwen3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:dec597ea8db0b1d787b40ebb39c37cf28e19d2350cac2b3dc6f6e4a21d6498f5

Observation 7ff97028-1fd7-4a5f-a61d-4d5f6b572086 · outbound

This paper cites Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:a555f11e53cc72374f6578c13c2d0ee0e29e31b448325febddcfab1be5e94a06

Observation 004c5abb-9941-4284-85b1-43c7f583578b · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:788975891a2e68581b5ec733de466da86045e3aef86f5c4518c115b7896cccfa

Observation 88526c1a-bb01-437b-b255-0f591bb5aaa7 · outbound

This paper cites Wildchat: 1M chatgpt interaction logs in the wild.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Wildchat: 1M chatgpt interaction logs in the wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:0ae618aab266609771ffdedd870f1d0c64f568a83036d3e17f605fd841e907b3

Observation 252353ac-9d46-432b-a971-9ed7b3db867e · outbound

This paper cites Xing, Hao Zhang, Joseph E.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Xing, Hao Zhang, Joseph E

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:879d3a77c40a15273c9cef67a970e45a735197b3295355e07be0e637074ccd6e

Observation 68d79dff-b85a-4d60-ae93-6e56fe66c968 · outbound

This paper cites , and the criterion might be.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL , and the criterion might be

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:dcc0e284cf8f47205d1bc5478fe2045cc1cf82bd72b27bea8b3980af900f23ac

Pith citing papers

No inbound Pith citation observations are available.