Pith. sign in

Paper Citation Record · LEDGER

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2508.16741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16741 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:15:24.398771Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1efc04f9-7977-44a6-8ba2-92ab331fc131 · outbound

This paper cites PRL: Prompts from Reinforcement Learning.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning PRL: Prompts from Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.100117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.100117Z digest=sha256:ac7046169c34488febc3b20bdc05850ab982898aa08281d575cb34f58427823c

Observation ae1e4285-087a-4fd4-a2da-5791b2ed3363 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.139898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.139898Z digest=sha256:739191c5cf3b0bb78c4b314bc8f553a029bf129a51a4f758e5ce4feec9ac2c57

Observation 579e80bc-d222-41cb-9496-958a6861aefe · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.202120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.202120Z digest=sha256:6df51bb9fb5475d53b26f71b6025b664b1d340cbe18238bc00c0186e167c0f3c

Observation ce242bad-4f0b-43f4-807d-e9271fb98361 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.290301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.290301Z digest=sha256:b1b07ecc95b59638cc0d27838f44d1d07e1ffabed6beab18ae0870aa8c0ce721

Observation 6ba090e3-3d41-40cd-bfd6-9d292e8913b4 · outbound

This paper cites Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.339311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.339311Z digest=sha256:30bbcb993461eefa1371c1cd1e8798634ae467cdf9dc47456d2ed4ca82eec4a0

Observation e8d1db19-62bc-4a99-a464-7831a5a7df94 · outbound

This paper cites an unresolved cited work.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:15:24.541148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:15:24.398771Z digest=sha256:a9b45a9ea0335ddd5c2726a9a15ecc65a7de868153daf8afce863047192de2e1

Pith citing papers

No inbound Pith citation observations are available.