Pith. sign in

Paper Citation Record · LEDGER

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2508.16741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16741 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:15:24.398771Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1efc04f9-7977-44a6-8ba2-92ab331fc131 · outbound

This paper cites PRL: Prompts from Reinforcement Learning.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning PRL: Prompts from Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.100117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.100117Z digest=sha256:24bf8990162a7569af5ce48fc8e692a1423895785c417f0a71d6af67fa464099

Observation ae1e4285-087a-4fd4-a2da-5791b2ed3363 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.139898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.139898Z digest=sha256:e661e47c77da27da0aa63374773eab0810a7f5231a14600b71a014686ee88897

Observation 579e80bc-d222-41cb-9496-958a6861aefe · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.202120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.202120Z digest=sha256:2c7ee426016e70374e92ff80df8b5f597fc005b19a88a3d97af9b778a9cf7f18

Observation ce242bad-4f0b-43f4-807d-e9271fb98361 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.290301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.290301Z digest=sha256:7dfa9d085e00bce97e8009451cf82aae4ffb104e6fab2f6207f95a55ef05a1dc

Observation 6ba090e3-3d41-40cd-bfd6-9d292e8913b4 · outbound

This paper cites Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:15:24.339311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:15:24.339311Z digest=sha256:e1bb8e600c587bb50305cb90745365ffad0116113c57625c9399306e3f1a7b7a

Observation e8d1db19-62bc-4a99-a464-7831a5a7df94 · outbound

This paper cites an unresolved cited work.

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:15:24.541148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:15:24.398771Z digest=sha256:8887126992b4d9910c9b33e95d1b7cedc85d0df279ac344b5e0d61c81bee6615

Pith citing papers

No inbound Pith citation observations are available.