Pith. sign in

Paper Citation Record · LEDGER

TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2407.16574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.16574 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:06.580832Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:14:10.944002Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1be80ad1-8f70-43a1-9af2-c61bda13c2d9 · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.443264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.443264Z digest=sha256:92598bc5d16a9eefac046d7901c42d4f52e42ff75e19a5c9900065659e823de1

Observation 106e2763-17de-4a95-bce3-91c8dbafc6cf · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.580832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.580832Z digest=sha256:c84203b19d5b4f4e4f499f82506dd6afbc88462c9684a8c0a6f5d468914b0ae0

Observation 8a80d37e-8a54-43cf-8c36-634b7706e89c · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.576255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.576255Z digest=sha256:1f0c0cef141be8a582071752da4f563e0ef2ec9ea9daf66818d983476fd03936

Observation 70da07b1-0cf1-44f1-ba5c-0c3613c1b6ab · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.365581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.365581Z digest=sha256:7ae8a7c19eed279c3f3bce7d1a0c9368e90fe1a9a7e9db4f619bdef1b57ed87d

Observation b795ecb2-01f1-4d2c-9ace-c13414b4b50b · inbound

SGPO: Self-Generated Preference Optimization based on Self-Improver cites this paper.

SGPO: Self-Generated Preference Optimization based on Self-Improver TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:49:12.744748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:49:12.744748Z digest=sha256:ae8fd88c907389b2bbdbb780693ba4977f8191ede268854fe953f067aaed5fbc

Observation cf5c47bb-8e85-47b3-a615-562ec3d23ee3 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.945445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:3ef641c002bcc0efcec8471dc1dc707d29eff1e7ad1202bfc5a0f0c1147fbdbc

Observation 9426f27c-1d05-4acf-8f71-d3236dfb0862 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:46.007242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:46.007242Z digest=sha256:e5d9903bc69ecc9688699fdd316fedc949b4a4bf0eabb99b06d81ded20f9093c

Observation 69356ac8-eeec-4bed-9ebf-82ed49ed2f7a · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.503407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.503407Z digest=sha256:bf37dca8a8cdabb8a3f16a2b8c4302e004e5a65dc55edcf2ada96470166a56af