Pith. sign in

Paper Citation Record · LEDGER

Self-Generated Critiques Boost Reward Modeling for Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2411.16646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16646 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:25:52.136038Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T12:25:43.079840Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d8777be-64db-4555-bd5a-c50a6fe447bf · inbound

In Context Learning and Reasoning for Symbolic Regression with Large Language Models cites this paper.

In Context Learning and Reasoning for Symbolic Regression with Large Language Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:53:21.144572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T18:50:40.378720Z digest=sha256:f5e84cfb5c895738de0eb3e0a987ed6572d99ecd036a2abf10ceaabdb437184c

Observation 199020c2-2a2a-4000-8eb8-0509df32500a · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 278

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.406469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:8ccacc7242d24d5fe5f0136d59b73fb255dd4d34485558ca1443a3a1df202cc4

Observation 2bf10488-a08d-454a-a18f-e96d366e2c8e · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.136038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.136038Z digest=sha256:ea2fcfb6e99da95687c1ce8018d045cb07e1cb52583c7473da68afe6aa6560b2

Observation d4229abe-225d-4a77-996c-286beb3372a5 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.481768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.481768Z digest=sha256:469c2f23f6c9bfd6fe575048f8d247f6e81e8ddd43f4512560220a661756188a

Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.445894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.445894Z digest=sha256:eebf08b5044304b17d912b360768ece0807d63c77b9315e5a6f7dcb7b84fbdea

Observation a978274c-3e04-4b41-9561-e696d770c2d0 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.139918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.139918Z digest=sha256:23f9bae731a10057ea6be67c8d2ed66b71ac6810d221e1ad25841121f46af840

Observation 422ef1bf-506e-4adc-bdba-de513b822b03 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.294496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.294496Z digest=sha256:4e13890a4248e58988bd539747ed9ec1157a02815fb7a60c9a511b90f84baac4

Observation 58c58590-3bd8-48e6-b25b-a0b1448d2c79 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.461763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:be5efe4fb2f48582065ffa9ed6a4cdf8d244ea999cbb9b099f662ff7dbfd32f7

Observation 0dfaad6f-0025-45e8-99fb-15e8ced842ba · inbound

Test-Time Verification for Text-to-SQL via Outcome Reward Models cites this paper.

Test-Time Verification for Text-to-SQL via Outcome Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:25:43.081624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T02:10:13.533450Z digest=sha256:b519c8437200aef3cd16ab16556a0d46d2874990e9a314cb09a3011b7dc97118