Pith. sign in

Paper Citation Record · LEDGER

Value-Free Policy Optimization via Reward Partitioning

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.13702.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13702 v4

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:49.368412Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:20.906356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d1e438d-cf68-4314-9294-e3aba8e89a37 · outbound

This paper cites Rrhf: Rank responses to align language models with human feedback,.

Value-Free Policy Optimization via Reward Partitioning Rrhf: Rank responses to align language models with human feedback,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.821253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.229111Z digest=sha256:b15b164a275ee65e48dd8d74b0aed72783bb1aa6ccdfd163e4f382aa7fe4e924

Observation 2be3843c-13a7-4a96-be30-6551e33c15e0 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Value-Free Policy Optimization via Reward Partitioning RLHF Workflow: From Reward Modeling to Online RLHF

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.233090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.233090Z digest=sha256:488bd473ea0198f49f42440ddc664d5daca04ef146f2870b939003b4eaed3bce

Observation c6740c56-52c9-4774-bd8a-e41a14d51045 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,.

Value-Free Policy Optimization via Reward Partitioning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.811145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.236673Z digest=sha256:17c5a63b3c3242513682f20fdbd34fe1411ce499559182fe441e6ba0005b0dc4

Observation 0364f190-1d28-4505-b39d-323425002d42 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Value-Free Policy Optimization via Reward Partitioning Direct preference optimization: Your language model is secretly a reward model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.241247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.241247Z digest=sha256:df46c1da9e4a5ca2204159c6769da0eaf327266571633b1a7951f08157cd2caa

Observation 170bbaa3-093d-4e6d-a9ea-1ab59660b784 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment,.

Value-Free Policy Optimization via Reward Partitioning Generalized preference optimization: A unified approach to offline alignment,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.795028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.244822Z digest=sha256:7b2d084050cfb40d7b7791218bedc395d5eb51ef007dc2476ae5b68b12b37ea5

Observation 848849fa-a948-449d-976a-015b8f7fb801 · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

Value-Free Policy Optimization via Reward Partitioning Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.248057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.248057Z digest=sha256:d07559b6f5258e4e19bfe8bc757c701186403bd99f44f0b16f093fe34fba8df6

Observation ffe4e4d3-d6bc-4aa1-8966-aed06214946b · outbound

This paper cites Model alignment as prospect theoretic optimization,.

Value-Free Policy Optimization via Reward Partitioning Model alignment as prospect theoretic optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.785448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.251658Z digest=sha256:6862dce054e914b3440db0baf97d1faf51cb1bd69015648261181d33a73d6110

Observation 6699bbdc-bc91-4330-9be9-e3013b84bd57 · outbound

This paper cites Instruction tuning for large language models: A survey,.

Value-Free Policy Optimization via Reward Partitioning Instruction tuning for large language models: A survey,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.254776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.254776Z digest=sha256:eeace2a6d1fcb6f746fab55ee9b9a4bcc1bd53e52198ac09c6b1fec5e49484c0

Observation fb0b05c5-085b-4d2c-b061-a4849a9ca341 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Value-Free Policy Optimization via Reward Partitioning Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.259019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.259019Z digest=sha256:692f38e82c6d1720a0d0c1d0c60ede66190714b9d5a9ab3002208c8651211075

Observation ff63cb40-f6da-4a37-bc7b-9a7815d04dff · outbound

This paper cites Trust region policy optimization,.

Value-Free Policy Optimization via Reward Partitioning Trust region policy optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.775562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.263025Z digest=sha256:0b73f742d573f5743180e9134d08fd2f57fb660aa166cc7692c8cd57355b2946

Observation 3b6d7598-5178-4239-ba4a-ae55ba16e2ab · outbound

This paper cites Training language models to follow instructions with human feedback,.

Value-Free Policy Optimization via Reward Partitioning Training language models to follow instructions with human feedback,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.266419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.266419Z digest=sha256:314984259ec93aa5069c8ec221cfea266e879af7e48f56b0978b1b9373a9e42c

Observation bc503428-b3a4-4786-8dab-a4ce942ed59c · outbound

This paper cites GPT-4 Technical Report.

Value-Free Policy Optimization via Reward Partitioning GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.269741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.269741Z digest=sha256:2b31e92d45a82e6feb6ca23c599d8bea7b5dcaa11cc4dbddd5ebf3fe3e012036

Observation 41a6f198-8d6c-4b8d-bc1c-481155b2fd4c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Value-Free Policy Optimization via Reward Partitioning The claude 3 model family: Opus, sonnet, haiku,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.759152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.272989Z digest=sha256:8697dea9bb0309cb8cd07d1b74d33cee43692c3e18e971df976d71032136265a

Observation 3056ed27-b9f2-4632-b3e7-ae35c551fc55 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,.

Value-Free Policy Optimization via Reward Partitioning Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.748634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.277532Z digest=sha256:51a8c6bd8185c0ae7434bc47534d1332572623d4774a629665b4d0c570bf5dc0

Observation cff56b3c-ca0c-458e-92b1-3cd249e28eca · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models,.

Value-Free Policy Optimization via Reward Partitioning Guiding pretraining in reinforcement learning with large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.738544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.280587Z digest=sha256:ed15dd7c133b6c107d3f4f265f3ce28f14c5f65d3a7f3a13d871cbad7212ca99

Observation 0d3ab8c7-82c6-472a-85b1-493e848b096b · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Value-Free Policy Optimization via Reward Partitioning RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.283692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.283692Z digest=sha256:262cf1d68eaaa08c30fd8931aab2eab3634b5ee4f9d8195e489fbbb1fe36ce9c

Observation 8e48b80a-f4a5-4b8c-a791-f04e80f8181f · outbound

This paper cites CREAM: consistency regularized self-rewarding language models,.

Value-Free Policy Optimization via Reward Partitioning CREAM: consistency regularized self-rewarding language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.727405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.287664Z digest=sha256:5cbac317d749a7e135513d826988a057263278662f793451c09ca2185a80c20e

Observation 0267a425-915e-45a6-a8be-60f29fa65428 · outbound

This paper cites Self-rewarding language models,.

Value-Free Policy Optimization via Reward Partitioning Self-rewarding language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.717989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.290911Z digest=sha256:a328b0caa752da7dfb3eaeaebcbff219fb2e64edb891865bc579a29e5f3db670

Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.293738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.293738Z digest=sha256:97cc7d7103ccb052b10bad5fb5415b25f52d1d25b2069dd3b4e5b866fa186c41

Observation 15bae0be-35c2-44a2-bd98-8b0a8a551dbd · outbound

This paper cites The Llama 3 Herd of Models.

Value-Free Policy Optimization via Reward Partitioning The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.297393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.297393Z digest=sha256:80bee7253431f07832079e4574f8acd8048522a9ddc7fdfa1b71295ed681c1ee

Observation 163ec618-6992-4fcf-ae7d-625c5dd4cf0a · outbound

This paper cites Qwen3 Technical Report.

Value-Free Policy Optimization via Reward Partitioning Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.300277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.300277Z digest=sha256:1f1e60b9215d635f876104baa7fc77e735017d33cfbe1b2e2645f88a2899bad3

Observation 2acf63e6-e1f2-4458-9c6f-427069ec6d62 · outbound

This paper cites Nemotron-4 340B Technical Report.

Value-Free Policy Optimization via Reward Partitioning Nemotron-4 340B Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.303322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.303322Z digest=sha256:2f635bb59296630a4926c7b71be20e5c2c48a9ddf1bcd8e588c00d4e184f4f96

Observation 4d4fdb80-5a78-4520-b10c-85b41aeb9bdf · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Value-Free Policy Optimization via Reward Partitioning Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.306523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.306523Z digest=sha256:e76fc67baffbac3cb51b006628a09e9703dd6ddf501afec9dad7c01198681f02

Observation b0633016-44d7-4956-a4d3-a56c310edf3e · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences,.

Value-Free Policy Optimization via Reward Partitioning A general theoretical paradigm to understand learning from human preferences,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.702083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.309895Z digest=sha256:f66e7c49848f36f8f6a905fbc337864c3f5bbeedffbc0772bef7ed04b1c68727

Observation f496c456-0247-4c12-b42d-338ec07bae57 · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty,.

Value-Free Policy Optimization via Reward Partitioning Advances in prospect theory: Cumulative representation of uncertainty,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.692625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.312840Z digest=sha256:8a3584c7cf1b3dd3054634b4091cb053148a48276b3a95b8bafd26647f62ac24

Observation efb414d2-4669-4b7d-a5eb-e8b65446ba0b · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Value-Free Policy Optimization via Reward Partitioning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.315854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.315854Z digest=sha256:f94e0dc37588bd6d5273b07fb6a467968f2891725565a4e240ed0c73dcd90deb

Observation 7262b461-5e87-474b-85dd-4960bc1d8d80 · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models,.

Value-Free Policy Optimization via Reward Partitioning Alpacaeval: An automatic evaluator of instruction-following models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.683149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.319398Z digest=sha256:42539785b85343be2dc2d4cc9d40ba2a058a5b1b2a65b1f34b09a75ff26c7a2d

Observation 29968375-32d7-49df-bf73-acb155c2e158 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Value-Free Policy Optimization via Reward Partitioning Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.322892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.322892Z digest=sha256:f8751a41afc093a567e334a35548fae72ee834bf62f9dbda80b54fbbb04177cc

Observation 3b78be03-5ee5-44bc-9e33-ffbaca2aee97 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Value-Free Policy Optimization via Reward Partitioning Instruction-Following Evaluation for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.326677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.326677Z digest=sha256:7c35718e539df3afd48995ac1386d774643bb1f0e8756a0639bb927003a1b976

Observation 2b07c684-5b46-4e78-b412-79e35815b2c1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Value-Free Policy Optimization via Reward Partitioning Training Verifiers to Solve Math Word Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.330155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.330155Z digest=sha256:85fe9c7765f65fe1d518073e67843dff1c4b276673db9ad82a73467a413a1600

Observation 17fce266-5c00-4cb7-852f-8059491a6f1c · outbound

This paper cites Scaling up models and data with t5x and seqio,.

Value-Free Policy Optimization via Reward Partitioning Scaling up models and data with t5x and seqio,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.667368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.333178Z digest=sha256:7f32a8f393c48562b246e2e1b29e396a406d5da73b5dc2adfab1dc2f2725bf2d

Observation 7fb1bfa7-74ea-4b4e-93ff-1c7f6ca8f999 · outbound

This paper cites Mistral 7b,.

Value-Free Policy Optimization via Reward Partitioning Mistral 7b,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.336652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.336652Z digest=sha256:69fa66dad80e26432fd7c55a9a0a6dbb56e8239c70ef0f75d23a91dba35f8194

Observation ab1ceb4b-8ef4-47cf-9ba4-9e85d1442e9d · outbound

This paper cites Llama: Open and efficient foundation language models,.

Value-Free Policy Optimization via Reward Partitioning Llama: Open and efficient foundation language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.650316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.340728Z digest=sha256:ddfccde3f18d6e99e053fc7e1ef4a7dd39b5858a4c142697d37ebfae13ca149a

Observation edba6570-de0e-4a9b-9180-4ac294105a90 · outbound

This paper cites Qwen2 Technical Report.

Value-Free Policy Optimization via Reward Partitioning Qwen2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.344252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.344252Z digest=sha256:0cf22136dc69ebe66ba9e0ec5f1cd37360388af40d0539b82b221c69754d4efa

Observation 526d418b-dd00-4c0e-8f9d-32f81c3ca79a · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Value-Free Policy Optimization via Reward Partitioning BERTScore: Evaluating Text Generation with BERT

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.347457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.347457Z digest=sha256:71faa6e78307ebaaae68be30e7a5437ba27efa5ac90eece8383d209bdd98399f

Observation 762e49c1-11ac-4aac-bf6e-ea267e0d46be · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Value-Free Policy Optimization via Reward Partitioning Rouge: A package for automatic evaluation of summaries,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.350864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.350864Z digest=sha256:200b38e0cc2c8184498913b341673f38dea0f19c8009a99bf16a7d1a1e90ec02

Observation 1c2fe268-4e4a-4beb-85cc-8fae2cd7f522 · outbound

This paper cites A diversity-promoting objective function for neural conversation models,.

Value-Free Policy Optimization via Reward Partitioning A diversity-promoting objective function for neural conversation models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.634832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.354796Z digest=sha256:6deb119b1965443f08854c2bac513d5655111ec27474383b869278ab8bf19f83

Observation c84dfb56-bb9d-4b13-8610-5d333745a163 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Value-Free Policy Optimization via Reward Partitioning A Survey on LLM-as-a-Judge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.358212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.358212Z digest=sha256:1e8d3aaa25004eec2b1bd73ce268df89eab673a63c46bcd20c1ff08eff7a14ef

Observation 0c8e74f6-81cf-4fd3-9c7a-df18274cf243 · outbound

This paper cites GPT-4o System Card.

Value-Free Policy Optimization via Reward Partitioning GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.361713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.361713Z digest=sha256:521c2dd7212406fcbc2a103874b13e27703ebe11b554fddd4a361228398d007a

Observation d8c59ea9-5c99-45a5-bbd7-2c6f3921eecc · outbound

This paper cites Claude 3.5 sonnet.

Value-Free Policy Optimization via Reward Partitioning Claude 3.5 sonnet

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:49.623894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:31:49.365419Z digest=sha256:26b5cfa56f3f4a1cd5a315a09a1b9c862cc81bd8702ddfdea746c7db004077bd

Observation 9a806028-570f-490e-bdc4-e44b439ba7ac · outbound

This paper cites Decoupled weight decay regularization,.

Value-Free Policy Optimization via Reward Partitioning Decoupled weight decay regularization,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.368412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.368412Z digest=sha256:19f7f4cbf225eb95be1674f26231fbc9d636786b7c428483c4e7b1e41ff24709

Pith citing papers

Observation 58037a65-8d50-481a-8f54-57238aecf6dd · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Value-Free Policy Optimization via Reward Partitioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.906356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.906356Z digest=sha256:93f0ab3e778fcfc511848386c5a3b488f6cf177ab73f6d67ac1cf8cb28c46377