Pith. sign in

Paper Citation Record · LEDGER

A Survey on Human Preference Learning for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.11191.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11191 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:52:37.405393Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb3b3895-c27c-4a72-9c09-74f8b2d8a384 · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate A Survey on Human Preference Learning for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.053077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.053077Z digest=sha256:500d77289c70d6f2db783b12e8ffcf54ebe3a17d2cedc24cbf50acd42b850e8a

Observation 04a03cee-d78f-47f2-80ae-d053c20ed67a · inbound

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives cites this paper.

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives A Survey on Human Preference Learning for Large Language Models

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.624269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.624269Z digest=sha256:8757ccefba08c3f3144c9e132ec383162263038454557e96298a994d8bee5cb2

Observation 145d718e-7b02-4278-98be-c7ed15b2070a · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs A Survey on Human Preference Learning for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:28.561126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:28.561126Z digest=sha256:8a395bc82a23fb6196ba0c84edd2a44e24d857819b03f18c3e35922dac96cbe5

Observation 96dae525-4015-46d6-ae74-723592f18dbd · inbound

Listwise Preference Alignment Optimization for Tail Item Recommendation cites this paper.

Listwise Preference Alignment Optimization for Tail Item Recommendation A Survey on Human Preference Learning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:41:38.324872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:41:38.324872Z digest=sha256:d89c8e3575b9557d27230883a95b758ab9ca6de33e14418220ca5c9d37760868

Observation fae028fc-2e86-47dd-bfe7-e01e9f7aec74 · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention A Survey on Human Preference Learning for Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:56:50.563426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:7929da67d5953031a763786d5f1052e30040c0550c7badd3ec6f8b16949a0521

Observation 2a89d976-dec3-48ac-bd83-84bc6cebfbf4 · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial A Survey on Human Preference Learning for Large Language Models

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:31.993618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:31.993618Z digest=sha256:90d5aa83ff8c2fb10a82d73c8b386856e66a9688f9cea359df399240c20bdb02

Observation c595a2f4-a866-47f8-9ec6-70a2fc92d731 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Survey on Human Preference Learning for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:31.560030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:31.560030Z digest=sha256:5d718cf45eeb423850d7d486c98c8669deaae1d928da6ddf0a4e4d5799e01f18

Observation 383cc46c-307d-4d06-8d98-0e4f5274ac29 · inbound

Can LLMs Make (Personalized) Access Control Decisions? cites this paper.

Can LLMs Make (Personalized) Access Control Decisions? A Survey on Human Preference Learning for Large Language Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:34:04.998514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:33:18.457421Z digest=sha256:ed6e505d492ed8ff683bbe8fdf4899200cf15ed309c77f72e0abe372757ee219

Observation a4db0292-b962-4748-b391-63151e5eddc8 · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge A Survey on Human Preference Learning for Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.769687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:54555704cac4a2664b4b231ae54a12e7810322bab2747b3ceda9bfac7cfb2e78

Observation 1458483e-a9e0-4b9c-b38d-aa888a30bd95 · inbound

Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents cites this paper.

Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents A Survey on Human Preference Learning for Large Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:37:25.788630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T18:47:36.189582Z digest=sha256:090151df619efac1173f93842a6c9620b2469b40b8efec6bb65f96cea9fd70ff

Observation 460c0d77-475f-4250-8534-24d8d2f6d150 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards A Survey on Human Preference Learning for Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.129924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:bfbb28b4e8e36218eb50e3be928e74f79ab9dc7dfc98804a73db070e40b241cf

Observation 6d16f15f-5418-4857-a0cf-53588e09ffba · inbound

Internal Pluralism and the Limits of Pairwise Comparisons cites this paper.

Internal Pluralism and the Limits of Pairwise Comparisons A Survey on Human Preference Learning for Large Language Models

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-02T09:03:14.928978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:03:14.928978Z digest=sha256:f9e4de1ccbbe1041599588a9b1151213fba5fce02ac209df76ea825206d628ae

Observation 7fe7ba90-5138-4f16-8f00-5575d2f34a96 · inbound

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning cites this paper.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning A Survey on Human Preference Learning for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.405393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.405393Z digest=sha256:cd7387986e6bf58315194079e98898b8dd7fe34193815fb4f42ccf1d57ec7edc