Pith. sign in

Paper Citation Record · LEDGER

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment

As of 14 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2502.04354.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04354 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:47:17.656530Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:46.666638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:34:25.763556Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 174bb4b2-35d2-4388-948b-939e6546f868 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:18.104986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.483686Z digest=sha256:fed09cf6e03dd516c4c29153503afa9437f37fa60af103731f3c6083db9e0944

Observation 574f42ca-56cb-4807-8531-c3e721a74159 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.488100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.488100Z digest=sha256:e2c6f0c7c0cfe2982e0b79418dffca6b241ed51d4ad7a03782a0619799a37370

Observation d0ff8392-f026-409e-b95a-8d437f5e521a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.492470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.492470Z digest=sha256:c27a2abda31fb0290a87f505986bd16b9e958127002ffeb38c655a87fcee2d96

Observation 421a4be3-99bf-4c32-88da-4b2d20da9c7f · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.497296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.497296Z digest=sha256:b4a52453cfeb712cddfa902cb855fcd58ee9414339af5025813072a60de4a741

Observation def6c9fb-8d3f-4735-8a6f-57c21c1d186c · outbound

This paper cites and Broderick, T.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment and Broderick, T

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:18.089022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.502024Z digest=sha256:c60c8162b6ce8ef48abd1cd6b538170792243fffc1525a282189cd1d4d80ce65

Observation 73be46cb-f446-40c5-bedd-122907673054 · outbound

This paper cites and Broderick, T.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment and Broderick, T

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:18.079532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.506476Z digest=sha256:0dbf90a07bd38a7d8c95b6f6785c525c86e30db71ecd8f00d0ba82ab672de2f2

Observation c3a7be71-321d-4e9c-a0f1-b9e1e4c87fcc · outbound

This paper cites and Verdinelli, I.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment and Verdinelli, I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:18.069294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.511315Z digest=sha256:1c45096eda510134edef7de729e8cac0d5145bf322170863491655179a51112b

Observation fdfe3e72-a9f0-48fd-a50f-37861c61d6bf · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.515446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.515446Z digest=sha256:1ab65366cc6fd00c80c224c526a9ad0d21d20cd7b0d56c049b9f835f8b18a50d

Observation 8d478326-31d9-4d69-ae24-51535ff72f5e · outbound

This paper cites I., Santos, S.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment I., Santos, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:18.054172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.518741Z digest=sha256:7729db38195198bcbd1864930e73ab586a68cde928d341df8aecb2532ff02d2f

Observation e538dc2d-97f4-4c23-8dd9-ab195e63b329 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.522690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.522690Z digest=sha256:0c519d5f15dd503dda861d403227fcf76e3054e0dc12603d746d7f6acd6cb722

Observation 0647d1c7-36b5-4203-bac7-7901b6c74ce9 · outbound

This paper cites and Lindley, D.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment and Lindley, D

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:18.044631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.526764Z digest=sha256:29802238257ed5e791e22cb85abbecf564f399ace61d045e12bbc4b234cc817f

Observation f369319e-ccf3-45c4-84da-8814aa6852dd · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.530335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.530335Z digest=sha256:40a22eebbc1ee6bc5750b511885d77d0c64fd5d4af46eb021aa36e499ffa401d

Observation fabc0822-4643-490c-bc48-f3d97a5d7106 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment RLHF Workflow: From Reward Modeling to Online RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.534445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.534445Z digest=sha256:3208fe0f50f7e5a830787a82aeefd2191eebc0b7f2a8436ce00123ab190ebd19

Observation 7222187e-abbb-4a67-bbde-25f6e9312ecc · outbound

This paper cites The Llama 3 Herd of Models.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.538302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.538302Z digest=sha256:bfdf7abb870a229f8dfee8f3c1baf4104ebb19ff6ecd0b06f6d533010322f7d1

Observation b9202e05-ea5c-4e0c-9332-e853598b3078 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:18.032853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.541845Z digest=sha256:fe10eec3183f87ce7f743bcf4c0d93eabcc365d3a0d37259a0e8f78eab7539c2

Observation eb7c576b-ffe9-4c0b-842d-b2cfb32a03f7 · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.545110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.545110Z digest=sha256:5034ce1949e6df0fe5ac2d736cef1ba3ef55cfaecdb48d2f9798739dbafb2e7a

Observation a118c3d7-72b8-488b-a82e-5df92d1e56ba · outbound

This paper cites Bayesian Active Learning for Classification and Preference Learning.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Bayesian Active Learning for Classification and Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.548632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.548632Z digest=sha256:0632f93986a96eacc912fe4ca1b350f96797c19902374a3a818370e2b0f67e62

Observation 9867a51b-9551-426f-8040-bc6394a2b309 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:18.013715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.552135Z digest=sha256:ddc1f0993dafa3b73b851b0d8eb46986c47901a3ebed236f72be60c8598ca304

Observation a7a444ad-3e2a-4116-b84a-7a03e2cea1a5 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:18.002625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.555638Z digest=sha256:09db6ddf11a90be844305222ae58b3d58ee2917dd6e657abea9334423243c7e4

Observation 44e5298d-977f-4d99-9569-d0ba02dd4071 · outbound

This paper cites Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:17.992376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.558917Z digest=sha256:a0fd0e62e74bfc32543cc3bbe5b3d30a980925247d7ab84782572520624a491f

Observation 05800bff-4fb0-469b-abc4-9d7c6776d3bc · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.562132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.562132Z digest=sha256:8168f88e33363925d4ae22dd04bbaaeb7be61b5dcf053461fa22dd866c74eedb

Observation 9f60a0c9-e24e-48f1-9307-f3c5565b4fa2 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.565775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.565775Z digest=sha256:b0abaf96a660866a37af846b77e7c68116aafea52f2b4055280627c0dd868fef

Observation 19a4e2fc-69d7-4c01-b1a9-f1fe15219fd6 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.569691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.569691Z digest=sha256:35696743c210e34461f23afda464e304431e62676b3b0f152f23902b8ad893c0

Observation 74abd0b1-95f9-4f5a-8379-137c39ff5c3c · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.981824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.573706Z digest=sha256:14c9d196f961750295d6dade353a67a127e8d8854d2eb9cb1f5f85328fa394a6

Observation 4216feaf-9aff-4e0b-b0be-172324f8c281 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.972166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.577086Z digest=sha256:39183c818aac375c0774e45af44b53389fffbecb7b52de2cd50fd41836afe5d4

Observation 4f263f93-3b0b-4669-90f3-6a7a378ae7f7 · outbound

This paper cites Active Preference Learning for Large Language Models.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Active Preference Learning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.580916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.580916Z digest=sha256:88884dd8a39ca7357260c1be388e153d295f505e4278b4617f94df19c94bf8d7

Observation 03b7d3ca-8995-4aaa-b5ea-ee8e23a7e105 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.962404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.584746Z digest=sha256:19984cb150929caa61c587c06831cc497c11d560c3fcdc9dee796c40da22da7a

Observation ad69f174-c741-4499-9a5d-ebdfd45bbb82 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.952310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.588213Z digest=sha256:14214bbbb5b52faf52441bfff10e1ea727e313119da4413c01ee358febb61d7c

Observation 8b5965a5-7467-42f9-bcd3-6b122384601f · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.942255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.592077Z digest=sha256:74b3d6104e1bf122cded65eb54ad46d6486b8dd23625b530c3c6e3912dfb6c38

Observation 33633452-c62a-4b11-be23-3328cb1c3c85 · outbound

This paper cites Active Learning for Convolutional Neural Networks: A Core-Set Approach.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Active Learning for Convolutional Neural Networks: A Core-Set Approach

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.595834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.595834Z digest=sha256:e2b6d252eccb7b520d14224b7041281085e0aa5eabdd6825795d3c8dde9264c7

Observation 630da295-44db-4efc-95eb-386572fc8886 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.599848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.599848Z digest=sha256:61363a305b2f85530f308b3ee35a1c5b0859eb2e92dccbc9ef4c252a84ee7205

Observation 63a778b0-0f9f-4b9d-96c1-f592b4fc4b51 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.603665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.603665Z digest=sha256:26103d43d83b0fe8f64f90d8b238f3419a78406116c76043e4911307b8e61b03

Observation e3e6f97b-f54d-4bf2-b662-4252f89b9899 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.918111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.607299Z digest=sha256:843065f2f7e03667540de5853cf975feea5de60f5d12bb5eeb23e719419bb98d

Observation e4c1d78b-fcf4-4cdd-a783-98b08dbba329 · outbound

This paper cites Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.611380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.611380Z digest=sha256:990b2bf3d5621e26b8b25a39df52500df4488eec99cab16f67a2dfbf09da975b

Observation 6b591223-f09c-4cc6-aca1-575fccf4d264 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Gemma: Open Models Based on Gemini Research and Technology

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.615382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.615382Z digest=sha256:69f1887b707029a6c7cb8af1ebb0a60af48da19dd0ade611831aebd4784e4b10

Observation 50417d67-ee5a-4c7d-a818-9f41f644a5dc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.619197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.619197Z digest=sha256:2b3cf7b13406e322f9208ddc021fb11bdc04c6d6f90440f54bebe487eec98888

Observation 9b8e3b49-8a77-4747-8ce7-b0eb02a0a7c5 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:47:17.907633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.622985Z digest=sha256:0b0bf5daee6b138ee722715c314529f286ca4a0b0100957caf87e2e7a6fa4722

Observation ca588570-0b37-4682-85fb-bc775174bb3e · outbound

This paper cites and Chris Glaze, B.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment and Chris Glaze, B

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:47:17.896575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T11:47:17.626782Z digest=sha256:efe978be163cb3ab4af29a6dcb9eface930b7e23f0799c9e9f7010e4e9ea3563

Observation c0bdc399-e989-4d4e-92bf-b379aed4f951 · outbound

This paper cites an unresolved cited work.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.630417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.630417Z digest=sha256:da94399580d4452230caf60db4a196d1556f3235be3724be4157ef2017bad4f5

Observation 791982b8-e588-4887-bc77-7e90d866e91f · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.634354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.634354Z digest=sha256:1e30a633d02c58bf1818b8b5fd91de1d18576d238485803f96648873ffdcdffa

Observation f269a302-d70b-4d9f-af17-0cb8f115d568 · outbound

This paper cites Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.638227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.638227Z digest=sha256:035a5e681f3027054b33903d45daa3386fe55a2b7d9a45a2466bcac8fc59ee67

Observation dd4ef48f-0cae-4d75-8fc2-22714ee90684 · outbound

This paper cites MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.641981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.641981Z digest=sha256:b6029edd07afea6038ef481c4da63231b9e62a3f0dbe54bb8557622695b8f697

Observation 68877fec-7230-405e-b79d-0453e484ba40 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.645871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.645871Z digest=sha256:4aca094fecf909a04151b23f02bd33fa4c2c53881415db46d97246ff654e5065

Observation 3ed37742-e476-45a4-a5c3-fec5ba2729d7 · outbound

This paper cites Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.649370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.649370Z digest=sha256:4d6b3ce90d9e2a0e2f2ef73207b0167e34ff87386843f261144e31dfab25150d

Observation b6ce621d-1846-4d13-b133-bd353dcd43ae · outbound

This paper cites Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.652900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.652900Z digest=sha256:0a4452b5e9f60f48421653f0334d3bacf06a8161191f7d1733a11be30af97a81

Observation a9dd4566-6569-4195-97c2-5c64f7de45c9 · outbound

This paper cites Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.656530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.656530Z digest=sha256:3bbf8f71ab9cca21280e939fcdefceff724921efed29cdf05d353e502b17d736

Pith citing papers

Observation 13b03f5f-35e5-4c53-8118-8a89fa713ed7 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Reviving The Classics: Active Reward Modeling in Large Language Model Alignment

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:46.666638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:46.666638Z digest=sha256:3d7c2f6f39734184e3558150068676203d82f4e61e92b146a745b03764979af7

Observation 6f256eb6-64ed-4697-934f-42d9e1153fb3 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reviving The Classics: Active Reward Modeling in Large Language Model Alignment

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:34:25.767096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:34:25.182215Z digest=sha256:0684e05bb69bd03c4b7935681f3af81cc1d3b9320333f7f42feb2242ac0dcd44