Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:32.708316Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2507.07375.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:32.708316Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd50d613-d633-4991-8971-8b128e83bfaf · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary A survey on evaluation of large language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4481ec-ba23-4c3d-b178-99c05edaff16 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Using an llm to help with code understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0001cbad-5f0e-480a-8c10-ab04577d1bc4 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary A Survey of Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c80e4a-ccfe-41b7-ab00-f763bfd2ca55 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a11b13-7de4-405b-b82f-73cda5024c44 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary A Survey on Large Language Models for Code Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 916dd231-eeeb-4be9-8b3f-46bbd95effa9 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Nguyen, Quang Pham, and Nghi D
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e740f70-b253-4026-90f8-c7df0600ac8c · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Rational decision-making agent with learning internal utility judgment
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 349be0f5-8bb4-44e0-a2ab-341f001069a5 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary How Far are LLMs from Real Search? A Comprehensive Study on Efficiency, Completeness, and Inherent Capabilities
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 77e7bd53-557f-4cd7-973b-dded99fde893 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Safe RLHF: Safe reinforcement learning from human feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 69708228-e735-41ca-9311-dc586d66fc1f · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Data-adaptive Safety Rules for Training Reward Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f6c85d-63e3-48c7-817c-802b866accaa · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Large Language Model Alignment: A Survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ebcb40-f4e9-4f43-b29b-46b60683ab82 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Catastrophic failure of LLM unlearning via quantization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0ad21580-69cc-445f-aa0b-1c0e961d356a · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1fea6c-175a-4353-82c2-3ac45f6e46bc · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Training language models to follow instructions with human feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7edfe9-52ea-4f00-8d2a-2cbe30e2a777 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82fcb283-bb09-4b1a-9049-a6cfe95800a0 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e15905dd-f349-45d5-99cc-214640b79ca9 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b6565b-6ce5-4490-8d0f-4d062136a615 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Scaling laws for reward model overoptimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5c8744-c079-420c-8e47-fd3f29ceef97 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Regularizing hidden states enables learning generalizable reward model for LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 72d56fc1-8a58-440d-b9d7-9e5b9069b774 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reinforced Self-Training (ReST) for Language Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f5ebdff-d09b-4348-a0a6-29c5c232a9a6 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary RAFT: Reward ranked finetuning for generative foundation model alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe264a2-4475-4269-a0e8-7d2a74d1a5dc · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary BoNBon alignment for large language models and the sweetness of best-of-n sampling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e5ff2773-3700-4367-8456-dbbaae253d92 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward model ensembles help mitigate overoptimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2eceda81-b0f9-4057-8cee-6fd326be5e8f · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814af997-c937-4d52-ae31-db71d34520ce · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1327b6b7-25c7-43f0-9cc9-2622a2a96aa8 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward-Robust RLHF in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608ab5de-8333-4751-8d8c-66a9fe2ffdbd · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary WARM: On the Benefits of Weight Averaged Reward Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f17fb1-6713-426e-a30e-02e9f03beac8 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Mitigating reward overoptimization via lightweight uncertainty estimation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 15f94c9c-a771-4795-8152-1f2b9969f667 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Confronting reward model overoptimization with constrained RLHF
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fd5e772f-2917-46fd-8d86-01b375aaf96e · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Provably mitigating overoptimization in RLHF: Your SFT loss is implicitly an adversarial regularizer
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0c6ff2cd-dc93-4461-8ad5-5cef5b5228ba · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2670194a-a4ac-4f5c-9ff8-66c87f5a78b3 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494dedb6-ce66-4401-a30c-8228dcc88c96 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Odin: Disentangled reward mitigates hacking in rlhf
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d0ba9d10-7228-4b8c-a064-a26c851f1978 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Mitigating reward over- optimization in RLHF via behavior-supported regularization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c8f808c6-1a69-4724-93a9-71a55743c510 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7d77ce26-f716-45ca-bb5b-e2e43879b913 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4d3b46b3-7012-40a1-ac29-651d3c92f037 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Ultrafeedback: Boosting language models with high-quality feedback
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6344dd9b-6b44-4d70-b083-52e5626041df · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Prometheus: Inducing fine-grained evaluation capability in language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1478dc94-e347-43a8-92da-4b5205119c2f · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary From generation to judgment: Opportunities and challenges of llm-as-a-judge
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96f1ab52-5e08-4ef2-b706-13e379c3b536 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary A Survey on LLM-as-a-Judge
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea44deb3-73ba-46be-8e99-2794c093ef3d · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9f64425d-4872-4be7-a1e0-8c8d38b36f77 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Smith, and Hannaneh Hajishirzi
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40709b58-5608-47ca-9a4d-baea0722a49d · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary RM-bench: Benchmarking reward models of language models with subtlety and style
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1dc1b6a7-b1a6-4575-8604-bfd8e6e670a3 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Rank analysis of incomplete block designs: I
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79436223-3dc2-4bc3-9038-6c364252e5ef · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2287ce39-6a6c-4864-8fb8-e7df4d8ab03a · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b72a5710-a58b-4182-82b3-2cbee629ab2d · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Learning to summarize with human feedback
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f5d77c5-b60c-443c-8d50-71c554792d2e · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Pairwise proximal policy optimization: Language model alignment with comparative RL
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec020fda-7a9c-441e-8d44-2e09e60de4c1 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward shaping to mitigate reward hacking in rlhf
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2be653-9f15-48ff-adac-e8727d1f82d8 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16cdb465-159a-4ae4-b6a3-f0fb8839c954 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Gemma: Open Models Based on Gemini Research and Technology
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1560d3f1-fc0a-4acf-8370-8e9b79abaaec · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward Model Ensembles Help Mitigate Overoptimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd38fb7b-20d0-4846-8fab-b52ad404bc98 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cd64a5-987a-4f44-983c-2810c38f8a4d · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27fdfeb5-0952-486b-a2b7-02ddb351d597 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 205cb3df-66f4-4bad-8e25-0bb8575dc0c0 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary The Llama 3 Herd of Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96ce776-6d65-4285-b50a-1ad17fcbb364 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary RLHF Workflow: From Reward Modeling to Online RLHF
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ae3f94-581a-41da-8b50-e41d5176c51d · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Rethinking reward model evaluation: Are we barking up the wrong tree? In The Thirteenth International Conference on Learning Representations, 2025
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4c137bbb-d842-4347-acf9-3bff36a85cbd · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary What makes a reward model a good teacher? an optimization perspective
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1f0d2a-e239-4670-9539-70369bca1236 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Mitigating the Alignment Tax of RLHF
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918f8fce-fff4-4e9b-a382-65f69b82e97e · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Rethinking reward modeling in preference-based large language model alignment
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 19f775c6-3d58-40b9-b8df-4c1f5703f672 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Starling-7b: Improving llm helpfulness & harmlessness with rlaif
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 50cbdeac-a90e-489e-b035-53ad11c6274f · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Nemotron-4 340B Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a7e88c4-c5a8-42ca-bc98-9dae587c9d73 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Fine-grained human feedback gives better rewards for language model training
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4885c4c4-937c-4e54-8052-92db587a203a · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f377fcb-db82-4eea-b21d-cbae3044e640 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Chatbot arena: An open platform for evaluating llms by human preference
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16812e59-4a34-4979-a196-8c672c4d8ea9 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary WebGPT: Browser-assisted question-answering with human feedback
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ade06b-4789-40c3-a681-cd6ab4084a6c · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150a17bd-89d2-431a-a845-643f1bb027d4 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0051497e-7fd1-4459-bdf7-3d22cd2c7d92 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99066ce0-bf66-45e2-bff9-babc6bf79869 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c2a574-043f-4123-b3fc-d777dbdac109 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary RRM: Robust reward model training mitigates reward hacking
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 666bbd5d-4176-40a4-904b-9923fbbc7e61 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa9054b3-3428-413d-b2ae-2c0022526db6 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9a15a556-a843-4b42-99f9-210db24ee158 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Convexity, classification, and risk bounds
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d2ff92-4fe3-4ed2-9710-3c5df2a99dfa · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Fundamentals of statistical signal processing: estimation theory
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6a67e93-2552-4e32-97b0-f01c49869cd0 · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary ARGS: Alignment as Reward-Guided Search
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7b5efa-c7a1-4ef0-ad2d-3ba0f589c4ab · outbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Transformers: State- of-the-art natural language processing
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
No inbound Pith citation observations are available.