Pith. sign in

Paper Citation Record · LEDGER

Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2406.12845.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12845 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:31:53.198146Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.398618Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65ddd036-d75f-47de-afe4-99952039e0d8 · inbound

Adaptive Decoding via Latent Preference Optimization cites this paper.

Adaptive Decoding via Latent Preference Optimization Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T20:31:53.198146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:31:53.198146Z digest=sha256:ab97afccea89e3533e723a90be2082ea7a7ed802bd30e83b5dffef5335aa3768

Observation 6586546e-2a46-48dc-8144-e1e528e25074 · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.933909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.933909Z digest=sha256:d38ded98731d4d1d7461b9c3b113b561839e5efc896e9875dcf112bf7cac6447

Observation 179382e3-09a0-4fee-93bb-aa3b7174946c · inbound

Interpreting Language Reward Models via Contrastive Explanations cites this paper.

Interpreting Language Reward Models via Contrastive Explanations Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:27.906346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:27.906346Z digest=sha256:5191c47c784e327d744150bf2ace7c83b08a6d2e26d74d6da7fa11a2903f391a

Observation ae6ea1ec-a58e-4b27-bf42-0d54fce4fbca · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.399577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.399577Z digest=sha256:76b0a489571f3cf3910589e32e3e9e1c699d2690caece335bbac08cfc43b06a6

Observation 4bbd1de7-89bb-4245-845b-a3ce80ee060e · inbound

Boosting LLM via Learning from Data Iteratively and Selectively cites this paper.

Boosting LLM via Learning from Data Iteratively and Selectively Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:38:36.946910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:38:36.946910Z digest=sha256:fcdc63b6fa19b6774c504cb151def8e236dfe3cddf99c9851ea321364dbb2277

Observation 76dd9384-d843-481b-9c36-4ba97b5fbc21 · inbound

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis cites this paper.

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:26:46.653799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:26:46.653799Z digest=sha256:980bb6800c8273f584f9db806890c92c1b054756de8bcdc98d80203d2c0d25dc

Observation ff6c558d-bff6-401c-8d5d-ec1eeff439c6 · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:09.084843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:09.084843Z digest=sha256:fa398ee692b6b9b8304a80a8d1842fd9d2fd7808886ed44a19d8d0b632a1195c

Observation e7092d19-35bc-4643-861b-ceaa4b504d6c · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.059884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.059884Z digest=sha256:3b29e46355cc065dcc082c328d254a26da8d22c29fd3366a891074a04025c0fd

Observation 589cc8e1-0d44-499e-99c3-9300818f6f27 · inbound

Data-adaptive Safety Rules for Training Reward Models cites this paper.

Data-adaptive Safety Rules for Training Reward Models Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.202305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.202305Z digest=sha256:e298280649265e71624e01f9e494d55b71e9088ef8b9cae35efe4e8889246f65

Observation 59b02dbf-8291-4708-9a20-58cfbc4b8699 · inbound

Mixture of Experts (MoE): A Big Data Perspective cites this paper.

Mixture of Experts (MoE): A Big Data Perspective Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-10T18:56:37.826261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:56:37.826261Z digest=sha256:9d972c482defbb477dc503c2355b794d3d27cc36d561df8c455bfc5b06b0c76f

Observation b1bdb19a-0c51-4eb1-9ac9-686627d8924a · inbound

Atla Selene Mini: A General Purpose Evaluation Model cites this paper.

Atla Selene Mini: A General Purpose Evaluation Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.821883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.821883Z digest=sha256:ffe41e49602c38ddffbb071c4a584a6479e45c5c365e397a685c9d847f7027f2

Observation 6b6b7f60-0c5e-4956-a978-5c3389a3891c · inbound

R.I.P.: Better Models by Survival of the Fittest Prompts cites this paper.

R.I.P.: Better Models by Survival of the Fittest Prompts Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T23:02:23.407554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T23:02:23.407554Z digest=sha256:e7ce3057a3f602f970fa07051042138be0be5d0d4f0e03a934e63295f4ef8eeb

Observation 338590b3-d24c-4faf-bf56-9d251942cf9c · inbound

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming cites this paper.

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:59.967922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:09:59.967922Z digest=sha256:315d886ed05f9c5a852d1cb9c078a72f36117c212d580fe9e1a22d0cd8f56e42

Observation 7323d348-b0fc-4847-9203-2a850d839a04 · inbound

MOSLIM:Align with diverse preferences in prompts through reward classification cites this paper.

MOSLIM:Align with diverse preferences in prompts through reward classification Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:05.677821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:05.677821Z digest=sha256:a59f5750172f1f7311795043bfe8983e19d740315349fd60ffea19f5b0882bca

Observation 4be40304-6593-490b-9633-e738ab1320a2 · inbound

Accelerating RLHF Training with Reward Variance Increase cites this paper.

Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.775977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.775977Z digest=sha256:20205fdd1ad4f1868115d2c9d798d063d1b57aad55288a0b4048a24d5524d8fc

Observation 60e0b91e-3536-448c-ad4b-885071fab307 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.873818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:8e2499f856f5dde5612444211a5d32af1d27b296755a575bbcbfb5b614a4cf7e

Observation 875c2e27-052c-4e7f-bbbd-d4b8646455d2 · inbound

Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding cites this paper.

Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:59.868635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:59.868635Z digest=sha256:1aa2efd2b377c3e589bee2dc75ed614f5bc123fc6efc9e9994e030481b87c827

Observation e82a1853-8285-4461-8166-fcdc7998f07b · inbound

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning cites this paper.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.079485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.079485Z digest=sha256:87c39a098c4390d211ffa1177bacb03cd665c4b00774a01d99428be69d147bda

Observation cceac6b1-0107-4437-8266-c51db615d615 · inbound

BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute cites this paper.

BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:05:28.911043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:05:28.911043Z digest=sha256:055af4283d0730eec25da67839c63b3a129addf41d29e15f1d4b3366cfd3e08b

Observation 849056ab-8732-4605-8bed-aa8161486bed · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.760921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.760921Z digest=sha256:5ff66da0b3036c3a69dbccb14c3ea8a121a8acd0d7e62960e4b48e946818d939

Observation b06ce073-c8ad-4ce4-bb51-450be7245415 · inbound

T-POP: Test-Time Personalization with Online Preference Feedback cites this paper.

T-POP: Test-Time Personalization with Online Preference Feedback Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:12.590452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:12.590452Z digest=sha256:7e6ecb72e358e900a1e2865c7c964dfc21c7350beebc25adbbd4373616582327

Observation 5daa66ef-e109-4667-939c-3f89ab5c0f62 · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:47.225695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:47.225695Z digest=sha256:a95acd6fd4797419e20e633cc1b5e800d9a7ea04368db598d0b9c4f9cd59224b

Observation 73b6f605-1e24-4871-bb6d-233259723df4 · inbound

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency cites this paper.

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:51.396297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T10:38:09.875786Z digest=sha256:6353d98d9b9238bff64b48f9e1cd585ccb75adba7e1384ebba157f7a85999529

Observation 78c93807-1306-4273-b3fb-328262e3caf1 · inbound

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient cites this paper.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.698135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:809f5e6797ffede9aaa73620e72bd742ec2cfd2e739f1409251ae7f2043d1ed3

Observation aa5406f5-3998-43dc-a48e-2dcc17ae23a3 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:48.664158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T20:07:07.548384Z digest=sha256:38a6da0268ff5bd9af2e9f8d022934539076ce01c1644256f2f9ed5a275d11f5

Observation ef04ec64-c2e6-464d-91e8-d7ff1db14216 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.620457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T06:25:16.650306Z digest=sha256:a4d09828bf20bdc73a81261cdb6aad699a6b63a842e4e54a635ea4dabaa1b955

Observation 684ed106-daed-4f8e-9b11-24a1aced04e7 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T02:30:35.617309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:30:35.617309Z digest=sha256:683bebdcb711fcb7dc260c5e1111bd703929e3c47f9551e73393dd12e23193ba

Observation 881263e6-a814-4432-ba11-50152e4a4d62 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.595220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:1e2b6a92be91075317543b8ca26ca0288fac13fbcb266011d5f759d55daa0f7c

Observation afc53f3c-b12b-409c-94b1-edb69dece7a5 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.918870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:a2170f3db7a886fb4110250810df422deaf8b3412f27ce0d34b74f4ed81c2e03

Observation aae0c89a-f6f3-440e-8496-99ba88e4727b · inbound

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning cites this paper.

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.400308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T04:49:28.598430Z digest=sha256:386863436adc0f7bf56c46a588e30704b129bf9efd62f97f7ef98636d3906e50

Observation 863bc362-198e-4315-9319-54c2924a9a3c · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.496824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:25.496824Z digest=sha256:61b44638a85d8a09e7af3f28cd343d48c6886dae9e7c7d06185f6109a41c3e85