Pith. sign in

Paper Citation Record · LEDGER

Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2405.17931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.17931 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:58:42.998808Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T06:25:27.959448Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9da7accf-d0c6-4a84-83b5-c3497a848cf6 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.597229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:1adb4611047e2c1275090131b4e337e38fe252e38d654a49856a670c6dd486d9

Observation 4a149c74-bc1f-48bd-bcf7-914f3eab9554 · inbound

Continual SFT Matches Multimodal RLHF with Negative Supervision cites this paper.

Continual SFT Matches Multimodal RLHF with Negative Supervision Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.998808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.998808Z digest=sha256:057d55e4b7279c7abaf7a9055ef5df43f12459cc362bcfeca14137edbddd074a

Observation f40e3127-5121-4f19-9b5f-901994342f6c · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.962311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:8009aff8502b4a900bd6ab623f8df02d61c95e2051a35d8dd40947e45c370a6f

Observation e5654da8-c50f-447f-94f5-15762b3e3d02 · inbound

CareBot: A Pioneering Full-Process Open-Source Medical Language Model cites this paper.

CareBot: A Pioneering Full-Process Open-Source Medical Language Model Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:28:07.903886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:28:07.903886Z digest=sha256:8e7865beeaf1cda2e86c6db94a0e25dfc05f0bdb9e95a1b8c299d35099121d58

Observation cd2b8c06-7a06-490c-919e-87236fc078f3 · inbound

BPO: Revisiting Preference Modeling in Direct Preference Optimization cites this paper.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.563139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.563139Z digest=sha256:bfc361302fe5b1b14bd5a5be4e9e4b86cca2f84fba33400d1e73dd91bd7edc9f

Observation 158db7b5-0ecb-4273-b7e3-4f4eb3b32069 · inbound

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging cites this paper.

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:36.108422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:08:36.108422Z digest=sha256:e8ef2ea659d2317ca76036ff1d6fa8a78852763b35e1367c3dca51c024a8b55b

Observation 9d51499f-dddf-4b90-bac9-be034ba83a96 · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 240

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:59.798792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:13af64cb4a462ea5b1b651cf7d8b16e47ccc551e8c7ec63deedf1be6e3dd1905

Observation 37544946-5c88-46e8-94cd-9a0dfcf70f97 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 252

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:a6696943bf3a5df849e84494f868d31a5be70ac13594871754c75bb9bc8edb52

Observation 63836574-52a5-499e-bf51-ddf136e7e96f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 253

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.653546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.653546Z digest=sha256:276461f75b0f9f3315edd99b1adee1b5c79b960552ba76469f21a03cccc94f6f

Observation 03c19841-5f0e-4102-9393-f561e2107b15 · inbound

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning cites this paper.

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T02:17:28.447082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:17:28.447082Z digest=sha256:e92dd8b229d940e21a20d2be52265e9a80dac081eda26538bb0c9eb5d2469638