Pith. sign in

Paper Citation Record · LEDGER

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

As of 19 August 2026, this Paper Citation Record lists 100 of 167 outbound references and 0 inbound Pith citation observations for arXiv:2608.03092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03092 v1

Coverage vector

measured 100 of 167 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.027500Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 167 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b44c3712-e4fb-4d25-b042-b058b40d0c96 · outbound

This paper cites Approximating.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Approximating

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.692427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.692427Z digest=sha256:271a70dacf7647ea8166ce14960bf0c469d97f84566a576dc9eac759eb3474ef

Observation 9d978763-eb09-4404-ac3a-ea6d751c60d3 · outbound

This paper cites Thinking Machines Lab: Connectionism , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Thinking Machines Lab: Connectionism , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.697244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.697244Z digest=sha256:279089107421cfa0c555de97abe5ed76ab7b72af79ae547703e0e791753d8797

Observation 3bc3d340-d0ad-46f6-886e-05ad422f23b4 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.701026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.701026Z digest=sha256:004fa30af8a83c86f315420bef57331d2c8951527a051cfe17dbd645582386d9

Observation 3f930e07-0f31-4a88-86bd-77b11984480b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.705018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.705018Z digest=sha256:e41e470e4386296b96e5e225285a10aa4e72ae0f0eaa21c104e8031ac6901f4c

Observation 60b9bebc-e6c2-4376-a935-351d30f38d02 · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.708487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.708487Z digest=sha256:c83a941cbcd0dbd3e2f81d616e2eba46bce16a4e0a592c97fd145a1c4029da23

Observation b4b0f2f6-ed7b-432d-b120-1e6766403ba6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.711697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.711697Z digest=sha256:a7191a20d81d653fbaba545906999b12a647f38501a36944842347690d1e7048

Observation be70df40-b56d-4ea7-a077-17a1b2023331 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.715648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.715648Z digest=sha256:96ec9ffa21ccbc49b88bc6580cc9f2fac24acb04de55eeb4fd4b04ed46dbaf56

Observation f0787fa9-408b-4eb8-9531-1affac13a384 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.719267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.719267Z digest=sha256:e06d7362a4507519cc787a0cec78fe3363eaa46ee24e4e6ca749791811680730

Observation a31bca12-0bbb-4470-bd94-c9265e587d90 · outbound

This paper cites NIPS Deep Learning and Representation Learning Workshop , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation NIPS Deep Learning and Representation Learning Workshop , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.722363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.722363Z digest=sha256:a238b06fc20f482846e439e7536f6b2f2edddd680cd1d1b75e09155799ee8414

Observation 1c332071-41cd-4d27-a618-1b40ac4a1c61 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.725125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.725125Z digest=sha256:e6f6ff4bf2fcc414c94f65fb2130dcd8f78de62b5f3afa6a8176eca1e6eb5dca

Observation 66007b7a-c580-4046-af05-70c4019a05ec · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.728125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.728125Z digest=sha256:128a51142b5d614572ca246188b130a365e0050e713e7568ef30ebf93bfe91aa

Observation 5e772e42-ca45-4703-972f-f65ccf585771 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.731053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.731053Z digest=sha256:54d054dabbaf6782a215e2927fb8049ed1b684afaa32f4a3059b77b0543fb8d7

Observation 58fb0731-3699-499a-973e-7019f2ddc652 · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.734494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.734494Z digest=sha256:75fae2ac71f6d79c7ce12bf3078027ade1debb704fce79f0bfaa1ffe827ab156

Observation 301e8b44-978d-47bb-8ef6-e89e0d9de9b6 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.737879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.737879Z digest=sha256:0baf88df6278ceb1fa413e524439b45e5507de151835e100cc91e585bc7a0fcf

Observation ea2d3af1-68aa-4862-81bf-184310a1320e · outbound

This paper cites 2026 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.740548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.740548Z digest=sha256:93a2c26d7ea167e5010db7d2754cb930b1aca79a33db1eec879637c3e59b34b5

Observation 7bffd4d0-e002-4b5a-8291-bc548fcdbf54 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.744202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.744202Z digest=sha256:a3b7aaed1db3f5b2b0eecf64e3cf6276112acf4cb2819138612226d1a965bc78

Observation 1e160db9-9907-4243-9747-5cd4c6284f54 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.748401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.748401Z digest=sha256:8de32dd5a8e032a9536d2c10fee8f644d7a82f3c8634a55d0673d132e01ce715

Observation 9981f875-77c8-4e0f-a65f-c41b1036962d · outbound

This paper cites Language Models are Super.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Language Models are Super

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.751873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.751873Z digest=sha256:a25ee0366d8737e4e49ece6ab81dfd6e5c5a172b5f73f932376ef5dda9f6e20b

Observation f8609371-1e60-447a-afec-bf88a774d8f6 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.755666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.755666Z digest=sha256:984f8ae48799dc0975184389bf4c738e76927f4ff38a0bf1a1c0e9894d0062e2

Observation d1c3a387-c06f-4e35-98e1-48f32b788dae · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning (ICML) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 39th International Conference on Machine Learning (ICML) , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.759548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.759548Z digest=sha256:4d973f187fc5642ba547a00d03ced9cb0a7b2f0b5ad0406194ed429a4be94e45

Observation 1bb26a02-3136-4743-be38-7b335f5a47f4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Constitutional AI: Harmlessness from AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.762569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.762569Z digest=sha256:463306eab2b78a1ed2e52d741e7c0baeaa385854fcb1606e68807cba8616b99e

Observation e55ff134-a8c9-4fef-be53-5822847b22cf · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.765987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.765987Z digest=sha256:b88a6237549143a48e1b682b4f20396138c52bf3243411ac41324847a2c704fb

Observation f469fe5c-7c2f-4362-a77f-f9cb0ec5a8f0 · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.769326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.769326Z digest=sha256:d78a2b91239b02d5641c66667037dc5efdea84b735d616ea26e068dc39891e61

Observation 2a04f551-0c78-48aa-bcd0-2e3b1650c66a · outbound

This paper cites Hashimoto , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Hashimoto , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.772522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.772522Z digest=sha256:6b56c6ebb78dfb484a57c439daf093011ba9c0166e73f6096f3d8dabb0131362

Observation 8a5ba99a-6b97-4694-a945-a9535aee84ac · outbound

This paper cites Arithmetic Control of.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Arithmetic Control of

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.775564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.775564Z digest=sha256:bc55395e66b092b936c89563ab63657a9b6e8659bc08fc45d7d2d9f4c6c04010

Observation 9c6e0cdb-3ef0-409f-b06d-7d3837124cd8 · outbound

This paper cites Rewarded soups: towards.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Rewarded soups: towards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.779097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.779097Z digest=sha256:2f053c16181a8134fd7bea9f86d53e353a2d4cdb95ea715a18715dc392d8f418

Observation 5911d5ed-af19-4c20-99d2-f54f52ae8c3a · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.782265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.782265Z digest=sha256:34dbab839895a0483fb511a30e37f7ea1a4f0c3bc279ce07648e9d050498308c

Observation 55315f55-352e-45fb-92d4-c8072d293dd1 · outbound

This paper cites Back to Basics: Revisiting.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Back to Basics: Revisiting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.785597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.785597Z digest=sha256:38f8a68789fbb7e248074a02225fdb5df218b834cb5e114fd9fa2479329f9c1f

Observation e845bf90-36b6-4458-8001-5b3997453d2c · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.788532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.788532Z digest=sha256:d2ff193a4ba495722be35ba857471c07dd574f05d5e1e224d42839c9db560178

Observation f21dcfbf-6f20-4b6c-b3b0-a49a4d561e47 · outbound

This paper cites 2606.16771 , archivePrefix=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2606.16771 , archivePrefix=

Reference 30

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:57:54.687922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.791392Z digest=sha256:17933a3f654eba203daff70dde0d52c38ed2d05109af7cdb713247d69e2a3394

Observation a56a4a33-f71d-4464-8afd-e78f8f8071ad · outbound

This paper cites DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:57:54.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.794931Z digest=sha256:d13149acc10c6c5e85ce8c354a8bbd4ddc0e2c3652cd27092dca5370495920b6

Observation 947a9e8c-6c31-4e97-b98e-94b3b52ea8b0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.799215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.799215Z digest=sha256:e8bcd8e6c00dedf38123f36dde3fd9de1c992bde477f2404a5409bca778fb72c

Observation 56257440-42c7-4d8e-8768-c59fd0c926f5 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.802430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.802430Z digest=sha256:88d92b3c11f9793e0d4eab20fee0b2803c97d5871363546e2d504cfd640aab06

Observation 60362f9e-c83b-4f2e-b4c4-f50e93c95758 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.806791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.806791Z digest=sha256:c62de190bc9d5f1a888a46b6472486bfda70b039940412a69bf058cb2865d019

Observation 1355564d-ac98-4810-a57f-0b906cb5eace · outbound

This paper cites Gonzalez and Hao Zhang and Ion Stoica , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Hao Zhang and Ion Stoica , booktitle=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.809737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.809737Z digest=sha256:5a39e3904188a733d41a0886c42eb6469f665bb3c587bcd1578d4b02bb107a33

Observation 2dab2f60-a705-448b-a87e-589b931b20b9 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.813058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.813058Z digest=sha256:15f8ea6b4df369e77a4d54628893ca7330778d680862cf95c815039f620700b5

Observation b3971191-ef85-4ff8-b1c6-f11cee07edff · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.816389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.816389Z digest=sha256:8843918a2c9732b7e89c844e4fee48296293e3ce94cd51208d845253defeba3f

Observation eb7f4ffe-1f1a-4a0c-9444-fbd9de0231a0 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.820297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.820297Z digest=sha256:2159a80fa5b940bc4271c52f4d5ce69a0b3481d7aa172e5086c6263b6481b03e

Observation d237aab3-d2f0-45e2-8d98-a83959f16682 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.823262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.823262Z digest=sha256:2d0086b9ef09dd257f259c26c0fd8c6b9e3f75422e05ee2821e9fe0b8d831ebb

Observation a9c73846-1a10-4f50-9319-e3fb17d0e450 · outbound

This paper cites SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:57:54.619457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.826235Z digest=sha256:0be62b304895d05944328cf6f46e8c8519005ea1bb9b96dcf5cb8af9a5f7e512

Observation 5f1707d8-0451-4278-be2d-ea5f5f65f148 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.831378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.831378Z digest=sha256:3865942a96153e2356a8e9ce29baac011d32486432f1957eb9891c6b8f67b25d

Observation 7100e852-0073-42bc-9aef-0a21d22ff9ef · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.836395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.836395Z digest=sha256:817c9fe86af7c7b30fe19cb402c84c96a663f8acb968dc5988a7890375f53657

Observation 458571d0-4753-49b3-bc5e-b90ebef85fe3 · outbound

This paper cites WARP: On the Benefits of Weight Averaged Rewarded Policies.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.839991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.839991Z digest=sha256:302a6656f63163baf54d0e5983f0626b2803e87a95a407d1848d62e9db6c8415

Observation b74d5406-0524-46d2-8a75-34d0a3b1ce8c · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: ACL 2024 , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.843008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.843008Z digest=sha256:c394b6fd394572cd3606cd236ac63aea87d39ed3364bcdf1c0b8e89d49aca8bf

Observation 52d38c34-105f-4271-84ec-e8728d765d87 · outbound

This paper cites Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.846073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.846073Z digest=sha256:83c7f250bb3d7b4bcc1528778e85990255c53171d9b8a2e20e25e167b98fa75c

Observation 4a920b9b-be2f-48cc-8f55-8369bda27d7d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.848979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.848979Z digest=sha256:4a107adc4fb09596082448f396c510df667094088469d56b6c592402217c9480

Observation 42b17cf8-95dd-4af2-a4f0-cf5d12562fa7 · outbound

This paper cites 2025 , doi=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , doi=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.851928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.851928Z digest=sha256:ded5ce42c63b7c86eed6660925c659aa9357fdd1ea82c0cf583fa36454d61946

Observation 910b124d-5a82-4792-b21f-92dba7479cf3 · outbound

This paper cites 2021 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.855105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.855105Z digest=sha256:474e52e576f797735e55d505680ebe8c99c7315e60f4e3f9ab2149105fc0cc7a

Observation b66bf340-d024-4cef-9b55-5130159e6940 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Measuring Mathematical Problem Solving With the

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.858442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.858442Z digest=sha256:754342d8623c213f96393d2746823679c3ccd0a404a86df0ac57d75e3fc54fac

Observation 7d649c44-756f-4c84-8759-1a43d032246e · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.861265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.861265Z digest=sha256:e38801db73aa42138cf0157fcf3dd9cff4d0fcec06b006720460a8a956f1b1d5

Observation 739164c7-0b3f-434f-b1dd-9d7647fb01c4 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.864101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.864101Z digest=sha256:c3f85a1b6aa77ac0916280dfbee31322de2111b506c11feeca40672ff26fc8a4

Observation 933ed797-10b5-4678-bba9-08d2b4935b00 · outbound

This paper cites International Conference on Machine Learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Machine Learning , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.867436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.867436Z digest=sha256:c5828c8bd34ff1c24618d956f100f292e5780779ec7a7f05a1986be6202e0269

Observation 4c6525fa-6a6e-43e7-b798-c2b3bd17038b · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.871011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.871011Z digest=sha256:5d2c9f9947b2fa9328c4f06ce9282ebaac2d8611cacdc0f51d5c66b63c9966b5

Observation 53692ec1-9980-46a6-ac36-d69d18c923d9 · outbound

This paper cites 2026 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.875132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.875132Z digest=sha256:e3b5ad684b2506115f1fb663207fe9e7629b98db10d61d80321077efe1f27d78

Observation bc3c0ad2-cb95-434c-b0b0-3abcff808756 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.878335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.878335Z digest=sha256:ce218c4cd18f2ceb0e7e973cdf642f250aa481ef16af1df679f5feb475da3f6b

Observation 1506ac23-0a71-46b5-a02c-d10d8f928cfd · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.882201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.882201Z digest=sha256:b87f30234a1cd4441594c72b2c943ea2fd2a6c21fc8cc446485e02ac4cae749d

Observation d606a231-f2ae-4509-8597-e32f247ea1ee · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Ion Stoica , booktitle=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.885627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.885627Z digest=sha256:43a78b452a020c1440802b5313dce87a03f0b772eccdcb8314c4149e9a5f736c

Observation 2370deaa-773b-479e-b087-935cb248bd00 · outbound

This paper cites Machine Learning , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Machine Learning , volume=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.889122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.889122Z digest=sha256:ce150462b92b7e180bf13620bcc94ea9ce6ad9b4b59dba3e2610e0236d6f3c87

Observation 4003c66c-0810-457c-b427-a3be5aa1dc7d · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.892306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.892306Z digest=sha256:89355b3512f4ea2f43470abb07f250d0a498591f9d3fd395109c1625b16659bb

Observation 96673f63-9191-467a-80d1-a86c82d45f75 · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2024 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the association for computational linguistics: EMNLP 2024 , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.895395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.895395Z digest=sha256:500abf9a9e03bd51fcd4ea4fae25d37408c76509c1075c203ddc4f18e9590ae3

Observation fb436b5f-bffc-44f3-8219-b28e115e373f · outbound

This paper cites Everyone Deserves A Reward: Learning Customized Human Preferences.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Everyone Deserves A Reward: Learning Customized Human Preferences

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.898499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.898499Z digest=sha256:53a663677f923a6fc15aac0f7a65ecb0400c929525b68dbacd089e46ac262d78

Observation 17f84292-f9a6-42c0-938a-43315dd4d88a · outbound

This paper cites Kimi-VL Technical Report.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Kimi-VL Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.902060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.902060Z digest=sha256:8f270d44d002903804d0cdbfcdf76ec80c6d4cad27e01e33aead729ba8b2d722

Observation 5c547009-05c7-4ae9-8842-3864d34f42c6 · outbound

This paper cites Advances in neural information processing systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.905649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.905649Z digest=sha256:a530ce6291450ff1cf51f34590906082b520208a2e698f98fe8e51d37f918a72

Observation e7d3048e-1d8c-4e85-be5b-7b2308f1e963 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.909368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.909368Z digest=sha256:e35456274cd8fb9d4add4f00a11dd8941156a77bb343be97174d9df2225622aa

Observation 37c9e7be-5e37-4de4-a6ae-018e2e77da91 · outbound

This paper cites 2000 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2000 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.912159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.912159Z digest=sha256:547994563e8fc0502605a4052827c54954419559af6c54b231c8226841af945d

Observation ed62110f-f4e8-4f0a-8f43-9a5398d88936 · outbound

This paper cites 1997 , publisher=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 1997 , publisher=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.916002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.916002Z digest=sha256:fad65f51b3ff391028a5ff7eb1202dc2244a4500e91648624c6c07c070583fb8

Observation 0a830bc8-5b09-4916-afbd-fc7199199291 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.919633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.919633Z digest=sha256:3a7946456ad6e43714e3cc50326ca2635ddec2d03c0990c6e8b6484e498d3a7d

Observation 74fe22e8-2260-483f-be44-6ee4dc7feb15 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.922992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.922992Z digest=sha256:b0a76622318a4e7ace8978f301353167e883728cb4f43d42d73aca868a48a7c1

Observation ff4ac2d3-d624-4216-bc39-52052533cbe1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Representation Learning with Contrastive Predictive Coding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.926780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.926780Z digest=sha256:cd8972d8db34827019baf6497dace4db2c991f8cde6b477906afb291b52a2cf2

Observation e8595768-a948-4f78-a4aa-69e7e79a006a · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.930464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.930464Z digest=sha256:908eb29a1875a3ccb0cede2972af3c243cdf88d5206ba9ca2325aff631b4517a

Observation 3b0a12de-5c06-49c9-b42c-d1066999ad54 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.933734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.933734Z digest=sha256:2d2c74279015a715ca562cf6705c094fc7c98e22ae78eeac75068e2692b166ca

Observation efceb99c-e966-482d-88be-1bee64ba9100 · outbound

This paper cites Advances in neural information processing systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.937179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.937179Z digest=sha256:9fe3f29a0bc2bd468bde11e4ca779ff6565ecc7703f67452635c20cbe5e14279

Observation e1911532-51fd-466a-9c76-2074d31e8c6b · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.940237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.940237Z digest=sha256:fbbc1457260c9acf2a26352bd456f7bd8950b52929c266e22181c2ac25b50694

Observation 09f95b76-65fa-442f-af2a-decc09c41a21 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Towards Understanding Sycophancy in Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.943369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.943369Z digest=sha256:e53b06efb34e5a5f2bb97b044ecb1ac168b179648205ee4d2fb92b2dac3a5eb1

Observation f904c68f-6879-4859-8ce7-4640b3670e3c · outbound

This paper cites Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.946749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.946749Z digest=sha256:753e719806cb614b65816a3dcb47b0874bf3777550fe5629121ef4613bb976b4

Observation 5ef0411b-074d-420d-b9bb-4b3e3314c8c2 · outbound

This paper cites Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.950455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.950455Z digest=sha256:9f84635094453adc20e2439f8b7fb614f17a9ac71ef4fad2810639363104f689

Observation 90036455-6a28-4aa9-8e78-d0208b1910d2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.953858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.953858Z digest=sha256:2dc88a83316183d18b40b8e2f94a5ab001032741352a65799244df54e1ac62f4

Observation dcdda4da-491a-432d-a3a1-4b8f546c4c3d · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.957789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.957789Z digest=sha256:e0a70985a9f78a8b1e75cfc5273c49cb762c40311da5974d0df9f9d04fd245cf

Observation d8724f47-7509-4199-9099-2e2d18c147e5 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.961334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.961334Z digest=sha256:aeeecd21249f2e811a9721b66e1f14fd54eaaa90ddaee174c2e28d87863d3ea3

Observation 783e67f2-7ba2-4956-96bd-cd9d5f43d6f6 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.964603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.964603Z digest=sha256:66e8708d53bd5037f123e98fb29bccce790b4a143aba6e67eeab1e7ebec56de5

Observation 283518ca-0f74-43ac-9f73-68e4f8110187 · outbound

This paper cites GPT-4o System Card.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GPT-4o System Card

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.967415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.967415Z digest=sha256:f65acb280fe504162cba4d89296f7f75edf47e4bbc93c465acd63b3b5c2dbb0d

Observation c6c8126d-f837-483d-947e-fbd78482a8dd · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.970403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.970403Z digest=sha256:7bf629827f7132dc24e80ed7e1c436cddd564735474f940a88e8c63ea52a4711

Observation 81a4ea9b-7e5b-4972-8c5a-333a96a32c0f · outbound

This paper cites 2023 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2023 , eprint=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.973324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.973324Z digest=sha256:7908b41cf35327a43d154531d33e9e3b97debfaffcdaffe7979bf9e796387d87

Observation 8f68f78d-087d-498d-8e1d-b1b9609f78ce · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.976058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.976058Z digest=sha256:f8c15f246988063d0f24a9793b735f5447e097989c0a1d61439d8f146282f790

Observation cd26a059-5d3c-4d27-828a-4c16ff5ba1a5 · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.979341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.979341Z digest=sha256:89459321770bef153bb04afd5a652482338ad78f1e964f5494e567c980367a6d

Observation 2c52bc28-2c74-4fa0-9f45-d1f6cf46fb42 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.982198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.982198Z digest=sha256:4e084ca4876207cb504ab21248c758e4c11d75193a217f6624a3d37578ab734e

Observation 8d7d7669-d088-4acb-b031-e45a426084a8 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.985088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.985088Z digest=sha256:6aa1422cc802de3abe99ece988ea4acfa16b118898cdb36f030c7eb1a4826d4e

Observation 6d1257dc-eee2-430c-9915-3f920322fb5b · outbound

This paper cites 2021 , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , booktitle=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.988450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.988450Z digest=sha256:cf185f9a88769499bc00e3497292614ce7ff089051a1d3ba7ff708dde93dd266

Observation 98600c1f-5a41-4af8-ae8b-0c0c8d4267af · outbound

This paper cites 2019 , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , booktitle=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.991165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.991165Z digest=sha256:325925dd24d8aa10d08c5e5b72be8bce6b43e94671f027f0b84c52e70505123c

Observation eaeff446-e299-471c-bef7-f32774ea6311 · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.994950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.994950Z digest=sha256:a9ceaa9388d65bd7abc9c23239a01dabdd8d4c32adab7447bc30c254e2d04830

Observation c9157bab-d55e-46d8-972f-18644528a140 · outbound

This paper cites The method of paired comparisons , author=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The method of paired comparisons , author=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.998943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.998943Z digest=sha256:032468c793a47e8a8efd29debc0724fdb3fc779bcd79bf7bfec61aef6349ae49

Observation 3671c0d6-78d4-49cf-9a2f-e02ec65f551a · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.001963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.001963Z digest=sha256:b95e49b08aae5a4b83c47a638fd91c544c771a64d4b15a77d8ff3550135bc18f

Observation 6db5d0e8-e58a-437c-85d2-fc3a6e5e6409 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.005398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.005398Z digest=sha256:8dff5dff7138998cfc83f03ae334e087bdb72d96439fbbd33be23ece32482b1e

Observation 8a3c386e-3536-4b72-9788-809f86b0f89b · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.008261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.008261Z digest=sha256:8ec81a269e5524a17684f59ca709597b3252e9bc0c6c8ae97c7563f7f43ef450

Observation 6661f9bc-b9bc-4d85-a316-01b58d24a03f · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.011160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.011160Z digest=sha256:605545e541b2d2b7bc7a54b98bda88eb562fd903abbdb04ea802d84104915242

Observation f7c4114e-0800-4c37-9b5e-fc611cca1038 · outbound

This paper cites Instruction Tuning with GPT-4.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Instruction Tuning with GPT-4

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.014237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.014237Z digest=sha256:0cf5a99f147a4be657aa78eb88f2b83f6e642e80759b1a8fa9c013679821eae0

Observation be4702e3-bdb3-4135-b817-af21f4a2d47a · outbound

This paper cites 2021 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.017797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.017797Z digest=sha256:f32e8c01c97dec8383b82ec48a844d8ae461bdab8f3dfc47031f892ecdb8bbee

Observation 1884369b-ba39-47b9-86d7-cc2fd09e02fb · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.021547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.021547Z digest=sha256:386bb0a2b05381e74fec38ca7fd2c0f5f378cc702f2b9b6a7eb589182824f75e

Observation a58c936d-b789-4294-8d0b-8761a1b6fc64 · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.024618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.024618Z digest=sha256:69ce1864c2be496cd4733a22e466d172ee98fe80de66e651aaad8e0a0dbad2e0

Observation 9d8a1c50-2af2-488e-b1cd-26f9a98fee5a · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.027500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.027500Z digest=sha256:838ffc3737dae5cf0f166acc058ad273e1b1c945ff822b56992a01417eafb130

Pith citing papers

No inbound Pith citation observations are available.