Pith. sign in

Paper Citation Record · LEDGER

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

As of 20 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 4 inbound Pith citation observations for arXiv:2509.22047.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22047 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:07.633786Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:01:22.596796Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a3726d08-fc94-415a-b2fb-927bf1841e39 · outbound

This paper cites URL: " 'urlintro :=.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.483220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.483220Z digest=sha256:0302a6b70e0f3aa335b109042c3197674ff761a00be7f8bcefce67962b92bdb7

Observation 66e50fcd-521d-457e-9ccd-3ec88077a324 · outbound

This paper cites write newline.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.488719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.488719Z digest=sha256:44bb1556eb0fc8a5e9cca1336f1cf47a35800b6d3b63d1c0dabf2741fc071e3a

Observation 9a190610-c58f-4f4b-88be-babca967dd9d · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.493350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.493350Z digest=sha256:2a2492c0dc1a550acdb6fb846081f3f360031988e1b0142cc962ad20b97b8240

Observation 48324b8d-8d99-46a1-ad87-f61a763f4819 · outbound

This paper cites Concrete Problems in AI Safety.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Concrete Problems in AI Safety

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.499209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.499209Z digest=sha256:0c8ac27a6b7d2e3b72529a018fff0a9551d9ceb17adfb369e8afddc55b64f4a9

Observation cc09a5fc-0f1d-4b49-a885-e643a163392e · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:08.054890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.503886Z digest=sha256:540dd9812fb314f701ef80236726a7c10788d8c1c0ce35324152f6aab3bb7874

Observation 760ba96c-ecea-41e6-a84b-43d6098f94d1 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:08.040279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.508835Z digest=sha256:d8864cb4a9c713436b739deb6341db6126329a8f7f41501e90592de977486276

Observation 4c476ffd-6f71-47da-9b68-d7ec0b870b1e · outbound

This paper cites Alegre, Ann Nowe, Ana Bazzan, El Ghazali Talbi, Gr\' e goire Danoy, and Bruno C.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Alegre, Ann Nowe, Ana Bazzan, El Ghazali Talbi, Gr\' e goire Danoy, and Bruno C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:08.023659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.513194Z digest=sha256:88de75a8de4498b306730a1171a7f12cc1df30efe331c5a83bf41298d33f076d

Observation c99a6e90-0a57-4def-a4a1-a03e316fb8e0 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:08.007318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.517788Z digest=sha256:118ce3d36abf2f9f308c4e5c65f0bab9ba72f2db521dda13fd99a67d1fb22b57

Observation da0ff5f9-e48f-4e8f-8f72-50680748e675 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.521973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.521973Z digest=sha256:628cf57f8bb7e43a432f4c4ac4ff55c40d2e63069d155b0d44abdcfc7e1c06a5

Observation 4129b698-f09a-4d28-8c40-669c6998dace · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.527109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.527109Z digest=sha256:326136863795d2f8eaab50bbab4928b1762133a4da5a3e02a961579807c50809

Observation a123be4e-6328-4b17-89c9-46ee84a7d3a8 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.984262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.531530Z digest=sha256:8a5df908e61d4dc8c9a90639aa1d5c2093c0b512a05bc05f856792d50699a0fa

Observation 0cdecf5c-ab15-49e9-90d5-0e3b2c68a704 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.972283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.536619Z digest=sha256:83dc3132eb72931e5121127d78e6b51eca4534a668a650536223f409ed8124e6

Observation 7cb53439-b37c-46ac-a318-7ddb1aed75c9 · outbound

This paper cites The Llama 3 Herd of Models.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.541356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.541356Z digest=sha256:5a801b65bc095ab669473fc38fd6b501b3b4c9bddec6f1a6438f8f626ca3760d

Observation ee9dc9f2-ed03-451f-b8e3-6b8b7c43b9b7 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.958926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.545720Z digest=sha256:b54156b8791c1accea761ec8be56d013e65d96e9c44940fe47e0f5d01652a8a6

Observation e810e80a-1e2c-41ed-9040-29ecc7dfa711 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.945528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.550638Z digest=sha256:07439fe82e7484cc5644d20d6cf8b1fae72855c882ed9ffb035ae0f47343580e

Observation 3cde26d9-8acb-466c-b53a-8b8d9b873093 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.554403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.554403Z digest=sha256:71a83f9c60a371e10657b6c8319b66383bb48d032c4f61cba0177c970bd62419

Observation e021a919-49f1-4e17-a82c-b6afdea50566 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.932670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.558156Z digest=sha256:f416ca4a0322708eca0768f36d9fe6d3a499c880701aae44c6cfc422592b4800

Observation d8f6086f-edbb-444a-ad56-e6c81bfe27ab · outbound

This paper cites DeepSeek-V3 Technical Report.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems DeepSeek-V3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.562204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.562204Z digest=sha256:2c602ee093091dd2fe680f5094083cf47e9fbbc269113c4eafcfd7d112e61a3c

Observation d896acf7-ec33-4129-aec3-2e9a403b4625 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Understanding R1-Zero-Like Training: A Critical Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.566529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.566529Z digest=sha256:2e940a87afccce9d9ee667b93d822ea9c84a72d506c46d204c4e928446e6e28a

Observation 20c222c6-1e6b-4426-a944-ae9d3b266aa9 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.920052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.572304Z digest=sha256:abec6b33f71cc7b50d27e2269eade683695bca6425510d057eb374509f7e4492

Observation e23f06e1-02aa-418a-9697-73c43e8af065 · outbound

This paper cites Bradley Knox, Chelsea Finn, and Scott Niekum.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Bradley Knox, Chelsea Finn, and Scott Niekum

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:07.905287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.576241Z digest=sha256:9bc8bf38105b9481d094d6895c1fbd1601076a16e35503cdff5cdc77972b9d35

Observation eb683d59-e733-45c2-b54a-fd10b31088b4 · outbound

This paper cites Magistral.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Magistral

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.580795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.580795Z digest=sha256:9927ba718b090c47cad6c81feadfd46af6250cc517a4967134392e2f40736252

Observation 514b1e38-43e0-4bd5-a606-2a4150296a47 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.584960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.584960Z digest=sha256:c96f5142256b78343498f68d4a94af8b351ff2d56146a8c0e2302abafcbf6813

Observation 1a6ce09f-1bb0-455f-a487-bf80d7821883 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.589161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.589161Z digest=sha256:8d8f414124578ac6b1095c5221565be51cfc33505834e860d58b530436588e66

Observation 6771ef64-3ba7-43f3-b4e3-647754b8e4b1 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.887315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.593278Z digest=sha256:96f6ed06fc3168a02108a4187e7b33cbc8dc6d0708f492842b93d70de5ad4a38

Observation 14c1781f-c34e-4d4d-ae53-b8619391bfc8 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.872941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.596945Z digest=sha256:9bac14008feaee4d69d201fc408837c49de6d2bd12a2a554ed5e8aa88372ed94

Observation 9d1b11e7-5133-4d5a-b163-084082d8ac14 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.858387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.600648Z digest=sha256:3c4ad994a2db5fcd70bec880f7befdf48612d9622c91fb73a08dc2e0b92c15e9

Observation 6d7cbe3a-90be-465c-9e47-0fa424104b69 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.605435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.605435Z digest=sha256:276f16ed69c80fa78f87377ca403217b3ff89469f9a528564f28a324ccb333b3

Observation 968d5be7-7c4c-4c75-bfca-aba6ac703d52 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.609159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.609159Z digest=sha256:0a83e3c747077450e171428ad63f29a28ecc8708e91cd3012e22638d04abff78

Observation be2d50b4-5f92-4b7b-9444-e13a3ae5bbe5 · outbound

This paper cites BNPO: Beta Normalization Policy Optimization.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems BNPO: Beta Normalization Policy Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.612952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.612952Z digest=sha256:7fd43041ebe5b20f9668ae9c343a54951e982ef75eaa9b2fcf1e16a530735e0f

Observation bd5654c1-11e8-495d-aa80-613796a2a4f0 · outbound

This paper cites Qwen3 Technical Report.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Qwen3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.616957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.616957Z digest=sha256:b724ea58abe82fc023189ed7b5ba80a403acfafd0ec9089395625cbb46d7b443

Observation d1c491fc-cc26-41cc-820f-ab52ce302f17 · outbound

This paper cites Group Sequence Policy Optimization.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Group Sequence Policy Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.621278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.621278Z digest=sha256:b6b62097e160df718f7f6b591594ba0508800646dd1ba51000dd920077f45b8b

Observation a1e784c5-8604-4ff7-aa4a-d9fa2cb420a2 · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:07.836779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:49:07.625503Z digest=sha256:b24a13e6b1287dd392d3e832b370b48c24241364b18c129bdf9e0ca92edfe7ef

Observation ec15d00e-3a95-4f00-8c22-c978c59cd61b · outbound

This paper cites an unresolved cited work.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.629420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.629420Z digest=sha256:215a9a681478cf4c7eaf23c18e42e011a380523102bbf771ad8bea90f64d84a8

Observation adc5925d-4d6f-4cfe-acad-35db3d87ebff · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Fine-Tuning Language Models from Human Preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.633786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.633786Z digest=sha256:225ce27526041e9a47fa90152aa0c422728c0e855485f19ce6f2fbae5fbdb8a0

Pith citing papers

Observation 48dbd911-a145-4704-a5fb-8372531e2bb0 · inbound

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair cites this paper.

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:00.114209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:21:43.166304Z digest=sha256:dd745fcb094ca7e46b409c534bc91c78793296ffd4fe2025b84a8fc5e8ebca31

Observation 7406bf45-c10e-4699-a55d-e05fd9c9b6e3 · inbound

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition cites this paper.

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:00.114209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T18:44:14.878564Z digest=sha256:0ca9832480e5dd1735e9830437b4416b41a4fe2f6cf3955533c12a4561aa7753

Observation 6a1c1148-0f65-4735-b945-e1745b9a9113 · inbound

CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation cites this paper.

CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:00.114209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T14:26:38.264174Z digest=sha256:4a6d71a35998e8e11b34626f74e207e04a4dfa0634055dba7e09aa27f2757250

Observation 5e56dbc7-b428-44ad-acbf-d3500a46a56b · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:22.596796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:22.596796Z digest=sha256:c4a51272920956c4a2f3f48f847820298f7bfc22195fca174377d2b6065dad13