Pith. sign in

Paper Citation Record · LEDGER

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

As of 18 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 3 inbound Pith citation observations for arXiv:2504.20157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20157 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:40:09.533305Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.049711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:07:56.781634Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84014ee3-0a43-46f5-b326-d6970b87b155 · outbound

This paper cites Ahmadian, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ahmadian, C

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.128421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.128421Z digest=sha256:4b466639443a4d53b9576ac3da05f5953fb258c8a6db100e7c8ef55759ebf358

Observation 4306b0ca-f5f8-4992-8f57-d901d427a5b4 · outbound

This paper cites Amodei, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Amodei, C

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.884168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.135897Z digest=sha256:a437517c24e414393cc5a921c3f69adde274ab3283cfea3e1208239acdee5ee0

Observation a224da52-802f-42ba-b81f-db6ed05552aa · outbound

This paper cites Concrete Problems in AI Safety.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.141124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.141124Z digest=sha256:932f5235bce813e9377383b66143cc2dbaf132bdb62b0026a5774883ee2330d0

Observation 556e3ce8-83e2-4569-969f-8804d3cc9109 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.146832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.146832Z digest=sha256:e51b153d8f273f6a3296abe9bd18ce53fac8cb151605ac0ad3579faa5109242c

Observation 10bd2885-16aa-4cf9-9b8a-fa048e6f08a9 · outbound

This paper cites Buckley, T.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Buckley, T

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.868225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.152614Z digest=sha256:2790237da4c7cf8286b59b4797a83110e1189182f7a1a1176510d11ed0ebfdf8

Observation 91c60746-0dce-4ca3-a41b-b26db3800892 · outbound

This paper cites ODIN: Disentangled Reward Mitigates Hacking in RLHF.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.158251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.158251Z digest=sha256:18335fad99c51555ac5e267f3d185715b991038f6b51fb26995b68675ba3132c

Observation ae0345da-19bc-4bd4-81e0-f040ef5bffa9 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.852486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.164194Z digest=sha256:fb9163f81badcabe2b55667e5de173328ca89d79a45ce4ca22267b77462305b6

Observation 50b31a4f-c1f1-43c9-a793-b8ba75fb26f4 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.169269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.169269Z digest=sha256:7287ae7749e850dc0f88906568530e62728f3a3c6b8cd1db057d638836170804

Observation 32ef45af-468c-4eac-8426-7e477df43f2b · outbound

This paper cites Coste, U.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Coste, U

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.836698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.174620Z digest=sha256:b0947a6c01a3ccbd5a6d3fbbdfa36d091a9462b930d4d9df717b42fad54e6da3

Observation 836b619c-f099-4806-bf18-aa63d08c53d9 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.820404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.179775Z digest=sha256:3734a3b0e33b3d2023447c4ecfa2290af95a0d6eb5583302b20f25cbc7db47ed

Observation ae088c77-19a0-410f-835a-817d8ce11952 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.184730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.184730Z digest=sha256:34cd70f5ccb8780a203959165a040f1419ef4b6cc158131ac72ede379ea82421

Observation 66984877-6b63-4282-ba84-e1fa7d93d77a · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.189989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.189989Z digest=sha256:4620d1359eafd90faf57aacdddb4d9247d2d8038c77f7ae401dd0ce9b0c4267d

Observation 192bcb4f-0014-4344-86d5-96a3f0016df2 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.804793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.195069Z digest=sha256:aa10c9e5a8552850e7d57f7e45762fd3bc770728e50c1079ea461b590aa094af

Observation f2d9a440-6136-4c47-9383-4967c4f8086f · outbound

This paper cites Efklides.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Efklides

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.788673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.199747Z digest=sha256:834546797247e328d267deb855d5dd8fea5f788b14c55af7205f04593155c73a

Observation e1af0d2c-812f-41fd-9a01-dd322fcab31e · outbound

This paper cites Eisenstein, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Eisenstein, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.772380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.204248Z digest=sha256:5aeab19f69df5f420087641640ba0a254d5ed31af7e15c2a2ed728f77f2d6c5b

Observation f4431f81-cd12-44f6-bb02-e79f62bbd885 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models KTO: Model Alignment as Prospect Theoretic Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.208936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.208936Z digest=sha256:c9da3495a17fbc76e57f447a2e909d4f658f96aea7df24fedfd7c619d6f17e60

Observation d57dbda2-dc2e-4c8b-8c38-bf3675930d1a · outbound

This paper cites Everitt, M.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Everitt, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.754664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.213859Z digest=sha256:1b54f910e49201c04132555ecc793344887a67fb985c642f908ccd346cd61d41

Observation 69bd380b-f029-45cd-b87c-8e663f6923cc · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.738502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.219104Z digest=sha256:0fb6a42230339247a226e5f4141e6250c7c9fa28b6353ce752d9b4c5674b6f45

Observation 1088a2dc-d7d6-4b3d-aafb-dbb2db24837c · outbound

This paper cites The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.223889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.223889Z digest=sha256:8810be04be7f7d3c6f3fb8e842161818b726c83ae39bc20f26ca45c218d9cdec

Observation 622f78c8-129d-48ff-a38f-179609261787 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.228950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.228950Z digest=sha256:dc04b9c3a4ebd18f36d83761030d8159c8f4a76b887e4377dc720edd2b3e4497

Observation 8c964a05-3372-4c03-81cd-062246a087ae · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.722228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.233843Z digest=sha256:c94cdadc4063f87106853fe00e3f3e5cb6742d040d286a393b119a44ffedb739

Observation 972e2a40-d167-4af7-9275-05ca3f9ee066 · outbound

This paper cites Hamner, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Hamner, J

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.706274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.238853Z digest=sha256:b7e241b08592741e03fb08d075adfe9636b9a4f935d6497a16b07078246b4c51

Observation 9a360b64-3cbf-43bd-8b13-02fa13c14d9c · outbound

This paper cites Hendrycks, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Hendrycks, C

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.689981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.243702Z digest=sha256:562e9f23757e46f651b378cfd65c00486ba32e3f311c38c836102875047e5bc8

Observation 4b777d64-80d6-4460-95d2-25b7b5664f22 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.249327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.249327Z digest=sha256:1577c0af7c418886c4de0b434c7a3403b579e3c8a973aa8b7415dac2affb3270

Observation 4d4d5521-3aff-4dd4-b14c-0c29b73a4b73 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.254238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.254238Z digest=sha256:0016f39fb77b48f2f882cdaa8212aac38d344e0223d9ee363191d96d85a55644

Observation a1ad9198-c9d1-4078-ab47-40aee5bc693b · outbound

This paper cites Kornilova and V.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Kornilova and V

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.259920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.259920Z digest=sha256:feefd72911d381c32ea4082adc00577ff65abbf8b380f895f0d67a3a8dd45bf4

Observation f38cfb96-0bc9-4eea-a2aa-15bd1d42b31b · outbound

This paper cites Krakovna, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Krakovna, J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.265135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.265135Z digest=sha256:b9ba0825adc151456e2aa2f968791a4ef11f308d1818892277251506c4c94156

Observation 433df192-df45-420f-9bb0-346297af1b6e · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.663066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.270126Z digest=sha256:cbb16ea84220b988f4347b946d479478475ef9a95856ca690afb836b62de9f3f

Observation 5f8948f6-c1b9-43c5-a354-fc31c7dbf8ca · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.647596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.274936Z digest=sha256:9863ea7607eec5caa75c701efcae94cb5081b8cef3527946602bc265180731a7

Observation 65d092b0-0a66-4b03-80bc-38d1630052be · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.284787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.284787Z digest=sha256:2e1a8b17be6c1b2d4d377159439a10916e445176be15c9f292fba3ef65a01b8e

Observation 5fb6f3bd-3dd9-4bce-8e61-4dd24a6dfc98 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.619821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.289547Z digest=sha256:dc24dc107f501ee92aa8b3574ff6954a23a82f9cb8f35edd780d8644a110f075

Observation 2b6b871e-def0-4485-8e08-94be64156b19 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.604313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.294328Z digest=sha256:465699a72e8c221fe7b98d1e0b39bd52b693d521dd68db7a5c097e8829dee184

Observation 2a05bfa3-4df1-45af-a08b-56cb3225c0e6 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.300373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.300373Z digest=sha256:77c307d26deffbf2e7e56765f2d955d112ab1e7a7658d89a6efaad83d4166326

Observation 35f4bdc6-319e-43c5-9303-46afe503f3c0 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.588317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.305542Z digest=sha256:9f2919c65e2d7b1330a97ddc2281a5b1199a557ebb78716debad7828f2e2c2d2

Observation 8facfa69-c929-44de-be2e-537ca11185c1 · outbound

This paper cites Lourie, R.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Lourie, R

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.572366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.310275Z digest=sha256:394d88a33f31202ca23ffb69e88ddd3ae0a29c1e91b4ec880afa86576e09654f

Observation a83fbe96-8807-4a2b-83df-d972bad9ed09 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.556775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.315294Z digest=sha256:a76cc3886cd518e1e65583eb78860826841960ddf6eaabc80ca74b4abba90b21

Observation a9e4a206-746d-4a89-bc50-5f3bbb6f1d3e · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.540436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.319942Z digest=sha256:e7ac316d7e7a648eeffa90425012137f3e0b76cdbd3874fb757d105a8e1d0afd

Observation 7b1aaf42-37b4-471c-ae6c-3b964cf90611 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.524150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.324685Z digest=sha256:8e88ec8bedab850cf1f0c01f395632c25772b35d3031d3fe65592e1479417e54

Observation 527e7dc8-f5f6-4d90-8a32-520471788536 · outbound

This paper cites Metcalfe and N.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Metcalfe and N

Reference 40

Resolution
verified exact
doi, observed 2026-08-16T05:40:09.598519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.329272Z digest=sha256:5b73b32b030dabe8bf026d508f92f6c860472eb7f8881c6c58960f5dc5ffb524

Observation a58260a0-9af0-45cf-8590-43ca5d4e1164 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.506102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.334058Z digest=sha256:dd063e702d75c4de3932a56dd335dbe00bac18b1d8d0b0ad96fa0fb718ae5bab

Observation 2294f5e1-732b-40e7-910a-3e74daaa4d19 · outbound

This paper cites The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.338975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.338975Z digest=sha256:9fd1f4706e5ee6bfc57f68afa07eb9b445c111abec743b549b0856283eec147c

Observation 68bc6e34-2ae4-4707-a5da-c381600b1957 · outbound

This paper cites Ouyang, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.488096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.344202Z digest=sha256:0ba42ed86a6a33d30f301dae98693179b04f08d3c81a49c2ad96b5cba2e38b0c

Observation 1671728b-5791-4991-8af3-1404c6cb7c89 · outbound

This paper cites Ouyang, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.471645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.349041Z digest=sha256:4adae783eaadda5088c070da57c92cb7506504f8abbc0eddf0df218821b68a74

Observation b11e6543-9da9-4bbd-a492-8597831bab04 · outbound

This paper cites Ouyang, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.454776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.353973Z digest=sha256:95ef5f3cabd4ca0bd886498d9343dbafbf25d416467ada1fc6a850f877694b4f

Observation bcb71774-95aa-431c-bfb5-00ae317dee4e · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.437918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.358880Z digest=sha256:e66feeb6c0f2a67f33b1b98a0e870205206bcb58910e4592bb8d532b568cca7f

Observation b014d8c6-a3ac-4f19-95ae-579bfc4cf962 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.363769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.363769Z digest=sha256:e54acf84451714de13587aa0e8863f18317c5386bf140cb58fd5fadc020b4226

Observation cd4b6a38-51b6-4861-8ce8-b0683ac52398 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.422198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.369179Z digest=sha256:e278ff521e3a41f7f2196f4e823e3127937ff434c4095dcbcefdcbc7e5a95fdf

Observation cdf3ab89-eaf0-420f-a33d-d487e4804a56 · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.374182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.374182Z digest=sha256:d5786e2b22ba6db3159c6308d908049136dd5e5df7bcc5f2d7c47ea31eeefc0d

Observation 11df0c3d-e5ee-40f6-9677-7e0438431df9 · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.379491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.379491Z digest=sha256:5cf62d4917dcb283df52066fbb5358a9c6d4d1c70fca9263353785151a8a451c

Observation b76ba7ae-bc15-47c3-ad44-1e5664272e99 · outbound

This paper cites Perez, S.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Perez, S

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.384702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.384702Z digest=sha256:370dd322f2b86591a462ebbeb13d526dcd46e2737508cb00cf6f26ffb601b498

Observation b0d625eb-b103-4a56-b950-355b7800b92d · outbound

This paper cites Qwen2.5 Technical Report.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.389446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.389446Z digest=sha256:3033a82bf7d37451086d9a3ebb8517f5fbdb06d84db0e9a7f2e04ca155a66a84

Observation 8b2ac55e-fa62-4471-a7a6-955281384036 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.394787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.394121Z digest=sha256:ff604cd0d631c86c21415e3e3639286e8c559f4e774bd62e7533713faffd0dad

Observation ce64f43f-f56f-4ceb-a41b-946b213c5fc1 · outbound

This paper cites Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.399353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.399353Z digest=sha256:2223f66d9da9b8af3cada4a7abbcb5b1023a940a806f2a0f7dd23274a3257b12

Observation 9ac2ff3d-b2bc-488d-a9f0-6c7e6e41f189 · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Verbosity Bias in Preference Labeling by Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.404427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.404427Z digest=sha256:35a1cdb0c42b4e77747fe25250b3f686e59d0a6c373f4036a2af1161d991352c

Observation 1ffa1d05-eec1-4e60-afd8-78932d7724dc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.414460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.414460Z digest=sha256:b3639fa6773c6baf6cf646e1283d47ff79f2a7c9b0e556a07c6172936670f82a

Observation 945a2281-cd88-478f-a281-cabf37e7366b · outbound

This paper cites Sharma, M.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Sharma, M

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.378237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.419402Z digest=sha256:25ae5eac7910c000961e788620ee64a47a54b5c6f4553104b979257476c11f7a

Observation 0af90142-28ef-42f0-af13-5cf3c4d47bfc · outbound

This paper cites Singhal, T.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Singhal, T

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.361213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.424140Z digest=sha256:1a86809488f0d47dca6a2f5a7a6813f0cd27d0d5af6a937e11490bc557d4bdc1

Observation e811439f-75ea-4476-9b60-b3b631382a62 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.342362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.429286Z digest=sha256:f5a01b378af8dd53bee8ce360af070c0ed3c1f915c697e4bf69b2127fceb6836

Observation 22eb0913-acdf-4a0e-b0da-3ae9645d77bd · outbound

This paper cites Stiennon, L.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Stiennon, L

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.324517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.434243Z digest=sha256:61a99c117437d918042fb32afef9762e6c0dbb41ddee51ddae17fa8b0271f177

Observation d0391c6e-1b24-4948-9b00-a8c2744aa454 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.306564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.439183Z digest=sha256:2140e6e599bea17469e6cd8b4b3a78ecc0a8452719dc5a7e0fc7bce3087b52b9

Observation ede0c600-1e5e-4546-8b8e-15db8f5c538f · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.290637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.443899Z digest=sha256:2747a266a3847a9826701ade6d84dc422e585e1d53ddfbb43c0dd2951fb58f05

Observation 78c37b0b-b751-4d8c-812b-8f2c0eb3be84 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.274336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.448407Z digest=sha256:ec2c02fe824f722020fef5374b9dee93e9bbefca98071ce60bf0de2db885750b

Observation 88e587cd-4c3b-4ad9-93cb-8ef92bdc6926 · outbound

This paper cites von Werra, Y.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models von Werra, Y

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.453207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.453207Z digest=sha256:34fa018f6913b63e56675d4f65028861064ed24e618ae58801587560534eeef1

Observation ceb0dfb0-e5f0-41c0-b2a2-4499233940d3 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.248070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.458205Z digest=sha256:bd55fecd0517fde0e192466bf1986ec640fb387345cfab1b09c59b4da5f7845c

Observation 81672bab-2540-4b7e-a93a-9d2bdb0a2356 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.231946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.462785Z digest=sha256:cef8dd31e48c6fbbf4591e66e33e7cd6d26ee7670362b88fa8a7aa43d01fa486

Observation 8087916d-e8cb-4716-a863-e32a161a2792 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.467576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.467576Z digest=sha256:546c61d6404e0f476a2fb8516f1a00af0678001f4c67478e21c8ffeff5d73d8b

Observation 9725a20f-ecbb-4e9d-b395-b5b8eff154be · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.215955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.472576Z digest=sha256:41fd9ce3dac20515d96c61689204cd5bd43f3a4231ff89692332b724f12832c0

Observation d1fea079-aa42-499f-9b4d-ffc4df35466f · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.477323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.477323Z digest=sha256:5fdf5d6990831b65deea143946ad696872d3419d5f0df0329f16577661b04592

Observation e727195b-d96b-45c0-b113-21d25c98ee16 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.483189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.483189Z digest=sha256:863e32a185988a65ac359e17db66e1ac289efe04e1505165fba4cd72acb104d1

Observation 30273e3a-f987-4b7d-9c49-dd0e691a5f71 · outbound

This paper cites Qwen2 Technical Report.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Qwen2 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.488443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.488443Z digest=sha256:066297d2f00ca78c8897be4988fa9a453b9c2028354a8cc519106f3da45c71d3

Observation 3a049343-5f77-4b12-ba85-9963fe750d34 · outbound

This paper cites Self-Rewarding Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Self-Rewarding Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.493614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.493614Z digest=sha256:f02614a38785a1a8e50c07b3d4af9dfb7a304d197b27d3e4ba0bdf4e7b067cd4

Observation f871c059-471c-480a-a2df-30632b8255e1 · outbound

This paper cites Zhang, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Zhang, C

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.499053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.499053Z digest=sha256:c1962ed81347268d58a823deba2044c22f77fc981355a333535c76666aebcb2b

Observation 16fbac64-21f1-4c7b-be01-69dcc5635eaa · outbound

This paper cites Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.503684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.503684Z digest=sha256:9ccdaf6100afb1a547e9e97d01edce17fa9f182a95110f64103f5cc79c8ebaf1

Observation 91dd55db-cff6-4a6f-b309-b58a9a88d1c6 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.508737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.508737Z digest=sha256:d642a5e83b3ca60151237ccf278de0aae863e8fcadddbbaf9378dd2cf67bd6f8

Observation 28dd2998-59d5-438a-8394-5222b09e945c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Fine-Tuning Language Models from Human Preferences

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.518689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.518689Z digest=sha256:560a781bc720df7fd7b5077813e0f62745e4fb1381fa1e2b7eaccd60063c4682

Observation 6aa3b570-ab53-40e2-bba8-cbe2d4dd9940 · outbound

This paper cites @esa (Ref.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models @esa (Ref

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.523386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.523386Z digest=sha256:f4e2d91a6d8c3a50438325bc272a5e1e0e8053f5a11b2dbd008c5c6092205f71

Observation 725ee1b4-4440-437b-a44b-19c7a9fe4673 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.528530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.528530Z digest=sha256:787ae12a7aea8e6018139a9e4e1a550ea6f67fc5f2d5cba4d2cebf0c050323d9

Observation 87e22da6-d4ce-4a22-b621-5a20094e4d05 · outbound

This paper cites A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.533305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.533305Z digest=sha256:86ebaaa95c6462ac7eab27dcac8e0ea694b639aa844b9f74ef22a72f598b681d

Pith citing papers

Observation 4a11841a-1d3f-494b-b92f-976e3a3625a3 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.785000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:be07811e32187d4ffc278092267aa80f56c2ac5ed7ef9a444e56eec4fa1e0ab2

Observation 514714e3-6952-4251-a15e-d7d09cae203f · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 222

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:09.180660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:09.180660Z digest=sha256:2651059803c25f8456600bf3c289da75ce4edc548e47445d4a24c70872ab48d7

Observation 7112acee-2383-4d63-9ac7-fbb30715048d · inbound

Improving Generalization Robustness of Multimodal RLVR cites this paper.

Improving Generalization Robustness of Multimodal RLVR Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.049711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.049711Z digest=sha256:43721f28032f52a1543d45940281573d2dc9636b5757bb36653b1b520b3255a0