Pith. sign in

Paper Citation Record · LEDGER

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

As of 14 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.25659.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25659 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:52:02.517746Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved50
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4e5bf94-41ca-4fb9-ac42-97162986f8ed · outbound

This paper cites Advances in neural information processing systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in neural information processing systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.281764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.281764Z digest=sha256:795a33f6212afec9e41d8ec5fb0754e6eba3bc180c27285eb380ea9c827003fb

Observation bea19546-555e-4372-899e-faad9d654d27 · outbound

This paper cites Advances in neural information processing systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in neural information processing systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.287843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.287843Z digest=sha256:bee25ffa688ed671e5c266cb65b39a5e52e26c4f0621acddec94b74ce805b842

Observation 031f23cc-6ff5-4363-88f5-4a3d83f983e0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.292292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.292292Z digest=sha256:b75407ea1ccc582492c936d61b7a0f0d6f5c5b6a449e190cd1b78f8d902f2960

Observation cf5f8084-4cd4-4beb-9bb4-a4af4091f56a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.296787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.296787Z digest=sha256:1c2ee965896bdbcf227afcba79760d5085458c0b2911148ace3db42b4483926a

Observation 029dbbdc-9bc9-4769-91fe-1a7632f31b9e · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.301364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.301364Z digest=sha256:f2459c324c290061ee46b1b32325260abb907f595f9e2b95921ba1aa49fd6066

Observation 14cc2846-ea2e-46d3-9b93-9cfa3adf2951 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.305972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.305972Z digest=sha256:723f4def127c670af5acb9dfca4d00fd48367b585f073e409cd90d6d4a83991b

Observation 37c13eec-aa26-4575-87fb-3084278f5d59 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.310961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.310961Z digest=sha256:3d40495f771231f36c2c863f5f6ab27cbbca62c6207c2240c82e18d55163ca34

Observation 406497d2-7e78-4100-b276-e90df5358498 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.315552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.315552Z digest=sha256:06db544f7b8e0efb0e413b99540769d37c121708360eb525fd7373d1de72d5f0

Observation aceaeceb-60e9-490c-b384-bcbd9465245a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.319800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.319800Z digest=sha256:5199eab92428c0c1e66441d48b838ea21567d4272172f7dc84d8e0bc6ce4094e

Observation 34dae7ad-51d8-4702-aaec-601b67a86ddf · outbound

This paper cites NeurIPS 2025 Workshop on Efficient Reasoning , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization NeurIPS 2025 Workshop on Efficient Reasoning , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.324700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.324700Z digest=sha256:90cda6f063ebca758d5a6d2d4f83e1b564dba1ecc40d2825fde5f8112d150f4a

Observation 63151e3e-2ef8-4355-b3a7-d6476766c662 · outbound

This paper cites International Conference on Learning Representations , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.329319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.329319Z digest=sha256:bcc01ab3f4899631ae8f277e7d78bb79941e6ed598b844b7de8a0379cbd8e499

Observation 9c41472d-e77c-4984-9a62-0c9dc467881a · outbound

This paper cites International Conference on Learning Representations , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.333486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.333486Z digest=sha256:88fef3a3607f0d635286fff96f3cc3dc1d8bbd97c5972cd849e8c9b45a416488

Observation 0f59b9b5-f567-4b8a-9de0-ec738d5f571e · outbound

This paper cites International Conference on Learning Representations , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.337688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.337688Z digest=sha256:136d8dfa30e53a83cdabc38e68ac36311779416838c474fdadab90fcd48661a6

Observation 156cd443-06da-4e8b-9036-9fdb0d85d21f · outbound

This paper cites International Conference on Learning Representations , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization International Conference on Learning Representations , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.341921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.341921Z digest=sha256:d0cf08f4e25c01ed5419215efb1277f539c768d74ebc3d0a86da74ab48f2e905

Observation 1e4c47de-57bd-4b61-95e2-ec02196a1e89 · outbound

This paper cites Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques , pages =.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques , pages =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.346262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.346262Z digest=sha256:a2bc92be152f0eca18874d55cfca22486039fc8a1faba1a6d90a8ca9d69fea98

Observation 23d706cb-a763-44b6-9d35-af692cb66590 · outbound

This paper cites Machine Learning , volume =.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Machine Learning , volume =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.350560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.350560Z digest=sha256:8785141d3c5f0c9a231af84bbc130ec5070cd2605ef483cdbb4c8895df15ef38

Observation d368c68d-316b-41d8-b882-bed880420339 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.355009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.355009Z digest=sha256:35f9d020cb26bab005ae3652b89643439d0b8677a10678eb59ececcc298d6f5e

Observation 190e7da1-0325-4fa0-b071-885335bda75a · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.359613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.359613Z digest=sha256:4a7fe90bfaf0e48ba6951149e88f3a84eb60a1749f5f263fb047eb4591d8ce5c

Observation 6f3ea0cf-2511-4549-8d46-98bbe0eec267 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.364021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.364021Z digest=sha256:4e74406da61103dcc5df2f211e71e2d655cc306118e871d566ebaad4e68772c0

Observation a2065375-3491-475d-9581-1a15bbd250c4 · outbound

This paper cites From Generation to Judgment: Opportunities and Challenges of.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization From Generation to Judgment: Opportunities and Challenges of

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.368769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.368769Z digest=sha256:f815c8fedd59d8dfb10d5d8ab1da8cd0baf7cdeee27bd12244a42a7c0adf467b

Observation b870122e-ed2f-431a-ad5d-a06161f4a113 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.373417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.373417Z digest=sha256:c4e01f1f9f6148ca91838cdb2025f08ffeea4be9e0649721c76a91e43c76011f

Observation 7fcc690d-0329-4642-a660-b6d0647ddd8b · outbound

This paper cites Computational Linguistics , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Computational Linguistics , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.378348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.378348Z digest=sha256:9d6d6b0370124abe7431b33408ee3ff3059b3d375effa92f54e3c38ef59d1e2b

Observation ec8f0b2b-8204-4f70-9228-e9241234c450 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Constitutional AI: Harmlessness from AI Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.382901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.382901Z digest=sha256:2bd992306fe6e5fe0168cfb1bbb110419190ddbb15bfc16b7394fa7bffe509c5

Observation 8f63742b-9734-4084-9bc8-1ea3a7a3d81c · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.388245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.388245Z digest=sha256:56a7e880fb633e90d875b01d014a95578dec4b04b136518174443bb8dc0bf0bf

Observation 35c04bb7-a48e-4f65-b89a-eb4165b2b3d0 · outbound

This paper cites Probabilistic Attribution For Large Language Models.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Probabilistic Attribution For Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.393126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.393126Z digest=sha256:b21cffacdd197eac5bc1fca1800bc2583bc6b939e9fb7fae5e03f4f4f3b14408

Observation 89b284e0-dce9-4b4a-ac38-72c9538955f3 · outbound

This paper cites arXiv preprint arXiv:2512.23457 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2512.23457 , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.398455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.398455Z digest=sha256:4d98eb3c3dc0155c5ce0448f995dbf975b5c8a9ad55a0fd166af2ae8300bcbc1

Observation c2b0cc37-69f0-4e28-9e7c-54d4f13ed042 · outbound

This paper cites arXiv preprint arXiv:2511.10507 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2511.10507 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.403060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.403060Z digest=sha256:17068dc14cc34a26e4ee76df5d240c434076c77c98c25d41a2eef4aa4cb62580

Observation 0df0d396-1817-46a0-bea1-0533f77635f5 · outbound

This paper cites arXiv preprint arXiv:2509.19199 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2509.19199 , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.407471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.407471Z digest=sha256:a961b3191ed9c76cb59b6e02132b81121360cf27614cecd2564ba27456dec8af

Observation 4f318d49-7162-4062-9395-5e70f20c5c66 · outbound

This paper cites HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.411714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.411714Z digest=sha256:9b7111e2104479460ba7a8716f7537f025b3bd0d336127e70773fcecad991e59

Observation e72dedb9-1533-41b7-a921-bf81ed9c770e · outbound

This paper cites arXiv preprint arXiv:2603.08754 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2603.08754 , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.416701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.416701Z digest=sha256:5ce6296d9ad15c94f28c199fbea6e5f211154ef1183ce41b8737f99a3dbdbc22

Observation 647ab12c-787a-448b-a255-14fc95088eab · outbound

This paper cites arXiv preprint arXiv:2601.08430 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2601.08430 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.421730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.421730Z digest=sha256:7f3efdff898c28f5edd68318197fff71a11fb664a66a792be1e7385b049cdcd6

Observation 11a5e376-83be-4e00-9059-fc6f8fb46862 · outbound

This paper cites Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.426284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.426284Z digest=sha256:6a934f8d4259969733025ef30f6ad7bef30a22c8769fe8ce6bd41518c1ae5e93

Observation 104e30d9-1c2e-4632-abe1-f7ec2a78079c · outbound

This paper cites arXiv preprint arXiv:2508.16949 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2508.16949 , year=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.430980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.430980Z digest=sha256:899f2f8129222649e0291dee54599c04b557cc3652b73c2c2eef3ddd360fb087

Observation 7d82b2ed-2b5b-4b33-8bfc-2f0cbbea53c8 · outbound

This paper cites Self-Distilled RLVR.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Self-Distilled RLVR

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.435481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.435481Z digest=sha256:6718f000ad9c7e542ad0a023b607148b09e799c810971285811a48772e804654

Observation 7f85128b-f8cb-4594-8b9c-a692bf7224ca · outbound

This paper cites an unresolved cited work.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.440062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.440062Z digest=sha256:f2c25df56877a11c7b10c07cea36303ccbe90e2962a739486e1e0f3a3edf0a4c

Observation 6523b596-bfc5-436d-8191-23449ca69c77 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.444230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.444230Z digest=sha256:20a36b22765c0eef9bd669bbd4c7e134fac930e94f3df5d0276552259c55ba95

Observation 61fdaf94-133b-48ec-91c5-1134e1a81d18 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Instruction-Following Evaluation for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.448730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.448730Z digest=sha256:8b260fe91adce004f1e2f5cfefe090fcdb81903df58c27bd01924c1baca9ea34

Observation 8f26d11b-1a5c-4d90-842e-fa713e36e00b · outbound

This paper cites MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.453431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.453431Z digest=sha256:7757ebca72a794b58cd468a86d498c3f1366c8906a253e50b47c533710c5f86b

Observation 2ed24d3b-e25f-4c7f-860c-7829083f532e · outbound

This paper cites Group Sequence Policy Optimization.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Group Sequence Policy Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.458662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.458662Z digest=sha256:7d83a6f1f974d0f9fe98edad9bf2a140a27f35d7f4e548750742c5c981bfde96

Observation 395710e0-aa83-40c0-ab23-0e61267842fc · outbound

This paper cites The Llama 3 Herd of Models.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization The Llama 3 Herd of Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.463478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.463478Z digest=sha256:b7e22d637b97ca39e689809883b456af42ef026e682e0e2f2b3ff7191444bd8e

Observation 40ba7d64-5a45-4870-97dc-af7f0c83f7ad · outbound

This paper cites Qwen3 Technical Report.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.467945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.467945Z digest=sha256:ed443122a26159cef500144be3c4501bea7286e68a7fb7071f625da854a6f571

Observation 1c5bfd7a-f17f-48ca-88a5-d674ff3233ad · outbound

This paper cites Qwen2.5 Technical Report.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Qwen2.5 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.472574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.472574Z digest=sha256:36cbd9eec6400c1ffbb7cef9aa55dd01380bdf2bbb78304a4b191a33e71b59c3

Observation 4b357547-316a-48f5-a73c-4ec5b56ddc10 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.477551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.477551Z digest=sha256:45ae37bb099b018e71cb56c6b2e2c90ec968b2bcfb964733a63e7f309d2205e2

Observation 2dc21fb3-ccbe-4119-b851-cc5e0a8819cf · outbound

This paper cites Proximal Policy Optimization Algorithms.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.482243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.482243Z digest=sha256:01c7b1f4bedddd7d00526155a98fb4a18b2b62b7f7a0dd9bf1dd5c67074fb3b1

Observation dfeeab2b-5a52-4fe3-847c-679898dae623 · outbound

This paper cites Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.486709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.486709Z digest=sha256:96767434746f9bbaebd49b71e4e6e6aba229af6112d3eeedc7f268b83402d3fa

Observation 0b8610db-4088-4402-8f66-f47e301b87b4 · outbound

This paper cites arXiv preprint arXiv:2510.00194 , year=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization arXiv preprint arXiv:2510.00194 , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.491356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.491356Z digest=sha256:59e8c07062b78243bcc87261e95ffd684fe42e4bbc5a2d3bef0af05a07a1b42b

Observation 9d0c09e4-a157-4e5d-8442-b0d670056806 · outbound

This paper cites OpenAI o1 System Card.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization OpenAI o1 System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.496398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.496398Z digest=sha256:00ce9e4c90128f3d9498efc7177af5612ee73f4560182025fd8c02ec404682b4

Observation 1c1d62ca-9229-4710-aed5-547a1c4786c4 · outbound

This paper cites Nature , volume=.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Nature , volume=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.501775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.501775Z digest=sha256:5b71ec40480e0a5626b585dd66bc803ca670ed5f746dd31e848cf06426f11921

Observation 8027817f-d979-4dfd-9b7e-7779049f0268 · outbound

This paper cites an unresolved cited work.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Unresolved cited work

Reference 49

Resolution
parse uncertain
no resolver link, observed 2026-08-01T01:52:02.506654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.506654Z digest=sha256:d7391bafcaed99a302d3595bea880647af2c6b8f7bf8e6639f8cb16fe1db60a6

Observation cb902165-c989-4696-a123-a16c86dcfbe4 · outbound

This paper cites Proceedings of the Thirty-Eighth.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the Thirty-Eighth

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.512266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.512266Z digest=sha256:1aa13a424fad576c949675222cb9fe31bf142fa84bd4b21ca1120ffe7c41378c

Observation 292ad71d-5ead-4706-aeb1-4d505edcaef5 · outbound

This paper cites Proceedings of the Thirty-Eighth.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Proceedings of the Thirty-Eighth

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:02.517746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:52:02.517746Z digest=sha256:dcce3330aeac84e7b417cafedf3b71a99a450f7cb0c0d264902f011910e13779

Pith citing papers

No inbound Pith citation observations are available.