{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2020:ZYXQIT4PR5T462S7RCS5YTGJFD","short_pith_number":"pith:ZYXQIT4P","schema_version":"1.0","canonical_sha256":"ce2f044f8f8f67cf6a5f88a5dc4cc928f5270ec52565eda1708830b8ceb67ec0","source":{"kind":"arxiv","id":"2006.05990","version":1},"attestation_state":"computed","paper":{"title":"What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Anton Raichuk, L\\'eonard Hussenot, Manu Orsini, Marcin Andrychowicz, Marcin Michalski, Matthieu Geist, Olivier Bachem, Olivier Pietquin, Piotr Sta\\'nczyk, Raphael Marinier, Sertan Girgin, Sylvain Gelly","submitted_at":"2020-06-10T17:59:03Z","abstract_excerpt":"In recent years, on-policy reinforcement learning (RL) has been successfully applied to many different continuous control tasks. While RL algorithms are often conceptually simple, their state-of-the-art implementations take numerous low- and high-level design decisions that strongly affect the performance of the resulting agents. Those choices are usually not extensively discussed in the literature, leading to discrepancy between published descriptions of algorithms and their implementations. This makes it hard to attribute progress in RL and slows down overall progress [Engstrom'20]. As a ste"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2006.05990","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2020-06-10T17:59:03Z","cross_cats_sorted":["stat.ML"],"title_canon_sha256":"4af9e13ef012a5295aab85ec338c0dc010f1dd620c2fd9cb26ed743690db76c5","abstract_canon_sha256":"2311d2c1f9c7b47d9516703ff319e0cba2bac1424021cb5bd6c033eb91d65b58"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T01:09:18.729891Z","signature_b64":"xoXVFn02S/62uPV/Xss7dUy7n4oZl9Az1JC4BlAkNSfqC6BkjRqhYrzmjjMtTrYZg0zbvA1cZWP+hO7wgn4BAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"ce2f044f8f8f67cf6a5f88a5dc4cc928f5270ec52565eda1708830b8ceb67ec0","last_reissued_at":"2026-07-05T01:09:18.729408Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T01:09:18.729408Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Anton Raichuk, L\\'eonard Hussenot, Manu Orsini, Marcin Andrychowicz, Marcin Michalski, Matthieu Geist, Olivier Bachem, Olivier Pietquin, Piotr Sta\\'nczyk, Raphael Marinier, Sertan Girgin, Sylvain Gelly","submitted_at":"2020-06-10T17:59:03Z","abstract_excerpt":"In recent years, on-policy reinforcement learning (RL) has been successfully applied to many different continuous control tasks. While RL algorithms are often conceptually simple, their state-of-the-art implementations take numerous low- and high-level design decisions that strongly affect the performance of the resulting agents. Those choices are usually not extensively discussed in the literature, leading to discrepancy between published descriptions of algorithms and their implementations. This makes it hard to attribute progress in RL and slows down overall progress [Engstrom'20]. As a ste"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2006.05990","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2006.05990/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2006.05990","created_at":"2026-07-05T01:09:18.729467+00:00"},{"alias_kind":"arxiv_version","alias_value":"2006.05990v1","created_at":"2026-07-05T01:09:18.729467+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2006.05990","created_at":"2026-07-05T01:09:18.729467+00:00"},{"alias_kind":"pith_short_12","alias_value":"ZYXQIT4PR5T4","created_at":"2026-07-05T01:09:18.729467+00:00"},{"alias_kind":"pith_short_16","alias_value":"ZYXQIT4PR5T462S7","created_at":"2026-07-05T01:09:18.729467+00:00"},{"alias_kind":"pith_short_8","alias_value":"ZYXQIT4P","created_at":"2026-07-05T01:09:18.729467+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":18,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.07756","citing_title":"The Importance of Encoder Choice:A Tabular-Image Study","ref_index":168,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23993","citing_title":"Learning to Trigger: Reinforcement Learning at the Large Hadron Collider","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2606.17199","citing_title":"PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08665","citing_title":"Hint Tuning: Less Data Makes Better Reasoners","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2606.23993","citing_title":"Learning to Trigger: Reinforcement Learning at the Large Hadron Collider","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30719","citing_title":"When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30719","citing_title":"When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2201.03544","citing_title":"The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18591","citing_title":"Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation","ref_index":98,"is_internal_anchor":false},{"citing_arxiv_id":"2604.26126","citing_title":"Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2108.03298","citing_title":"What Matters in Learning from Offline Human Demonstrations for Robot Manipulation","ref_index":67,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11375","citing_title":"TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11473","citing_title":"TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08665","citing_title":"Hint Tuning: Less Data Makes Better Reasoners","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2604.26126","citing_title":"Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18978","citing_title":"Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14142","citing_title":"From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2301.04104","citing_title":"Mastering Diverse Domains through World Models","ref_index":13,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD","json":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD.json","graph_json":"https://pith.science/api/pith-number/ZYXQIT4PR5T462S7RCS5YTGJFD/graph.json","events_json":"https://pith.science/api/pith-number/ZYXQIT4PR5T462S7RCS5YTGJFD/events.json","paper":"https://pith.science/paper/ZYXQIT4P"},"agent_actions":{"view_html":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD","download_json":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD.json","view_paper":"https://pith.science/paper/ZYXQIT4P","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2006.05990&json=true","fetch_graph":"https://pith.science/api/pith-number/ZYXQIT4PR5T462S7RCS5YTGJFD/graph.json","fetch_events":"https://pith.science/api/pith-number/ZYXQIT4PR5T462S7RCS5YTGJFD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD/action/storage_attestation","attest_author":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD/action/author_attestation","sign_citation":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD/action/citation_signature","submit_replication":"https://pith.science/pith/ZYXQIT4PR5T462S7RCS5YTGJFD/action/replication_record"}},"created_at":"2026-07-05T01:09:18.729467+00:00","updated_at":"2026-07-05T01:09:18.729467+00:00"}