Pith. sign in

Paper Citation Record · LEDGER

Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2403.03950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.03950 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:32:44.345803Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68eb5b5f-310f-4f76-8ad4-8c2f50cec93c · inbound

Chronos: Learning the Language of Time Series cites this paper.

Chronos: Learning the Language of Time Series Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:27:23.487475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T08:27:23.298009Z digest=sha256:17b15d7d2ab8698b5bbed475f33693a028151851b3ea2aebfa0c1865df79ddf9

Observation 829d6796-a05e-4a0b-8d27-9a4f3b1b4acf · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:04:10.485375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:fda0cf4a65bcf7c47a65c5f15ab9aa0229b82811140019ac6b313df0656d8e90

Observation 57709a0f-590d-45ac-9b61-08887d3f6b20 · inbound

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior cites this paper.

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.424708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T02:38:26.196000Z digest=sha256:d8e0a93a8579474ca2f1bc6ed0f3ba4b12423092543f0bc51c37bbcb98217367

Observation fa4b0738-8482-4e9b-94ea-f186d525c781 · inbound

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior cites this paper.

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:55:33.154530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:53:48.604436Z digest=sha256:cb8c9183757269e2eeb4c88c2132264116a7252d054416dd60b5edb5e9d9021e

Observation 5d9a0d03-ad46-430e-a542-e611bd7328eb · inbound

D2 Actor Critic: Diffusion Actor Meets Distributional Critic cites this paper.

D2 Actor Critic: Diffusion Actor Meets Distributional Critic Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T07:35:29.400283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:31:27.330951Z digest=sha256:efc25e1626591aa0bf6e04e707cf1a8aaad70a8c82165a8d7403d441e3e35464

Observation 544e6d5b-cf72-4159-be1f-56330955559d · inbound

Value Flows cites this paper.

Value Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:28.459658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:28.459658Z digest=sha256:f44882f1b1d5a4a6e383572287a8d8c666810913500485f12b57a41242e37df4

Observation 2e14a882-24c4-442f-8c64-ba290c397f12 · inbound

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning cites this paper.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 843

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:49.758326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:49.758326Z digest=sha256:ec8d24d784cd4cf9f2623ad0b7b30b0a2d3c439abec34f83af4ed3b48d1d7f07

Observation 94690d6b-e86f-467c-b64f-00310c11c066 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:37:13.913128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:36:09.272019Z digest=sha256:f8a7a1e92b31e1ddb2c1476fe8caed3f1586b7609a0adece91b502b7930da695

Observation 94defd6d-5c96-4f7e-98bd-d0d42e4ee0a4 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:15.129472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:15.129472Z digest=sha256:7e6b748ebc1d011d251ab4e73b681d8142afd5bed08aea1374313aa6d6ac8f2c

Observation d3fc7f7f-8fe7-4ec1-a38f-88e78e34dfe7 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:36:17.834646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:f15c1874fb4323043cac73221f038d98d4336a0c22f16f66eca4722787085c1e

Observation 3bc30c35-a270-4039-aac2-5cb11aaa1904 · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:27.856199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:0bcc1f86576735632814ca636a7ddc7bdc35e98b9768054797ef96ea004c5d32

Observation b1523808-7122-45ad-97e8-60040f337875 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.858264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:8d225f339421d7128f3e0943b20ffc0272e655b79d419b37fc771e40a386c0c9

Observation d4073d5f-6457-45fb-bc1c-42240b54a772 · inbound

Hierarchical Behaviour Spaces cites this paper.

Hierarchical Behaviour Spaces Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:01:12.805850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:33:48.527621Z digest=sha256:f590fd8d261b17b9304c9323a92c1c7475e42e4c14018d86e0cd29db0e498fc1

Observation 83ddd3f9-3c3d-4670-a4d4-9280249664c5 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.701398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:b8dc85631caa2cbad953aabfb80c3620fd9ab4635a15836afd377b50ffacf131

Observation ac178a96-50c7-432e-87f1-73fda8433b49 · inbound

Survival Reinforcement Learning: Toward Scalable Self-Supervised RL cites this paper.

Survival Reinforcement Learning: Toward Scalable Self-Supervised RL Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:22:46.545028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:22:27.775787Z digest=sha256:d834885384aef0ae25c72566921cc617dbad0708e2a906916ddf83160aaa27d5

Observation 3ae4e93a-643f-4d7d-9153-0cd30820f91e · inbound

Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning cites this paper.

Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:31.233703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T18:12:00.111067Z digest=sha256:8ac8e174b912f1a46ff466f466bac3f811fe47541a65806a6cc43fcae66d6f77

Observation 5a0a0036-5fe6-4d6a-86ca-3f9a6224d9b1 · inbound

Superhuman AI for Generals.io Using Self-Play Reinforcement Learning cites this paper.

Superhuman AI for Generals.io Using Self-Play Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:45.477702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:03:49.551408Z digest=sha256:a2a62d56eae278924e5e702661efecf42ab37ec64ef72a8159718fe98ded0a58

Observation e00edf4c-2546-4721-8b4c-d7689ed94a6a · inbound

World Value Models for Robotic Manipulation cites this paper.

World Value Models for Robotic Manipulation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:40:00.236783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T23:33:50.854261Z digest=sha256:28fbbbf40edb457feee01b58cf3737c7f7d086563d977dcf4c96b6dd65e9e0ab

Observation 35f039f2-de60-4110-8433-5dad67dea679 · inbound

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation cites this paper.

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:03:05.132513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T03:55:47.911210Z digest=sha256:824c4bc36cabc7d7f04616cbea4b519e59f1e101fe3194e0ce83111308c08d28

Observation 6640e43d-9c1e-471c-b3fb-589ad26aba89 · inbound

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation cites this paper.

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T11:25:18.198611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:25:18.198611Z digest=sha256:1e510ff986119b48c001139c4ba018fc5afb2a2f014c290429d9d3cca3a6a690

Observation 55b73d7d-5c57-4140-a3a8-1e9cf64de5e8 · inbound

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation cites this paper.

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T07:17:21.581451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:17:21.581451Z digest=sha256:0a0bc76bcdfc175431cb2ad2d10558bc33f0a5c2b0207568e1bab727241d7ac0

Observation 703baa5e-1d4f-4cce-8642-aff5faab71a4 · inbound

Relative Value Learning cites this paper.

Relative Value Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:32:00.021033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:32:00.021033Z digest=sha256:51069880e75b9520dec7da73b30f288e5981b77b0f5e0461d0e48101f92528db

Observation 7d547984-a1a0-4dc0-995e-8ca771626df5 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:44.345803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:44.345803Z digest=sha256:0d00f22e0274f4530194f0aad7d7e0fd3e78a39aaa430a1741b1c3f554ab0c83