Pith. sign in

REVIEW 3 cited by

Deep Reinforcement Learning in Parameterized Action Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1511.04143 v5 pith:DPCG3BFZ submitted 2015-11-13 cs.AI cs.LGcs.MAcs.NE

classification cs.AIcs.LGcs.MAcs.NE
keywords actiondeeplearningparameterizedcontinuousreinforcementagentbest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that deep neural networks are capable of approximating both value functions and policies in reinforcement learning domains featuring continuous state and action spaces. However, to the best of our knowledge no previous work has succeeded at using deep neural networks in structured (parameterized) continuous action spaces. To fill this gap, this paper focuses on learning within the domain of simulated RoboCup soccer, which features a small set of discrete action types, each of which is parameterized with continuous variables. The best learned agent can score goals more reliably than the 2012 RoboCup champion agent. As such, this paper represents a successful extension of deep reinforcement learning to the class of parameterized action space MDPs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hybrid TD3: Overestimation Bias Analysis and Stable Policy Optimization for Hybrid Action Space

    cs.RO 2026-03 conditional novelty 5.0 of 10

    A weighted clipped Q-learning target that averages over discrete action choices makes TD3-style training stable for hybrid-action robot manipulation.

  2. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A convolutional residual VQ-VAE action tokenizer trained on over 100x more data than prior work improves OpenVLA success rates and inference speed on several manipulation tasks.

  3. Integrated Automated Car Following and Lane-changing control based on a Parametrized Deep Q-network with Hybrid Action Space

    eess.SY 2026-07 conditional novelty 4.0 of 10

    P-DQN with a hybrid discrete-continuous action space integrates CAV car-following and lane-changing and beats MOBIL+IDM on comfort and inverse-TTC in four simulated scenarios.

Pith tools