Pith. sign in

REVIEW 3 cited by

iQRL -- Implicitly Quantized Representations for Sample-efficient Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02696 v1 pith:S3HYM3XN submitted 2024-06-04 cs.LG

classification cs.LG
keywords learningrepresentationcontrollatentreinforcementcontinuousimplicitlyiqrl
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning representations for reinforcement learning (RL) has shown much promise for continuous control. We propose an efficient representation learning method using only a self-supervised latent-state consistency loss. Our approach employs an encoder and a dynamics model to map observations to latent states and predict future latent states, respectively. We achieve high performance and prevent representation collapse by quantizing the latent representation such that the rank of the representation is empirically preserved. Our method, named iQRL: implicitly Quantized Reinforcement Learning, is straightforward, compatible with any model-free RL algorithm, and demonstrates excellent performance by outperforming other recently proposed representation learning methods in continuous control benchmarks from DeepMind Control Suite.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProDVI: Programmatic Dynamics Priors for Value Network Initialization

    cs.LG 2026-08 conditional novelty 6.0 of 10

    LLM-generated dynamics programs, used only to pretrain a value network's state-action encoder, improve sample efficiency of model-free RL on continuous control tasks.

  2. Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A model-free visual RL agent that jointly trains latent self-prediction and next-observation prediction, mediated by two adapters, improves aggregate DMControl scores over prior methods.

  3. Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

    cs.LG 2025-10 conditional novelty 5.0 of 10

    Simplicial embeddings — group-wise softmax feature layers — improve sample efficiency and final performance of FastTD3, FastSAC, and PPO across continuous- and discrete-control benchmarks at no meaningful runtime cost.

Pith tools