Pith. sign in

REVIEW 11 cited by

Noisy Networks for Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1706.10295 v3 pith:KVWDVK2L submitted 2017-06-30 cs.LG stat.ML

classification cs.LGstat.ML
keywords agentexplorationnoisynetnoiseweightsaddedaddsadvancing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration. The parameters of the noise are learned with gradient descent along with the remaining network weights. NoisyNet is straightforward to implement and adds little computational overhead. We find that replacing the conventional exploration heuristics for A3C, DQN and dueling agents (entropy reward and $\epsilon$-greedy respectively) with NoisyNet yields substantially higher scores for a wide range of Atari games, in some cases advancing the agent from sub to super-human performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 390 citations worldwide. Full citation record

  1. Prompt-Driven Exploration

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Prompt-Driven Exploration refines language prompts from rollout videos via a VLM, enabling RL to escape zero-reward VLA and LLM policies where action-space noise fails.

  2. Skillful joint probabilistic weather forecasting from marginals

    cs.LG 2025-06 conditional novelty 7.0 of 10

    FGN, a neural weather model trained only on per-location forecast scores, produces more accurate global ensemble forecasts than GenCast and captures realistic spatial correlations.

  3. How Should We Meta-Learn Reinforcement Learning Algorithms?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A systematic comparison of black-box evolution, neural and symbolic distillation, and LLM-based proposal for meta-learning RL algorithms yields practical recommendations: warm-started LLM proposal is sample-efficient,...

  4. Meta-learning how to Share Credit among Macro-Actions

    cs.LG 2025-06 conditional novelty 6.0 of 10

    MASP meta-learns a similarity matrix over macro-actions and regularizes Q-values so that similar actions move together, improving exploration and performance in augmented-action-space RL.

  5. Hadamax Encoding: Elevating Performance in Model-Free Atari

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Hadamax, a Hadamard-product and max-pooling encoder, improves PQN's median human-normalized Atari-57 score by about 80% with no algorithmic changes.

  6. When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design

    physics.optics 2026-07 conditional novelty 5.5 of 10

    Dueling DQN is the only tested value-based RL variant that reliably improves seven-variable PCSEL designs under a matched 83-call FDTD budget across four seeds.

  7. A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Tunable energy landscapes whose thermal averages equal sigmoid, softmax, and matrix-vector products can, in principle, form the basis of a low-energy analog computer, with a superconducting double-well device as a fir...

  8. Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Bayesian deep Q-learning exhibits a cold posterior effect, caused partly by misspecified Gaussian priors, and Laplace or meta-learned priors improve performance.

  9. HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning

    cs.AI 2025-05 conditional novelty 5.0 of 10

    HAVA weights RL rewards by an agent reputation that falls when norms are violated, letting written safety rules and learned social norms be combined in one policy.

  10. Quantum Reinforcement Learning by Adaptive Non-local Observables

    quant-ph 2025-07 conditional novelty 4.0 of 10

    Adaptive non-local observables, jointly trained with variational circuit parameters, improve DQN and A3C reinforcement learning agents on simulated benchmark tasks relative to fixed Pauli-measurement baselines.

  11. Designing Adaptive Algorithms Based on Reinforcement Learning for Dynamic Optimization of Sliding Window Size in Multi-Dimensional Data Streams

    cs.LG 2025-07 reject novelty 4.0 of 10

    A DQN-based reinforcement learning agent that dynamically chooses sliding window sizes is claimed to improve classification accuracy on multi-dimensional streams, but the reported evaluation is inconsistent and not re...

Pith tools