citation dossier

Fast and accurate deep network learning by exponential linear units (elus)

Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015 · arXiv 1511.07289

20Pith papers citing it

20reference links

cs.LGtop field · 7 papers

UNVERDICTEDtop verdict bucket · 14 papers

This arXiv-backed work is queued for full Pith review when it crosses the high-inbound sweep. That review runs reader · skeptic · desk-editor · referee · rebuttal · circularity · lean confirmation · RS check · pith extraction.

read on arXiv PDF

why this work matters in Pith

Pith has found this work in 20 reviewed papers. Its strongest current cluster is cs.LG (7 papers). The largest review-status bucket among citing papers is UNVERDICTED (14 papers). For highly cited works, this page shows a dossier first and a bounded explorer second; it never tries to render every citing paper at once.

representative citing papers

Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their Gradients

cs.LG · 2026-05-03 · unverdicted · novelty 8.0

Floating-point neural networks with automatic differentiation can represent arbitrary floating-point functions and their gradients under mild conditions.

Building Normalizing Flows with Stochastic Interpolants

cs.LG · 2022-09-30 · conditional · novelty 8.0

Normalizing flows are constructed by learning the velocity of a stochastic interpolant via a quadratic loss derived from its probability current, yielding an efficient ODE-based alternative to diffusion models.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling

cs.SD · 2026-05-11 · unverdicted · novelty 7.0

AffectCodec is an emotion-guided neural speech codec that preserves emotional cues during quantization while maintaining semantic fidelity and prosodic naturalness.

GravityGraphSAGE: Link Prediction in Directed Attributed Graphs

cs.LG · 2026-05-10 · unverdicted · novelty 7.0

GravityGraphSAGE adapts GraphSAGE with a gravity-inspired decoder to outperform prior graph deep learning methods on directed link prediction across citation networks and 16 real-world graphs.

ReActor: Reinforcement Learning for Physics-Aware Motion Retargeting

cs.RO · 2026-05-07 · unverdicted · novelty 7.0

ReActor jointly optimizes motion retargeting and RL policy training with an approximate gradient to generate physically consistent robot motions from human references using only sparse body correspondences.

Neuro-Symbolic ODE Discovery with Latent Grammar Flow

cs.LG · 2026-04-17 · unverdicted · novelty 7.0

Latent Grammar Flow discovers ODEs by placing grammar-based equation representations in a discrete latent space, using a behavioral loss to cluster similar equations, and sampling via a discrete flow model guided by data fit and constraints.

High Fidelity Neural Audio Compression

eess.AS · 2022-10-24 · accept · novelty 7.0

EnCodec is an end-to-end trained streaming neural audio codec that uses a single multiscale spectrogram discriminator and a gradient-normalizing loss balancer to achieve higher fidelity than prior methods at the same bitrates for 24 kHz mono and 48 kHz stereo audio.

Rethinking Attention with Performers

cs.LG · 2020-09-30 · unverdicted · novelty 7.0

Performers approximate full-rank softmax attention in Transformers via FAVOR+ random features for linear complexity, with theoretical guarantees of unbiased estimation and competitive results on pixel, text, and protein tasks.

Dream to Control: Learning Behaviors by Latent Imagination

cs.LG · 2019-12-03 · accept · novelty 7.0

Dreamer learns to control from images by imagining and optimizing behaviors in a learned latent world model, outperforming prior methods on 20 visual tasks in data efficiency and final performance.

Searching for Activation Functions

cs.NE · 2017-10-16 · conditional · novelty 7.0

Automated search discovers Swish activation f(x) = x * sigmoid(βx) that improves top-1 ImageNet accuracy over ReLU by 0.9% on Mobile NASNet-A and 0.6% on Inception-ResNet-v2.

Wide Residual Networks

cs.CV · 2016-05-23 · accept · novelty 7.0

Wide residual networks achieve higher accuracy and faster training than very deep thin residual networks by increasing width and decreasing depth, setting new state-of-the-art results on CIFAR, SVHN, and ImageNet.

Training continuously-coupled reconfigurable photonic chips with quantum machine learning

quant-ph · 2026-05-11 · unverdicted · novelty 6.0

A black-box machine learning technique trains continuously-coupled photonic waveguide arrays to implement target unitaries using limited single- and two-photon measurements without requiring detailed internal models.

Geometric Monomial (GEM): a family of rational 2N-differentiable activation functions

cs.LG · 2026-04-23 · unverdicted · novelty 6.0

GEM is a new family of C^{2N}-smooth rational activation functions with variants that achieve performance on par with or exceeding GELU on ResNet, GPT-2, and BERT benchmarks.

Fast neural network surrogate for multimodal effective-one-body gravitational waveforms from generically precessing compact binaries

gr-qc · 2026-04-15 · unverdicted · novelty 6.0

Neural network surrogate approximates precessing compact binary gravitational waveforms up to 1000x faster than the base EOB model with validated accuracy.

Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery

cs.CV · 2026-05-05 · unverdicted · novelty 5.0

LAGCD inserts residual linear adapters into each ViT block plus a distribution alignment loss to improve generalized category discovery by increasing model flexibility while reducing bias between seen and novel classes.

Universal Smoothness via Bernstein Polynomials: A Constructive Approximation Approach for Activation Functions

cs.AI · 2026-05-04 · unverdicted · novelty 5.0

BerLU constructs a C1-differentiable activation with Lipschitz constant 1 via Bernstein polynomial approximation, showing better performance and efficiency than baselines on image classification with ViTs and CNNs.

Constraints on the baryon density from fast radio bursts using a non-parametric reconstruction of the Hubble parameter

astro-ph.CO · 2026-05-03 · conditional · novelty 5.0

FRB dispersion measures combined with non-parametric H(z) reconstruction yield Ω_b h² = 0.02236 ± 0.00090, agreeing with BBN and Planck CMB to within 0.05%.

A sound-horizon-free measurement of the Hubble constant from DESI DR2 baryon acoustic oscillations using artificial neural networks

astro-ph.CO · 2026-04-27 · unverdicted · novelty 5.0

Neural network reconstruction of DESI DR2 BAO, SNe Ia, and cosmic chronometer data gives H0 = 71.5 ± 2.2 km s^{-1} Mpc^{-1} without sound horizon input.

Testing $\Lambda$CDM with ANN-Reconstructed Expansion History from Cosmic Chronometers

astro-ph.CO · 2026-04-24 · unverdicted · novelty 4.0

The ANN-reconstructed Hubble parameter H(z) from cosmic chronometers aligns with Lambda CDM predictions within uncertainties.

Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input

cs.RO · 2026-04-21 · unverdicted · novelty 4.0

Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.

citing papers explorer

Showing 20 of 20 citing papers.

Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their Gradients cs.LG · 2026-05-03 · unverdicted · none · ref 36
Floating-point neural networks with automatic differentiation can represent arbitrary floating-point functions and their gradients under mild conditions.
Building Normalizing Flows with Stochastic Interpolants cs.LG · 2022-09-30 · conditional · none · ref 9
Normalizing flows are constructed by learning the velocity of a stochastic interpolant via a quadratic loss derived from its probability current, yielding an efficient ODE-based alternative to diffusion models.
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cs.SD · 2026-05-11 · unverdicted · none · ref 68
AffectCodec is an emotion-guided neural speech codec that preserves emotional cues during quantization while maintaining semantic fidelity and prosodic naturalness.
GravityGraphSAGE: Link Prediction in Directed Attributed Graphs cs.LG · 2026-05-10 · unverdicted · none · ref 23
GravityGraphSAGE adapts GraphSAGE with a gravity-inspired decoder to outperform prior graph deep learning methods on directed link prediction across citation networks and 16 real-world graphs.
ReActor: Reinforcement Learning for Physics-Aware Motion Retargeting cs.RO · 2026-05-07 · unverdicted · none · ref 46
ReActor jointly optimizes motion retargeting and RL policy training with an approximate gradient to generate physically consistent robot motions from human references using only sparse body correspondences.
Neuro-Symbolic ODE Discovery with Latent Grammar Flow cs.LG · 2026-04-17 · unverdicted · none · ref 43
Latent Grammar Flow discovers ODEs by placing grammar-based equation representations in a discrete latent space, using a behavioral loss to cluster similar equations, and sampling via a discrete flow model guided by data fit and constraints.
High Fidelity Neural Audio Compression eess.AS · 2022-10-24 · accept · none · ref 7
EnCodec is an end-to-end trained streaming neural audio codec that uses a single multiscale spectrogram discriminator and a gradient-normalizing loss balancer to achieve higher fidelity than prior methods at the same bitrates for 24 kHz mono and 48 kHz stereo audio.
Rethinking Attention with Performers cs.LG · 2020-09-30 · unverdicted · none · ref 118
Performers approximate full-rank softmax attention in Transformers via FAVOR+ random features for linear complexity, with theoretical guarantees of unbiased estimation and competitive results on pixel, text, and protein tasks.
Dream to Control: Learning Behaviors by Latent Imagination cs.LG · 2019-12-03 · accept · none · ref 9
Dreamer learns to control from images by imagining and optimizing behaviors in a learned latent world model, outperforming prior methods on 20 visual tasks in data efficiency and final performance.
Searching for Activation Functions cs.NE · 2017-10-16 · conditional · none · ref 3
Automated search discovers Swish activation f(x) = x * sigmoid(βx) that improves top-1 ImageNet accuracy over ReLU by 0.9% on Mobile NASNet-A and 0.6% on Inception-ResNet-v2.
Wide Residual Networks cs.CV · 2016-05-23 · accept · none · ref 5
Wide residual networks achieve higher accuracy and faster training than very deep thin residual networks by increasing width and decreasing depth, setting new state-of-the-art results on CIFAR, SVHN, and ImageNet.
Training continuously-coupled reconfigurable photonic chips with quantum machine learning quant-ph · 2026-05-11 · unverdicted · none · ref 61
A black-box machine learning technique trains continuously-coupled photonic waveguide arrays to implement target unitaries using limited single- and two-photon measurements without requiring detailed internal models.
Geometric Monomial (GEM): a family of rational 2N-differentiable activation functions cs.LG · 2026-04-23 · unverdicted · none · ref 3
GEM is a new family of C^{2N}-smooth rational activation functions with variants that achieve performance on par with or exceeding GELU on ResNet, GPT-2, and BERT benchmarks.
Fast neural network surrogate for multimodal effective-one-body gravitational waveforms from generically precessing compact binaries gr-qc · 2026-04-15 · unverdicted · none · ref 92
Neural network surrogate approximates precessing compact binary gravitational waveforms up to 1000x faster than the base EOB model with validated accuracy.
Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery cs.CV · 2026-05-05 · unverdicted · none · ref 39
LAGCD inserts residual linear adapters into each ViT block plus a distribution alignment loss to improve generalized category discovery by increasing model flexibility while reducing bias between seen and novel classes.
Universal Smoothness via Bernstein Polynomials: A Constructive Approximation Approach for Activation Functions cs.AI · 2026-05-04 · unverdicted · none · ref 17
BerLU constructs a C1-differentiable activation with Lipschitz constant 1 via Bernstein polynomial approximation, showing better performance and efficiency than baselines on image classification with ViTs and CNNs.
Constraints on the baryon density from fast radio bursts using a non-parametric reconstruction of the Hubble parameter astro-ph.CO · 2026-05-03 · conditional · none · ref 25
FRB dispersion measures combined with non-parametric H(z) reconstruction yield Ω_b h² = 0.02236 ± 0.00090, agreeing with BBN and Planck CMB to within 0.05%.
A sound-horizon-free measurement of the Hubble constant from DESI DR2 baryon acoustic oscillations using artificial neural networks astro-ph.CO · 2026-04-27 · unverdicted · none · ref 34
Neural network reconstruction of DESI DR2 BAO, SNe Ia, and cosmic chronometer data gives H0 = 71.5 ± 2.2 km s^{-1} Mpc^{-1} without sound horizon input.
Testing $\Lambda$CDM with ANN-Reconstructed Expansion History from Cosmic Chronometers astro-ph.CO · 2026-04-24 · unverdicted · none · ref 64
The ANN-reconstructed Hubble parameter H(z) from cosmic chronometers aligns with Lambda CDM predictions within uncertainties.
Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input cs.RO · 2026-04-21 · unverdicted · none · ref 34
Sparsely gated MoE policies double the success rate of a real Unitree Go2 quadruped on large-obstacle parkour versus matched-active-parameter MLP baselines while cutting inference time compared with a scaled-up MLP.

Fast and accurate deep network learning by exponential linear units (elus)

why this work matters in Pith

fields

years

verdicts

representative citing papers

citing papers explorer