Pith. sign in

REVIEW 5 cited by

Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.09586 v1 pith:C5L6G4HG submitted 2019-09-12 cs.NE cs.CLcs.LG

classification cs.NEcs.CLcs.LG
keywords understandingwelllongmemorynetworksneuralpublicationsrecurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN) are one of the most powerful dynamic classifiers publicly known. The network itself and the related learning algorithms are reasonably well documented to get an idea how it works. This paper will shed more light into understanding how LSTM-RNNs evolved and why they work impressively well, focusing on the early, ground-breaking publications. We significantly improved documentation and fixed a number of errors and inconsistencies that accumulated in previous publications. To support understanding we as well revised and unified the notation used.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neural Inhibition Improves Dynamic Routing and Mixture of Experts

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Neural inhibition gating on MoE router inputs improves a synthetic digit/squares benchmark by about four points over plain MoE, but the language-model evidence is unreliable.

  2. Likelihood-Free Estimation for Spatiotemporal Hawkes processes with missing data and application to predictive policing

    cs.LG 2025-02 reject novelty 5.0 of 10

    A WGAN with an exact Hawkes simulator as its generator estimates spatiotemporal Hawkes parameters from thinned crime data, improving hotspot prediction on simulated Bogota data.

  3. Automating the Deep Space Network Data Systems; A Case Study in Adaptive Anomaly Detection through Agentic AI

    cs.LG 2025-08 reject novelty 4.0 of 10

    An internship report integrates reconstruction-based deep learning, Q-learning, and a Mistral LLM into an agentic workflow for DSN anomaly detection, without reporting any performance metrics.

  4. Joint Flow And Feature Refinement Using Attention For Video Restoration

    cs.CV 2025-05 conditional novelty 4.0 of 10

    JFFRA jointly refines optical flow and frame features in an iterative attention-based loop, reporting gains up to 1.62 dB over prior video restoration methods.

  5. Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models

    eess.AS 2025-02 conditional novelty 4.0 of 10

    Combining knowledge distillation with l0 or low-rank pruning improves compressed RNN-T ASR, and joint pruning with fine-tuning gives 8.9% and 13.4% relative WER gains over baseline.

Pith tools