REVIEW 5 cited by
Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN) are one of the most powerful dynamic classifiers publicly known. The network itself and the related learning algorithms are reasonably well documented to get an idea how it works. This paper will shed more light into understanding how LSTM-RNNs evolved and why they work impressively well, focusing on the early, ground-breaking publications. We significantly improved documentation and fixed a number of errors and inconsistencies that accumulated in previous publications. To support understanding we as well revised and unified the notation used.
Forward citations
Cited by 5 Pith papers
-
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
Neural inhibition gating on MoE router inputs improves a synthetic digit/squares benchmark by about four points over plain MoE, but the language-model evidence is unreliable.
-
Likelihood-Free Estimation for Spatiotemporal Hawkes processes with missing data and application to predictive policing
A WGAN with an exact Hawkes simulator as its generator estimates spatiotemporal Hawkes parameters from thinned crime data, improving hotspot prediction on simulated Bogota data.
-
Automating the Deep Space Network Data Systems; A Case Study in Adaptive Anomaly Detection through Agentic AI
An internship report integrates reconstruction-based deep learning, Q-learning, and a Mistral LLM into an agentic workflow for DSN anomaly detection, without reporting any performance metrics.
-
Joint Flow And Feature Refinement Using Attention For Video Restoration
JFFRA jointly refines optical flow and frame features in an iterative attention-based loop, reporting gains up to 1.62 dB over prior video restoration methods.
-
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
Combining knowledge distillation with l0 or low-rank pruning improves compressed RNN-T ASR, and joint pruning with fine-tuning gives 8.9% and 13.4% relative WER gains over baseline.
Discussion (0). Continue with ORCID to comment.