Pith. sign in

REVIEW 3 cited by

LSTM: A Search Space Odyssey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1503.04069 v2 pith:TSE6BMAT submitted 2015-03-13 cs.NE cs.LG

classification cs.NEcs.LG
keywords lstmvariantsnetworksarchitecturecomponentshyperparametersrecognitionresults
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Several variants of the Long Short-Term Memory (LSTM) architecture for recurrent neural networks have been proposed since its inception in 1995. In recent years, these networks have become the state-of-the-art models for a variety of machine learning problems. This has led to a renewed interest in understanding the role and utility of various computational components of typical LSTM variants. In this paper, we present the first large-scale analysis of eight LSTM variants on three representative tasks: speech recognition, handwriting recognition, and polyphonic music modeling. The hyperparameters of all LSTM variants for each task were optimized separately using random search, and their importance was assessed using the powerful fANOVA framework. In total, we summarize the results of 5400 experimental runs ($\approx 15$ years of CPU time), which makes our study the largest of its kind on LSTM networks. Our results show that none of the variants can improve upon the standard LSTM architecture significantly, and demonstrate the forget gate and the output activation function to be its most critical components. We further observe that the studied hyperparameters are virtually independent and derive guidelines for their efficient adjustment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Models as Knowledge Bases?

    cs.CL 2019-09 accept novelty 7.0 of 10

    BERT stores relational knowledge extractable via cloze queries without fine-tuning and matches supervised baselines on open-domain QA tasks.

  2. Quantity doesn't buy quality syntax with neural language models

    cs.CL 2019-08 conditional novelty 6.0 of 10

    More training data and larger LSTM hidden layers yield diminishing returns on subject-verb agreement accuracy, and GPT and BERT sometimes score below LSTMs trained on far less data.

  3. Multimodal Techniques for Malware Classification

    cs.CR 2025-01 conditional novelty 3.0 of 10

    Stacking header and section models gives a small accuracy gain over single models, but the gain lacks statistical validation and the dataset is small and imbalanced.

Pith tools