{"work":{"id":"04430f4e-b270-479c-9dcd-bee723164789","openalex_id":"https://openalex.org/W2792764867","doi":"10.48550/arxiv.1803.01271","arxiv_id":"1803.01271","raw_key":null,"title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling","authors":null,"authors_text":"Shaojie Bai, J. Zico Kolter, Vladlen Koltun","year":2018,"venue":"cs.LG","abstract":"For most deep learning practitioners, sequence modeling is synonymous with recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should one use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. The models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple convolutional architecture outperforms canonical recurrent networks such as LSTMs across a diverse range of tasks and datasets, while demonstrating longer effective memory. We conclude that the common association between sequence modeling and recurrent networks should be reconsidered, and convolutional networks should be regarded as a natural starting point for sequence modeling tasks. To assist related work, we have made code available at http://github.com/locuslab/TCN .","external_url":"https://arxiv.org/abs/1803.01271","cited_by_count":4345,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1803.01271","created_at":"2026-05-10T01:04:50.357825+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling","render_title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling"},"hub":{"state":{"work_id":"04430f4e-b270-479c-9dcd-bee723164789","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":133,"external_cited_by_count":4345,"distinct_field_count":23,"first_pith_cited_at":"2019-06-21T04:35:14+00:00","last_pith_cited_at":"2026-07-09T09:16:36+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T01:39:18.274403+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":9},{"context_role":"baseline","n":3},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":9},{"context_polarity":"baseline","n":3},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling","claims":[{"claim_text":"For most deep learning practitioners, sequence modeling is synonymous with recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should one use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. The models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple conv","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"a hierarchical downsample-convolve-interact architecture to capture dynamic temporal dependencies at different temporal resolutions of time series data. Inspired by the idea of masked convolution [129], Wavenet [130] introduces causal convolution and dilated causal convolution to model long-range temporal causality. Similar to Wavenet, Temporal Convolutional Networks (TCN) [131] uses a stack of dilated convolutional kernels with progressively enlarged dilation factors to achieve a large receptiv","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"finance [1], power systems [2], climate science [3], among others. With increasing data availability and computational power, purely data-driven and machine learning-aided TSM has become an active research area [4]. Various machine learning architectures have been adopted in the literature for effective TSM, including Convolu- tional/Recurrent Neural Networks (CNN/RNN) [5], Trans- formers [6], and deep state-space models (SSMs) [7]. Al- though there has been a surge of Transformer-based solution","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"training, alignment, hard sync, finalize), executed on 8× NVIDIA H200 GPUs for ∼54 GPU-hours. Full optimizer and architecture hyperparameters are deferred to Appendix B. Downstream baselines.For downstream evaluation (§5.4), we compare against four standard strong baselines from the clinical-ML literature: Ridge regression, LightGBM [29], LSTM [30], and TCN [31], trained on raw clinical features. Curriculum ablation variants.For curriculum ablation (§5.2), we compare CLIN-JEPA against four train","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Visual-to-auditory Spatial processing of auditory signals, i.e. sound source lo- calization, is partially performed in the right occipital cortex, more speciﬁcally, in the dorsal and lateral ventral parts [39, 116], for CB. The inferior parietal lobule 23 seems to mitigate the auditory spatial information for more demanding computation towards the occipital areas [116]. Similarly, Collignon and colleagues [45] reported activations of the right cuneous and the right occipital gyrus in CB. These r","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"depth before they fall within the effective receptive field, and such depth can make optimization more challenging. Previous CNN architectures have explicitly manipulated inductive biases in order to address multiscale representations and shift-invariance. To capture long- range and multiscale dependencies, models often use causal dilated convolutions [15] or parallel multi-branch kernels [7,16]. Conversely, local shift-invariance can be preserved via pre-decimation low-pass filtering [17], whil","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Since LRDs are perhaps the foremost challenge for sequence models, all standard model families such as continuous-time models (CTMs), RNNs, CNNs, and Transformers include many specialized variants designed to address them. Modern examples include orthogonal and Lipschitz RNNs [ 1, 13] to combat vanishing gradients, dilated convolutions to increase context size [ 3, 28], and an increasingly vast family of eﬃcient Transformers that reduce the quadratic dependence on sequence length [ 8, 22]. Despi","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":3,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-07-01T17:01:26.792299+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"18ec2b92-75b6-4979-a90c-74e6200070ad","orcid":null,"display_name":"Shaojie Bai"},{"id":"af830e8c-1c58-4458-915c-56d208aed1b5","orcid":null,"display_name":"J. Zico Kolter"},{"id":"d731164d-961a-4471-a43a-a362bc515a4d","orcid":null,"display_name":"Vladlen Koltun"}]},"error":null,"updated_at":"2026-07-01T17:01:27.126691+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:59:53.666769+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"A Time Series is Worth 64 Words: Long-term Forecasting with Transformers","work_id":"d6d0a3ac-d695-4de0-ba2d-4e1d31ac8359","shared_citers":5},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":4},{"title":"Hochreiter and J","work_id":"c3b0bfa7-6764-45f1-a40d-45baaee9d22c","shared_citers":4},{"title":"A., Lines, J., Flynn, M., Large, J., Bostrom, A","work_id":"79852b21-480f-48b1-9483-33c25c5001ab","shared_citers":3},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":3},{"title":"WaveNet: A Generative Model for Raw Audio","work_id":"05682736-8137-4a5f-a5b6-1127e99b0041","shared_citers":3},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":2},{"title":"Are Transformers Effective for Time Series Forecasting?Proceedings of the AAAI Conference on Artificial Intelligence, 37(9): 11121–11128","work_id":"a543feff-80bc-42c9-91c4-7a95410ba01e","shared_citers":2},{"title":"arXiv preprint arXiv:2405.14616 , year=","work_id":"7045d4c0-99f6-4e00-9067-99a0e5b2247e","shared_citers":2},{"title":"Etsformer: Exponential smoothing transformers for time-series forecasting","work_id":"d6e7f964-94a6-432e-9ec3-7b27b3a841a3","shared_citers":2},{"title":"Gaussian Error Linear Units (GELUs)","work_id":"0466fd22-03a1-4a61-af0a-a900e77bb023","shared_citers":2},{"title":"Graph wavenet for deep spatial-temporal graph model- ing.arXiv preprint arXiv:1906.00121","work_id":"adb5a4a7-33f0-446b-99e0-025d57226403","shared_citers":2},{"title":"In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp","work_id":"1c2af073-b162-4cd0-a5cd-d8aade23b9b5","shared_citers":2},{"title":"iTransformer: Inverted Transformers Are Effective for Time Series Forecasting","work_id":"c1ea0b21-a74c-4315-a48c-4ef730cec588","shared_citers":2},{"title":"Neural Machine Translation by Jointly Learning to Align and Translate","work_id":"d831e763-d530-4029-a65c-ac595d82cb2a","shared_citers":2},{"title":"On the Opportunities and Risks of Foundation Models","work_id":"a18039e9-928d-47c9-a836-32656a71bf71","shared_citers":2},{"title":"The Kinetics Human Action Video Dataset","work_id":"c8a3de61-cfd3-4aeb-bcf7-a0372c015748","shared_citers":2},{"title":"","work_id":"0ceeea7a-0aaf-49c3-ad94-fa0a2a466cfd","shared_citers":1},{"title":"11 Published as a conference paper at ICLR 2023 Sana Tonekaboni, Danny Eytan, and Anna Goldenberg","work_id":"d8c8f3df-f2bd-4b9f-b4c9-837f46c1b27a","shared_citers":1},{"title":"18 Published as a conference paper at ICLR 2023 Models PatchTST DLinear FEDformer Autoformer InformerFine-tuning Lin","work_id":"b5255845-db05-48e4-8a1d-5bf96772e934","shared_citers":1},{"title":"2001 , issn =","work_id":"ad5a46e1-a70d-4090-9f16-57d9d5b7e9e4","shared_citers":1},{"title":"2007 , issn =","work_id":"da12b3e1-15ba-48fe-86a1-26cb0eae540d","shared_citers":1},{"title":"2010, Astronomy & Astrophysics, 511, L2, doi: 10.1051/0004-6361/200913456","work_id":"9fe7f1e0-ab87-44ec-b266-f2d505ae30aa","shared_citers":1},{"title":"2012, Solar Physics, 279, 317, doi: 10.1007/s11207-012-0037-5","work_id":"1a1d2b0a-7875-4d98-b273-2087f3da484a","shared_citers":1}],"time_series":[{"n":1,"year":2021},{"n":1,"year":2022},{"n":1,"year":2023},{"n":39,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:59:45.540571+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:00:11.362829+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling","claims":[{"claim_text":"For most deep learning practitioners, sequence modeling is synonymous with recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should one use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. The models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple conv","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"a hierarchical downsample-convolve-interact architecture to capture dynamic temporal dependencies at different temporal resolutions of time series data. Inspired by the idea of masked convolution [129], Wavenet [130] introduces causal convolution and dilated causal convolution to model long-range temporal causality. Similar to Wavenet, Temporal Convolutional Networks (TCN) [131] uses a stack of dilated convolutional kernels with progressively enlarged dilation factors to achieve a large receptiv","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"finance [1], power systems [2], climate science [3], among others. With increasing data availability and computational power, purely data-driven and machine learning-aided TSM has become an active research area [4]. Various machine learning architectures have been adopted in the literature for effective TSM, including Convolu- tional/Recurrent Neural Networks (CNN/RNN) [5], Trans- formers [6], and deep state-space models (SSMs) [7]. Al- though there has been a surge of Transformer-based solution","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"training, alignment, hard sync, finalize), executed on 8× NVIDIA H200 GPUs for ∼54 GPU-hours. Full optimizer and architecture hyperparameters are deferred to Appendix B. Downstream baselines.For downstream evaluation (§5.4), we compare against four standard strong baselines from the clinical-ML literature: Ridge regression, LightGBM [29], LSTM [30], and TCN [31], trained on raw clinical features. Curriculum ablation variants.For curriculum ablation (§5.2), we compare CLIN-JEPA against four train","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Visual-to-auditory Spatial processing of auditory signals, i.e. sound source lo- calization, is partially performed in the right occipital cortex, more speciﬁcally, in the dorsal and lateral ventral parts [39, 116], for CB. The inferior parietal lobule 23 seems to mitigate the auditory spatial information for more demanding computation towards the occipital areas [116]. Similarly, Collignon and colleagues [45] reported activations of the right cuneous and the right occipital gyrus in CB. These r","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"depth before they fall within the effective receptive field, and such depth can make optimization more challenging. Previous CNN architectures have explicitly manipulated inductive biases in order to address multiscale representations and shift-invariance. To capture long- range and multiscale dependencies, models often use causal dilated convolutions [15] or parallel multi-branch kernels [7,16]. Conversely, local shift-invariance can be preserved via pre-decimation low-pass filtering [17], whil","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Since LRDs are perhaps the foremost challenge for sequence models, all standard model families such as continuous-time models (CTMs), RNNs, CNNs, and Transformers include many specialized variants designed to address them. Modern examples include orthogonal and Lipschitz RNNs [ 1, 13] to combat vanishing gradients, dilated convolutions to increase context size [ 3, 28], and an increasingly vast family of eﬃcient Transformers that reduce the quadratic dependence on sequence length [ 8, 22]. Despi","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":3,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-07-01T17:01:27.129834+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling","claims":[{"claim_text":"For most deep learning practitioners, sequence modeling is synonymous with recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should one use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. The models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple conv","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:00:06.867178+00:00"}},"summary":{"title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling","claims":[{"claim_text":"For most deep learning practitioners, sequence modeling is synonymous with recurrent networks. Yet recent results indicate that convolutional architectures can outperform recurrent networks on tasks such as audio synthesis and machine translation. Given a new sequence modeling task or dataset, which architecture should one use? We conduct a systematic evaluation of generic convolutional and recurrent architectures for sequence modeling. The models are evaluated across a broad range of standard tasks that are commonly used to benchmark recurrent networks. Our results indicate that a simple conv","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"A Time Series is Worth 64 Words: Long-term Forecasting with Transformers","work_id":"d6d0a3ac-d695-4de0-ba2d-4e1d31ac8359","shared_citers":5},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":4},{"title":"Hochreiter and J","work_id":"c3b0bfa7-6764-45f1-a40d-45baaee9d22c","shared_citers":4},{"title":"A., Lines, J., Flynn, M., Large, J., Bostrom, A","work_id":"79852b21-480f-48b1-9483-33c25c5001ab","shared_citers":3},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":3},{"title":"WaveNet: A Generative Model for Raw Audio","work_id":"05682736-8137-4a5f-a5b6-1127e99b0041","shared_citers":3},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":2},{"title":"Are Transformers Effective for Time Series Forecasting?Proceedings of the AAAI Conference on Artificial Intelligence, 37(9): 11121–11128","work_id":"a543feff-80bc-42c9-91c4-7a95410ba01e","shared_citers":2},{"title":"arXiv preprint arXiv:2405.14616 , year=","work_id":"7045d4c0-99f6-4e00-9067-99a0e5b2247e","shared_citers":2},{"title":"Etsformer: Exponential smoothing transformers for time-series forecasting","work_id":"d6e7f964-94a6-432e-9ec3-7b27b3a841a3","shared_citers":2},{"title":"Gaussian Error Linear Units (GELUs)","work_id":"0466fd22-03a1-4a61-af0a-a900e77bb023","shared_citers":2},{"title":"Graph wavenet for deep spatial-temporal graph model- ing.arXiv preprint arXiv:1906.00121","work_id":"adb5a4a7-33f0-446b-99e0-025d57226403","shared_citers":2},{"title":"In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp","work_id":"1c2af073-b162-4cd0-a5cd-d8aade23b9b5","shared_citers":2},{"title":"iTransformer: Inverted Transformers Are Effective for Time Series Forecasting","work_id":"c1ea0b21-a74c-4315-a48c-4ef730cec588","shared_citers":2},{"title":"Neural Machine Translation by Jointly Learning to Align and Translate","work_id":"d831e763-d530-4029-a65c-ac595d82cb2a","shared_citers":2},{"title":"On the Opportunities and Risks of Foundation Models","work_id":"a18039e9-928d-47c9-a836-32656a71bf71","shared_citers":2},{"title":"The Kinetics Human Action Video Dataset","work_id":"c8a3de61-cfd3-4aeb-bcf7-a0372c015748","shared_citers":2},{"title":"","work_id":"0ceeea7a-0aaf-49c3-ad94-fa0a2a466cfd","shared_citers":1},{"title":"11 Published as a conference paper at ICLR 2023 Sana Tonekaboni, Danny Eytan, and Anna Goldenberg","work_id":"d8c8f3df-f2bd-4b9f-b4c9-837f46c1b27a","shared_citers":1},{"title":"18 Published as a conference paper at ICLR 2023 Models PatchTST DLinear FEDformer Autoformer InformerFine-tuning Lin","work_id":"b5255845-db05-48e4-8a1d-5bf96772e934","shared_citers":1},{"title":"2001 , issn =","work_id":"ad5a46e1-a70d-4090-9f16-57d9d5b7e9e4","shared_citers":1},{"title":"2007 , issn =","work_id":"da12b3e1-15ba-48fe-86a1-26cb0eae540d","shared_citers":1},{"title":"2010, Astronomy & Astrophysics, 511, L2, doi: 10.1051/0004-6361/200913456","work_id":"9fe7f1e0-ab87-44ec-b266-f2d505ae30aa","shared_citers":1},{"title":"2012, Solar Physics, 279, 317, doi: 10.1007/s11207-012-0037-5","work_id":"1a1d2b0a-7875-4d98-b273-2087f3da484a","shared_citers":1}],"time_series":[{"n":1,"year":2021},{"n":1,"year":2022},{"n":1,"year":2023},{"n":39,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"af830e8c-1c58-4458-915c-56d208aed1b5","orcid":null,"display_name":"J. Zico Kolter","source":"manual","import_confidence":0.72},{"id":"18ec2b92-75b6-4979-a90c-74e6200070ad","orcid":null,"display_name":"Shaojie Bai","source":"manual","import_confidence":0.72},{"id":"d731164d-961a-4471-a43a-a362bc515a4d","orcid":null,"display_name":"Vladlen Koltun","source":"manual","import_confidence":0.72}]}}