REVIEW 4 major objections 6 minor 233 references
Survey on Deep Neural Networks in Speech and Vision Systems
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This survey's central claim is that it is one of the most comprehensive accounts to date of deep learning for vision and speech, spanning architectures, benchmark results, industrial systems, and deployment on phones and embedded devices.
desk verdict A well-organized survey with a useful mobile-deployment angle, but citation errors in the summary tables undermine its reference value as published. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the survey's organizing taxonomy: deep architectures sorted by learning paradigm (convolutional, recurrent, generative, attention-based, and machine-searched), each mapped to its signature applications and benchmark results, and then re-examined under hardware constraints. What carries the argument is the pairing of every architecture family with concrete resource costs on the deployment side, so that the survey can claim coverage along both a software axis and a hardware axis. The load-bearing instruments are the comparative tables: they assert specific, checkable numbers such as AlexNet's 17.0% top-5 error on ImageNet, a roughly seven-fold parameter reduction for MobileNets at about 1% accuracy loss, and memory cuts from 59.1 MB to 3.2 MB for a compressed speech recognizer.
What would settle it
Check the survey's benchmark tables against the original papers: if a reader can list significant 2019-era vision or speech systems, architecture families, or record results that the survey omits, or finds reported numbers misattributed, the comprehensiveness claim fails. One concrete discrepancy is already visible: Table I's ResNet row is keyed to a reference that the reference list assigns to a speech-synthesis paper, while the running text keys ResNet to a different reference entirely.
Extended reading notes
Core claim
The paper's claim, stated in its abstract and conclusion, is that it delivers one of the most comprehensive surveys of the latest developments in intelligent vision and speech systems, from both software and hardware perspectives. On the software side, it traces the architecture families—convolutional networks, deep belief networks and stacked autoencoders, variational autoencoders, generative adversarial networks, flow models, recurrent networks and LSTMs, attention mechanisms, and neural architecture search—and ties each to its state-of-the-art results in tasks such as ImageNet classification, face, action, and pose recognition, speech recognition, and image generation. On the hardware side, it catalogs the techniques that shrink these models onto resource-restricted platforms: pruning, quantization, structured matrices, low-rank compression, and efficient designs such as depthwise-separable MobileNets, together with measured memory and energy trade-offs. It then argues that these systems are already reshaping behavioral science, intelligent transportation, and precision medicine, while naming data hunger, computational cost, and black-box interpretability as the limits that remain.
Load-bearing premise
The survey's claim of comprehensiveness rests on the assumption that its informal selection of papers and the numbers it reports faithfully represent the state of the art at the time, since the paper states no search strategy, inclusion criteria, or date range that would let a reader check the coverage.
Editorial extensions
If this is right
- If the survey's coverage holds, a newcomer can use it as a first-level map of where vision and speech deep learning stood in 2019, including dominant architectures, benchmark numbers, and the main industrial systems.
- The survey's emphasis on resource-constrained deployment implies that the next wave of intelligent vision and speech systems will be defined less by raw accuracy and more by the ability to run within limited memory, battery, and compute budgets.
- Its catalog of small-footprint results—keyword spotting, mobile speech recognition, and compact CNNs with order-of-magnitude parameter reductions—implies that a growing range of vision and speech applications can run on-device rather than in the cloud.
- The emerging-applications sections argue that automated behavioral analysis, intelligent transportation, and precision medicine will be early high-impact adopters of these systems.
- The survey's stated limitations—data hunger, computational burden, and black-box opacity—imply that progress in small-data learning, 3D and 4D data handling, and hardware-software co-design will determine how widely the technology spreads.
Reading between the lines
- The survey's 2019 taxonomy ends with attention as an emerging alternative to recurrence in sequence tasks; the same pattern it documents is what later carried attention-based machinery into vision, so the survey's framing helps a reader understand that migration.
- Because the survey states no inclusion criteria, its tables are best treated as pointers to the original papers rather than audited figures; a reader who needs decision-grade numbers should verify each entry before relying on it.
- The hardware-side tension the survey documents—accurate models are too large to deploy, efficient models lose accuracy—implies that future benchmarks will increasingly report accuracy per unit of memory or energy rather than accuracy alone.
- The survey's comprehensiveness claim is time-stamped: the field moved quickly after 2019, yet the architecture-to-application structure the survey builds is portable and can be re-populated with newer results as a living map.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of deep neural network research as applied to vision and speech systems. It covers network architectures (CNNs, generative models, RNNs, attention), their use in vision and speech applications, deployment on resource-constrained mobile and embedded platforms, and emerging application areas such as behavioral science, transportation, and medicine. The paper's central claim, stated in the abstract and in Section 7, is that it is one of the most comprehensive surveys of the latest developments in intelligent vision and speech applications from both software and hardware perspectives.
Significance. If accurate, this survey would provide a broad and useful entry point for researchers and practitioners, covering architectures, applications, hardware constraints, and emerging frontiers in a single document. The organization is logical and the scope is ambitious. However, the paper's value as an authoritative reference is currently undermined by several concrete citation errors in its central summary tables and in the main text. These errors are not peripheral: they occur in exactly the tables and sentences a reader would rely on to compare state-of-the-art results. The paper also provides no search strategy or inclusion criteria to support its comprehensiveness claim, which further weakens confidence in its coverage. With a systematic reference audit and methodological clarification, the survey could serve its intended purpose, but in its present form it cannot be recommended as a reliable reference.
major comments (4)
- [Table I] Table I, ResNet row: the table attributes 'ResNet [71] – Microsoft 2015' with a 4.70% top-5 error, but reference [71] is Prenger et al., 'Waveglow: A flow-based generative network for speech synthesis,' not a ResNet paper. The text (Section 3.1) instead cites ResNet as [97], but [97] is Wu et al., 'Wider or deeper: Revisiting the resnet model for visual recognition,' which is not the original He et al. ResNet paper either. A reader cannot verify the claimed state-of-the-art error rate from the cited source, and this error appears in a core summary table that the survey asks readers to trust.
- [Table III] Table III, 'Autoencoder/DBN [128]' row: reference [128] is Sainath et al., 'Deep convolutional neural networks for LVCSR' (ICASSP 2013), a paper about deep convolutional networks for speech recognition, not about autoencoders or deep belief networks. The row also lists a 15.5% word error rate on English Broadcast News while the cited paper's experiments concern a different setup. This misattribution in a table that summarizes state-of-the-art speech models materially weakens the survey's reliability.
- [Section 3.2] Section 3.2, paragraph on LSTM training: the sentence 'the development of long short-term memory (LSTM) networks that use special hidden units known as "gates" to retain memory over longer portions of a sequence' is cited to [40], which is Alam et al., 'Novel hierarchical Cellular Simultaneous Recurrent Neural Network for object detection' (IJCNN 2015), an object-detection paper with no LSTM gating content. The same paragraph later attributes statements about memory-network stories to [99] (Huang et al., DenseNet) and about translation difficulties to [101] (Tan and Le, EfficientNet), neither of which discusses those topics. These citation mismatches in the main text suggest a systematic reference-integrity problem.
- [Sections 1 and 7] The abstract and Sections 1 and 7 claim the paper is 'one of the most comprehensive surveys' of vision and speech deep learning, but the paper provides no search strategy, inclusion criteria, database list, or date range for the literature covered. The selection appears to be informal, and the citation errors documented above indicate that the covered literature is not always accurately represented. Without a stated methodology, the comprehensiveness claim cannot be independently checked, and the current reference errors further undermine it.
minor comments (6)
- [Table II] Table II: 'Tomson et al.' should be 'Tompson et al.'
- [Section 3.3] Section 3.3: the CIFAR-10 description says 'there are 10 classes with 60,000 images each'; CIFAR-10 has 60,000 images in total (6,000 per class), so this needs rewording.
- [Section 3.3] Section 3.3: 'TIMT' is a typo for TIMIT.
- [Section 3.1] Section 3.1: 'Neural Architecture Search (described in Section 2.8)' should refer to Section 2.9, which is the Neural Architecture Search section.
- [Table III footnote] Table III footnote: 'PERPEPLEXITY' is a typo for 'PERPLEXITY.'
- [Figures 3, 4, and 6] Figures 3, 4, and 6 are difficult to read in the PDF; higher-resolution images would improve clarity.
Circularity Check
No circularity: this is a literature survey with no fitted parameters, no derivations, and no predictive claims; its claims summarize external published work rather than reducing to the paper's own inputs.
full rationale
The paper is a narrative review of deep learning architectures, applications, and hardware constraints for vision and speech. It contains no mathematical derivation chain, no fitted parameters, no experimental predictions, and no uniqueness theorem imported from prior work. The central claim, that the paper is 'one of the most comprehensive surveys,' is a scope and quality assertion about literature coverage, not a result derived from its own definitions or equations. A few references are self-citations by the authors, but none is load-bearing: for example, reference [40] is the authors' own object-detection paper cited as background for LSTM gating, which is a descriptive citation rather than an argument that depends on that self-citation for its force. The skeptic's concern about mis-attributed references in Tables I and III is a correctness and reliability issue about the survey's representation of the external literature, not a circularity issue: the survey's content is still drawn from outside sources and is not equivalent by construction to its inputs. Therefore, the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The cited papers are accurately represented and their reported performance numbers are correctly transcribed.
- ad hoc to paper The unsystematic selection of papers is representative of the state of the art in vision and speech deep learning as of 2019.
- standard math Background knowledge of neural network concepts is assumed.
Cite this review
Pith. "Pith review of Survey on Deep Neural Networks in Speech and Vision Systems." pith.science (2026). https://pith.science/paper/5XVTLN7L
@misc{pith2026190807656,
author = {Pith},
title = {Pith review of: Survey on Deep Neural Networks in Speech and Vision Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/5XVTLN7L}},
note = {Machine review of arXiv:1908.07656}
}
read the original abstract
This survey presents a review of state-of-the-art deep neural network architectures, algorithms, and systems in vision and speech applications. Recent advances in deep artificial neural network algorithms and architectures have spurred rapid innovation and development of intelligent vision and speech systems. With availability of vast amounts of sensor data and cloud computing for processing and training of deep neural networks, and with increased sophistication in mobile and embedded technology, the next-generation intelligent systems are poised to revolutionize personal and commercial computing. This survey begins by providing background and evolution of some of the most successful deep learning models for intelligent vision and speech systems to date. An overview of large-scale industrial research and development efforts is provided to emphasize future trends and prospects of intelligent vision and speech systems. Robust and efficient intelligent systems demand low-latency and high fidelity in resource-constrained hardware platforms such as mobile devices, robots, and automobiles. Therefore, this survey also provides a summary of key challenges and recent successes in running deep neural networks on hardware-restricted platforms, i.e. within limited memory, battery life, and processing capabilities. Finally, emerging applications of vision and speech across disciplines such as affective computing, intelligent transportation, and precision medicine are discussed. To our knowledge, this paper provides one of the most comprehensive surveys on the latest developments in intelligent vision and speech applications from the perspectives of both software and hardware systems. Many of these emerging technologies using deep neural networks show tremendous promise to revolutionize research and development for future vision and speech systems.
Figures
Reference graph
Works this paper leans on
-
[71]
Waveglow: A flow-based generative network for speech synthesis,
R. Prenger, R. Valle, and B. Catanzaro, "Waveglow: A flow-based generative network for speech synthesis," in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019: IEEE, pp. 3617-3621
2019
-
[97]
Wider or deeper: Revisiting the resnet model for visual recognition,
Z. Wu, C. Shen, and A. v. d. Hengel, "Wider or deeper: Revisiting the resnet model for visual recognition," arXiv preprint arXiv:1611.10080, pp. 1-19, 2016
arXiv 2016
-
[98]
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,
K. He, X. Zhang, S. Ren, and J. Sun, "Delving deep into rectifiers: Surpassing human-level performance on imagenet classification," in Proceedings of the IEEE international conference on computer vision, 2015, pp. 1026-1034
2015
-
[128]
Deep convolutional neural networks for LVCSR,
T. N. Sainath, A.-r. Mohamed, B. Kingsbury, and B. Ramabhadran, "Deep convolutional neural networks for LVCSR," in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on, 2013: IEEE, pp. 8614-8618
2013
-
[40]
Novel hierarchical Cellular Simultaneous Recurrent neural Network for object detection,
M. Alam, L. Vidyaratne, and K. M. Iftekharuddin, "Novel hierarchical Cellular Simultaneous Recurrent neural Network for object detection," in Neural Networks (IJCNN), 2015 International Joint Conference on, 12-17 July 2015 2015, pp. 1-7, doi: 10.1109/IJCNN.2015.7280480
-
[99]
Densely connected convolutional networks,
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, "Densely connected convolutional networks," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4700-4708
2017
-
[101]
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,
M. Tan and Q. V. Le, "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks," arXiv preprint arXiv:1905.11946, 2019
arXiv 1905
-
[1]
Driver inattention monitoring system for intelligent vehicles: A review,
Y. Dong, Z. Hu, K. Uchimura, and N. Murayama, "Driver inattention monitoring system for intelligent vehicles: A review," 2011 2010, vol. 12, 2 ed., pp. 596-614, doi: 10.1109/TITS.2010.2092770
Show all 233 references
-
[2]
Video-based lane estimation and tracking for driver assistance: Survey, system, and evaluation,
J. C. McCall and M. M. Trivedi, "Video-based lane estimation and tracking for driver assistance: Survey, system, and evaluation," vol. 7, ed, 2006, pp. 20-37
2006
-
[3]
A Review of Computer Vision Techniques for the Analysis of Urban Traffic,
N. Buch, S. a. Velastin, and J. Orwell, "A Review of Computer Vision Techniques for the Analysis of Urban Traffic," IEEE Transactions on Intelligent Transportation Systems, vol. 12, no. 3, pp. 920-939, 2011, doi: 10.1109/TITS.2011.2119372
2011
-
[4]
Looking at Humans in the Age of Self-Driving and Highly Automated Vehicles,
E. Ohn-Bar and M. M. Trivedi, "Looking at Humans in the Age of Self-Driving and Highly Automated Vehicles," IEEE Transactions on Intelligent Vehicles, vol. 1, no. 1, pp. 90-104, 2016, doi: 10.1109/TIV.2016.2571067
2016
-
[5]
End to End Learning for Self-Driving Cars,
M. Bojarski et al., "End to End Learning for Self-Driving Cars," arXiv:1604, pp. 1-9, 2016. [Online]. Available: http://arxiv.org/abs/1604.07316
2016 arXiv
-
[6]
Lane-Change Detection Based on Vehicle-Trajectory Prediction,
H. Woo et al., "Lane-Change Detection Based on Vehicle-Trajectory Prediction," IEEE Robotics and Automation Letters, vol. 2, no. 2, pp. 1109-1116, 2017, doi: 10.1109/LRA.2017.2660543
2017
-
[7]
Single-pedestrian detection aided by two-pedestrian detection,
W. Ouyang, X. Zeng, and X. Wang, "Single-pedestrian detection aided by two-pedestrian detection," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 9, pp. 1875-1889, 2015, doi: 10.1109/TPAMI.2014.2377734
2015
-
[8]
Deep Architecture for Traffic Flow Prediction: Deep Belief Networks With Multitask Learning,
W. Huang, G. Song, H. Hong, and K. Xie, "Deep Architecture for Traffic Flow Prediction: Deep Belief Networks With Multitask Learning," IEEE Transactions on Intelligent Transportation Systems, vol. 15, no. 5, pp. 2191-2201, 2014, doi: 10.1109/TITS.2014.2311123
2014
-
[9]
Capturing Car-Following Behaviors by Deep Learning,
X. Wang, R. Jiang, L. Li, Y. Lin, X. Zheng, and F.-Y. Wang, "Capturing Car-Following Behaviors by Deep Learning," IEEE Transactions on Intelligent Transportation Systems, pp. 1-11, 2017, doi: 10.1109/TITS.2017.2706963
2017
-
[10]
Deep Learning for Reliable Mobile Edge Analytics in Intelligent Transportation Systems: An Overview,
A. Ferdowsi, U. Challita, and W. Saad, "Deep Learning for Reliable Mobile Edge Analytics in Intelligent Transportation Systems: An Overview," ieee vehicular technology magazine, vol. 14, no. 1, pp. 62-70, 2019
2019
-
[11]
Brain tumor segmentation with Deep Neural Networks,
M. Havaei et al., "Brain tumor segmentation with Deep Neural Networks," Medical Image Analysis, vol. 35, pp. 18-31, 2017, doi: 10.1016/j.media.2016.05.004
2017 doi
-
[12]
Multimodal Neuroimaging Feature Learning for Multiclass Diagnosis of Alzheimer's Disease,
S. Liu et al., "Multimodal Neuroimaging Feature Learning for Multiclass Diagnosis of Alzheimer's Disease," IEEE Transactions on Biomedical Engineering, vol. 62, no. 4, pp. 1132-1140, 2015, doi: 10.1109/TBME.2014.2372011
2015
-
[13]
Deep biomarkers of human aging: Application of deep neural networks to biomarker development,
E. Putin et al., "Deep biomarkers of human aging: Application of deep neural networks to biomarker development," Aging, vol. 8, no. 5, pp. 1021-1033, 2016, doi: 10.18632/aging.100968
2016 doi
-
[14]
An end-to-end computer vision pipeline for automated cardiac function assessment by echocardiography,
R. C. Deo et al., "An end-to-end computer vision pipeline for automated cardiac function assessment by echocardiography," CoRR, 2017
2017
-
[15]
A review of smart homes—Past, present, and future,
M. R. Alam, M. B. I. Reaz, and M. A. M. Ali, "A review of smart homes—Past, present, and future," IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), vol. 42, no. 6, pp. 1190-1203, 2012
2012
-
[16]
Personal virtual assistant,
R. S. Cooper, J. F. McElroy, W. Rolandi, D. Sanders, R. M. Ulmer, and E. Peebles, "Personal virtual assistant," ed: Google Patents, 2011
2011
-
[17]
Application of data mining techniques in customer relationship management: A literature review and classification,
E. W. Ngai, L. Xiu, and D. C. Chau, "Application of data mining techniques in customer relationship management: A literature review and classification," Expert systems with applications, vol. 36, no. 2, pp. 2592-2602, 2009
2009
-
[18]
A review on application of data mining techniques to combat natural disasters,
S. Goswami, S. Chakraborty, S. Ghosh, A. Chakrabarti, and B. Chakraborty, "A review on application of data mining techniques to combat natural disasters," Ain Shams Engineering Journal, pp. 1-14, 2016
2016
-
[19]
Vision based hand gesture recognition for human computer interaction: a survey,
S. S. Rautaray and A. Agrawal, "Vision based hand gesture recognition for human computer interaction: a survey," Artificial Intelligence Review, vol. 43, no. 1, pp. 1-54, 2015
2015
-
[20]
Deeppose: Human pose estimation via deep neural networks,
A. Toshev and C. Szegedy, "Deeppose: Human pose estimation via deep neural networks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 1653-1660
2014
-
[21]
Joint training of a convolutional network and a graphical model for human pose estimation,
J. J. Tompson, A. Jain, Y. LeCun, and C. Bregler, "Joint training of a convolutional network and a graphical model for human pose estimation," in Advances in neural information processing systems, 2014, pp. 1799-1807
2014
-
[22]
Safety and security in smart cities using artificial intelligence—A review,
S. Srivastava, A. Bisht, and N. Narayan, "Safety and security in smart cities using artificial intelligence—A review," in Cloud Computing, Data Science & Engineering-Confluence, 2017 7th International Conference on, 2017: IEEE, pp. 130-133. 19
2017
-
[23]
Detecting depression severity from vocal prosody,
Y. Yang, C. Fairbairn, and J. F. Cohn, "Detecting depression severity from vocal prosody," IEEE Transactions on Affective Computing, vol. 4, no. 2, pp. 142-150, 2013, doi: 10.1109/T-AFFC.2012.38
2013 doi
-
[24]
Speech and prosody characteristics of adolescents and adults with high-functioning autism and Asperger syndrome,
L. D. Shriberg, R. Paul, J. L. McSweeny, A. Klin, D. J. Cohen, and F. R. Volkmar, "Speech and prosody characteristics of adolescents and adults with high-functioning autism and Asperger syndrome," Journal of Speech, Language, and Hearing Research, vol. 44, no. 5, pp. 1097-1115...
2001 doi
-
[25]
Survey on speech emotion recognition: Features, classification schemes, and databases,
M. El Ayadi, M. S. Kamel, and F. Karray, "Survey on speech emotion recognition: Features, classification schemes, and databases," Pattern Recognition, vol. 44, no. 3, pp. 572-587, 2011, doi: 10.1016/j.patcog.2010.09.020
2011 doi
-
[26]
Evaluating deep learning architectures for Speech Emotion Recognition,
H. M. Fayek, M. Lech, and L. Cavedon, "Evaluating deep learning architectures for Speech Emotion Recognition," Neural Networks, vol. 92, pp. 60-68, 2017, doi: 10.1016/j.neunet.2017.02.013
2017 doi
-
[27]
Deep learning for robust feature generation in audiovisual emotion recognition,
Y. Kim, H. Lee, and E. M. Provost, "Deep learning for robust feature generation in audiovisual emotion recognition," 2013, pp. 3687- 3691, doi: 10.1109/ICASSP.2013.6638346. [Online]. Available: http://ieeexplore.ieee.org/lpdocs/epic03/wrapper.htm?arnumber=6638346
2013
-
[28]
A fast learning algorithm for deep belief nets,
G. E. Hinton, S. Osindero, and Y.-W. Teh, "A fast learning algorithm for deep belief nets," Neural computation, vol. 18, no. 7, pp. 1527-1554, 2006
2006
-
[29]
Learning multiple layers of representation,
G. E. Hinton, "Learning multiple layers of representation," Trends in cognitive sciences, vol. 11, no. 10, pp. 428-434, 2007
2007
-
[30]
Comparison of deep neural networks to spatio-temporal cortical dynamics of human visual object recognition reveals hierarchical correspondence,
R. M. Cichy, A. Khosla, D. Pantazis, A. Torralba, and A. Oliva, "Comparison of deep neural networks to spatio-temporal cortical dynamics of human visual object recognition reveals hierarchical correspondence," Scientific reports, vol. 6, pp. 1-13, 2016, Art no. 27755
2016
-
[31]
Deep hierarchies in the primate visual cortex: What can we learn for computer vision?,
N. Kruger et al., "Deep hierarchies in the primate visual cortex: What can we learn for computer vision?," IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 8, pp. 1847-1871, 2013
2013
-
[32]
Deep learning in neural networks: An overview,
J. Schmidhuber, "Deep learning in neural networks: An overview," Neural networks, vol. 61, pp. 85-117, 2015
2015
-
[33]
Imagenet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, "Imagenet classification with deep convolutional neural networks," in Advances in neural information processing systems, 2012, pp. 1097-1105
2012
-
[34]
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, "Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion," Journal of Machine Learning Research, vol. 11, no. Dec, pp. 3371-3408, 2010
2010
-
[35]
Generative adversarial nets,
I. Goodfellow et al., "Generative adversarial nets," in Advances in neural information processing systems, 2014, pp. 2672-2680
2014
-
[36]
Auto-encoding variational bayes,
D. P. Kingma and M. Welling, "Auto-encoding variational bayes," arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[37]
Density estimation using real nvp,
L. Dinh, J. Sohl-Dickstein, and S. Bengio, "Density estimation using real nvp," arXiv preprint arXiv:1605.08803, 2016
2016 arXiv
-
[38]
A critical review of recurrent neural networks for sequence learning,
Z. C. Lipton, J. Berkowitz, and C. Elkan, "A critical review of recurrent neural networks for sequence learning," arXiv preprint arXiv:1506.00019, pp. 1-38, 2015
2015 arXiv
-
[39]
Attention is all you need,
A. Vaswani et al., "Attention is all you need," in Advances in neural information processing systems, 2017, pp. 5998-6008
2017
-
[41]
Restricted Boltzmann machines for collaborative filtering,
R. Salakhutdinov, A. Mnih, and G. Hinton, "Restricted Boltzmann machines for collaborative filtering," in Proceedings of the 24th international conference on Machine learning, 2007: ACM, pp. 791-798
2007
-
[42]
Deep boltzmann machines,
R. Salakhutdinov and G. Hinton, "Deep boltzmann machines," in Artificial Intelligence and Statistics, 2009, pp. 448-455
2009
-
[43]
Extracting deep bottleneck features using stacked auto-encoders,
J. Gehring, Y. Miao, F. Metze, and A. Waibel, "Extracting deep bottleneck features using stacked auto-encoders," in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on, 2013: IEEE, pp. 3377-3381
2013
-
[44]
Extracting and composing robust features with denoising autoencoders,
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, "Extracting and composing robust features with denoising autoencoders," in Proceedings of the 25th international conference on Machine learning, 2008: ACM, pp. 1096-1103
2008
-
[45]
Learning hierarchical representations for face verification with convolutional deep belief networks,
G. B. Huang, H. Lee, and E. Learned-Miller, "Learning hierarchical representations for face verification with convolutional deep belief networks," in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, 2012: IEEE, pp. 2518-2525
2012
-
[46]
Investigation of deep boltzmann machines for phone recognition,
Z. You, X. Wang, and B. Xu, "Investigation of deep boltzmann machines for phone recognition," in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, 2013: IEEE, pp. 7600-7603
2013
-
[47]
Random Deep Belief Networks for Recognizing Emotions from Speech Signals,
G. Wen, H. Li, J. Huang, D. Li, and E. Xun, "Random Deep Belief Networks for Recognizing Emotions from Speech Signals," Computational intelligence and neuroscience, vol. 2017, pp. 1-9, 2017
2017
-
[48]
A research of speech emotion recognition based on deep belief network and SVM,
C. Huang, W. Gong, W. Fu, and D. Feng, "A research of speech emotion recognition based on deep belief network and SVM," Mathematical Problems in Engineering, vol. 2014, pp. 1-7, 2014
2014
-
[49]
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations,
H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng, "Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations," in Proceedings of the 26th annual international conference on machine learning, 2009: ACM, pp. 609-616
2009
-
[50]
Attribute2image: Conditional image generation from visual attributes,
X. Yan, J. Yang, K. Sohn, and H. Lee, "Attribute2image: Conditional image generation from visual attributes," in European Conference on Computer Vision, 2016: Springer, pp. 776-791
2016
-
[51]
An uncertain future: Forecasting from static images using variational autoencoders,
J. Walker, C. Doersch, A. Gupta, and M. Hebert, "An uncertain future: Forecasting from static images using variational autoencoders," in European Conference on Computer Vision, 2016: Springer, pp. 835-851
2016
-
[52]
A hybrid convolutional variational autoencoder for text generation,
S. Semeniuta, A. Severyn, and E. Barth, "A hybrid convolutional variational autoencoder for text generation," arXiv preprint arXiv:1702.02390, 2017
2017 arXiv
-
[53]
Expressive speech synthesis via modeling expressions with variational autoencoder,
K. Akuzawa, Y. Iwasawa, and Y. Matsuo, "Expressive speech synthesis via modeling expressions with variational autoencoder," arXiv preprint arXiv:1804.02135, 2018
2018 arXiv
-
[54]
Generative adversarial text to image synthesis,
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, "Generative adversarial text to image synthesis," arXiv preprint arXiv:1605.05396, 2016
2016 arXiv
-
[55]
Photo-realistic single image super-resolution using a generative adversarial network,
C. Ledig et al., "Photo-realistic single image super-resolution using a generative adversarial network," arXiv preprint, 2017
2017
-
[56]
Conditional generative adversarial nets,
M. Mirza and S. Osindero, "Conditional generative adversarial nets," arXiv preprint arXiv:1411.1784, 2014
2014 arXiv
-
[57]
High-resolution image synthesis and semantic manipulation with conditional gans,
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, A. Tao, J. Kautz, and B. Catanzaro, "High-resolution image synthesis and semantic manipulation with conditional gans," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 8798-8807
2018
-
[58]
Adversarial feature learning,
J. Donahue, P. Krähenbühl, and T. Darrell, "Adversarial feature learning," arXiv preprint arXiv:1605.09782, 2016. 20
2016 arXiv
-
[59]
Large scale adversarial representation learning,
J. Donahue and K. Simonyan, "Large scale adversarial representation learning," in Advances in Neural Information Processing Systems, 2019, pp. 10541-10551
2019
-
[60]
Towards principled methods for training generative adversarial networks,
M. Arjovsky and L. Bottou, "Towards principled methods for training generative adversarial networks," arXiv preprint arXiv:1701.04862, 2017
2017 arXiv
-
[61]
NIPS 2016 tutorial: Generative adversarial networks,
I. Goodfellow, "NIPS 2016 tutorial: Generative adversarial networks," arXiv preprint arXiv:1701.00160, 2016
2016 arXiv
-
[62]
Spectral normalization for generative adversarial networks,
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, "Spectral normalization for generative adversarial networks," arXiv preprint arXiv:1802.05957, 2018
2018 arXiv
-
[63]
Wasserstein generative adversarial networks,
M. Arjovsky, S. Chintala, and L. Bottou, "Wasserstein generative adversarial networks," in International conference on machine learning, 2017, pp. 214-223
2017
-
[64]
Improved training of wasserstein gans,
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, "Improved training of wasserstein gans," in Advances in neural information processing systems, 2017, pp. 5767-5777
2017
-
[65]
Least squares generative adversarial networks,
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, "Least squares generative adversarial networks," in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 2794-2802
2017
-
[66]
Generating Diverse High-Fidelity Images with VQ-VAE-2,
A. Razavi, A. v. d. Oord, and O. Vinyals, "Generating Diverse High-Fidelity Images with VQ-VAE-2," arXiv preprint arXiv:1906.00446, 2019
1906 arXiv
-
[67]
Nice: Non-linear independent components estimation,
L. Dinh, D. Krueger, and Y. Bengio, "Nice: Non-linear independent components estimation," arXiv preprint arXiv:1410.8516, 2014
2014 arXiv
-
[68]
Glow: Generative flow with invertible 1x1 convolutions,
D. P. Kingma and P. Dhariwal, "Glow: Generative flow with invertible 1x1 convolutions," in Advances in Neural Information Processing Systems, 2018, pp. 10215-10224
2018
-
[69]
Wavenet: A generative model for raw audio,
A. v. d. Oord et al., "Wavenet: A generative model for raw audio," arXiv preprint arXiv:1609.03499, 2016
2016 arXiv
-
[70]
Pixel recurrent neural networks,
A. v. d. Oord, N. Kalchbrenner, and K. Kavukcuoglu, "Pixel recurrent neural networks," arXiv preprint arXiv:1601.06759, 2016
2016 arXiv
-
[72]
SEGAN: Speech enhancement generative adversarial network,
S. Pascual, A. Bonafonte, and J. Serra, "SEGAN: Speech enhancement generative adversarial network," arXiv preprint arXiv:1703.09452, 2017
2017 arXiv
-
[73]
Speech Enhancement for Noise-Robust Speech Synthesis Using Wasserstein GAN}},
N. Adiga, Y. Pantazis, V. Tsiaras, and Y. Stylianou, "Speech Enhancement for Noise-Robust Speech Synthesis Using Wasserstein GAN}}," Proc. Interspeech 2019, pp. 1821-1825, 2019
2019
-
[74]
End-to-end sequence labeling via bi-directional lstm-cnns-crf,
X. Ma and E. Hovy, "End-to-end sequence labeling via bi-directional lstm-cnns-crf," arXiv preprint arXiv:1603.01354, 2016
2016 arXiv
-
[75]
On the properties of neural machine translation: Encoder-decoder approaches,
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio, "On the properties of neural machine translation: Encoder-decoder approaches," arXiv preprint arXiv:1409.1259, 2014
2014 arXiv
-
[76]
LSTM: A search space odyssey,
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber, "LSTM: A search space odyssey," IEEE transactions on neural networks and learning systems, vol. 28, no. 10, pp. 2222-2232, 2016
2016
-
[77]
Recurrent models of visual attention,
V. Mnih, N. Heess, and A. Graves, "Recurrent models of visual attention," in Advances in neural information processing systems, 2014, pp. 2204-2212
2014
-
[78]
Learning to combine foveal glimpses with a third-order Boltzmann machine,
H. Larochelle and G. E. Hinton, "Learning to combine foveal glimpses with a third-order Boltzmann machine," in Advances in neural information processing systems, 2010, pp. 1243-1251
2010
-
[79]
On learning where to look,
M. A. Ranzato, "On learning where to look," arXiv preprint arXiv:1405.5488, 2014
2014 arXiv
-
[80]
Learning where to attend with deep architectures for image tracking,
M. Denil, L. Bazzani, H. Larochelle, and N. de Freitas, "Learning where to attend with deep architectures for image tracking," Neural computation, vol. 24, no. 8, pp. 2151-2184, 2012
2012
-
[81]
Draw: A recurrent neural network for image generation,
K. Gregor, I. Danihelka, A. Graves, D. J. Rezende, and D. Wierstra, "Draw: A recurrent neural network for image generation," arXiv preprint arXiv:1502.04623, 2015
2015 arXiv
-
[82]
Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific Contexts,
K. Fu, J. Jin, R. Cui, F. Sha, and C. Zhang, "Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific Contexts," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 12, pp. 2321-2334, 2017, doi: 10.1109/T...
2017
-
[83]
Generating images from captions with attention,
E. Mansimov, E. Parisotto, J. L. Ba, and R. Salakhutdinov, "Generating images from captions with attention," arXiv preprint arXiv:1511.02793, 2015
2015 arXiv
-
[84]
Neural turing machines,
A. Graves, G. Wayne, and I. Danihelka, "Neural turing machines," arXiv preprint arXiv:1410.5401, pp. 1-26, 2014
2014 arXiv
-
[85]
Effective approaches to attention-based neural machine translation,
M.-T. Luong, H. Pham, and C. D. Manning, "Effective approaches to attention-based neural machine translation," arXiv preprint arXiv:1508.04025, pp. 1-11, 2015
2015 arXiv
-
[86]
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, "Listen, attend and spell: A neural network for large vocabulary conversational speech recognition," in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on, 2016: IEEE, pp. 4960- 4964
2016
-
[87]
Skeleton-based action recognition using spatio-temporal LSTM network with trust gates,
J. Liu, A. Shahroudy, D. Xu, A. C. Kot, and G. Wang, "Skeleton-based action recognition using spatio-temporal LSTM network with trust gates," IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 12, pp. 3007-3021, 2018
2018
-
[88]
Self-attention generative adversarial networks,
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, "Self-attention generative adversarial networks," arXiv preprint arXiv:1805.08318, 2018
2018 arXiv
-
[89]
Neural architecture search with reinforcement learning,
B. Zoph and Q. V. Le, "Neural architecture search with reinforcement learning," arXiv preprint arXiv:1611.01578, 2016
2016 arXiv
-
[90]
Darts: Differentiable architecture search,
H. Liu, K. Simonyan, and Y. Yang, "Darts: Differentiable architecture search," arXiv preprint arXiv:1806.09055, 2018
2018 arXiv
-
[91]
Progressive neural architecture search,
C. Liu et al., "Progressive neural architecture search," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 19-34
2018
-
[92]
Learning hierarchical features for scene labeling,
C. Farabet, C. Couprie, L. Najman, and Y. LeCun, "Learning hierarchical features for scene labeling," IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 8, pp. 1915-1929, 2013
1915
-
[93]
Going deeper with convolutions,
C. Szegedy et al., "Going deeper with convolutions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1-9
2015
-
[94]
Imagenet large scale visual recognition challenge,
O. Russakovsky et al., "Imagenet large scale visual recognition challenge," International Journal of Computer Vision, vol. 115, no. 3, pp. 211-252, 2015. 21
2015
-
[95]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," arXiv preprint arXiv:1409.1556, pp. 1-14, 2014
2014 arXiv
-
[96]
Visualizing and understanding convolutional networks,
M. D. Zeiler and R. Fergus, "Visualizing and understanding convolutional networks," in European conference on computer vision, 2014: Springer, pp. 818-833
2014
-
[100]
Squeeze-and-excitation networks,
J. Hu, L. Shen, and G. Sun, "Squeeze-and-excitation networks," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132-7141
2018
-
[102]
Cosface: Large margin cosine loss for deep face recognition,
H. Wang et al., "Cosface: Large margin cosine loss for deep face recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5265-5274
2018
-
[103]
Arcface: Additive angular margin loss for deep face recognition,
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, "Arcface: Additive angular margin loss for deep face recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 4690-4699
2019
-
[104]
Temporal pyramid pooling based convolutional neural networks for action recognition,
P. Wang, Y. Cao, C. Shen, L. Liu, and H. T. Shen, "Temporal pyramid pooling based convolutional neural networks for action recognition," IEEE Trans. Circuits and Systems for Video Technology, vol. 27, no. 12, pp. 2613-2622, 2017
2017
-
[105]
Contextual action recognition with r* cnn,
G. Gkioxari, R. Girshick, and J. Malik, "Contextual action recognition with r* cnn," in Proceedings of the IEEE international conference on computer vision, 2015, pp. 1080-1088
2015
-
[106]
Enhanced computer vision with microsoft kinect sensor: A review,
J. Han, L. Shao, D. Xu, and J. Shotton, "Enhanced computer vision with microsoft kinect sensor: A review," IEEE transactions on cybernetics, vol. 43, no. 5, pp. 1318-1334, 2013
2013
-
[107]
Exploiting deep residual networks for human action recognition from skeletal data,
H.-H. Pham, L. Khoudour, A. Crouzil, P. Zegers, and S. A. Velastin, "Exploiting deep residual networks for human action recognition from skeletal data," Computer Vision and Image Understanding, vol. 170, pp. 51-66, 2018
2018
-
[108]
Deep progressive reinforcement learning for skeleton-based action recognition,
Y. Tang, Y. Tian, J. Lu, P. Li, and J. Zhou, "Deep progressive reinforcement learning for skeleton-based action recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5323-5332
2018
-
[109]
Deep convolutional neural networks for human action recognition using depth maps and postures,
A. Kamel, B. Sheng, P. Yang, P. Li, R. Shen, and D. D. Feng, "Deep convolutional neural networks for human action recognition using depth maps and postures," IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2018
2018
-
[110]
Pictorial structures for object recognition,
P. F. Felzenszwalb and D. P. Huttenlocher, "Pictorial structures for object recognition," International journal of computer vision, vol. 61, no. 1, pp. 55-79, 2005
2005
-
[111]
3d human pose estimation in the wild by adversarial learning,
W. Yang, W. Ouyang, X. Wang, J. Ren, H. Li, and X. Wang, "3d human pose estimation in the wild by adversarial learning," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5255-5264
2018
-
[112]
Real-time 3D hand pose estimation with 3D convolutional neural networks,
L. Ge, H. Liang, J. Yuan, and D. Thalmann, "Real-time 3D hand pose estimation with 3D convolutional neural networks," IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 4, pp. 956-970, 2019
2019
-
[113]
Densepose: Dense human pose estimation in the wild,
R. Alp Güler, N. Neverova, and I. Kokkinos, "Densepose: Dense human pose estimation in the wild," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7297-7306
2018
-
[114]
Detect globally, refine locally: A novel approach to saliency detection,
T. Wang et al., "Detect globally, refine locally: A novel approach to saliency detection," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 3127-3135
2018
-
[115]
Progressive attention guided recurrent network for salient object detection,
X. Zhang, T. Wang, J. Qi, H. Lu, and G. Wang, "Progressive attention guided recurrent network for salient object detection," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 714-722
2018
-
[116]
A deep-learning based feature hybrid framework for spatiotemporal saliency detection inside videos,
Z. Wang, J. Ren, D. Zhang, M. Sun, and J. Jiang, "A deep-learning based feature hybrid framework for spatiotemporal saliency detection inside videos," Neurocomputing, vol. 287, pp. 68-83, 2018
2018
-
[117]
Pyramid dilated deeper convlstm for video salient object detection,
H. Song, W. Wang, S. Zhao, J. Shen, and K.-M. Lam, "Pyramid dilated deeper convlstm for video salient object detection," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 715-731
2018
-
[118]
Learning by tracking: Siamese CNN for robust target association,
L. Leal-Taixé, C. Canton-Ferrer, and K. Schindler, "Learning by tracking: Siamese CNN for robust target association," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016, pp. 33-40
2016
-
[119]
Fast online object tracking and segmentation: A unifying approach,
Q. Wang, L. Zhang, L. Bertinetto, W. Hu, and P. H. Torr, "Fast online object tracking and segmentation: A unifying approach," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 1328-1338
2019
-
[120]
Deep Reinforcement Learning for Subpixel Neural Tracking,
T. Dai et al., "Deep Reinforcement Learning for Subpixel Neural Tracking," in International Conference on Medical Imaging with Deep Learning, 2019, pp. 130-150
2019
-
[121]
Unpaired image-to-image translation using cycle-consistent adversarial networks,
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, "Unpaired image-to-image translation using cycle-consistent adversarial networks," in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2223-2232
2017
-
[122]
Semantic image inpainting with deep generative models,
R. A. Yeh, C. Chen, T. Yian Lim, A. G. Schwing, M. Hasegawa-Johnson, and M. N. Do, "Semantic image inpainting with deep generative models," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5485-5493
2017
-
[123]
Context encoders: Feature learning by inpainting,
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, "Context encoders: Feature learning by inpainting," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2536-2544
2016
-
[124]
Image inpainting for irregular holes using partial convolutions,
G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro, "Image inpainting for irregular holes using partial convolutions," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 85-100
2018
-
[125]
VideoFlow: A flow-based generative model for video,
M. Kumar et al., "VideoFlow: A flow-based generative model for video," arXiv preprint arXiv:1903.01434, 2019
1903 arXiv
-
[126]
Strategies for training large scale neural network language models,
T. Mikolov, A. Deoras, D. Povey, L. Burget, and J. Černocký, "Strategies for training large scale neural network language models," in Automatic Speech Recognition and Understanding (ASRU), 2011 IEEE Workshop on, 2011: IEEE, pp. 196-201
2011
-
[127]
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,
G. Hinton et al., "Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups," IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82-97, 2012. 22
2012
-
[129]
Long short-term memory recurrent neural network architectures for large scale acoustic modeling,
H. Sak, A. Senior, and F. Beaufays, "Long short-term memory recurrent neural network architectures for large scale acoustic modeling," in Fifteenth Annual Conference of the International Speech Communication Association, 2014, pp. 338-342
2014
-
[130]
Deep long short-term memory networks for speech recognition,
J.-T. Chien and A. Misbullah, "Deep long short-term memory networks for speech recognition," in Chinese Spoken Language Processing (ISCSLP), 2016 10th International Symposium on, 2016: IEEE, pp. 1-5
2016
-
[131]
The Microsoft 2017 conversational speech recognition system,
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke, "The Microsoft 2017 conversational speech recognition system," in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2018: IEEE, pp. 5934-5938
2017
-
[132]
State-of-the-art speech recognition with sequence-to-sequence models,
C.-C. Chiu et al., "State-of-the-art speech recognition with sequence-to-sequence models," in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018: IEEE, pp. 4774-4778
2018
-
[133]
Improved training of end-to-end attention models for speech recognition,
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, "Improved training of end-to-end attention models for speech recognition," arXiv preprint arXiv:1805.03294, 2018
2018 arXiv
-
[134]
Memory networks,
J. Weston, S. Chopra, and A. Bordes, "Memory networks," arXiv preprint arXiv:1410.3916, pp. 1-15, 2014
2014 arXiv
-
[135]
Improved semantic representations from tree-structured long short-term memory networks,
K. S. Tai, R. Socher, and C. D. Manning, "Improved semantic representations from tree-structured long short-term memory networks," arXiv preprint arXiv:1503.00075, pp. 1-11, 2015
2015 arXiv
-
[136]
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,
Y. Wu et al., "Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation," arXiv preprint arXiv:1609.08144, pp. 1-23, 2016
2016 arXiv
-
[137]
Deep visual-semantic alignments for generating image descriptions,
A. Karpathy and L. Fei-Fei, "Deep visual-semantic alignments for generating image descriptions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3128-3137
2015
-
[138]
Automatic speech emotion recognition using recurrent neural networks with local attention,
S. Mirsamadi, E. Barsoum, and C. Zhang, "Automatic speech emotion recognition using recurrent neural networks with local attention," in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017: IEEE, pp. 2227-2231
2017
-
[139]
3-D convolutional recurrent neural networks with attention model for speech emotion recognition,
M. Chen, X. He, J. Yang, and H. Zhang, "3-D convolutional recurrent neural networks with attention model for speech emotion recognition," IEEE Signal Processing Letters, vol. 25, no. 10, pp. 1440-1444, 2018
2018
-
[140]
Adversarial auto-encoders for speech based emotion recognition,
S. Sahu, R. Gupta, G. Sivaraman, W. AbdAlmageed, and C. Espy-Wilson, "Adversarial auto-encoders for speech based emotion recognition," arXiv preprint arXiv:1806.02146, 2018
2018 arXiv
-
[141]
Deep audio-visual speech recognition,
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, "Deep audio-visual speech recognition," IEEE transactions on pattern analysis and machine intelligence, 2018
2018
-
[142]
Zero-shot keyword spotting for visual speech recognition in-the-wild,
T. Stafylakis and G. Tzimiropoulos, "Zero-shot keyword spotting for visual speech recognition in-the-wild," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 513-529
2018
-
[143]
The CIFAR-10 dataset,
A. Krizhevsky, V. Nair, and G. Hinton, "The CIFAR-10 dataset," online: http://www. cs. toronto. edu/kriz/cifar. html, vol. 55, 2014
2014
-
[144]
Microsoft coco: Common objects in context,
T.-Y. Lin et al., "Microsoft coco: Common objects in context," in European conference on computer vision, 2014: Springer, pp. 740- 755
2014
-
[145]
The unmanned aerial vehicle benchmark: Object detection and tracking,
D. Du et al., "The unmanned aerial vehicle benchmark: Object detection and tracking," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 370-386
2018
-
[146]
Vision meets drones: A challenge,
P. Zhu, L. Wen, X. Bian, H. Ling, and Q. Hu, "Vision meets drones: A challenge," arXiv preprint arXiv:1804.07437, 2018
2018 arXiv
-
[147]
Phone recognition on the TIMIT database,
C. Lopes and F. Perdigao, "Phone recognition on the TIMIT database," Speech Technologies/Book, vol. 1, pp. 285-302, 2011
2011
-
[148]
Voxceleb: a large-scale speaker identification dataset,
A. Nagrani, J. S. Chung, and A. Zisserman, "Voxceleb: a large-scale speaker identification dataset," arXiv preprint arXiv:1706.08612, 2017
2017 arXiv
-
[149]
Neural Machine Translation
Stanford. "Neural Machine Translation." https://nlp.stanford.edu/projects/nmt/ (accessed
-
[150]
The fifth'CHiME'Speech Separation and Recognition Challenge: Dataset, task and baselines,
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, "The fifth'CHiME'Speech Separation and Recognition Challenge: Dataset, task and baselines," arXiv preprint arXiv:1803.10609, 2018
2018 arXiv
-
[151]
LRS3-TED: a large-scale dataset for visual speech recognition,
T. Afouras, J. S. Chung, and A. Zisserman, "LRS3-TED: a large-scale dataset for visual speech recognition," arXiv preprint arXiv:1809.00496, 2018
2018 arXiv
-
[152]
Google Brain Team's Mission
Google. "Google Brain Team's Mission." https://ai.google/research/teams/brain/ (accessed
-
[153]
Facebook AI Research (FAIR)
Facebook. "Facebook AI Research (FAIR)." https://research.fb.com/category/facebook-ai-research-fair/ (accessed
-
[154]
Facebook’s Perfect, Impossible Chatbot,
T. Simonite, "Facebook’s Perfect, Impossible Chatbot," MIT Technology Review. [Online]. Available: https://www.technologyreview.com/s/604117/facebooks-perfect-impossible-chatbot/
-
[155]
Cognitive Toolkit
Microsoft. "Cognitive Toolkit." https://docs.microsoft.com/en-us/cognitive-toolkit/ (accessed
-
[156]
Achieving human parity in conversational speech recognition,
W. Xiong et al., "Achieving human parity in conversational speech recognition," arXiv preprint arXiv:1610.05256, pp. 1-13, 2016
2016 arXiv
-
[157]
Cortana
Microsoft. "Cortana." https://www.microsoft.com/en-us/cortana (accessed
-
[158]
Specification FAQ
I. T. Association, "Specification FAQ." [Online]. Available: http://www.infinibandta.org/content/pages.php?pg=technology_faq
-
[159]
Deep speech 2: End-to-end speech recognition in english and mandarin,
D. Amodei et al., "Deep speech 2: End-to-end speech recognition in english and mandarin," in International Conference on Machine Learning, 2016, pp. 173-182
2016
-
[160]
Deep Learning AI
NVIDIA. "Deep Learning AI." https://www.nvidia.com/en-us/deep-learning-ai/ (accessed
-
[161]
"Watson." https://www.ibm.com/watson/ (accessed
IBM. "Watson." https://www.ibm.com/watson/ (accessed
-
[162]
Apple Machine Learning Journal
A. Inc. "Apple Machine Learning Journal." https://machinelearning.apple.com/ (accessed
-
[163]
Amazon Machine Learning
A. W. Services. "Amazon Machine Learning." https://aws.amazon.com/sagemaker (accessed
-
[164]
Engineering More Reliable Transportation with Machine Learning and AI at Uber
U. Engineering, "Engineering More Reliable Transportation with Machine Learning and AI at Uber." [Online]. Available: https://eng.uber.com/machine-learning/
-
[165]
Machine Learning Offers a Path to Deeper Insight
Intel, "Machine Learning Offers a Path to Deeper Insight." [Online]. Available: https://www.intel.com/content/www/us/en/analytics/machine-learning/overview.html
-
[166]
“Your Word is my Command
J. Schalkwyk et al., "“Your Word is my Command”: Google Search by Voice: A Case Study," in Advances in Speech Recognition: Springer, 2010, pp. 61-90
2010
-
[167]
Small-footprint keyword spotting using deep neural networks,
G. Chen, C. Parada, and G. Heigold, "Small-footprint keyword spotting using deep neural networks," in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on, 2014: IEEE, pp. 4087-4091. 23
2014
-
[168]
Convolutional neural networks for small-footprint keyword spotting,
T. N. Sainath and C. Parada, "Convolutional neural networks for small-footprint keyword spotting," in Sixteenth Annual Conference of the International Speech Communication Association, 2015, pp. 1478-1482
2015
-
[169]
Query-by-example keyword spotting using long short-term memory networks,
G. Chen, C. Parada, and T. N. Sainath, "Query-by-example keyword spotting using long short-term memory networks," in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on, 2015: IEEE, pp. 5236-5240
2015
-
[170]
Accurate and compact large vocabulary speech recognition on mobile devices,
X. Lei, A. W. Senior, A. Gruenstein, and J. Sorensen, "Accurate and compact large vocabulary speech recognition on mobile devices," in Interspeech, 2013, vol. 1, pp. 662-665
2013
-
[171]
On-demand language model interpolation for mobile speech input,
B. Ballinger, C. Allauzen, A. Gruenstein, and J. Schalkwyk, "On-demand language model interpolation for mobile speech input," in Interspeech, 2010, pp. 1812-1815
2010
-
[172]
Unary data structures for language models,
J. Sorensen and C. Allauzen, "Unary data structures for language models," in Twelfth Annual Conference of the International Speech Communication Association, 2011, pp. 1425-1428
2011
-
[173]
Small-footprint high-performance deep neural network-based speech recognition using split-VQ,
Y. Wang, J. Li, and Y. Gong, "Small-footprint high-performance deep neural network-based speech recognition using split-VQ," in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on, 2015: IEEE, pp. 4984-4988
2015
-
[174]
Model Compression Applied to Small-Footprint Keyword Spotting,
G. Tucker, M. Wu, M. Sun, S. Panchapagesan, G. Fu, and S. Vitaladevuni, "Model Compression Applied to Small-Footprint Keyword Spotting," in INTERSPEECH, 2016, pp. 1878-1882
2016
-
[175]
Deep feature-based face detection on mobile devices,
S. Sarkar, V. M. Patel, and R. Chellappa, "Deep feature-based face detection on mobile devices," in Identity, Security and Behavior Analysis (ISBA), 2016 IEEE International Conference on, 2016: IEEE, pp. 1-8
2016
-
[176]
Deep learners benefit more from out-of-distribution examples,
Y. Bengio et al., "Deep learners benefit more from out-of-distribution examples," in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 2011, pp. 164-172
2011
-
[177]
Face-based active authentication on mobile devices,
M. E. Fathy, V. M. Patel, and R. Chellappa, "Face-based active authentication on mobile devices," in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on, 2015: IEEE, pp. 1687-1691
2015
-
[178]
Mobio database for the ICPR 2010 face and speech competition,
C. McCool and S. Marcel, "Mobio database for the ICPR 2010 face and speech competition," Idiap, 2009
2010
-
[179]
Mobilenets: Efficient convolutional neural networks for mobile vision applications,
A. G. Howard et al., "Mobilenets: Efficient convolutional neural networks for mobile vision applications," arXiv preprint arXiv:1704.04861, 2017
2017 arXiv
-
[180]
Redundancy-Reduced MobileNet Acceleration on Reconfigurable Logic for ImageNet Classification,
J. Su et al., "Redundancy-Reduced MobileNet Acceleration on Reconfigurable Logic for ImageNet Classification," Cham, 2018: Springer International Publishing, in Applied Reconfigurable Computing. Architectures, Tools, and Applications, pp. 16-28
2018
-
[181]
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,
S. Han, H. Mao, and W. J. Dally, "Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding," arXiv preprint arXiv:1510.00149, pp. 1-14, 2015
2015 arXiv
-
[182]
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients,
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou, "Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients," arXiv preprint arXiv:1606.06160, 2016
2016 arXiv
-
[183]
An early resource characterization of deep learning on wearables, smartphones and internet-of-things devices,
N. D. Lane, S. Bhattacharya, P. Georgiev, C. Forlivesi, and F. Kawsar, "An early resource characterization of deep learning on wearables, smartphones and internet-of-things devices," in Proceedings of the 2015 International Workshop on Internet of Things towards Applications, ...
2015
-
[184]
Multi-digit number recognition from street view imagery using deep convolutional neural networks,
I. J. Goodfellow, Y. Bulatov, J. Ibarz, S. Arnoud, and V. Shet, "Multi-digit number recognition from street view imagery using deep convolutional neural networks," arXiv preprint arXiv:1312.6082, pp. 1-13, 2013
2013 arXiv
-
[185]
Can deep learning revolutionize mobile sensing?,
N. D. Lane and P. Georgiev, "Can deep learning revolutionize mobile sensing?," in Proceedings of the 16th International Workshop on Mobile Computing Systems and Applications, 2015: ACM, pp. 117-122
2015
-
[186]
Deepx: A software accelerator for low-power deep learning inference on mobile devices,
N. D. Lane et al., "Deepx: A software accelerator for low-power deep learning inference on mobile devices," in Information Processing in Sensor Networks (IPSN), 2016 15th ACM/IEEE International Conference on, 2016: IEEE, pp. 1-12
2016
-
[187]
Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015) Database,
N. Evans, Z. Wu, J. Yamagishi, and T. Kinnunen, "Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015) Database," 2015
2015
-
[188]
Reading digits in natural images with unsupervised feature learning,
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, "Reading digits in natural images with unsupervised feature learning," in NIPS workshop on deep learning and unsupervised feature learning, 2011, vol. 2011, no. 2, p. 5
2011
-
[189]
Structured transforms for small-footprint deep learning,
V. Sindhwani, T. Sainath, and S. Kumar, "Structured transforms for small-footprint deep learning," in Advances in Neural Information Processing Systems, 2015, pp. 3088-3096
2015
-
[190]
Pan, Structured matrices and polynomials: unified superfast algorithms
V. Pan, Structured matrices and polynomials: unified superfast algorithms. Springer Science & Business Media, 2012
2012
-
[191]
Learning natural language inference with LSTM,
S. Wang and J. Jiang, "Learning natural language inference with LSTM," arXiv preprint arXiv:1512.08849, pp. 1-10, 2015
2015 arXiv
-
[192]
Shufflenet: An extremely efficient convolutional neural network for mobile devices,
X. Zhang, X. Zhou, M. Lin, and J. Sun, "Shufflenet: An extremely efficient convolutional neural network for mobile devices," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6848-6856
2018
-
[193]
Affect detection: An interdisciplinary review of models, methods, and their applications,
R. A. Calvo and S. D'Mello, "Affect detection: An interdisciplinary review of models, methods, and their applications," IEEE Transactions on Affective Computing, vol. 1, no. 1, pp. 18-37, 2010, doi: 10.1109/T-AFFC.2010.1
2010 doi
-
[194]
Recognizing facial expression: machine learning and application to spontaneous behavior,
M. S. Bartlett, G. Littlewort, M. Frank, C. Lainscsek, I. Fasel, and J. Movellan, "Recognizing facial expression: machine learning and application to spontaneous behavior," in Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on, 2005, ...
2005
-
[195]
Spoken Emotion Recognition Using Deep Learning,
E. M. Albornoz, M. Sánchez-Gutiérrez, F. Martinez-Licona, H. L. Rufiner, and J. Goddard, "Spoken Emotion Recognition Using Deep Learning," Springer, Cham, 2014, pp. 104-111
2014
-
[196]
Video affective content analysis: a survey of state of the art methods,
S. Wang and Q. Ji, "Video affective content analysis: a survey of state of the art methods," IEEE Transactions on Affective Computing, vol. 6, no. 4, pp. 1-1, 2015, doi: 10.1109/TAFFC.2015.2432791
2015
-
[197]
A review of the use of computational intelligence in the design of military surveillance networks,
M. G. Ball, B. Qela, and S. Wesolkowski, "A review of the use of computational intelligence in the design of military surveillance networks," in Recent Advances in Computational Intelligence in Defense and Security: Springer, 2016, pp. 663-693
2016
-
[198]
Automatic handgun detection alarm in videos using deep learning,
R. Olmos, S. Tabik, and F. Herrera, "Automatic handgun detection alarm in videos using deep learning," Neurocomputing, vol. 275, pp. 66-72, 2018
2018
-
[199]
Towards reading hidden emotions: A comparative study of spontaneous micro-expression spotting and recognition methods,
X. Li et al., "Towards reading hidden emotions: A comparative study of spontaneous micro-expression spotting and recognition methods," IEEE Transactions on Affective Computing, 2017
2017
-
[200]
Facial Action Coding System - Manual and Investigator’s Guide. FACS,
P. Ekman, Friesen, W. V., & Hager, J. C. , "Facial Action Coding System - Manual and Investigator’s Guide. FACS," Research Nexus,
-
[201]
The faces of engagement: Automatic recognition of student engagement from facial expressions,
J. Whitehill, Z. Serpell, Y. C. Lin, A. Foster, and J. R. Movellan, "The faces of engagement: Automatic recognition of student engagement from facial expressions," IEEE Transactions on Affective Computing, vol. 5, no. 1, pp. 86-98, 2014, doi: 10.1109/TAFFC.2014.2316163
2014
-
[202]
Characterizing consumer emotional response to sweeteners using an emotion terminology questionnaire and facial expression analysis,
K. A. Leitch, S. E. Duncan, S. O'Keefe, R. Rudd, and D. L. Gallagher, "Characterizing consumer emotional response to sweeteners using an emotion terminology questionnaire and facial expression analysis," Food Research International, vol. 76, pp. 283-292, 2015, doi: 10.1016/j.f...
2015 doi
-
[203]
Artificial intelligence and behavioral economics,
C. F. Camerer, "Artificial intelligence and behavioral economics," in Economics of Artificial Intelligence: University of Chicago Press, 2017
2017
-
[204]
A Feasibility Study of Autism Behavioral Markers in Spontaneous Facial, Visual, and Hand Movement Response Data,
M. D. Samad, N. Diawara, J. L. Bobzien, J. W. Harrington, M. A. Witherow, and K. M. Iftekharuddin, "A Feasibility Study of Autism Behavioral Markers in Spontaneous Facial, Visual, and Hand Movement Response Data," IEEE Transactions on Neural Systems and Rehabilitation Engineer...
2018
-
[205]
Computational Analysis of Deep Visual Data for Quantifying Facial Expression Production,
M. Leo et al., "Computational Analysis of Deep Visual Data for Quantifying Facial Expression Production," Applied Sciences, vol. 9, no. 21, p. 4542, 2019
2019
-
[206]
A pilot study to identify autism related traits in spontaneous facial actions using computer vision,
M. D. Samad, N. Diawara, J. L. Bobzien, C. M. Taylor, J. W. Harrington, and K. M. Iftekharuddin, "A pilot study to identify autism related traits in spontaneous facial actions using computer vision," Research in Autism Spectrum Disorders, vol. 65, pp. 14-24, 2019
2019
-
[207]
Autonomous Driving
Audi. "Autonomous Driving." https://www.audi.com/en/experience-audi/mobility-and-trends/autonomous-driving.html (accessed
-
[208]
All Tesla Cars Being Produced Now Have Full Self-Driving Hardware
Tesla. "All Tesla Cars Being Produced Now Have Full Self-Driving Hardware." https://www.tesla.com/blog/all-tesla-cars-being- produced-now-have-full-self-driving-hardware (accessed
-
[209]
Data-Driven Multi-step Demand Prediction for Ride-Hailing Services Using Convolutional Neural Network,
C. Wang, Y. Hou, and M. Barth, "Data-Driven Multi-step Demand Prediction for Ride-Hailing Services Using Convolutional Neural Network," in Science and Information Conference, 2019: Springer, pp. 11-22
2019
-
[210]
Map Enhanced Route Travel Time Prediction using Deep Neural Networks,
S. Das et al., "Map Enhanced Route Travel Time Prediction using Deep Neural Networks," arXiv preprint arXiv:1911.02623, 2019
1911 arXiv
-
[211]
Deeppool: Distributed model-free algorithm for ride-sharing using deep reinforcement learning,
A. Alabbasi, A. Ghosh, and V. Aggarwal, "Deeppool: Distributed model-free algorithm for ride-sharing using deep reinforcement learning," arXiv preprint arXiv:1903.03882, 2019
1903 arXiv
-
[212]
Medical error—the third leading cause of death in the US,
M. Daniel and M. A. Makary, "Medical error—the third leading cause of death in the US," Bmj, vol. 353, no. i2139, p. 476636183, 2016
2016
-
[214]
A deep neural network to enhance prediction of 1-year mortality using echocardiographic videos of the heart,
A. Ulloa et al., "A deep neural network to enhance prediction of 1-year mortality using echocardiographic videos of the heart," arXiv preprint arXiv:1811.10553, 2018
2018 arXiv
-
[215]
Deep learning applications in ophthalmology,
E. Rahimy, "Deep learning applications in ophthalmology," Current opinion in ophthalmology, vol. 29, no. 3, pp. 254-260, 2018
2018
-
[216]
Detection and diagnosis of dental caries using a deep learning-based convolutional neural network algorithm,
J.-H. Lee, D.-H. Kim, S.-N. Jeong, and S.-H. Choi, "Detection and diagnosis of dental caries using a deep learning-based convolutional neural network algorithm," Journal of dentistry, vol. 77, pp. 106-111, 2018
2018
-
[217]
Dermatologist-level classification of skin cancer with deep neural networks,
A. Esteva et al., "Dermatologist-level classification of skin cancer with deep neural networks," Nature, vol. 542, no. 7639, p. 115, 2017
2017
-
[218]
Deep learning enables reduced gadolinium dose for contrast‐enhanced brain MRI,
E. Gong, J. M. Pauly, M. Wintermark, and G. Zaharchuk, "Deep learning enables reduced gadolinium dose for contrast‐enhanced brain MRI," Journal of Magnetic Resonance Imaging, vol. 48, no. 2, pp. 330-340, 2018
2018
-
[219]
Deep-learning cardiac motion analysis for human survival prediction,
G. A. Bello et al., "Deep-learning cardiac motion analysis for human survival prediction," Nature machine intelligence, vol. 1, no. 2, p. 95, 2019
2019
-
[220]
Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: Is the problem solved?,
O. Bernard et al., "Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: Is the problem solved?," IEEE transactions on medical imaging, vol. 37, no. 11, pp. 2514-2525, 2018
2018
-
[221]
Automated Gleason grading of prostate cancer tissue microarrays via deep learning,
E. Arvaniti et al., "Automated Gleason grading of prostate cancer tissue microarrays via deep learning," Scientific reports, vol. 8, 2018
2018
-
[222]
Lung pattern classification for interstitial lung diseases using a deep convolutional neural network,
M. Anthimopoulos, S. Christodoulidis, L. Ebner, A. Christe, and S. Mougiakakou, "Lung pattern classification for interstitial lung diseases using a deep convolutional neural network," IEEE transactions on medical imaging, vol. 35, no. 5, pp. 1207-1216, 2016
2016
-
[223]
Advanced machine learning in action: identification of intracranial hemorrhage on computed tomography scans of the head with clinical workflow integration,
M. R. Arbabshirani et al., "Advanced machine learning in action: identification of intracranial hemorrhage on computed tomography scans of the head with clinical workflow integration," npj Digital Medicine, vol. 1, no. 1, p. 9, 2018
2018
-
[224]
A data augmentation methodology for training machine/deep learning gait recognition algorithms,
C. C. Charalambous and A. A. Bharath, "A data augmentation methodology for training machine/deep learning gait recognition algorithms," arXiv preprint arXiv:1610.07570, pp. 1-12, 2016
2016 arXiv
-
[225]
Understanding data augmentation for classification: when to warp?,
S. C. Wong, A. Gatt, V. Stamatescu, and M. D. McDonnell, "Understanding data augmentation for classification: when to warp?," arXiv preprint arXiv:1609.08764, pp. 1-6, 2016
2016 arXiv
-
[226]
Transfer learning using computational intelligence: a survey,
J. Lu, V. Behbood, P. Hao, H. Zuo, S. Xue, and G. Zhang, "Transfer learning using computational intelligence: a survey," Knowledge- Based Systems, vol. 80, pp. 14-23, 2015
2015
-
[227]
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,
Y. Gal and Z. Ghahramani, "Dropout as a Bayesian approximation: Representing model uncertainty in deep learning," in international conference on machine learning, 2016, pp. 1050-1059
2016
-
[228]
Towards bayesian deep learning: A survey,
H. Wang and D.-Y. Yeung, "Towards bayesian deep learning: A survey," arXiv preprint arXiv:1604.01662, pp. 1-17, 2016
2016 arXiv
-
[229]
Deep learning,
Y. LeCun, Y. Bengio, and G. Hinton, "Deep learning," nature, vol. 521, no. 7553, p. 436, 2015
2015
-
[230]
Mastering the game of Go with deep neural networks and tree search,
D. Silver et al., "Mastering the game of Go with deep neural networks and tree search," nature, vol. 529, no. 7587, p. 484, 2016
2016
-
[231]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural computation, vol. 9, no. 8, pp. 1735-1780, 1997
1997
-
[232]
Densecap: Fully convolutional localization networks for dense captioning,
J. Johnson, A. Karpathy, and L. Fei-Fei, "Densecap: Fully convolutional localization networks for dense captioning," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4565-4574
2016
-
[233]
Histogram of gradients of time–frequency representations for audio scene classification,
A. Rakotomamonjy and G. Gasso, "Histogram of gradients of time–frequency representations for audio scene classification," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 23, no. 1, pp. 142-153, 2014
2014
-
[2002]
Available: https://doi.org/10.1016/j.msea.2004.04.064
[Online]. Available: https://doi.org/10.1016/j.msea.2004.04.064. 24
2004 doi
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.