{"id":"4148234d-fbfd-4083-a878-5e5f8b87debf","arxiv_id":"1908.07656","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A broad survey of deep learning architectures and systems for vision and speech, with an emphasis on mobile deployment and emerging applications, containing no new results.","lead":"This paper surveys deep learning for vision and speech, covering network designs, applications, and running models on phones and other limited hardware. It is a reference summary with no new research findings, useful mainly to readers who want a broad map of the field as of 2019.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central 'comprehensive' claim rests on the accuracy of its summary tables, and multiple table entries mis-cite their sources (e.g., Table I attributes ResNet to reference [71], which is Waveglow), so the survey's reference value is in question.","rationale":"The reader's weakest assumption is that the survey's informal paper selection is representative and that the cited works and reported numbers are accurately represented. My independent reading confirmed this is the load-bearing weakness: the central claim is a claim of comprehensiveness, and comprehensiveness can only be evaluated if the selection is described and if the cited entries are trustworthy. Multiple concrete citation mismatches appear in the headline tables, including ResNet being assigned to a Waveglow reference, a CNN paper being labeled as an autoencoder/DBN, and an LSTM description citing an object-detection paper. These errors are directly in the survey's comparative reference apparatus, so they undermine the claimed reference value more than ordinary typos would. The concern is real but addressable: a careful audit and correction of the tables and citations could restore the survey's utility. Therefore, the verdict CONDITIONAL is appropriate, consistent with the reader's assessment. I do not see a separate load-bearing concern beyond the citation/representation issue, so no stronger verdict change is warranted.","tokens_in":34379,"tokens_out":3548,"duration_ms":36364,"concrete_test":"Audit every row of Tables I, III, V, and VI by resolving each bracketed reference to its bibliography entry and verifying that the architecture, dataset, metric, and result match the cited source. In particular, resolve Table I's 'ResNet [71]' entry against reference [71] (Waveglow) and against the canonical ResNet publication; resolve Table III's 'Autoencoder/DBN [128]' against reference [128]; and verify Table I's AlexNet 17.0% against Krizhevsky et al. If any row in any table misattributes its source, the survey needs a corrections pass before the comprehensiveness claim can be accepted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and Section 7 claim the paper is 'one of the most comprehensive surveys on the latest developments in intelligent vision and speech applications.' That claim is only as strong as the survey's accurate representation of the cited literature. The paper gives no search strategy, inclusion criteria, or date range, so the comprehensiveness claim cannot be independently checked. Worse, internal spot-checks show concrete attribution errors in the summary tables that are the survey's main reference apparatus. In Table I, 'ResNet [71] - Microsoft 2015' points to reference [71], which is Prenger et al. 'Waveglow' (a speech synthesis model), not a ResNet image-classification paper; the correct ResNet paper is cited in the text as [97], but reference [97] is 'Wider or deeper' by Wu et al., not the original He et al. ResNet. In Table III, the row 'Autoencoder/DBN [128]' cites Sainath et al.'s deep convolutional neural network paper for LVCSR, which is neither an autoencoder nor a DBN. In Section 3.2, the text cites [40] for LSTM gating, but [40] is an object-detection paper. These are not typos in a peripheral section; they are mis-attributions in exactly the tables a reader would trust to compare state-of-the-art results. If representative entries cannot be relied on, the survey cannot serve as a comprehensive or authoritative reference, and the central claim fails unless the citations are corrected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of deep neural network research as applied to vision and speech systems. It covers network architectures (CNNs, generative models, RNNs, attention), their use in vision and speech applications, deployment on resource-constrained mobile and embedded platforms, and emerging application areas such as behavioral science, transportation, and medicine. The paper's central claim, stated in the abstract and in Section 7, is that it is one of the most comprehensive surveys of the latest developments in intelligent vision and speech applications from both software and hardware perspectives.","tokens_in":34680,"tokens_out":5879,"duration_ms":49403,"significance":"If accurate, this survey would provide a broad and useful entry point for researchers and practitioners, covering architectures, applications, hardware constraints, and emerging frontiers in a single document. The organization is logical and the scope is ambitious. However, the paper's value as an authoritative reference is currently undermined by several concrete citation errors in its central summary tables and in the main text. These errors are not peripheral: they occur in exactly the tables and sentences a reader would rely on to compare state-of-the-art results. The paper also provides no search strategy or inclusion criteria to support its comprehensiveness claim, which further weakens confidence in its coverage. With a systematic reference audit and methodological clarification, the survey could serve its intended purpose, but in its present form it cannot be recommended as a reliable reference.","major_comments":[{"comment":"Table I, ResNet row: the table attributes 'ResNet [71] – Microsoft 2015' with a 4.70% top-5 error, but reference [71] is Prenger et al., 'Waveglow: A flow-based generative network for speech synthesis,' not a ResNet paper. The text (Section 3.1) instead cites ResNet as [97], but [97] is Wu et al., 'Wider or deeper: Revisiting the resnet model for visual recognition,' which is not the original He et al. ResNet paper either. A reader cannot verify the claimed state-of-the-art error rate from the cited source, and this error appears in a core summary table that the survey asks readers to trust.","section":"Table I"},{"comment":"Table III, 'Autoencoder/DBN [128]' row: reference [128] is Sainath et al., 'Deep convolutional neural networks for LVCSR' (ICASSP 2013), a paper about deep convolutional networks for speech recognition, not about autoencoders or deep belief networks. The row also lists a 15.5% word error rate on English Broadcast News while the cited paper's experiments concern a different setup. This misattribution in a table that summarizes state-of-the-art speech models materially weakens the survey's reliability.","section":"Table III"},{"comment":"Section 3.2, paragraph on LSTM training: the sentence 'the development of long short-term memory (LSTM) networks that use special hidden units known as \"gates\" to retain memory over longer portions of a sequence' is cited to [40], which is Alam et al., 'Novel hierarchical Cellular Simultaneous Recurrent Neural Network for object detection' (IJCNN 2015), an object-detection paper with no LSTM gating content. The same paragraph later attributes statements about memory-network stories to [99] (Huang et al., DenseNet) and about translation difficulties to [101] (Tan and Le, EfficientNet), neither of which discusses those topics. These citation mismatches in the main text suggest a systematic reference-integrity problem.","section":"Section 3.2"},{"comment":"The abstract and Sections 1 and 7 claim the paper is 'one of the most comprehensive surveys' of vision and speech deep learning, but the paper provides no search strategy, inclusion criteria, database list, or date range for the literature covered. The selection appears to be informal, and the citation errors documented above indicate that the covered literature is not always accurately represented. Without a stated methodology, the comprehensiveness claim cannot be independently checked, and the current reference errors further undermine it.","section":"Sections 1 and 7"}],"minor_comments":[{"comment":"Table II: 'Tomson et al.' should be 'Tompson et al.'","section":"Table II"},{"comment":"Section 3.3: the CIFAR-10 description says 'there are 10 classes with 60,000 images each'; CIFAR-10 has 60,000 images in total (6,000 per class), so this needs rewording.","section":"Section 3.3"},{"comment":"Section 3.3: 'TIMT' is a typo for TIMIT.","section":"Section 3.3"},{"comment":"Section 3.1: 'Neural Architecture Search (described in Section 2.8)' should refer to Section 2.9, which is the Neural Architecture Search section.","section":"Section 3.1"},{"comment":"Table III footnote: 'PERPEPLEXITY' is a typo for 'PERPLEXITY.'","section":"Table III footnote"},{"comment":"Figures 3, 4, and 6 are difficult to read in the PDF; higher-resolution images would improve clarity.","section":"Figures 3, 4, and 6"}],"recommendation":"major_revision","confidential_remarks":"The large number of misattributed references suggests the reference list was not carefully checked against the bibliography. I recommend the editor require a systematic reference audit rather than case-by-case fixes. The lack of methodology for the comprehensiveness claim is also worth addressing. The paper has a reasonable structure, but its utility as an authoritative survey is currently compromised."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this survey is that it is a useful orientation map for someone entering vision/speech deep learning, but it is not a reliable reference as it stands. The organizational thread—software and hardware perspectives, with attention to resource-constrained deployment—is genuinely helpful. The architecture background (CNN, DBN, SAE, VAE, GAN, flow, RNN, attention, NAS) is presented clearly, and the emerging applications sections (behavioral, transportation, medicine) are broad and mostly accurate in narrative.\n\nThat said, the citation apparatus has real problems that the stress-test note correctly identifies. Table I attributes ResNet to reference [71], which is Waveglow—a speech synthesis paper. The text cites ResNet as [97], which is 'Wider or deeper' by Wu et al., not the original He et al. paper; He et al. appears only as [98] for a different claim. Table III lists 'Autoencoder/DBN [128]' but [128] is Sainath et al.'s deep CNN for LVCSR. In Section 3.2, LSTM gating is cited to [40], which is an object-detection paper by the first author. These are not peripheral typos; they are mis-attributions in exactly the tables a reader would trust for state-of-the-art numbers. The comprehensiveness claim in the abstract and conclusion also lacks a stated search strategy, inclusion criteria, or date range, so 'one of the most comprehensive surveys' is an unsupported overclaim.\n\nOn the positive side, the survey does a decent job of summarizing the evolution of models and highlighting mobile and embedded constraints, an angle many surveys neglect. The limitation sections (sample size, computational burden, interpretability, over-optimism) are sensible and balanced. The paper is not a random collection; it is organized with intent.\n\nFor a newcomer, the survey could be a starting point, but I would not hand it to a student without warning about the citations. The errors are fixable, and the underlying scope is legitimate. I would send it to peer review with a request for thorough revision and verification of every table entry, not desk-reject it.\n\nWould I cite it? Probably not, given the reference issues. Would I bring it to reading group? Maybe, as a cautionary example of why citation hygiene matters in surveys. But for a reliable reference, wait for a corrected version.","headline":"A well-organized survey with a useful mobile-deployment angle, but citation errors in the summary tables undermine its reference value as published.","tokens_in":35145,"tokens_out":2767,"would_cite":false,"duration_ms":25237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey's central claim is that it is one of the most comprehensive accounts to date of deep learning for vision and speech, spanning architectures, benchmark results, industrial systems, and deployment on phones and embedded devices.","keywords":["deep learning","computer vision","speech recognition","convolutional neural networks","recurrent neural networks","attention mechanisms","model compression","embedded systems"],"falsifier":"Check the survey's benchmark tables against the original papers: if a reader can list significant 2019-era vision or speech systems, architecture families, or record results that the survey omits, or finds reported numbers misattributed, the comprehensiveness claim fails. One concrete discrepancy is already visible: Table I's ResNet row is keyed to a reference that the reference list assigns to a speech-synthesis paper, while the running text keys ResNet to a different reference entirely.","tokens_in":34197,"feed_emoji":"🧠","tokens_out":15685,"duration_ms":131524,"temperature":0.7,"pith_summary":"This survey tries to be the go-to reference for deep learning in vision and speech as of 2019: one organized account of the architecture families, the applications they dominate, the industrial systems that deploy them, and the effort to shrink those models onto phones and embedded devices. A reader should care because the paper's value is precisely its breadth, connecting algorithmic evolution to benchmark numbers and to the memory, battery, and compute budgets that decide whether a model can leave the server. The survey's central contention is that deep learning won by replacing hand-engineered features with end-to-end trainable hierarchies, and that the next frontier is efficiency rather than raw accuracy. It further argues that these systems are already reshaping behavioral science, transportation, and medicine, while three weaknesses—data hunger, computational burden, and black-box opacity—limit how far they can go.","feed_headline":"Survey maps deep learning across vision, speech, and mobile hardware","feed_subtitle":"Architectures, benchmark results, phone-chip deployment, and emerging medical and transport uses in one reference.","key_machinery":"The central object is the survey's organizing taxonomy: deep architectures sorted by learning paradigm (convolutional, recurrent, generative, attention-based, and machine-searched), each mapped to its signature applications and benchmark results, and then re-examined under hardware constraints. What carries the argument is the pairing of every architecture family with concrete resource costs on the deployment side, so that the survey can claim coverage along both a software axis and a hardware axis. The load-bearing instruments are the comparative tables: they assert specific, checkable numbers such as AlexNet's 17.0% top-5 error on ImageNet, a roughly seven-fold parameter reduction for MobileNets at about 1% accuracy loss, and memory cuts from 59.1 MB to 3.2 MB for a compressed speech recognizer.","core_discovery":"The paper's claim, stated in its abstract and conclusion, is that it delivers one of the most comprehensive surveys of the latest developments in intelligent vision and speech systems, from both software and hardware perspectives. On the software side, it traces the architecture families—convolutional networks, deep belief networks and stacked autoencoders, variational autoencoders, generative adversarial networks, flow models, recurrent networks and LSTMs, attention mechanisms, and neural architecture search—and ties each to its state-of-the-art results in tasks such as ImageNet classification, face, action, and pose recognition, speech recognition, and image generation. On the hardware side, it catalogs the techniques that shrink these models onto resource-restricted platforms: pruning, quantization, structured matrices, low-rank compression, and efficient designs such as depthwise-separable MobileNets, together with measured memory and energy trade-offs. It then argues that these systems are already reshaping behavioral science, intelligent transportation, and precision medicine, while naming data hunger, computational cost, and black-box interpretability as the limits that remain.","pith_inferences":["The survey's 2019 taxonomy ends with attention as an emerging alternative to recurrence in sequence tasks; the same pattern it documents is what later carried attention-based machinery into vision, so the survey's framing helps a reader understand that migration.","Because the survey states no inclusion criteria, its tables are best treated as pointers to the original papers rather than audited figures; a reader who needs decision-grade numbers should verify each entry before relying on it.","The hardware-side tension the survey documents—accurate models are too large to deploy, efficient models lose accuracy—implies that future benchmarks will increasingly report accuracy per unit of memory or energy rather than accuracy alone.","The survey's comprehensiveness claim is time-stamped: the field moved quickly after 2019, yet the architecture-to-application structure the survey builds is portable and can be re-populated with newer results as a living map."],"forward_implications":["If the survey's coverage holds, a newcomer can use it as a first-level map of where vision and speech deep learning stood in 2019, including dominant architectures, benchmark numbers, and the main industrial systems.","The survey's emphasis on resource-constrained deployment implies that the next wave of intelligent vision and speech systems will be defined less by raw accuracy and more by the ability to run within limited memory, battery, and compute budgets.","Its catalog of small-footprint results—keyword spotting, mobile speech recognition, and compact CNNs with order-of-magnitude parameter reductions—implies that a growing range of vision and speech applications can run on-device rather than in the cloud.","The emerging-applications sections argue that automated behavioral analysis, intelligent transportation, and precision medicine will be early high-impact adopters of these systems.","The survey's stated limitations—data hunger, computational burden, and black-box opacity—imply that progress in small-data learning, 3D and 4D data handling, and hardware-software co-design will determine how widely the technology spreads."],"supporting_citations":[{"why":"Supplies the greedy layer-wise pretraining algorithm that the survey credits as the breakthrough enabling deep belief networks and the later deep learning wave.","marker":"[28]"},{"why":"The AlexNet ImageNet result that anchors the survey's CNN story and serves as the baseline for later architecture comparisons and mobile-deployment studies.","marker":"[33]"},{"why":"The Transformer attention model the survey cites to argue that attention-based processing has overtaken recurrent networks for sequential and context-based tasks.","marker":"[39]"},{"why":"The ImageNet dataset and challenge that the survey uses as the benchmark yardstick for its image classification narrative and tables.","marker":"[94]"},{"why":"The reference the survey uses for ResNet and its residual-block design, which it credits with enabling very deep CNNs that solve the vanishing-gradient problem.","marker":"[97]"},{"why":"The multi-institution acoustic modeling study the survey marks as the milestone that established deep networks in large-vocabulary speech recognition.","marker":"[127]"},{"why":"The Deep Speech 2 end-to-end system that anchors the survey's account of large-scale industrial speech recognition and its near-human results.","marker":"[159]"},{"why":"The MobileNets architecture with depthwise-separable convolutions that anchors the survey's account of efficient CNNs for mobile and embedded vision.","marker":"[179]"},{"why":"The three-stage pruning, quantization, and Huffman coding pipeline the survey uses to show how deep models can be compressed for low-memory, low-power deployment.","marker":"[181]"}],"fun_headline_variants":["From CNNs to edge: full-stack deep learning survey for vision and speech","One survey spans DNN architectures, hardware tricks, and vision-speech apps","Deep learning for sight and sound: software to silicon in a single guide","Vision and speech AI survey covers model families to mobile efficiency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim of comprehensiveness rests on the assumption that its informal selection of papers and the numbers it reports faithfully represent the state of the art at the time, since the paper states no search strategy, inclusion criteria, or date range that would let a reader check the coverage.","fun_headline_variants_meta":{"raw":{"variants":["From CNNs to edge: full-stack deep learning survey for vision and speech","One survey spans DNN architectures, hardware tricks, and vision-speech apps","Deep learning for sight and sound: software to silicon in a single guide","Vision and speech AI survey covers model families to mobile efficiency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000322,"raw_usage":{"total_tokens":1845,"prompt_tokens":1014,"completion_tokens":831,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":752}},"tokens_in":630,"tokens_out":831,"duration_ms":8013,"temperature":1.0,"reasoning_tokens":752,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:56:40.860895+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the survey's benchmark tables against the original papers: if a reader can list significant 2019-era vision or speech systems, architecture families, or record results that the survey omits, or finds reported numbers misattributed, the comprehensiveness claim fails. One concrete discrepancy is already visible: Table I's ResNet row is keyed to a reference that the reference list assigns to a speech-synthesis paper, while the running text keys ResNet to a different reference entirely.","supporting_citations":[],"review_version":1}