{"id":"009f0717-e9c2-4763-aa63-90763a1e61b9","arxiv_id":"2412.06044","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review that aggregates and compares major cloud providers' generative AI tools and services, with no new empirical results.","lead":"This paper reviews cloud platforms for building generative AI applications, comparing services from AWS, Azure, Google Cloud, IBM, Oracle, and Alibaba. It organizes compute, serverless, edge, storage, data, API, security, and cost offerings into tables, SWOT analyses, and recommendations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comparative tables contain verifiable factual errors (e.g., Table 3 reports 16 H100 GPUs for AWS P5 and Google A3; official specs list 8), so the review's central comparative claim is not dependable until audited.","rationale":"The paper is a scoping review, not an experimental study, and its central contribution is a comparative guide for practitioners. Read charitably, the claim is that the review maps current cloud services and critically compares them. For that claim to hold, the comparative tables must be accurate. The reader flagged reliance on vendor materials as the weakest assumption. I agree that is a risk, but I found a sharper and more decisive version: the tables contain transcription errors that are checkable against primary sources. Table 3's H100 counts are off by a factor of two for AWS and Google; Section 2.2.2 has unresolved '?' citation placeholders; Sections 4.5.4 through 4.5.6 are duplicated and mislabeled. These are not merely style issues; they affect the information practitioners would act on. A single corrected table would not rescue the argument unless all comparative data are audited. I do not think this warrants rejection because the review's organizational framework, supplementary matrices, and source distribution analysis are useful and corrigible. However, without a factual audit, the central claim of 'critical analysis' is not supportable. The reader's CONDITIONAL verdict already captures the need for revision, so my read does not move the verdict; it strengthens the conditions that should be attached to acceptance.","tokens_in":43408,"tokens_out":5757,"duration_ms":57385,"concrete_test":"Verify Table 3 against current official AWS EC2 P5 and Google Cloud A3 product pages, counting the maximum number of NVIDIA H100 GPUs per instance. If both are 8 rather than 16, the table is erroneous. Then re-audit all comparative tables (Tables 2 through 12) against official provider documentation and independent benchmarks; if more than a small number of entries fail, the central comparative claim should be downgraded or the paper revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is a comprehensive, critical comparison of cloud platforms for generative AI. That claim stands or falls on the accuracy of the comparative tables in Sections 4 and 5. Several entries are demonstrably wrong or unsupported. In Table 3, AWS EC2 P5 instances and Google Cloud A3 instances are each listed with '16 NVIDIA H100 GPUs'; current official product documentation specifies 8 H100 GPUs per instance for both. Table 2 similarly conflates instance families (AWS P4d/P3/G5) under a single 'Max GPU per Instance' column. The manuscript also contains missing-citation placeholders ('?') in Section 2.2.2, duplicated and mislabeled subsections (4.5.4/4.5.5/4.5.6), and references that do not support the statements they are attached to (e.g., [7] and [8] in Section 4.5.4). If a reader uses the tables to select a provider, a single wrong GPU count changes capacity planning. This is not a matter of vendor bias alone; the comparative conclusions are internally unreliable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a scoping review of cloud platforms for developing generative AI solutions, covering AWS, Microsoft Azure, Google Cloud, IBM Cloud, Oracle Cloud, and Alibaba Cloud, with additional sections on Databricks and Snowflake. The authors aim to critically analyze and compare provider offerings across compute, serverless, edge, storage, data, security, cost, and AI-specific services, and to provide practitioner guidance. The manuscript includes extensive tables, figures, supplementary material, SWOT analyses, and a bibliometric analysis of 255 references. The central claim is that the review provides a reliable comparative guide to selecting cloud providers for generative AI.","tokens_in":43616,"tokens_out":5409,"duration_ms":46868,"significance":"If the comparative content were accurate, the paper would be a useful broad map of the current cloud AI landscape for practitioners and researchers. The authors cover an unusually wide set of providers and service categories, and the supplementary matrices, SWOT analyses, and reference-distribution analysis are helpful organizational devices. The paper does not present new empirical measurements, but it could serve as a starting point for provider selection if the factual basis is repaired. However, the paper's value as a 'critical' comparative review depends on the correctness of its central tables and the independence of its sources; in its present form that value is substantially compromised by the errors and citation problems identified below.","major_comments":[{"comment":"Table 3 lists AWS EC2 P5 instances and Google Cloud A3 instances as each having 16 NVIDIA H100 GPUs; the official AWS and Google documentation specifies 8 H100 GPUs per instance for both. The accompanying text in Section 4.1.1 then concludes that 'AWS and Google Cloud lead in raw GPU capacity with their latest offerings supporting up to 16 NVIDIA H100 GPUs per instance,' so this error directly supports a load-bearing comparative claim. The authors should correct Table 3 and re-check all GPU counts in Tables 2 and 3 against vendor documentation.","section":"Table 3, §4.1.1"},{"comment":"Section 2.2.2 contains literal question marks in place of citation numbers in all five bullet points (e.g., 'cost-effectiveness ?', 'AWS EC2 P4d instances ?', 'Amazon SageMaker Autopilot ?'). A scoping review that describes key cloud capabilities for generative AI cannot leave its central claims uncited; this indicates an incomplete reference pass and needs to be fixed systematically. The Section 1.3 methodology promises a 'systematic and rigorous analytical approach,' which these placeholders contradict.","section":"Section 2.2.2"},{"comment":"Sections 5.2 and 5.3 are near-verbatim duplicates ('Security Aspects' and 'Deployment and Orchestration Challenges' contain the same text with different reference numbers), and Tables 8 and 9 are identical duplicated tables with different captions. In addition, Section 4.5.6 is titled 'Monitoring and Observability' but its content discusses interoperability and ethical AI, while Section 4.5.7 is titled 'Challenges and Future Directions' but discusses API-accessible services. These structural problems make the manuscript internally inconsistent and must be resolved before the review can be considered reliable.","section":"Sections 5.2, 5.3, 4.5.6, 4.5.7, Tables 8 and 9"},{"comment":"In Section 4.5.4, the claims about serverless inference and model compression are supported by references [7] and [8]; [7] is an image-captioning paper and [8] is the Arksey and O'Malley scoping-review methodology paper, neither of which supports the statements. Similar citation mismatches appear elsewhere (e.g., reference [54] is labeled 'Azure AI' but points to a Google Cloud URL). Because the paper's comparative claims are only as credible as its citations, the authors need to conduct a full citation audit.","section":"Section 4.5.4"},{"comment":"The paper's comparative conclusions depend heavily on provider-authored or vendor-adjacent sources: Section S5.1 reports that 56.1% of the 255 references are web resources and only 5.5% are cloud provider documentation, with vendor blogs counted inside the web category. A central example is the claim in Section 3.2.2 that Azure's 'exclusive partnership with OpenAI' allows it to 'take the lead among managed AI services,' which is supported by Microsoft's own Azure OpenAI page. For a review that promises to 'critically analyze' cloud platforms, the authors must either add independent third-party benchmarks and analyst assessments or substantially qualify conclusions that are based on vendor marketing material.","section":"Sections 3.2.2 and S5.1"}],"minor_comments":[{"comment":"The paragraph in Section 1.2 repeats the same statement about 60% of enterprises planning to integrate generative AI by 2025 twice in consecutive sentences; one occurrence should be removed.","section":"Section 1.2"},{"comment":"Table 2's 'Max GPU per Instance' column mixes different instance families (P4d, P3, G5) under one entry, which is misleading; Table 6 reports 'Max Throughput' values such as 50 Gbps for Azure Blob and 240 Gbps for Google Cloud Storage without clear definitions or sourcing. These tables need per-instance detail and cited units.","section":"Tables 2 and 6"},{"comment":"There are numerous typographical issues: 'T able' appears in table captions, 'F uture' appears in headings, 'Iransformer' appears in Figure 2, 'GAteway' appears in Figure 10, and Section 5.1 is titled 'Performance Analysis' but begins with text about security concerns. These should be corrected in a careful copyedit.","section":"Figures and headings"},{"comment":"The first paragraph of Section 3.2.2 contains a duplicated sentence fragment ('Microsoft Azure is the second-largest cloud provider, renowned for its enterprise-friendly environment and deep Microsoft Azure is the second-largest...'), which needs to be rewritten.","section":"Section 3.2.2"},{"comment":"Reference [54] is labeled 'Azure AI' but points to a Google Cloud URL, and several web references in the bibliography lack access dates; the reference list should be standardized.","section":"Reference list"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a draft rather than a submission-ready review. The duplicated sections, question-mark citations, and mislabeled headings suggest the final editing pass was not completed. Given the central claim is a comparative guide, I would ask the authors to correct the tables, fill the citations, and rework the source base before considering it for publication. If the authors cannot provide corrected GPU counts and independent sources, the paper's value as a critical review is questionable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a broad but uneven scoping review. It could be a useful first map of the cloud-genAI vendor landscape, but the central comparative tables contain verifiable factual errors and the evidence is heavily vendor-sourced, so no one should rely on it without a major audit.\n\nWhat it does well: it covers a lot of ground—compute, serverless, edge, storage, data, security, cost, and provider overviews, with SWOT analyses and a “Modern AI Stack” diagram. It pulls together a huge amount of vendor documentation into one place. The supplementary bibliometric breakdown is a nice extra. For someone who wants a quick orientation to what AWS, Azure, GCP, IBM, Oracle, Alibaba, Databricks, and Snowflake offer for generative AI, this saves time.\n\nThe problems are real. Table 3 lists 16 H100 GPUs for both AWS P5 and Google A3; official specs say 8. Table 2 conflates instance families under a single max-GPU column. That is not a peripheral detail—the paper’s whole purpose is comparative evaluation, and a wrong GPU count changes capacity planning. Section 2.2.2 has literal “?” placeholders instead of citations. Sections 5.2 and 5.3 are duplicated, as are Tables 8 and 9. Subsection headings 4.5.6 and 4.5.7 are mislabeled. The reference base is 56% web resources, and much of that is vendor marketing. The claim that Azure “took the lead” among managed AI services is backed by Microsoft’s own pages. The methodology is described in general terms but without a reproducible search protocol, so it falls short of a rigorous scoping review.\n\nWho is this for? A practitioner who wants a broad orientation and understands the tables need checking. Not for a decision-maker who will rely on the numbers. It could be a decent starting point after heavy revision.\n\nRecommendation: yes, send it to peer review—the scope is useful and the field needs maps like this. But I would expect major revisions: fact-check all tables, remove duplications, replace the “?” citations, add a methods appendix, and explicitly discuss the vendor-sourced evidence as a limitation. Without that, it is a draft, not a review.","headline":"A broad but sloppy scoping review of cloud platforms for generative AI: useful as a first orientation, but the comparative tables contain verifiable factual errors and the evidence is heavily vendor-sourced, so it needs major revision before anyone should rely on it.","tokens_in":44160,"tokens_out":2584,"would_cite":false,"duration_ms":24170,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scoping review argues that major cloud platforms offer converging but differently positioned toolchains for generative AI, and that the right choice depends on workload-specific needs.","keywords":["Generative AI","Cloud computing","Scoping review","High-performance computing","Serverless architectures","Edge computing","Vendor comparison","Cloud security"],"falsifier":"A standardized benchmark, running the same generative-AI model, dataset, and request pattern on AWS, Azure, Google Cloud, IBM, Oracle, and Alibaba with independent measurement of cost, latency, and throughput, would confirm or overturn claims such as Azure's lead in managed AI services and Google's TPU performance advantages.","tokens_in":43217,"feed_emoji":"☁️","tokens_out":5531,"duration_ms":52907,"temperature":0.7,"pith_summary":"This paper is a scoping review of cloud platforms for building generative AI applications. It argues that AWS, Microsoft Azure, Google Cloud, IBM, Oracle, and Alibaba Cloud each assemble a full generative-AI stack, covering compute, serverless, edge, storage, data, AI services, security, and cost, but position those pieces differently. The review claims Azure currently leads managed AI services through its OpenAI partnership, while AWS and Google Cloud lead in raw infrastructure, and IBM leads in responsible and enterprise AI; Oracle and Alibaba target specific enterprise and regional niches. If the review is right, an organization's cloud choice matters less at the level of basic capability and more at the level of fit, meaning which provider's strengths align with its scale, compliance, and cost constraints.","feed_headline":"Six cloud providers compared for generative AI buildouts","feed_subtitle":"AWS, Azure, Google, IBM, Oracle, and Alibaba each lead in a different slice; here is where each fits.","key_machinery":"The carrying mechanism is a comparative framework that evaluates providers along eight dimensions: high-performance computing and GPUs, serverless architectures, edge computing, storage and data lakes or warehousing, generative-AI development ecosystems, API-accessible AI services, security, and cost. Named in the paper as the 'comparative framework,' it organizes service matrices, capability tables, and SWOT analyses for AWS, Azure, Google Cloud, IBM, Oracle, and Alibaba Cloud into a single decision-oriented map. The framework's work is to turn a large collection of vendor offerings into comparable columns so that strengths and weaknesses can be weighed side by side, while the review uses Arksey and O'Malley's scoping-review methodology to bound the literature search.","core_discovery":"The paper's central claim is that the cloud ecosystem for generative AI can be systematically mapped and compared across service categories, and that this mapping reveals genuine, decision-relevant differences among providers. On the paper's own terms, the discovery is that every major provider now offers an end-to-end generative-AI stack, so competition has shifted from availability to emphasis: Azure's exclusive OpenAI partnership gives it the lead in managed AI services; AWS and Google Cloud dominate scalability, infrastructure, orchestration, and security; IBM leads in ethical and explainable AI; and Oracle excels in enterprise integration. The review presents these as comparative findings supported by service matrices, SWOT analyses, and market-share data, and it recommends that enterprises match provider strengths to their own needs, use hybrid or multi-cloud strategies to reduce lock-in, and conduct their own total-cost-of-ownership analysis before choosing.","pith_inferences":["Because most references are vendor web pages and provider documentation, the 'leadership' claims, such as Azure's managed-AI lead, are closer to vendor positioning than to independent measurement; neutral benchmarks would be needed to confirm them.","The review's structure implies a testable extension: run the same generative-AI workload, with the same model, data, and traffic pattern, on all six providers and compare latency, throughput, and cost, which would turn the comparative map into a quantitative ranking.","The inclusion of Databricks, Snowflake, and the 'Modern AI Stack' suggests the unit of comparison may soon shift from cloud provider to toolchain, with providers becoming one layer among GPU, vector database, orchestration, and observability vendors.","One consequence the authors leave implicit is that specific service names in the tables will date quickly, while the comparative dimensions themselves will remain useful for future updates."],"forward_implications":["Enterprises can use the paper's service matrices to shortlist cloud providers by the specific needs of a generative-AI project rather than by general reputation.","Azure's exclusive partnership with OpenAI implies that teams wanting frontier GPT models through a managed API will most often land on Azure, while teams wanting custom training at scale may start with AWS or Google Cloud.","The paper's recommendation of hybrid and multi-cloud strategies implies that vendor lock-in is best treated as a design constraint from the start, not as a later migration problem.","Because cost models differ by usage pattern, the paper's cost analysis says organizations should run their own total-cost-of-ownership study before committing to a provider.","Security, compliance, and explainability tools now exist across all major providers, so the differentiator is which provider's governance features match an organization's regulatory environment."],"supporting_citations":[{"why":"Supplies the scoping-review methodology the paper follows.","marker":"[8]"},{"why":"Provides the cloud-AI market size and CAGR figures that frame the review.","marker":"[93]"},{"why":"Supplies the overall cloud market-share percentages that anchor the provider ranking.","marker":"[94]"},{"why":"Provides the AI and ML market-share and usability claims for individual providers.","marker":"[95]"},{"why":"Documents Azure OpenAI's availability and the OpenAI partnership behind Azure's lead claim.","marker":"[96]"},{"why":"Supplies the Google Cloud TPU v4 performance comparison cited in the performance analysis.","marker":"[57]"},{"why":"Documents AWS EC2 P5 H100 hardware used in the compute comparison.","marker":"[106]"},{"why":"Documents Google Cloud A3 H100 instances used in the compute comparison.","marker":"[108]"}],"fun_headline_variants":["Cloud AI showdown: six providers, six strengths","Generative AI cloud race: from access to emphasis","AWS, Azure, Google, IBM, Oracle, Alibaba: AI stack comparison","Which cloud for generative AI? Scoping review points the way","Cloud platforms for generative AI: a comparative guide"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's comparative conclusions assume that the vendor documentation, marketing pages, and industry web resources that make up a majority of its 255 references are accurate, unbiased, and current enough to support judgments about which provider leads.","fun_headline_variants_meta":{"raw":{"variants":["Cloud AI showdown: six providers, six strengths","Generative AI cloud race: from access to emphasis","AWS, Azure, Google, IBM, Oracle, Alibaba: AI stack comparison","Which cloud for generative AI? Scoping review points the way","Cloud platforms for generative AI: a comparative guide"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1206,"prompt_tokens":942,"completion_tokens":264,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":182}},"tokens_in":558,"tokens_out":264,"duration_ms":3722,"temperature":1.0,"reasoning_tokens":182,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:03:26.043267+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A standardized benchmark, running the same generative-AI model, dataset, and request pattern on AWS, Azure, Google Cloud, IBM, Oracle, and Alibaba with independent measurement of cost, latency, and throughput, would confirm or overturn claims such as Azure's lead in managed AI services and Google's TPU performance advantages.","supporting_citations":[],"review_version":1}