{"id":"b462b7f4-5a52-4cea-9201-e1612f041a8f","arxiv_id":"2411.10408","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A grounded-theory review of 144 metaverse records produces three unvalidated readiness models covering user experience, business, and national competitiveness.","lead":"A systematic review of 144 academic and industry texts about the metaverse, coded into three conceptual models for user experience, business readiness, and national competitiveness. Why might a smart generalist read it? It offers a compact, multi-level checklist for thinking about how the metaverse could affect people, companies, and countries.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No reproducible audit trail for the n=144 corpus: the models' empirical grounding cannot be checked, and the coding categories may be predetermined by the search terms.","rationale":"The reader's weakest assumption is that the corpus is representative enough to ground the models, citing the missing PRISMA counts, absent search dates and query sets, non-reproducible LinkedIn sampling, and no included-record list. I agree with that identification: it is the most load-bearing weakness because the three models are explicitly presented as the result of analyzing n=144 records, and if the record set cannot be audited or reproduced, the factor categories and the models built from them lack a verifiable empirical foundation. My reading sharpens the concern by pointing to the leaked 'RAYYAN-INCLUSION' annotations in the reference list, which show that screening records exist but were not reported, and by noting that Table 1's search terms already encode the Micro-Meso-Macro structure, so the models' top-level organization is partly predetermined by the search strategy rather than purely emergent. These issues do not change the reader's conditional verdict: the paper can still be made acceptable by supplying the missing methodological audit trail and tempering the claim that the models are grounded-theory discoveries. The proposed test is therefore a reproducibility audit of the corpus and its screening counts, which would settle whether the empirical grounding holds.","tokens_in":28103,"tokens_out":3256,"duration_ms":36917,"concrete_test":"Publish the PRISMA 2020 flow diagram with exact counts (records identified, screened, excluded by criterion, duplicates removed, included), the full search strings, search dates, and the complete list of the 144 included records or a public dataset. Then independently re-run the searches using those stated terms and dates; if the same 144 records cannot be recovered, or the flow counts do not sum to 144, the corpus-based model derivation is not reproducible and the models' empirical grounding fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that three research models emerge from grounded theory analysis of n=144 white- and grey-literature records (Abstract; §3). For that claim to hold, the corpus must be retrieved reproducibly and the coding must genuinely derive from it. Section 3.2 reports 'first fifty results query searching on LinkedIn' and Google Scholar searches, but gives no search dates, no complete query strings, no PRISMA screening counts, and no list of included records. Figure 1's flow diagram cannot be inspected in the supplied text. The bibliography also contains leaked screening-tool annotations, e.g. 'RAYYAN-INCLUSION: \"Amir\"=¿\"Included\"' in references [46], [52], [56], [105], and [130], indicating that inclusion decisions were made in a tool but the documented screening trail was not reported. Without the full PRISMA flow, the claimed n=144 is unverifiable; and because LinkedIn results are personalized and Google Scholar results change over time, the corpus is not reproducible. If the retrieved record set is biased or incomplete, the six UX factors, four business factors, and four national factors that constitute the models will not generalize. A secondary concern is that Table 1's search terms are already organized by the three target levels ('User Experience (Micro)', 'Business Readiness (Meso)', 'Governance (Macro)'), so the three-model structure is built into the search strategy rather than purely emerging from grounded theory; this weakens the paper's claim of an emergent theoretical derivation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic literature review of 144 white- and grey-literature records, analyzed with grounded theory, to produce three conceptual models: a Metaverse User Experience Maturity Model, a Metaverse Business Readiness Model, and a Metaverse National Competitiveness Model. The authors frame the metaverse as a sociotechnical imaginary and organize the review by micro (user), meso (business), and macro (national) levels, with heuristic rubrics and checklists for practical use. The paper is qualitative and integrative, covering user-level factors such as accessibility, flexibility, gamification, interoperability, social interaction, and empowerment; business factors such as digital twin integration, marketing, product/service development, and resources; and national factors such as policy, service development, human capital readiness, and sociocultural influence.","tokens_in":28337,"tokens_out":3840,"duration_ms":39203,"significance":"If the synthesis is trustworthy, the three models and their associated assessment tools would be a useful contribution to HCI and information science, bridging academic and industry perspectives and giving designers, strategists, and policymakers a structured starting point. The paper is openly self-critical: it acknowledges the interpretive nature of grounded theory, the lack of validation of the proposed rubrics, and the need for inter-rater reliability. The practical appendices are detailed and could seed future empirical work. However, the significance is conditional on the credibility of the systematic review corpus and on whether the model structure genuinely derives from the literature rather than being imposed by the search design.","major_comments":[{"comment":"The corpus is not reproducible. The text states that the study collected 'first fifty results query searching on LinkedIn' and searched Google Scholar, but it does not provide search dates, complete query strings, database/platform details, or the PRISMA screening counts (identified, screened, excluded, included). The flow diagram in Figure 1 is not inspectable in the supplied text, and no list of the 144 included records is given. This is load-bearing because the three models are claimed to emerge from this corpus. The reference list contains leaked inclusion annotations such as 'RAYYAN-INCLUSION: \"Amir\"=¿\"Included\"' in refs. [46], [52], [56], [105], and [130], which suggest the screening was performed in a tool but the documented screening trail was not reported. Please provide a complete PRISMA flow table, the full search strings with dates, and an appendix listing the 144 records.","section":"§3.2–3.3, Figure 1"},{"comment":"The claim of grounded-theory emergence is undermined by the search-term design. The search terms in Table 1 are already grouped by the three target levels ('User Experience (Micro)', 'Business Readiness (Meso)', 'Governance (Macro)'), and Section 3.1 declares the micro-meso-macro analytical framework before data collection. The coding process in Section 3.4 then produces categories that map neatly onto those three levels. The paper does not explain how the grounded-theory analysis could have revised, rejected, or restructured the three-level typology. Please revise the presentation to acknowledge the hybrid deductive–inductive nature of the study, or provide evidence that the coding was open to alternative level structures.","section":"Table 1, §3.1, §3.4"},{"comment":"The national-competitiveness model is the least transparently derived. The text says 'The coding processes is presented in table??' and the actual reference to Table 6 is missing. The selective-coding step is described as using WEF terminologies of network readiness and competitiveness, but the open-to-axial-to-selective mapping is less clearly evidenced than in Tables 4 and 5. Additional coding examples or a more explicit trace from the codes to the final four national factors are needed to support the model as an empirical synthesis.","section":"§4.3, Table 6"},{"comment":"The LDA topic modeling of titles is reported but never connected to the grounded-theory analysis or the three research questions. Table 3 lists keyword sets for five topics, but there is no interpretation of these topics, no coherence measure, and no explanation of how they informed the models. As it stands, the LDA section is a disconnected descriptive addition; either integrate it into the analysis narrative or remove it.","section":"§4, Table 3"}],"minor_comments":[{"comment":"The cross-reference to 'table??' should be 'Table 6'.","section":"§4.3"},{"comment":"The text '144.2.2Marketing' appears to be a numbering artifact; the stray '14' should be removed.","section":"§4.2.2"},{"comment":"The leaked 'RAYYAN-INCLUSION' annotations in several reference entries, e.g. refs. [46], [52], [56], [105], and [130], should be removed from the reference list.","section":"References"},{"comment":"'Data availability: Not Applicable' is not appropriate for a systematic literature review. If the corpus cannot be shared, the authors should state this explicitly and justify the exception; otherwise, a supplementary list of the 144 records is expected.","section":"Declarations"},{"comment":"There are numerous typographical and spacing errors (e.g., 'sociotechnichal' in Section 2, 'guidedbyof', 'glasserian' in Section 5.4). The manuscript needs a full copyedit.","section":"Throughout"},{"comment":"The same paper by Bibri (2022) appears to be cited twice as separate references; this should be consolidated.","section":"References [13] and [14]"}],"recommendation":"major_revision","confidential_remarks":"The leaked RAYYAN annotations in the reference list, while probably an artifact of the screening workflow, are the kind of detail that could raise reviewer suspicion about reporting integrity; I would recommend the editor insist on a clean, reproducible PRISMA audit trail before further consideration. The manuscript's novelty is moderate for the journal, but the integrative three-level framing is a reasonable contribution if the method is made defensible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. This paper builds three maturity/readiness models (UX, business, national) from a mixed white/grey literature corpus, and the models themselves are sensible and well organized. But the empirical grounding for the claimed systematic review is not visible, and the three-level structure does not really 'emerge' from the data. Treat it as a thorough scoping synthesis, not as a validated grounded-theory derivation.\n\nWhat's good: the paper engages with existing maturity models (Radoff, Digitopia, Weinberger & Gross), spans academic and industry sources, and produces concrete heuristic rubrics and checklists that are explicitly flagged as unvalidated. The six UX factors, four business factors, and four national factors are a reasonable reading of the current discourse. The LDA topic model is a nice supplementary touch, and the Python script is provided.\n\nThe soft spots are real. The paper invokes PRISMA but reports no screening counts, no search dates, no complete query strings, and no list of the 144 included records. The LinkedIn sampling ('first fifty results') is not reproducible, and Google Scholar results are personalized; the corpus cannot be reconstructed. The reference list leaks 'RAYYAN-INCLUSION' annotations, which tells me screening was done in a tool but the audit trail was not reported. On top of that, Table 1's search terms are already grouped into User Experience (Micro), Business Readiness (Meso), and Governance (Macro), so the micro-meso-macro model structure is built into the search strategy. The grounded theory coding then reproduces that scaffolding. The paper acknowledges researcher bias and lack of inter-rater reliability in the Limitations section, which is honest, but the reporting of the corpus itself is the bigger problem.\n\nThe abstract also overstates readiness: the 'practical assessment tools' are heuristics and checklists that the authors themselves say are not yet validated.\n\nWho gets value: anyone wanting a compact multi-level checklist of metaverse factors for design, strategy, or policy brainstorming. Not a reliable empirical foundation for further systematic claims.\n\nOn peer review: I'd send it to referees, but with a clear request that the authors supply the full PRISMA flow, the exact queries, dates, and the complete list of included records. If that is not possible, they should drop the PRISMA framing and call it a scoping review.","headline":"Plausible tri-level metaverse readiness taxonomy, but the n=144 corpus is unreported and the level structure is prefigured by the search terms; useful as a scoping synthesis, not as a grounded-theory result.","tokens_in":28893,"tokens_out":2157,"would_cite":false,"duration_ms":20363,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 144 academic and industry records identifies six user-experience factors, four business-readiness factors, and four national-competitiveness factors, and packages them into three maturity-style models with practical…","keywords":["metaverse","systematic literature review","grounded theory","user experience","business readiness","national competitiveness","extended reality","maturity model"],"falsifier":"Re-run the review with a fully documented protocol: fixed search queries and dates, a complete screening log, and the full list of included records, then check whether the same six user factors, four business factors, and four national factors emerge. If a comparable corpus yields different factor categories, the models are corpus-dependent rather than universal.","tokens_in":27862,"feed_emoji":"🌐","tokens_out":7299,"duration_ms":61650,"temperature":0.7,"pith_summary":"This paper claims that the metaverse is best understood as a sociotechnical system that must be studied at three levels at once: the individual user, the business, and the nation. Through a systematic review of 144 academic and industry records and grounded-theory coding, the authors extract six user-experience factors (accessibility, flexibility, gamification, interoperability, social interaction, and user empowerment), four business-readiness factors (digital-twin integration, marketing, product/service development, and resources), and four national-competitiveness factors (policy, service development, human-capital readiness, and sociocultural influence). They assemble these factors into three research models—a Metaverse User Experience Maturity Model, a Metaverse Business Readiness Model, and a Metaverse National Competitiveness Model—each accompanied by a heuristic rubric or checklist intended for practical assessment. If the models are right, designers, business strategists, and policymakers gain concrete, multi-level tools for evaluating where they stand and what to build next as the metaverse develops. The models are presented as provisional: the paper explicitly says the rubrics are not yet validated and future expert interviews are planned.","feed_headline":"Three models map the metaverse from user to nation","feed_subtitle":"A systematic review of 144 studies distills UX, business, and policy factors into rubrics and checklists for readiness.","key_machinery":"The analytical engine is grounded theory applied to a mixed corpus of 144 records. Open coding tags discrete concerns in the literature; axial coding groups those concerns into candidate categories; selective coding consolidates them into the core categories that become the three models. The analysis is organized by a micro-meso-macro lens, with each level corresponding to one research question, and the resulting categories are then expressed as maturity-model rubrics: the UX heuristic rubric rates each user factor on a 1–5 scale, the business checklist offers yes/no readiness questions, and the national checklist organizes readiness into infrastructure, regulatory, economic, human-capital, cultural, and other policy areas. Topic modeling of titles is used as a secondary, descriptive view of the corpus's themes.","core_discovery":"On the paper's own terms, the central discovery is that the scattered literature on the metaverse converges on a small, ordered set of factors at each scale of analysis. At the individual level, adoption hinges on accessibility, flexibility, gamification, interoperability, social affordances, and user empowerment; at the business level, readiness hinges on digital-twin integration, marketing capability, product and service development, and resources; at the national level, competitiveness hinges on policy frameworks, service development, human-capital readiness, and sociocultural influence. Working from a corpus that deliberately mixes peer-reviewed and grey literature from industry, the authors code these factors using grounded theory and then join them into three layered models, asserting that the levels are mutually reinforcing: good user experience drives business innovation, business innovation drives national economic activity, and government policy shapes the conditions for both. The paper also proposes that these factors can be operationalized as maturity-style rubrics—a heuristic evaluation for UX, a self-assessment checklist for businesses, and a readiness checklist for nations—so that the models are usable before they are fully validated.","pith_inferences":["The models imply a causal chain—user experience to business innovation to national competitiveness to policy and back to user experience—that the review surfaces but does not test; modeling that loop with actual adoption data would be a natural next step.","The reliance on search-engine results that are personalized and time-varying means replicating the review with a frozen, archived corpus would determine how stable the factor lists are, and such a replication could also reveal whether the grey-literature proportion changes the conclusions.","Because the rubrics are presented as unvalidated, they could be converted into survey instruments and field-tested in specific sectors such as retail or education to see which items discriminate between high- and low-readiness organizations.","The national model's inclusion of cultural influence and virtual citizenship points toward a policy question the paper leaves open: how countries without strong cultural export industries should compete in the metaverse."],"forward_implications":["User experience researchers can use the six-factor UX maturity rubric to benchmark existing metaverse applications and identify which dimension, such as interoperability or user empowerment, is least mature.","Business strategists can use the four-factor readiness checklist to audit their organization for digital-twin integration, marketing channels, product/service innovation, and resource readiness before committing to metaverse initiatives.","Policymakers can use the national competitiveness checklist to compare readiness across infrastructure, regulation, economic development, human capital, and cultural influence, and to spot gaps in high-speed connectivity investment or workforce training.","Researchers can treat the three models as a variable catalog for designing studies that connect levels, such as testing whether a rise in user empowerment predicts business adoption of metaverse commerce.","The paper's planned expert interviews constitute a direct validation path for the models and their rubrics."],"supporting_citations":[{"why":"Supplies the grounded-theory literature-review procedure that the analysis follows.","marker":"[33]"},{"why":"Provides the systematic-review reporting guideline the screening adapts.","marker":"[34]"},{"why":"Supplies the micro-meso-macro analytical lens that organizes the three research questions.","marker":"[35]"},{"why":"Justifies including non-peer-reviewed grey literature in the corpus to capture industry perspectives.","marker":"[37]"},{"why":"Foundational source of the grounded-theory coding methodology used for analysis.","marker":"[39]"},{"why":"Provides the topic-modeling technique used for the descriptive overview of article titles.","marker":"[44]"},{"why":"Supplies the e-readiness framework tradition that the national competitiveness model extends.","marker":"[27]"},{"why":"Provides the adapted systematic-review procedure that the paper follows for data collection.","marker":"[31]"}],"fun_headline_variants":["144 metaverse studies distil into three roadmaps","Three layered models for metaverse adoption and policy","Metaverse maturity: UX, business, and national checklists","From individual to nation: a metaverse readiness triad"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 144 records gathered from Google Scholar and LinkedIn fairly represent the broader metaverse discourse, since the search steps are reported without screening counts or the full query set and search-engine results change over time.","fun_headline_variants_meta":{"raw":{"variants":["144 metaverse studies distil into three roadmaps","Three layered models for metaverse adoption and policy","Metaverse maturity: UX, business, and national checklists","From individual to nation: a metaverse readiness triad"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000273,"raw_usage":{"total_tokens":1631,"prompt_tokens":937,"completion_tokens":694,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":630}},"tokens_in":553,"tokens_out":694,"duration_ms":6862,"temperature":1.0,"reasoning_tokens":630,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:38:50.769289+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the review with a fully documented protocol: fixed search queries and dates, a complete screening log, and the full list of included records, then check whether the same six user factors, four business factors, and four national factors emerge. If a comparable corpus yields different factor categories, the models are corpus-dependent rather than universal.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies including non-peer-reviewed grey literature in the corpus to capture industry perspectives."},{"cited_title":"Nursing research 17(4), 364 (1968).Publisher:LWW [40]Remenyi,D.:GroundedTheory:AReaderforResearchers,Students,FacultyandOthers.ACPIPublishing,???(2013)","cited_arxiv_id":null,"evidence_quote":"Foundational source of the grounded-theory coding methodology used for analysis."}],"review_version":1}