{"id":"1d512dde-63ca-420e-b955-5d505caab8ed","arxiv_id":"2509.01018","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The field of visualization literacy is mapped into five contribution categories and four competency themes, showing that reading charts is well studied while critique, connection, and non-Western contexts are neglected.","lead":"This survey organizes 374 papers on visualization literacy into five research categories and four skill themes. It offers a shared vocabulary for defining the term and exposes gaps in what is measured and taught.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Keyword search omits critique/connection-related terminology; reported competency and non-Western gaps may be sampling artifacts.","rationale":"The reader's weakest assumption was that the 374-paper corpus might be incomplete or unrepresentative, especially for non-Western and non-computer-science venues. My concern is a more specific variant: the Boolean search string itself is likely to systematically exclude a substantial body of research on critical visualization literacy and on decision-making/connection competencies, because those papers often use terminology like 'misleading', 'deception', 'rhetoric', or 'decision-making' rather than 'literacy' or 'education'. Since the pilot coding was seeded by 'visualization literacy' keyword hits, the extracted keywords in Section 3.2 are unlikely to reach that literature. This directly undermines the specific empirical claim that critique and connection are under-studied, and the non-Western gap claim is similarly exposed by the English-language database selection. I partially agree with the reader because the general completeness concern is correct, but I pinpoint a concrete mechanism that a simple corpus release would not fully remediate—one would also need to validate search coverage against alternative terminology. The verdict remains CONDITIONAL: the paper is a valuable organizational contribution, but its headline gap analysis should be accepted only after the authors either demonstrate that the omitted terminology does not change the gap patterns or release an expanded corpus that incorporates such work. This is not a rejection because the taxonomy may still be useful even if some counts shift; it is a condition because the current evidence is insufficient to support the strong 'systematic gaps' claims.","tokens_in":38442,"tokens_out":4985,"duration_ms":59446,"concrete_test":"Run a supplementary search on the same five databases plus Google Scholar and regional indexes (e.g., SciELO, CNKI) using additional keyword families: (visualization OR graph OR chart) AND (critical OR misleading OR deceptive OR rhetoric OR persuasion OR decision-making OR storytelling OR comprehension), plus translated equivalents (e.g., 'alfabetización visual', '可视化素养', 'letramento visual'). Screen the new results with the paper's inclusion criteria, add eligible papers to the corpus, and recompute the competency and population gap analyses. If CRITIQUE/CONNECTION counts or non-Western representation increase materially, the reported gaps are sampling artifacts and the central claims must be revised.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central gap claims (Sections 5.2, 5.4, 6) depend on the corpus assembled in Section 3.2. The Boolean search string is: (visualization AND (literacy OR education OR teach* OR learn* OR student OR novice)) OR (graphicacy OR \"graph literacy\" OR \"graph comprehension\" OR \"graphical literacy\" OR \"graphical comprehension\" OR \"risk literacy\"). This string captures work explicitly using literacy/education vocabulary, but omits a large space of relevant work on critical evaluation of visualizations (e.g., terms like \"misleading\", \"deceptive\", \"critical\", \"rhetoric\", \"persuasion\") and on connecting visualizations to decisions or contexts (e.g., \"decision-making\", \"reasoning\", \"storytelling\") when those papers do not mention literacy/education/student/novice. Because the pilot corpus was assembled via \"visualization literacy\" keyword searches (Section 3.1), the keyword extraction likely inherits this narrow vocabulary. Consequently, the conclusions that CRITIQUE and CONNECTION are under-studied, and that non-Western contexts are nearly absent, may reflect the search frame rather than the field. The five databases (IEEE Xplore, ACM DL, DBLP, Scopus, Web of Science) are also English-centric, so non-English visualization literacy research (e.g., in Spanish, Chinese, Portuguese) could be systematically missed—an especially direct threat to the non-Western gap claim. This is load-bearing because the taxonomy and the reported gaps are the paper's main contribution, and they are only as valid as the corpus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This state-of-the-art report surveys visualization literacy research, synthesizing 374 papers collected through a five-database keyword search, manual inclusion tagging, and snowballing. The authors propose a two-dimensional taxonomy: five research categories (ontology, assessment, mechanisms, populiteracy, intervention) and four competency themes (consumption, construction, critique, connection). They use this taxonomy to characterize the field, reporting that consumption is the most frequently measured competency while critique and connection receive less attention, and identifying gaps in non-Western and accessibility-focused research. The report also offers an operationalization framework for visualization literacy grounded in application contexts and competencies.","tokens_in":38790,"tokens_out":6588,"duration_ms":76871,"significance":"If the claims are accurate, this survey provides a valuable organizing structure for a fragmented field and a practical vocabulary for defining and comparing studies. Strengths include the transparent four-stage methodology (pilot coding with reported Cohen's kappa, explicit search string, manual inclusion tagging with dual annotation, snowballing) and the publication of corpus statistics. The four competency themes are intuitive and likely to be adopted by the community. However, the central gap claims rest entirely on the composition of the corpus, and the corpus search strategy is narrowly scoped. The survey's value depends on addressing the sampling threats identified below.","major_comments":[{"comment":"The Boolean search string requires visualization AND (literacy OR education OR teach* OR learn* OR student OR novice), plus a fixed set of graph-literacy terms. It does not include terms such as 'misleading', 'deceptive', 'critical', 'persuasion', 'rhetoric', 'decision-making', 'reasoning', or 'storytelling' unless accompanied by the enumerated literacy/education vocabulary. Because the pilot corpus (Section 3.1) was assembled from 'visualization literacy' searches, the extracted keywords likely inherit this narrow frame. Therefore, the conclusion in Section 6 that CRITIQUE and CONNECTION are under-studied, and the assessment-gap discussion in Section 5.2, may reflect the search frame rather than the field. This is load-bearing. I recommend supplementing the search with competency-specific terms and re-tagging the additional papers, or alternatively re-framing the claims as 'within the v","section":"Section 3.2 / Section 6"},{"comment":"The five databases (IEEE Xplore, ACM DL, DBLP, Scopus, Web of Science) are all English-centric, and the search string is English-only. The claim in Section 6 that non-Western contexts are nearly absent (echoed from [Sol22]) is directly threatened by the lack of non-English indexes and non-English keyword equivalents. Without a supplementary search in non-English venues, or at least a clear qualification that the non-Western gap holds only for the English-language indexed literature, this conclusion is unsupported.","section":"Section 3.2 / Section 6"},{"comment":"The only quantified inter-coder reliability is the inclusion kappa of 0.633 (Section 3.1). The final assignment of papers to categories and competency themes is described as consensus-based, but no reliability statistic or detailed rubric for competency tagging is provided. Since the gap analysis and the counts in Figure 3 depend on these tags, the authors should either report reliability for the final category/competency coding, or justify why dual annotation with consensus is sufficient. This is relevant to the central quantitative distribution claims.","section":"Section 3.4"}],"minor_comments":[{"comment":"Several encoding artifacts appear, e.g., 'â ˘AIJ', 'â ˘A¸ S', 'KÃ˝ urner'. These should be cleaned in the final PDF/source.","section":"Text encoding"},{"comment":"The 374-paper corpus is not listed. Consider an appendix or online repository listing all included papers, enabling readers to audit the taxonomy and re-use the corpus.","section":"Appendix / reproducibility"},{"comment":"The 'Relevant Age Groups' row (Children, Adolescents) overlaps with the 'Students' grouping (e.g., primary school students). Clarify whether these are distinct population categories or a cross-cutting dimension.","section":"Table 2"},{"comment":"Reference [MHS15] is cited twice in close succession; coalesce or remove the duplicate citation.","section":"Section 5.4.1"},{"comment":"Add axis labels and consider a companion table or matrix showing competency coverage per category; the current UpSet plot shows category intersections but not the competency themes that are central to the gap claims.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"This is a useful survey with a strong methodological skeleton, but the corpus-level gap claims are vulnerable to the search-strategy critique. I would ask the authors to either expand the search and re-code, or explicitly restate the contribution as a taxonomy/framework rather than a field-wide gap analysis. As a STAR, the paper should also include a list of included papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the most comprehensive visualization literacy survey to date, and the two-axis taxonomy (research categories x competency themes) is a genuinely useful organizing device. The methodology is more transparent than most surveys in the field. But the stress test lands: the Boolean search string is built from literacy/education vocabulary, so the finding that critique and connection are under-studied might be partly a sampling artifact. Also, the 374-paper corpus and tag assignments are not shipped, which undercuts auditability.\n\nWhat's new: prior surveys covered 20–34 papers; this one screens 4,519 and includes 374. The five-category taxonomy (ontology, assessment, mechanisms, populiteracy, intervention) and the four competency themes (consumption, construction, critique, connection) provide a workable language for talking about gaps. The discussion of complementary literacies and the operationalization framework are thoughtful, not just list-making. The inclusion coding was dual-annotated with a reported kappa of 0.633, and the pilot-then-expand process is sensible.\n\nWhere the soft spots are: the stress-test concern is fair. The keyword string includes graphicacy, graph literacy, graph comprehension, etc., but omits terms like misleading, deceptive, critical, rhetoric, persuasion, decision-making, storytelling. Work on critical visualization or visualization reasoning that doesn't use literacy vocabulary will not be retrieved. Since the taxonomy and gap analysis are derived from this corpus, the claims that critique and connection are under-studied, and that non-Western contexts are nearly absent, may reflect the sampling frame rather than the field. The authors do acknowledge the non-Western limitation, but they don't interrogate whether their search strategy itself is the cause. The second issue is reproducibility: no accompanying list of included papers or annotations. For a survey that makes quantitative claims about category sizes, that's a real shortfall.\n\nNone of this sinks the paper. The taxonomy is plausible and the synthesis is valuable; the gaps are probably real even if the effect sizes are uncertain. But 'probably' isn't 'demonstrated.'\n\nWho this is for: anyone entering visualization literacy research, and anyone wanting a structured map of the field. It deserves a serious referee. I'd recommend acceptance conditional on releasing the corpus and tag data, and on softening or better supporting the gap claims given the search string's limits.","headline":"A useful, transparent STAR with a plausible taxonomy, but the search string likely biases the gap claims and the corpus isn't released.","tokens_in":39253,"tokens_out":2345,"would_cite":true,"duration_ms":27670,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey of 374 papers argues that visualization literacy research is best organized by five contribution types and four competency themes, and that this organization reveals systematic gaps—especially in critique, connection, accessibil","keywords":["visualization literacy","taxonomy","competency framework","state-of-the-art survey","assessment instruments","data visualization education","critical visualization literacy","populiteracy"],"falsifier":"A concrete check: rerun the same boolean keyword search against non-English and non-computer-science databases (e.g., ERIC, PsycINFO, CINAHL, SciELO, CNKI) and hand-code the results with the paper's inclusion criteria. If a substantial number of new papers in critique, connection, accessibility, or non-Western contexts emerges, the gap analysis would need revision. Alternatively, if the authors release the tagged corpus, an independent replication of the coding on a random subset could confirm the category counts and the claimed imbalances.","tokens_in":38354,"feed_emoji":"📊","tokens_out":4405,"duration_ms":52579,"temperature":0.7,"pith_summary":"The paper is a systematic survey of research on visualization literacy—the skills people need to read, make, question, and use data visualizations. It claims that the field can be organized by five kinds of research contributions (definitions, measurements, models of comprehension, population studies, and teaching interventions) crossed with four competency areas (consumption, construction, critique, connection). Applying this taxonomy to 374 papers shows the field is lopsided: most work measures how well people read charts, while critical evaluation and connecting charts to real-world knowledge are under-studied, and research outside Western contexts and on accessibility is nearly absent. The taxonomy gives researchers a shared vocabulary for saying exactly which skills they study and a map of where new work is needed.","feed_headline":"374-paper survey maps the gaps in visualization literacy research","feed_subtitle":"Reading charts is well studied; critiquing, connecting, and non-Western contexts are not.","key_machinery":"The central mechanism is the taxonomy itself. On one axis are five research categories derived from iterative open coding of the 374-paper corpus: ontology (definitions and scope), assessment (instruments for measuring literacy), mechanisms (models of user processes, barriers, and comprehension steps), populiteracy (descriptions of literacy within a specific population), and intervention (efforts to teach or improve literacy). On the other axis are four competency themes: consumption (reading and interpreting), construction (designing and implementing), critique (critical thinking about visualizations), and connection (relating visualizations to external knowledge and context). The taxonomy","core_discovery":"On the paper's own terms, the central claim is that visualization literacy research can be unified and usefully described by a two-dimensional taxonomy: five research categories (ontology, assessment, mechanisms, populiteracy, intervention) crossed with four competency themes (consumption, construction, critique, connection). The authors argue that visualization literacy should be operationalized as a context-dependent set of competencies rather than a single ability, and that doing so reveals systematic gaps. Consumption is heavily measured, interventions mostly target construction, critique and connection are rarely addressed across categories, and accessibility and non-Western populations","pith_inferences":["The same taxonomy could be applied to adjacent literacies (data literacy, graph literacy, risk literacy) to test whether their research portfolios are similarly lopsided.","If the field shifts toward critique and connection, current default instruments like VLAT may lose status, and rubric-based or qualitative assessments may become more central.","The near-absence of non-Western studies suggests chart conventions themselves may be partly culture-specific; consumption measures developed in Western contexts would need cross-cultural validation before being trusted elsewhere.","A testable extension: rerun the same search protocol on non-English databases and compare category and theme distributions; if critique and connection papers are plentiful there, the taxonomy's gaps are a sampling artifact rather than a property of the field."],"forward_implications":["Assessment developers can use the framework to state explicitly which competencies their instrument measures and to justify creating new tests for construction and critique.","Intervention designers can see that most teaching targets construction and basic reading, making teaching of critique and connection comparatively open ground.","Researchers who adopt the competency themes will be forced to make their operationalization of visualization literacy explicit, making studies more comparable across the field.","Populiteracy studies receive a checklist of near-empty populations: non-Western groups, people with visual impairments or dyslexia, and specific domain specialists such as policymakers.","The taxonomy is structured to be extended or revised, so future surveys can treat it as a living classification rather than a fixed endpoint."],"supporting_citations":[{"why":"Prior survey of 34 papers limited to evaluative approaches; this STAR positions itself as broader in scope and covering non-evaluative research.","marker":"[FJL22a]"},{"why":"Prior survey of 20 papers proposing three research themes; also the source of the observation that most visualization literacy research is Western.","marker":"[Sol22]"},{"why":"VLAT, the canonical consumption-focused assessment instrument, used to illustrate the dominance of consumption measurement in the assessment category.","marker":"[LKK17]"},{"why":"Boy et al.'s principled assessment with an item bank, a foundational contribution to the assessment category.","marker":"[BRBF14]"},{"why":"Carpenter and Shah's cognitive model of graph comprehension, grounding the mechanisms category.","marker":"[CS98]"},{"why":"CALVI, the critical-thinking assessment, evidence that critique-focused instruments exist but are rare.","marker":"[GCK23]"},{"why":"Camba et al. arguing that identifying deception is a core component of visualization literacy, supporting the ontology and critique discussion.","marker":"[CCB22]"},{"why":"Iguanodon, a game for improving visualization construction literacy, representative of the intervention category's focus on construction.","marker":"[ALE*24]"}],"fun_headline_variants":["Visualization literacy: critique and connection understudied","374 papers: visualization literacy gaps in critique, connection","Survey: visualization literacy research misses critical skills","Visualization literacy taxonomy reveals lopsided research","Study: visualization literacy focuses on reading, not critiquing"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The survey's findings rest on the 374-paper corpus being a representative sample of the field; if the keyword search, database choices, and manual tagging missed large bodies of relevant work—especially outside Western, English-language, computer-science-indexed venues—the reported gaps would be artifacts of the sampling frame rather than real properties of the literature.","fun_headline_variants_meta":{"raw":{"variants":["Visualization literacy: critique and connection understudied","374 papers: visualization literacy gaps in critique, connection","Survey: visualization literacy research misses critical skills","Visualization literacy taxonomy reveals lopsided research","Study: visualization literacy focuses on reading, not critiquing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000675,"raw_usage":{"total_tokens":2851,"prompt_tokens":629,"completion_tokens":2222,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":373,"completion_tokens_details":{"reasoning_tokens":2147}},"tokens_in":373,"tokens_out":2222,"duration_ms":18747,"temperature":1.0,"reasoning_tokens":2147,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:56:45.345004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: rerun the same boolean keyword search against non-English and non-computer-science databases (e.g., ERIC, PsycINFO, CINAHL, SciELO, CNKI) and hand-code the results with the paper's inclusion criteria. If a substantial number of new papers in critique, connection, accessibility, or non-Western contexts emerges, the gap analysis would need revision. Alternatively, if the authors release the tagged corpus, an independent replication of the coding on a random subset could confirm the category counts and the claimed imbalances.","supporting_citations":[],"review_version":1}