{"id":"129f2017-a678-406d-abb8-495a28662192","arxiv_id":"2608.00330","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Sketching and think-aloud critique assessments distinguish visualization expertise levels better than multiple-choice tests and capture partially distinct literacy skills.","lead":"This paper develops two new “qualitative” ways to test visualization literacy: asking people to talk through what is wrong with charts (critique) and to draw their own charts (sketching). In a study of 80 people, these tests separated beginners, students, and experts better than standard multiple-choice tests, suggesting they measure skills those tests miss.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on single-rater, non-blind grading of the main study; pilot inter-rater reliability does not cover the data that produces the group differences.","rationale":"The reader's weakest_assumption identifies exactly the risk I consider load-bearing. The abstract's central claim is comparative: sketching and critique show improved sensitivity over CALVI and Mini-VLAT. That comparison is produced from scores assigned by a single coder who recognized expert participants' voices. The inter-rater reliability reported in Sec. 4 was computed on pilot tasks and reflects consensus-building during rubric development; it does not provide evidence about scoring consistency on the main dataset, and it cannot detect systematic bias toward group labels. Section 7's admission makes this concrete. Because the assessment rubrics are holistic, a rater's expectation that experts should produce better sketches/critiques can plausibly influence scores across all rubric items. This would inflate exactly the pairwise differences in Fig. 8, especially expert vs. non-expert, and would also imply the correlations with CALVI/Mini-VLAT may be distorted. A blind multi-rater re-grade of at least the expert/student strata is therefore a decisive test. I did not find a separate internal inconsistency that would justify rejection; the statistical reporting is transparent and the non-significant student-crowd critique contrast is reported honestly. The appropriate verdict remains CONDITIONAL, contingent on the blind-grading check.","tokens_in":19036,"tokens_out":4987,"duration_ms":51457,"concrete_test":"Take all 17 expert and 21 student responses plus a random sample of 20 crowdworker responses from the main study; strip audio and replace with anonymized transcripts (or text-to-speech with voice masking), and have two independent raters who were not involved in rubric development and are blind to participant group grade all Sketching and Critique items using the final rubrics. Compute quadratic-weighted Cohen's κ on this sample and re-estimate the pairwise group differences (Welch t-tests / Bayesian IRT) from blinded scores. If the expert–student or expert–crowd CIs move toward zero or become non-significant, the central claim is not supported; if they remain, the grader-bias concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The key evidence for the central claim is the between-group differentiation shown in Figs. 5 and 8 (Sketching and Critique pairwise CIs). Those scores were assigned by the first author alone on main-experiment data. Inter-rater reliability (κ=0.81 critique, 0.76 sketching) was measured only on a pilot subset (69 critique tasks, 15 sketching tasks) during rubric alignment, as reported in Section 4 (Grading). Section 7 explicitly acknowledges that \"many participants' voices were recognizable by the first author during grading\" in the expert condition. Because the rubric is holistic and relies on subjective judgment, and because the rater knew the research hypotheses and could identify group membership, the large expert-vs-student/crowd gaps could be inflated by expectation effects. The pilot κ does not rule this out: it shows two aligned authors can agree after discussion, not that the main-study scores are unbiased. If blind multi-rater grading produced smaller group differences, the \"improved sensitivity over CALVI/Mini-VLAT\" conclusion would be substantially weaker. This is not an accusation of bad faith; it is a structural property of non-blind qualitative scoring, acknowledged by the authors. It is load-bearing because it affects every group-level quantitative claim, not just a secondary analysis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces two qualitative visualization literacy assessments—freeform sketching and think-aloud critique—and compares them against CALVI and Mini-VLAT in a within-subjects online study with 80 participants from three expertise groups (crowdworkers, students, and researchers). Using Bayesian IRT, the authors estimate item discrimination and latent abilities and report pairwise group contrasts. They find that both new assessments differentiate experts from non-experts and that they capture skills weakly to moderately correlated with existing tests. The paper argues that qualitative assessments complement multiple-choice measures, especially for distinguishing highly skilled individuals.","tokens_in":19260,"tokens_out":4960,"duration_ms":45418,"significance":"If the findings hold, the paper provides a useful step toward measuring higher-order visualization literacy, with practical guidance for instrument selection. Strengths include a transparent study design, Bayesian IRT with reported credible intervals, public study materials and code, and a sensitivity analysis for prior treemap exposure. The candid acknowledgment of grading bias in Section 7 is commendable. However, the central claim of improved sensitivity is weakened by single-rater, non-blind grading of the main study and the absence of a formal cross-assessment statistical comparison. The paper is likely to stimulate further research on qualitative literacy assessments, but the current evidence does not fully support the strong conclusions in the abstract.","major_comments":[{"comment":"The primary group-level results (Figs. 5, 8) are based on scores assigned by the first author alone on the main experiment. Pilot inter-rater reliability (κ=0.81 critique, 0.76 sketching) was computed on a small pilot subset (69 critique tasks, 15 sketching tasks) during rubric alignment, not on the main data. Section 7 explicitly acknowledges that 'many participants’ voices were recognizable by the first author during grading,' particularly in the expert condition. Because the rubric is holistic and the rater knew the hypotheses and group membership, the expert-vs-student and expert-vs-crowd gaps could be inflated. This is load-bearing for the paper’s central claim. Please blind the grading (e.g., use transcripts with voice masking), add a second independent rater on a sample of main data, or provide a sensitivity analysis bounding the possible bias.","section":"§4 Grading; §7 Limitations"},{"comment":"The abstract and Section 1 claim that both sketching and critique 'succeed at differentiating between groups of varying skill level.' However, the critique assessment does not significantly differentiate students from crowdworkers (Fig. 8: 0.39 [−0.05, 0.82]). Only expert vs. student and expert vs. crowdworker contrasts are significant. The claim should be qualified to avoid overstating the evidence. Similarly, the phrase 'improved sensitivity over both Mini-VLAT and CALVI' is only supported by a qualitative comparison of which CIs exclude zero; no direct statistical comparison of sensitivity across assessments is reported.","section":"§5 Results; Fig. 8; Abstract"},{"comment":"The conclusion that the new assessments exhibit improved sensitivity is inferred from separate Welch t-tests per assessment. No model or test directly compares the magnitude of group differences across assessments (e.g., an assessment-by-group interaction or a posterior distribution of the contrast difference). A formal comparison would strengthen the claim and provide a quantitative basis for the title and abstract.","section":"§5 Will Sketching...; Fig. 8"},{"comment":"The sensitivity analysis in Appendix A.1 only removes four students with prior exposure to the treemap. However, the critique stimuli include other widely publicized charts (the line chart, the map), and Section 7 notes that experts 'likely have seen some of the examples before.' Prior familiarity with specific charts could differentially boost expert scores. Please measure or ask about prior exposure for all stimuli, or at least broaden the sensitivity analysis, so that the assessments measure generalizable literacy rather than recognition of famous charts.","section":"§6 Participant Group Differences; §7 Limitations; Appendix A.1"}],"minor_comments":[{"comment":"The abstract contains 'Y et' with a stray space; should be 'Yet.'","section":"Abstract"},{"comment":"The caption reads 'If a test confidence interval does not overlap with the associated mean, then the test is significant.' This is inaccurate; significance is determined by whether the CI excludes zero, not the mean. Please correct.","section":"Fig. 8 caption"},{"comment":"The test name is spelled inconsistently: 'Calvi/Mini-VLAT' and 'CALVI' appear interchangeably. Use 'CALVI' consistently.","section":"§4 Study Methods"},{"comment":"The reference to 'Fig. 6' in the text occurs before the reader has context for the figure's layout; consider adding a pointer to the rubric table (Table 1) when discussing item discrimination.","section":"§5 Results"}],"recommendation":"major_revision","confidential_remarks":"The paper shares authors with prior work on CALVI and AVEC, but the self-citations are largely for study infrastructure and established rubrics; I do not see a disclosure problem. The main concern is the non-blind single-rater grading, which could bias the central group comparisons. If the authors can provide a blind re-grading of at least a subset or a convincing sensitivity analysis, the paper would be much stronger. The sample sizes for experts and students are small but not disqualifying given the transparent reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this paper is useful. It builds two qualitative assessments—sketch a visualization from data, critique flawed charts aloud—with concrete rubrics, and it shows they separate crowdworkers, students, and experts better than Mini-VLAT and CALVI in a mostly transparent design. It deserves peer review and likely encourages revision.\n\nWhat’s actually new: prior instruments are multiple-choice or constrained construction tasks. Freeform sketching and think-aloud critique with published rubrics are a real step, and the comparison is done within-subjects with Bayesian IRT, credible intervals, and a sensitivity analysis for treemap exposure. The qualitative findings—participants who only know bar charts, anchoring in critiques, the table-with-reasoning grading call—are valuable on their own. Credit also for shipping the study code and at least crowdworker-level supplemental data.\n\nThe soft spots, in proportion:\n\n1. The main grading was done by the first author alone, not blind, and the authors acknowledge that expert voices were recognizable. Pilot inter-rater reliability (κ≈0.81 critique, 0.76 sketching) came from a small aligned pilot, not from the main dataset. This is load-bearing because every between-group difference and most IRT results run through that one grader. I don’t think it sinks the paper—the gaps are large and the pattern is plausible—but it is exactly the kind of expectation effect that can inflate group differences. A blind multi-rater check on a sample of main data, especially the expert condition, should be a required revision.\n\n2. The expert and student samples are thin (n=17 and 21) for IRT-based claims. The authors report posterior SDs to argue precision is comparable across groups, which helps, but the expert–student pairwise contrasts are the least certain and one CI barely excludes zero.\n\n3. The CALVI comparison uses the 15 trick items with the normal items replaced by Mini-VLAT. That hybrid matches prior work, so it is defensible, but the paper should be careful not to imply it re-administered standard CALVI.\n\nThe stress-test concern about circularity is largely wrong: rubrics came from prior literature and pilot iteration, and this is a feasibility comparison, not a fitted prediction. The self-citations are to study infrastructure, which is fine.\n\nThe same transparency decision that protects participants—withholding expert and student data because voices are recognizable—makes independent verification harder. That is a real limitation, not a flaw in the analysis.\n\nBottom line: a solid step toward measuring higher-order visualization literacy. The assessments need external validation and independent grading before being treated as validated instruments, but the claim that open-ended critique and sketching capture something multiple-choice tests miss is credible. Send it to reviewers; the main thing I would ask for is a blinded multi-rater grading sample and a clearer statement of what the CALVI hybrid does and does not measure.","headline":"Worth engaging: the first real shot at open-ended sketching/critique assessments for visualization literacy, but the headline claim leans on single-rater, non-blind grading.","tokens_in":19773,"tokens_out":3185,"would_cite":true,"duration_ms":31206,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Open-ended sketching and critique assessments separate visualization skill levels that multiple-choice tests cannot, especially among experienced viewers.","keywords":["visualization literacy","qualitative assessment","sketching","critique","think-aloud","item response theory","ceiling effects","higher-order skills"],"falsifier":"Rescore all 80 participants' sketches and critiques with two independent raters who are blind to group and hear only anonymized transcripts; if the expert–student and student–crowdworker gaps no longer reach significance, the paper's central sensitivity claim would not hold.","tokens_in":18897,"feed_emoji":"✏️","tokens_out":5226,"duration_ms":47787,"temperature":0.7,"pith_summary":"The paper tries to show that visualization literacy is not fully captured by multiple-choice comprehension tests. It introduces two web-based qualitative assessments—thinking aloud while critiquing flawed charts, and sketching a visualization for a given dataset—and compares them with two established tests across crowdworkers, students, and visualization researchers. The results suggest both qualitative modalities differentiate between experience levels better than the established tests, with Mini-VLAT showing a ceiling even among crowdworkers. Because the new assessments correlate only moderately with existing ones, they appear to measure distinct, higher-order skills such as contextual judgment and design reasoning. A sympathetic reader would conclude that qualitative assessments are a promising complement for measuring high-end visualization literacy.","feed_headline":"Open-ended tests reveal visualization skill multiple-choice misses","feed_subtitle":"Thinking aloud and drawing charts separate expert from novice viewers where Mini-VLAT hits its ceiling.","key_machinery":"The carrying mechanism is the pair of five-point holistic rubrics applied to think-aloud transcripts and final sketches. Unlike early analytic checklists, the rubrics credit the reasoning behind a judgment—a participant can defend a 3D treemap well and score well—rather than ticking off expected features. For sketching, rubric categories cover visual hierarchy, base encoding, semantic validity, and text/guides; for critique, reading, contextualization, and judgment. A Bayesian ordinal item-response model then converts the rubric scores into latent ability estimates used for group comparisons and correlation analysis.","core_discovery":"The paper claims that both sketching and critique succeed at differentiating between groups of varying visualization skill and exhibit improved sensitivity over both Mini-VLAT and CALVI. Under the authors' rubrics—four categories for sketching and three for critique—sketching separates all three pairwise comparisons (experts vs students, students vs crowdworkers, experts vs crowdworkers), while critique separates experts from each other group but not students from crowdworkers. Mini-VLAT separates no pair and CALVI separates only experts from crowdworkers. The qualitative assessments correlate moderately with each other and with CALVI, and weakly with Mini-VLAT, which the authors read as evi","pith_inferences":["Editorial inference: If baseline visualization literacy has risen since VLAT's introduction, as the paper tentatively suggests, norms for older multiple-choice instruments may need revalidation in current populations before being used as study filters.","Editorial inference: The authors' acknowledged grader-bias risk implies a concrete test: if expert-student differences persist when grading is blind and voice-anonymized, the construct-validity case would be substantially stronger.","Editorial inference: Because correlation between reading- and writing-style assessments around 0.5 mirrors text literacy research, one might predict that combining sketching and critique with multiple-choice scores gives a more complete latent visualization-literacy factor than any single modality.","Editorial inference: The rubric's treatment of ambiguous sketches, where audio is used to infer implied labels, suggests an automated pipeline could grade transcripts plus final canvas state; a pilot study comparing automated grading to human rubric scores would test feasibility."],"forward_implications":["Multiple-choice-only assessment of visualization literacy should be reconsidered when the goal is to distinguish experienced individuals; Mini-VLAT's ceiling makes it unsuitable for that purpose.","The critique rubric appears transferable: similar discrimination across three very different charts suggests the method, not just the specific stimuli, carries the signal.","Sketching and critique can be used as a complement after a short screening test, an adaptive design the authors explicitly propose.","Researchers can reduce grading burden by selecting a subset of sketching tasks, such as keeping the tree task for expert populations, without much loss of discrimination.","Qualitative assessments open the door to studying how people construct transformations and design choices, not just whether they can read a chart."],"fun_headline_variants":["Critique and sketching expose visualization skill gaps","Open-ended viz tests beat multiple choice at expert levels","Sketching and critique reveal expert visualization literacy","Why multiple-choice visualization tests hit a ceiling"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the single grader's rubric scores on the main dataset were not biased toward the expert group; inter-rater reliability was checked only on a small pilot, and the grader could recognize many experts' voices.","fun_headline_variants_meta":{"raw":{"variants":["Critique and sketching expose visualization skill gaps","Open-ended viz tests beat multiple choice at expert levels","Sketching and critique reveal expert visualization literacy","Why multiple-choice visualization tests hit a ceiling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1108,"prompt_tokens":745,"completion_tokens":363,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":489,"tokens_out":363,"duration_ms":4125,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:42:21.194028+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rescore all 80 participants' sketches and critiques with two independent raters who are blind to group and hear only anonymized transcripts; if the expert–student and student–crowdworker gaps no longer reach significance, the paper's central sensitivity claim would not hold.","supporting_citations":[],"review_version":1}