{"id":"5c068aa8-f68d-421e-8b62-b218aed0ee0a","arxiv_id":"2608.06804","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In an exploratory study, fact-checkers using the FYI browser extension adopted AI-first, manual-first, and parallel workflows, using visualizations to audit AI verdicts.","lead":"A browser extension called FYI embeds AI-assisted fact-checking inside a reading pane, letting readers verify statistics in data-driven articles against the underlying dataset. In a 22-person study, readers combined AI and manual tools in three distinct workflows, with self-built charts used to audit AI conclusions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed workflow flexibility may be an artifact of a task designed so no single tool suffices; a control condition with single-tool-answerable claims would test this.","rationale":"The strongest claim has two premises: first, that the observed workflow archetypes and flexibility are genuine reader behaviors rather than experimental artifacts; second, that they generalize to FYI's stated target population. The reader's weakest_assumption concerned the second premise (sample representativeness). I found the more load-bearing of the two to be the first, because Section 4.2 explicitly states that the claims were designed to require combining multiple verification strategies. That design choice makes the absence of fixed-pipeline behavior almost inevitable, so the descriptive claim that 'participants did not follow a fixed pipeline' is partially built into the stimulus rather than discovered. The paper is commendably transparent about this and about the sample, and the released artifacts support reproducibility, but transparency does not remove the design-dependence of the central descriptive claim. A control condition with single-tool-answerable claims would directly test whether flexible multi-tool composition is a general strategy or a response to task demands. I did not make inter-rater reliability the headline concern because even perfect coding reliability would not resolve whether the observed patterns are task-induced. Likewise, the trust-calibration claim is qualitative but is tied to concrete behaviors (e.g., P13 abandoning Auto Check after an error) and is appropriately hedged as an exploratory finding. Because the reader already recommended CONDITIONAL, and this concern reinforces rather than overturns that judgment, the appropriate final verdict remains unchanged.","tokens_in":21703,"tokens_out":5730,"duration_ms":60600,"concrete_test":"Run a second condition (or a between-subjects control) with the same FYI prototype and same recruitment pool, but replace half of the embedded claims with claims that are verifiable by a single direct operation, such as one exact value in Table Explorer or one simple chart in Chart Builder. For those simple claims, measure the mean number of tools used per claim and the fraction of claims where Chart Builder was launched after Auto Check. If simple claims still average two or more tools and show the same visual-audit pattern, the archetype and flexibility findings are robust; if simple claims are typically settled with one tool, the 'flexibly compose' conclusion is an artifact of the forced multi-tool stimulus.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central descriptive claim—that readers 'actually fact-check' by flexibly composing AI and manual tools and by auditing AI with self-built charts—rests on behavior observed under an instrument that was deliberately arranged to make such composition necessary. Section 4.2 states that the six embedded claims were designed 'to require participants to combine multiple verification strategies rather than rely on any single tool or pathway.' Under that design, a single-tool pipeline is not a viable option, so observing multi-tool workflows, archetypes, and post-AI chart audits is partly guaranteed. The observed archetypes (T3–T5) and their fluidity therefore reflect the task's demands as much as (or instead of) a general reader behavior. A second, more internal version of the same concern: the 'visualization as primary auditing mechanism' result is likely gated by the sample's high visualization literacy (M=5.68/7) and by claims that require group comparisons and aggregations. The paper itself acknowledges this in Sec. 5.4 (P17 avoided Chart Builder) and Sec. 7 ('may limit generalizability'), so the concern is not a hidden flaw. But it is load-bearing: if the scenario had included simple, directly table-lookupable claims, or less visualization-literate readers, the archetype distribution and the centrality of Chart Builder could change substantially. The sample limitation in Secs. 4.1 and 7 is transparent, but the task-design issue in Sec. 4.2 is the less-defended premise for the 'flexible composition' part of the strongest claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents FYI, an open-source browser extension that embeds claim detection, four verification tools (Auto Check, AI Chat, Table Explorer, Chart Builder), and user-authored verdicts into the reading environment. Using FYI as an instrumented design probe, the authors ran an exploratory N=22 study in which participants fact-checked six embedded data claims in a movie-data article while thinking aloud. Based on interaction logs, transcripts, and questionnaires, they report three workflow archetypes (AI-first with manual confirmation; manual-first with AI supplement; parallel co-review), dynamic trust calibration driven by cross-tool convergence and inconsistency, and the use of self-built charts as a primary mechanism for auditing AI outputs. They derive four design implications for mixed-initiative fact-checking systems and release the prototype and study materials.","tokens_in":21977,"tokens_out":6209,"duration_ms":61603,"significance":"If the behavioral findings hold, the paper makes a useful empirical contribution to human-AI sensemaking and fact-checking research: it moves beyond system-capability papers to document real tool-composition behavior, and it provides a reusable open-source testbed with fine-grained interaction logging, prompts, and raw logs. The triangulation of interaction logs, think-aloud protocols, and interviews is a strength, as is the transparent reporting of verdict accuracy against researcher ground truth in Sec. 5.5. The main significance is as an exploratory design probe; the contribution is descriptive rather than confirmatory, and the implications (DI1–DI4) are tied to the observed behaviors in a way that should generalize only if the identified threats are addressed.","major_comments":[{"comment":"Section 4.2 states that the six embedded claims were designed 'to require participants to combine multiple verification strategies rather than rely on any single tool or pathway.' Because the stimulus was deliberately constructed so that no single tool suffices, the observation that participants composed multi-tool workflows (Sec. 5.3) and that three archetypes emerged is in part an artifact of the experimental instrument, not an independent fact about reader behavior. To support the central 'actually fact-check' claim, the authors should either include claims answerable by a single tool or direct table lookup, conduct a per-claim analysis showing that the archetypes and chart-auditing behavior also hold for simple claims, or explicitly re-scope the contribution to behavior under a task that demands multi-tool composition. This is load-bearing because the paper's main empirical result—flexible multi-tool composition—depends on it.","section":"Sec. 4.2"},{"comment":"The stated target population is 'laypersons with basic data and visualization literacy' (Sec. 1 and Sec. 4.1), but the sample reports high visualization literacy (M=5.68/7) and near-daily generative AI use (M=6.18/7). The conclusions that 'visualization serves as the primary auditing mechanism' (Sec. 5.4, T7) and that limited visualization literacy creates a 'usability–confidence paradox' (Sec. 6.3) are therefore likely gated by the sample's visualization skill; the paper itself notes that P17 avoided Chart Builder (Sec. 5.4) and acknowledges generalizability limits (Sec. 7). To make the population claim, the authors should either recruit a broader sample, stratify the analyses by data literacy, or reframe the contribution as an account of data-literate users' behavior with implications, rather than evidence, for laypersons.","section":"Sec. 4.1 and Sec. 7"},{"comment":"The three workflow archetypes (T3–T5) are central to the paper's contribution, but the method for assigning participants to archetypes is not reported. Table 4 lists participant IDs per theme, yet the paper does not state whether classification was done per participant across the whole session or per claim, what criteria or thresholds define each archetype, or whether the three coders agreed on these assignments. Figure 2 shows per-claim sequences while the text reports archetype counts (9/6/4) as session-level categories, making the taxonomy hard to audit. Please provide the operational coding scheme, a per-participant classification rule, and inter-rater agreement statistics, or present the archetypes as per-claim orientations rather than participant types.","section":"Sec. 5.3 and Table 4"}],"minor_comments":[{"comment":"There are unresolved citation placeholders: '[?]' appears in Sec. 2.1 for the statement that data claims implicitly refer to an underlying dataset, and '[?]' appears in Sec. 2.2 for the Thucy system, which corresponds to reference [53]. These should be replaced with proper citations.","section":"Sec. 2.1 and Sec. 2.2"},{"comment":"The caption says the figure covers 131 claims, while Sec. 5.1 reports 156 actively investigated claims; please clarify whether the figure excludes claims without tool-use sequences and state the inclusion criterion.","section":"Figure 2"},{"comment":"The thematic analysis section describes independent coding by three researchers but reports no inter-rater reliability measure; for a qualitative exploratory study this is acceptable, but a brief statement on consensus and stability of the final themes would strengthen the reporting.","section":"Sec. 4.4"},{"comment":"The accuracy comparison (n=79) includes only verdict instances for which researcher ground truth, Auto Check, and participant labels were all available; please clarify how the excluded instances were distributed and whether this selection affects the reported agreement rates.","section":"Sec. 5.5"},{"comment":"The design goals are clear, but DG2 states that 'web search provides external corroboration' while web search is implemented only as a toggle within AI Chat; the paper could note this placement earlier to avoid implying a fifth standalone tool.","section":"Sec. 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper fits TVCG's scope as a design-probe study and benefits from its open-source release and detailed logging. The main risk is overclaiming from a task deliberately designed to require multi-tool composition. I would not require a full replication, but the authors should provide the per-claim analysis separating simple from composite claims, clarify the archetype coding, and either broaden the sample or re-scope the population claims. The placeholder citations should also be fixed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a solid, honestly-reported design probe, and the central finding—that people flexibly compose AI and manual tools and audit AI with self-built charts—is plausible but partly pre-arranged by the study design. It deserves a serious referee, but the authors should soften the generalizing language and add a reliability check.\n\nWhat's genuinely new: an open-source, instrumented prototype (FYI) that spans the full detection–verification–determination pipeline, plus interaction logs and think-aloud data from 22 participants. The three workflow archetypes (AI-first with manual confirmation, manual-first with AI supplement, parallel co-review) and the dynamic trust-calibration pattern (convergence builds trust, inconsistency erodes it) are concrete and extend prior work on trust calibration. The paper also ships the artifact, prompts, logs, transcripts, and ground-truth labels—that's real reproducibility value.\n\nThe soft spots, in order of severity. First, the stimulus was designed so that no single tool would suffice (Sec. 4.2 states this plainly). Observing multi-tool workflows and flexible composition is partly a consequence of the task, not purely a discovery about reader behavior. The paper acknowledges this indirectly, but it doesn't defend why the result generalizes. A control condition with claims answerable by a single tool would have strengthened it. Second, the sample is 22 university-affiliated participants with high visualization literacy (5.68/7) and daily AI use (6.18/7); the paper labels them 'laypersons' but they're not typical readers. The limitations section is transparent about this, but the design implications (DI1–DI4) are written for a broader population. Third, the thematic coding has no inter-rater reliability metric; they say three researchers read independently and iteratively developed themes, but no kappa or similar. For a qualitative study this is a moderate gap, not fatal. Minor: trust dynamics are inferred partly from self-report, partly from logs, and the small N means the archetype counts (9/22, 6/22, 4/22) are fragile.\n\nNone of this sinks the paper. The central qualitative finding—that AI was treated as an initial guide rather than an authority, and that visualization served as an auditing mechanism—is consistent across quotes, logs, and the verdict agreement data. The design implications are reasonable for the studied population, and the open prototype is a real community resource.\n\nWho this is for: HCI and visualization researchers working on fact-checking, human-AI interaction, or mixed-initiative systems. It's exploratory, so don't expect causal claims. I'd send it to review, but request revisions on the generalizing language and the coding reliability. I'd also suggest a follow-up with a more diverse sample.\n\nRecommendation: engage seriously—accept with major revisions if the generalizing claims are tightened.","headline":"Useful design-probe study with open artifacts and honest limitations; the central claims are plausible but the task design partly pre-arranges what it observes.","tokens_in":22498,"tokens_out":1947,"would_cite":true,"duration_ms":20411,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Readers do not follow a fixed fact-checking pipeline: in a 22-person study they flexibly composed AI-first, manual-first, and parallel workflows, used self-built charts to audit AI conclusions, and calibrated trust by cross-tool agreement.","keywords":["data claims","fact-checking","human-AI interaction","trust calibration","visualization","design probe","large language models","mixed-initiative systems"],"falsifier":"Run the same FYI study with a sample matched to the target population (lower self-rated visualization literacy and sporadic generative-AI use) on articles that require multi-table joins: if most participants follow one fixed tool order, rarely open Chart Builder after Auto Check, or accept AI verdicts without inspecting the data, the three-archetype and visual-auditing claims fail to generalize.","tokens_in":1619,"feed_emoji":"📊","tokens_out":1641,"duration_ms":72962,"temperature":0.7,"pith_summary":"FYI is a browser extension that lets readers fact-check statistical statements in online articles against the underlying dataset, using four tools that range from full automation (an Auto Check pipeline and an AI chat) to hands-on inspection (a sortable table and a self-service chart builder). The paper uses FYI as an instrumented probe to answer a question prior systems have left open: when both AI and manual verification are available at once, how do people actually detect, verify, and judge data claims? In a 22-participant exploratory study, readers did not march through a fixed detection-verification-determination pipeline. Instead they composed three recurring workflow archetypes—AI-first with manual confirmation, manual-first with AI supplement, and parallel co-review—and they treated self-built visualizations as the primary mechanism for checking AI conclusions. The paper argues that trust in AI is calibrated dynamically, growing when independent tools converge and eroding when outputs disagree, and that future fact-checking systems should therefore treat AI as a starting point rather than an authority, while elevating visualization to a core verification capability.","feed_headline":"Readers fact-check by building charts to audit AI","feed_subtitle":"A 22-person study finds flexible AI-plus-manual workflows, with trust rising when separate tools agree.","key_machinery":"FYI is a browser extension that embeds the whole fact-checking workspace in a side panel beside the article under review. Its load-bearing design is a spectrum of four verification tools sharing one uploaded dataset: Auto Check (a four-step automated pipeline that streams an evidence chart and verdict), AI Chat (a multi-turn LLM with optional web search and client-side Python data analysis), Table Explorer (direct sorting and filtering of raw rows), and Chart Builder (a shelf-based chart authoring tool). A 25-type interaction logger records every user action as a timestamped event, so the reading session itself becomes observable data. This combination—multiple complementary modalities plus full provenance logging—is what allows the authors to characterize workflow archetypes and trust calibration rather than measuring only end-task accuracy.","core_discovery":"The paper's central empirical claim is that mixed-initiative data fact-checking is neither automation-led nor manual-led but a fluid, multi-modal activity. Analyzing 2,250 logged interactions, think-aloud protocols, and interviews from 22 participants, the authors identify three workflow archetypes—AI-first with manual confirmation (9/22), manual-first with AI supplement (6/22), and parallel co-review (4/22)—and show that individuals often moved between archetypes within a session. Chart Builder was the most-used per-claim tool (105 of 156 investigated claims), and in 55 instances participants deliberately opened it after Auto Check to audit the AI's conclusion; a self-built chart that aligned with or refuted a claim often served as the stopping criterion for a verdict, sometimes overriding AI outputs. Trust in AI shifted with experience: convergence of independent tools raised confidence, numerical inconsistencies eroded it, and one early AI error could collapse trust for the whole session. The authors also document a dominant 'AI initiates, human decides' reliance model, with Auto Check broadly accurate on straightforward claims (71% agreement with ground truth) but blind to the one claim requiring contextual judgment that the dataset could not settle, which participants were better at flagging as unverifiable.","pith_inferences":["If these workflow archetypes generalize, accuracy reporting for fact-checking tools should include tool-sequence analyses, not just verdict correctness, because a tool's value depends on how readers interleave it with other modalities.","The usability-confidence paradox suggests a concrete design test: giving low-literacy users AI-assisted chart suggestions should increase their willingness to audit AI outputs, and a controlled comparison of verdict confidence and override behavior could verify this.","The observed AI-on-AI cross-checks imply that future systems should distinguish 'grounded in raw data' from 'two language models happen to agree,' since agreement between LLMs is weaker evidence than a user-built chart derived from the dataset.","A natural extension is to run the same probe on finance or public-health articles whose claims require joins across tables; the plausible prediction is more manual-first behavior and more unverifiable verdicts, but the paper itself does not establish that."],"forward_implications":["Fact-checking systems should present AI verdicts as provisional hypotheses that invite human confirmation, rather than as definitive answers that require effort to override.","Visualization authoring should be treated as a primary auditing capability, with AI-assisted chart suggestions offered to users who cannot easily build charts themselves.","Interfaces should offer independent, freely composable tools instead of rigid step-by-step wizards, because participants adapted their workflow across claims and even within a single session.","Systems should expose intermediate reasoning, data queries, generated code, and confidence scores, since visible process supported trust calibration and selective override.","Designers of trust should expect calibration to be brittle: one inconsistent numeric output can outweigh many successes, so cross-tool consistency matters as much as average accuracy."],"supporting_citations":[{"why":"It supplies the canonical detection-verification-determination pipeline that frames the study's research questions and design.","marker":"[22]"},{"why":"It introduces the concept of data claims and an end-to-end automated fact-checking architecture that FYI positions itself against.","marker":"[18]"},{"why":"It provides the code-execution verification approach that Auto Check adapts for generating and running verification code.","marker":"[14]"},{"why":"It supplies the prior interactive system for contextualizing statistical statements with data exploration whose composite workflow findings this study extends with generative AI.","marker":"[33]"},{"why":"It establishes the paradigm of using charts to check text-caption consistency, which the visual-auditing finding extends to user-built charts.","marker":"[31]"},{"why":"It contributes the trust-calibration and user-reliance model used to interpret the dynamic trust findings.","marker":"[8]"},{"why":"It provides the think-aloud protocol used to capture participants' reasoning during verification.","marker":"[17]"},{"why":"It supplies the thematic analysis method through which the workflow and trust themes were derived.","marker":"[7]"}],"fun_headline_variants":["Readers audit AI fact-checks by building their own charts","AI-plus-manual workflows dominate data claim fact-checking","Trust in AI fact-checks hinges on tool convergence","Study: Flexible AI+manual workflows for fact-checking data"],"cache_read_input_tokens":24704,"weakest_assumption_plain":"The load-bearing premise is that the 22 study participants—university-affiliated, highly visualization-literate, and daily generative-AI users—behave like the lay readers with basic data and visualization literacy that FYI is designed for, when both groups fact-check a single data-driven article.","fun_headline_variants_meta":{"raw":{"variants":["Readers audit AI fact-checks by building their own charts","AI-plus-manual workflows dominate data claim fact-checking","Trust in AI fact-checks hinges on tool convergence","Study: Flexible AI+manual workflows for fact-checking data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1546,"prompt_tokens":1053,"completion_tokens":493,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":426}},"tokens_in":669,"tokens_out":493,"duration_ms":5009,"temperature":1.0,"reasoning_tokens":426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:28:34.123723+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same FYI study with a sample matched to the target population (lower self-rated visualization literacy and sporadic generative-AI use) on articles that require multi-table joins: if most participants follow one fixed tool order, rarely open Chart Builder after Auto Check, or accept AI verdicts without inspecting the data, the three-archetype and visual-auditing claims fail to generalize.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the prior interactive system for contextualizing statistical statements with data exploration whose composite workflow findings this study extends with generative AI."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the think-aloud protocol used to capture participants' reasoning during verification."}],"review_version":1}