{"id":"b2b26666-e9e8-4726-8c8d-0086184901ae","arxiv_id":"2607.25131","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Active Inference Fast and Slow agents for bar-chart averaging yield inspectable, parameter-tied failure modes (tick bias vs memory decay) as a framework for in silico visualization evaluation.","lead":"The paper builds Active Inference agents that simulate Fast and Slow chart reading on a two-bar average task, producing inspectable belief and fixation traces with distinct failure modes. It offers a path to test visualization designs in simulation before or alongside human studies.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The Fast/Slow contrast is confounded by asymmetric sensory precision and time budgets (bar_obs_sigma=0.0 vs pair_obs_sigma=0.22; T=10 vs T=7), so the \"mechanism-specific\" quantitative signatures may partly encode gifted parameters rather than strategy architecture.","rationale":"I read the paper in good faith as what it declares itself to be: a methodology proof-of-concept, explicitly not a human-validated model, with honest framing that the failure modes are \"expected consequences of the strategy definitions\" rather than discoveries. On that basis the reader's CONDITIONAL/HIGH verdict is well-calibrated, and the reader's weakest assumption (ecological validity of the hand-specified architectures) is a real but acknowledged limitation. My concern is distinct and internal: even granting the paper's own terms, the controlled-comparison claim in §1.2 is not met by the reported experiments, because the two agents differ in sensory precision and time budget in ways aligned with the direction of every headline contrast. This matters specifically because the paper's proposed empirical program treats the current quantitative curves (not just their qualitative existence) as fitting/falsification targets; unstable targets would compromise that program even before any human data arrive. The concern is cheap to settle — the authors release code, and the symmetrization re-runs are parameter changes, not new methods — so it strengthens rather than overturns the conditional verdict. I therefore recommend UNCHANGED (verdict remains CONDITIONAL), with the added explicit condition that signature robustness under symmetric sensory/time budgets be demonstrated alongside the reader's condition of empirical grounding. Partial agreement with the reader: same family of worry about hand-specification, but localized to internal validity of the comparison rather than transfer to human strategies.","tokens_in":26615,"tokens_out":3068,"duration_ms":96226,"concrete_test":"Re-run the baseline sweep (Fig. 5) and both perturbation experiments (Figs. 6–7) under symmetrized settings: (a) bar_obs_sigma = pair_obs_sigma at three shared values {0.0, 0.11, 0.22}; (b) a common trial horizon (e.g., T=10 for both, or equal per-bar time budgets); (c) memory-decay exposure normalized per stored value (scale the Slow model's per-step decay so that total decay per retained estimate matches the Fast model's single memory). Using the released code (github.com/NatLabRockies/AIF_for_Vis), if Slow baseline accuracy falls toward Fast levels, or the tick-bias crossover at segment_tick_anchor_bias≈0.5 vanishes or shifts materially, the reported signatures are parameter-sensitive and should be published as families over symmetrized settings before serving as falsification targets.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the Fast and Slow agents' distinct failure signatures (accuracy curves, failure maps, decay/bias sensitivities) follow from their architectural commitments and can therefore serve as falsifiable mechanistic hypotheses. Section 1.2 explicitly promises a \"matched parameterization\" enabling a \"controlled computational experiment.\" But Table 1 breaks that control in exactly the dimensions the results depend on. The Slow agent receives a noiseless coarse cue (bar_obs_sigma=0.0, i.e., the σ→0 deterministic snap of Eq. 56) while the Fast agent's gist cue carries σ=0.22; the Slow agent also gets a longer trial horizon (T=10 vs T=7). Two consequences follow. (a) The baseline gap (0.98 vs 0.92, Fig. 5) and the diagonal banded failure map (Fig. 8) are confounded: a noise-free bar-top anchor alone would yield near-ceiling accuracy and tick-aligned structure regardless of whether representation is bar-wise or compressed. (b) The perturbation curves used as \"quantitative signatures\" are not attributable to strategy alone: in the memory-decay sweep (Fig. 6) the Slow agent decays two memories per action (Eqs. 69–70) versus one for Fast, and operates under a different time budget; in the tick-bias sweep (Fig. 7) the crossing at segment_tick_anchor_bias≈0.5 depends on Slow's averaging cancellation plus its zero-noise coarse anchor. Because §§5–6 propose fitting precisely these curves and maps to human data, signatures that move under parameter symmetrization would make human-data matches absorbable by parameter tweaks and mismatches non-diagnostic. This does not sink the proof-of-concept contribution, but it undercuts the specific claim that the reported numbers are mechanism fingerprints rather than one point in a hand-tuned parameter family.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper proposes Active Inference as a framework for converting qualitative cognitive theories of visualization interpretation into executable process models, and demonstrates the translation on Padilla et al.'s dual-process account using a two-bar average-estimation task. Two discrete-state POMDP agents are constructed: a Fast (Type 1) agent maintaining a single compressed average memory with a gist–anchor–refine routine, and a Slow (Type 2) agent maintaining two bar-specific memories with sequential bar-wise estimation. Tick-salience bias is placed in the environment but not the agents' likelihood models, and memory decay is implemented as retention/forgetting/diffusion over memory states. Simulations show baseline accuracies of ~0.92 (Fast) vs ~0.98 (Slow), a diagonal-banded pairwise failure map, stronger Slow sensitivity to memory-decay sweeps, and stronger Fast sensitivity to segment tick-bias sweeps, plus inspectable belief and action traces. The authors frame the work strictly as a proof of concept and a call to parameterize, falsify, or refine such agents against human data.","tokens_in":27182,"tokens_out":4269,"duration_ms":34061,"significance":"If the framing holds up, this is a useful methodological contribution to the growing 'virtual viewer' literature: a fully specified, executable, and inspectable instantiation of a dual-process account of chart reading, with mechanism-specific failure signatures (accuracy curves, failure maps, step counts, belief traces) that are in principle falsifiable against human data. Strengths that deserve explicit credit: the appendices give a complete formal specification (cue likelihoods Eqs. 56–61, memory dynamics Eq. 13, full policy libraries Eqs. 91–110) at a level of detail rare in this venue; code is publicly released; the stimulus sweep is exhaustive (9610 runs per model at baseline); and the authors are unusually candid that the Fast/Slow decomposition is a modeling choice, that several results are expected consequences of the architecture rather than discoveries, and that the models are 'tentative mechanisms' pending empirical grounding. The paper does not overclaim predictive validity for human behavior. Its significance is as a scaffold and a call to action, and on those terms it is valuable.","major_comments":[{"comment":"§1.2 promises a 'matched parameterization' with parameters differing 'only in the architectural and temporal features required by their respective cognitive modes,' enabling a 'controlled computational experiment.' Table 1 does not deliver this control in the dimensions the results depend on. The Slow agent receives a deterministic coarse cue (bar_obs_sigma = 0.0, the σ→0 snap of Eq. 56) while the Fast agent's gist cue carries σ = 0.22, and the Slow agent receives a longer horizon (T=10 vs T=7). These two gifted parameters, not the representational architecture, may drive (a) the baseline gap (0.98 vs 0.92, Fig. 5) — a noise-free bar-top anchor alone yields near-ceiling accuracy regardless of whether memory is bar-wise or compressed — and (b) the diagonal banded failure structure of Fig. 8. Because §§5–6 propose fitting precisely these accuracy curves and failure maps to human data as 'q","section":"§1.2, Table 1, Figs. 5–8"},{"comment":"The perturbation curves are presented as quantitative signatures usable to parameterize or falsify the strategies against human data (§5, §6), but their quantitative content depends on unconstrained preference/effort parameters that are not themselves candidates for the mechanisms under test. The crossing of Fast and Slow accuracy at segment_tick_anchor_bias ≈ 0.5 (Fig. 7) and the accuracy optimum at decay rate ≈ 0.2 (Fig. 6) depend on C_fb = (0, −80, 12) (Eq. 35), c_nr = 0.35, γ = 8.0, mem_sigma = 5.0, and the segment_sigma = 0.045 refinement width. A sensitivity analysis varying these nuisance parameters over a plausible range — showing that the ordinal, mechanism-level signatures (Slow decays faster under ρ_mem sweeps; Fast decays faster under λ_tick sweeps) are robust while only crossing points move — is needed before these curves can serve as identifiable fitting targets. Without it","section":"§4.2–4.3, Figs. 6–7, Eq. (35), Table 1"},{"comment":"The degenerate all-report policy templates ([REPORT r ×4], Eqs. 101 and 110) are included in both policy libraries alongside report-ending templates, with report_action_instant = True and strong feedback preferences. The manuscript does not report how often degenerate or truncated policies are actually selected. Since step counts (Figs. 13–14) and the belief-trace interpretation (§A.2, Fig. 10) are offered as behavioral signatures, the authors should report the empirical distribution over selected template families per condition, and confirm that the deliberation-time effects in Figs. 13–14 are not artifacts of template-library composition (e.g., the Fast library containing shorter paths to report). This is a reporting gap rather than a suspected error, but it is load-bearing for the step-count and trace claims.","section":"Appendix D.3, Eqs. (101)/(110), Figs. 13–14"}],"minor_comments":[{"comment":"The Fig. 8 caption states 'Green regions indicate Fast model failure,' while the Appendix A.1 text states 'Red areas indicate bar pairs for which the Fast agent failed... more often than the Slow model.' Please reconcile.","section":"Fig. 8 caption vs. §A.1"},{"comment":"Table 1 lists the default mem_sigma = 0.25, while §4.2 states the decay experiment used mem_sigma = 5.0. Please clarify which values are 'defaults' and flag the override in the table or caption.","section":"Table 1 vs. §4.2"},{"comment":"Eq. (6) defines B_ij without an action index while the surrounding text describes B(u) as 'the transition matrix induced by action u'; the known-action operator form (Eqs. 25–29) later reintroduces action conditioning. A sentence reconciling the two notations would help readers.","section":"§2.2, Eq. (6)"},{"comment":"Typos/grammar: §1.2 'capture capture'; §3.2 'the hight of each bar'; 'course' used throughout for 'coarse' (course_tick_anchor_bias, 'course initial cue'); Fig. 6 and Fig. 7 captions 'random seems'; §A.2 'Padillaet al.' and 'starts by with the LOOK_PAIR action'; §5 'interpretable timeseries'; §6 'whenwhy the viewer responds'; §2.1 'an architecture that express both' (missing 'es'); Fig. 10 caption 'actions show as the series' and 'the believe average'.","section":"Various"},{"comment":"Fig. 8 visualizes only the accuracy difference between models. Since Slow accuracy is near ceiling, absolute per-pair accuracies (available in Fig. 9) would make the diagonal-band claim easier to assess independently of the difference scaling.","section":"Fig. 8"},{"comment":"Abstract and Fig. 1 caption contain spacing artifacts from PDF extraction/typesetting ('Type 1,Fast', 'forin silico', 'in silicoevaluation'); please check the final proofs.","section":"Abstract, Fig. 1"},{"comment":"The supplemental link is given as '<removed for review>' immediately followed by the live GitHub URL; given the non-anonymous author list this is presumably intentional, but the placeholder should be cleaned up.","section":"Supplemental Materials"}],"recommendation":"major_revision","confidential_remarks":"The manuscript carries a TVCG/VIS copyright line indicating prior publication; I have evaluated it as a standard submission on its merits. The author team includes the lead author of the framework being formalized (Padilla et al. [19]); this is disclosed by the author list itself and is not improper, but reviewers and readers should be aware the target theory and the model originate from overlapping groups. The parameter-symmetrization concern in my major comments is, in my view, the one issue that must be resolved before the \"matched controlled experiment\" framing can stand; it is fully addressable with additional simulation runs, and I would not want the paper rejected over it if the venue permits revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful, executable translation of Padilla et al. into Fast/Slow Active Inference agents on two-bar average estimation, with belief traces, failure maps, and parameter sweeps you can actually inspect. It does not claim to replace user studies, and it mostly stays inside that lane.\n\nWhat is new is the concrete factorization—compressed average memory vs two bar memories, staged LOOK_PAIR/BAR → TICK → SEGMENT → REPORT policies, environment-side tick bias the agent does not know about, and the quantitative signatures those choices produce. The appendices spell out A/B/C/D, memory decay, and report projection clearly enough to reproduce. Baseline accuracies (~0.92 vs ~0.98), memory-decay hurting Slow harder, and tick bias hurting Fast faster are coherent with the architectures they built. Code pointer and full math are real credit.\n\nSoft spots, in proportion. The “matched parameterization / controlled computational experiment” language is oversold. Table 1 gifts Slow a noiseless coarse cue (bar_obs_sigma=0) and a longer horizon (T=10 vs 7) while Fast gets pair_obs_sigma=0.22. They justify the noise as “eye-balled midpoint,” which is fair as a strategy difference, but then baseline gaps and some sweep shapes are not pure architecture fingerprints; they partly encode those free choices. The paper already says the failure modes are expected from the definitions—good—but §§5–6 still pitch those curves as the things to fit to humans, so confounded signatures will make matches easy and mismatches non-diagnostic until they symmetrize or fit jointly. No human data yet; circularity is acknowledged, not hidden. Hand-built policy library and symbolic cues are fine for a PoC and flagged in Limitations.\n\nMath and citations look solid: Friston/Active Inference, Padilla, graphical perception, and the virtual-viewer/scanpath line are in the right places; they are not reinventing saliency models and claiming novelty there.\n\nWho it is for: people who want mechanistic, falsifiable process models in viz/HCI, not designers hunting a drop-in evaluator tomorrow. I would bring it to a reading group that cares about cognitive modeling or evaluation methodology. It deserves serious referee time—revise the matched-parameter claim, stress-test under symmetrized noise/horizons, and keep the empirical program front and center. Engage; do not over-read the sweeps as discoveries about people.","headline":"Solid proof-of-concept that turns Padilla-style dual-process chart reading into runnable Active Inference agents with inspectable failure signatures—useful scaffold, not yet human-validated mechanism discovery.","tokens_in":28057,"tokens_out":616,"would_cite":true,"duration_ms":21989,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Active Inference turns dual-process chart reading into executable agents whose distinct failure modes can be simulated before user studies.","keywords":["Perception and Cognition","Visualization Design and Evaluation","Cognitive Modeling","Active Inference","Decision Making","Dual-process theory","Graphical perception"],"falsifier":"Run the same two-bar average task with humans and check whether error maps, step/fixation counts, and sensitivity to memory load versus tick salience match the Fast versus Slow predictions; systematic mismatch falsifies the chosen strategy decomposition or its parameters.","tokens_in":27744,"feed_emoji":"📊","tokens_out":929,"duration_ms":32511,"temperature":0.7,"pith_summary":"User studies tell us which charts work after they are built, but they rarely give a causal, predictive account of why a viewer succeeds or fails. This paper shows that a dual-process account of visualization judgment can be turned into runnable Active Inference agents that treat chart reading as sequential visual search: agents update beliefs and choose where to look by balancing uncertainty reduction against cognitive effort. As a proof of concept, a Fast agent estimates a compressed bar-pair average and a Slow agent estimates each bar then averages; the Fast agent is more hurt by tick-salience bias, the Slow agent by working-memory decay. Both produce inspectable traces—belief uncertainty, fixation sequences, and stimulus-specific error maps—so the hypothesized mechanisms become parameters that later human data can fit, refine, or falsify. The aim is earlier in silico evaluation of visualization efficacy, with empirical studies used to ground the simulations rather than only to grade finished designs.","feed_headline":"Simulated chart readers fail in human-like ways","feed_subtitle":"Fast and Slow agents predict tick bias vs memory decay before you run the user study","key_machinery":"Discrete-time Active Inference agents cast as POMDPs: hidden states hold task register and working-memory contents (one average memory vs two bar memories), observations are staged cues (relative position, axis tick, segment, feedback), and actions are chosen from a short policy library by minimizing expected free energy that trades epistemic information gain against effort and preference for correct reports.","core_discovery":"Active Inference can instantiate dual-process strategies for a two-bar average-estimation task as matched Fast (compressed average-first) and Slow (bar-wise sequential) agents, such that hypothesized human vulnerabilities—tick-salience bias for the Fast agent and working-memory decay for the Slow agent—appear as distinct quantitative signatures in accuracy, failure maps, step counts, and belief/fixation traces, providing a formal scaffold for mechanistic hypothesis testing about visualization interpretation.","pith_inferences":["If the parameter-fitting pipeline works, visualization toolkits could ship with default agent profiles (high tick bias, low memory retention) as automated design linters.","The same Fast/Slow split may expose when common encodings (stacked bars, dual axes) systematically push viewers into the more fragile strategy.","Mismatch between model and eye-tracking would most cleanly revise the observation modalities or policy library rather than abandon the free-energy objective.","Pairing an inspectable agent engine with a natural-language front end could let non-modelers queue in silico studies without writing POMDP matrices."],"forward_implications":["Designers could stress-test encodings in simulation for tick-anchoring or memory-load failure before committing to full user studies.","Empirical work shifts from only grading finished charts to also fitting, refining, or rejecting explicit process models.","Failure maps, belief trajectories, and fixation sequences become shared quantitative targets across model and human data.","A later hierarchical arbiter could switch between Fast and Slow policies when uncertainty exceeds a threshold, predicting when a design forces analytic effort.","Scanpath and virtual-viewer models can propose observation and action vocabularies for new tasks that Active Inference then mechanizes."],"fun_headline_variants":["Active Inference models Fast and Slow chart readers with human-like biases","Simulated Fast agents show tick-salience bias; Slow ones show memory decay","Dual-process Active Inference agents predict chart interpretation errors","In silico Fast vs Slow readers expose distinct bar-chart failure signatures","Executable Active Inference scaffold tests visualization error mechanisms"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hand-built Fast and Slow architectures, staged look-then-report routines, and restricted policy library are adequate stand-ins for how people actually read these charts; if humans use different strategies or cues, the simulated failure signatures will not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Active Inference models Fast and Slow chart readers with human-like biases","Simulated Fast agents show tick-salience bias; Slow ones show memory decay","Dual-process Active Inference agents predict chart interpretation errors","In silico Fast vs Slow readers expose distinct bar-chart failure signatures","Executable Active Inference scaffold tests visualization error mechanisms"]},"model":"grok-4.5","effort":"low","cost_usd":0.003669,"raw_usage":{"total_tokens":1212,"prompt_tokens":792,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":36688000,"prompt_tokens_details":{"text_tokens":792,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":350,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":792,"tokens_out":70,"duration_ms":7203,"temperature":1.0,"reasoning_tokens":350,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T00:28:22.171558+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same two-bar average task with humans and check whether error maps, step/fixation counts, and sensitivity to memory load versus tick salience match the Fast versus Slow predictions; systematic mismatch falsifies the chosen strategy decomposition or its parameters.","supporting_citations":[],"review_version":1}