{"id":"41c49dec-ccc2-4ee2-b04c-4796ae39fae7","arxiv_id":"2607.09018","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Higher variance in simultaneously shown AI-style design sets increases human selection of center-proximal designs in both preference and representativeness tasks.","lead":"People picking among AI design options favor the middle-looking ones more when the set is more varied. This matters because AI tools that show many options at once may quietly push users toward safer, average choices.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Stimulus construction confounds set variance with absolute extremity of options, so the high-variance effect may not isolate ensemble central-tendency bias.","rationale":"The reader's weakest_assumption correctly flags ecological validity of hand-constructed stimuli, but the more load-bearing internal threat is the confound between variance and absolute extremity that is baked into the continuum sampling itself. That confound directly threatens the causal attribution of the strongest claim (higher variance → more center selection via ensemble perception). The paper's own Discussion acknowledges extremeness aversion and compromise effect as complementary explanations yet offers no design that disentangles them; the proposed third condition would do so cleanly. Because the statistical contrasts are real and the manipulation check succeeded, the claim is not overturned, only its mechanistic interpretation is under-specified; CONDITIONAL remains appropriate, with the additional requirement that future work (or a revision) separate range from variance. Agreement with the reader is therefore partial: same overall verdict, different primary soft spot.","tokens_in":8341,"tokens_out":586,"duration_ms":7181,"concrete_test":"Construct a third matched condition that keeps the same absolute extremes as high-variance (includes -3 and +3) but reduces interior dispersion (e.g., {-3,-3,0,0,0',0,+3,+3} or {-3,-1,0,0,0',0,+1,+3}), re-run the preference and representativeness tasks with N≈50, and test whether center-selection rates remain elevated relative to the original low-variance set. If rates drop to low-variance levels once interior variance is reduced while extremes remain, the ensemble-variance interpretation is weakened; if rates stay high, the original claim is supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that higher set variance increases center-proximal selection via ensemble perception. High-variance sets are defined as {-3,-2,-1,0,0',+1,+2,+3} and low-variance as {-1,-1,0,0,0',0,+1,+1} (Method, stimulus construction). Thus high-variance sets uniquely contain the extreme poles (-3,+3). Extremeness aversion / compromise effect (Discussion cites Chernev 2004; Simonson & Tversky 1992) predict avoidance of those poles and therefore higher relative selection of the shared center items, independent of any ensemble mean extraction. Because the design never holds the option set's absolute range fixed while varying only dispersion around a matched centroid, the observed 56% vs 34% (preference) and 66% vs 38% (representativeness) contrasts cannot cleanly attribute the increase to ensemble mechanisms rather than classic choice-set composition effects. The paper itself notes these alternative accounts but does not experimentally separate them.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper claims that simultaneous multi-option presentation of AI-generated design variations can induce central tendency bias via ensemble perception: users preferentially select designs near the perceptual center of a set, and this bias strengthens when set variance is higher. In a within-subjects experiment (N=50), participants viewed controlled 8-poster grids under high-variance ({-3..+3}) versus low-variance (clustered near 0) conditions and selected once for aesthetic preference and once for set representativeness. Center-selection rates rose with variance in both tasks (preference 56% vs 34%, t(49)=3.51, p=0.001; representativeness 66% vs 38%, t(49)=4.34, p<0.001), with a successful perceived-diversity manipulation check. The authors conclude that multi-variation interfaces may constrain selection diversity even when generation diversity is high, and discuss interface mitigations.","tokens_in":8621,"tokens_out":946,"duration_ms":9187,"significance":"If the result holds, it identifies a concrete, interface-level cognitive bottleneck in human–AI co-creation that is currently under-studied relative to generation quality. The work usefully bridges ensemble-perception theory to evaluative selection of complex visual artifacts, reports a clean within-subjects design with counterbalancing, randomized grid positions, and a strong manipulation check, and surfaces a practical tension between output diversity and selection diversity. The dual-task structure (preference plus representativeness) and self-report reason data strengthen the claim that the bias is not merely a conscious typicality strategy. These contributions are relevant to HCI and design-tool research even if the precise mechanism remains partly ambiguous.","major_comments":[{"comment":"Method (stimulus construction) and Discussion: high-variance sets uniquely include the absolute poles (-3, +3) while low-variance sets do not. Extremeness aversion / compromise-effect accounts therefore predict elevated center selection under high variance without requiring ensemble mean extraction. The paper cites these alternatives but does not hold absolute range fixed while varying only dispersion around a matched centroid, so the reported contrasts (Table I; 56% vs 34%, 66% vs 38%) cannot cleanly attribute the increase to ensemble mechanisms. A load-bearing revision is either an additional condition that equates range or a substantially stronger experimental argument that the present design isolates ensemble coding.","section":null},{"comment":"Method and Limitations: stimuli are manually constructed posters ordered by three-rater holistic consensus along a sparse-to-busy continuum, not live model outputs. While control is understandable, this weakens the ecological claim that the results speak to Midjourney/DALL·E-style interfaces. The central applied claim (multi-variation AI interfaces constrain selection diversity) therefore rests on an untested transfer assumption that should be either tested with real generative outputs or more tightly scoped in the abstract and conclusions.","section":null}],"minor_comments":[{"comment":"Table I and Results: report effect sizes (e.g., Cohen’s d or equivalent) alongside the paired t-tests so readers can judge practical magnitude beyond p-values.","section":null},{"comment":"Figure 1 caption and Method: clarify whether the two near-centroid items (0 and 0') were visually distinct enough that participants treated them as separate options rather than near-duplicates; any recognition of duplication could itself bias center rates.","section":null},{"comment":"Self-reported reasons (Table II) are pooled across variance conditions; a brief breakdown by high vs low variance would help confirm that reason distributions remain stable under the manipulation.","section":null},{"comment":"Related Work / Discussion: a short note on whether positional randomization fully eliminates known grid-position biases (center-of-screen preference) would strengthen the design description.","section":null},{"comment":"Minor typography: “USERSTUDY”, “RELATEDWORK”, and similar concatenated headings should be spaced for readability.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core empirical pattern is clear and the design is careful on many dimensions; the main risk is over-claiming a specific ensemble mechanism when the variance manipulation is confounded with absolute extremity. If the authors can either add a range-matched control or substantially temper the mechanistic language and ecological claims, the paper becomes a solid HCI contribution. Fit for a serious HCI / design-cognition venue is good once the attribution gap is addressed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is a clear, readable user study: when people see eight poster designs at once, higher set variance raises center-proximal picks in both preference (56% vs 34%) and representativeness (66% vs 38%) tasks, with solid paired t-tests and a strong manipulation check. That is new data for the selection stage of generative-AI workflows, not just another generation-fixation paper.\n\nWhat they do well: within-subjects design, counterbalanced task order, randomized grid positions, two near-centroid items held constant, and explicit self-report reasons that look similar for center vs non-center choices. The citation base is appropriate (ensemble perception plus the compromise/extremeness literature they themselves flag). No circular stats; the dependent measure is independent of the manipulation.\n\nThe soft spot that matters is the one the stress-test flags. High-variance sets are {-3..+3}; low-variance sets are clustered around 0 with no poles. So the contrast confounds dispersion with the presence of extremes. Extremeness aversion or compromise effects predict exactly the same pattern without any ensemble mean extraction. The authors note the alternatives in Discussion but never hold range fixed while varying only variance around a matched centroid. Hand-crafted stimuli and modest N are secondary limitations they already own; the confound is the load-bearing one for the mechanism claim.\n\nStill, the directional finding is real and the interface implication (multi-option grids may shrink selection diversity even when generation diversity rises) is worth having on the table for HCI and tool designers. I would bring it to reading group as a methods discussion piece, cite the empirical rates if I am writing about selection interfaces, and send it to peer review. Referees should demand a cleaner variance-only control or an explicit test against the decision-level accounts, but the paper is already serious enough to deserve that conversation.","headline":"Clean within-subjects result that multi-option design sets with higher variance pull selections toward the center, but the high-variance condition also uniquely adds extreme poles, so ensemble perception is not cleanly isolated from classic extremeness aversion.","tokens_in":9139,"tokens_out":477,"would_cite":true,"duration_ms":5437,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When AI design tools show more varied options at once, people pick the middle ones more often, in both preference and typicality judgments.","keywords":["Image-generation AI","Human–AI collaboration","Ensemble perception","Design selection","Central tendency bias","Multi-option presentation","Creative workflows"],"falsifier":"Run the same preference and representativeness tasks with live outputs from commercial image models that truly vary in stylistic spread, or present the same high-variance options sequentially rather than simultaneously; if center-selection rates no longer rise with variance (or fall under sequential presentation), the interface-ensemble account is undermined.","tokens_in":9266,"feed_emoji":"🎨","tokens_out":818,"duration_ms":9356,"temperature":0.7,"pith_summary":"This paper claims that the common grid layout of image-generation AI tools can systematically pull human choices toward the visual center of the set. Drawing on ensemble perception—the automatic extraction of summary statistics from groups of objects viewed together—the authors argue that higher variance among simultaneous design alternatives strengthens that central pull. In a controlled experiment, participants saw high- and low-variance sets of eight poster designs and chose once for personal preference and once for which design best represented the set. Higher variance raised the rate of center-proximal selections in both tasks. The practical stakes are clear: systems built to expand creative options may, through the way those options are shown, shrink the diversity of what users actually select.","feed_headline":"More varied AI design grids pull people toward the middle","feed_subtitle":"Higher set variance raised center picks in both liking and typicality tasks, shrinking selection diversity","key_machinery":"Ensemble perception / central tendency bias: the automatic extraction of a set-level mean representation that anchors evaluation toward options near the perceptual center; the experiment manipulates set variance (high vs. low continua of poster designs) while holding the two center items constant across conditions.","core_discovery":"Higher variance in simultaneously presented sets of AI-style design variations increases selection of center-proximal designs, both when people choose what they like most and when they choose which design best represents the overall style of the set.","pith_inferences":["Default 2×2 or 2×4 grids in tools like Midjourney-style interfaces may be quietly producing more conservative final selections than the raw diversity of the model suggests.","The same bias could appear in other multi-option creative UIs (logo variants, layout systems, product configurators) whenever options share a visual continuum.","A/B tests that only measure generation diversity or click-through may miss a selection-stage bottleneck that flattens creative outcomes.","If sequential or subset presentation reduces the bias, product teams have a low-cost lever that does not require changing the underlying model."],"forward_implications":["Multi-option grids in generative tools can reduce selection diversity even when generation diversity is high.","Raising output variance alone may strengthen rather than weaken the pull toward average-looking options.","Self-reported reasons for liking a design may stay stable while actual choices still shift toward the set center.","Interface interventions that break simultaneous ensemble extraction (sequential presentation, visual emphasis on extremes) become a design lever for preserving selection diversity.","Evaluations of generative AI creative tools should treat selection-interface structure as a first-class variable alongside generation quality."],"fun_headline_variants":["Higher-variance AI design sets drive more center picks","Wider AI variation grids bias people to midpoint options","Central tendency bias rises with varied AI design grids","Bigger spreads of AI designs shrink selection diversity","More variance in AI grids pulls selections toward the middle"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim rests on the idea that hand-built poster sets ordered by holistic sparse-to-busy ranking, shown in grids, are a fair stand-in for the variance and presentation structure of real AI image-generation interfaces.","fun_headline_variants_meta":{"raw":{"variants":["Higher-variance AI design sets drive more center picks","Wider AI variation grids bias people to midpoint options","Central tendency bias rises with varied AI design grids","Bigger spreads of AI designs shrink selection diversity","More variance in AI grids pulls selections toward the middle"]},"model":"grok-4.5","effort":"low","cost_usd":0.002576,"raw_usage":{"total_tokens":878,"prompt_tokens":673,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":25760000,"prompt_tokens_details":{"text_tokens":673,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":131,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":673,"tokens_out":74,"duration_ms":2212,"temperature":1.0,"reasoning_tokens":131,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T00:57:48.703404+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same preference and representativeness tasks with live outputs from commercial image models that truly vary in stylistic spread, or present the same high-variance options sequentially rather than simultaneously; if center-selection rates no longer rise with variance (or fall under sequential presentation), the interface-ensemble account is undermined.","supporting_citations":[],"review_version":1}