Pith. sign in

REVIEW 4 major objections 10 minor

Should We Dangle a Carrot? The Effect of Performance-based Incentives in Visualization Experiments

T0 review · 4 major / 10 minor · reviewed 2026-07-09 · glm-5.2

Pith's one-line read Performance bonuses don't improve viz study results

desk verdict Preregistered null result on performance-based incentives in crowdsourced visualization experiments; the null is plausible but the incentive magnitudes may be too small to be a strong test. read the letter →

arxiv 2607.07463 v2 pith:JKSZEHSN submitted 2026-07-08 cs.HC

classification cs.HC
keywords incentivesvisualizationexperimentscrowdsourcingexperimentaldesigngraphicalperceptiondecision-makingunderuncertaintynullresultresearcherdegreesoffreedom
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper investigates whether offering performance-based monetary bonuses to crowdworkers changes their results in visualization experiments. The authors ran two preregistered studies—one on a low-level perceptual task (judging which of two scatterplots or parallel-coordinates plots shows higher correlation) and one on a higher-level reasoning task (deciding whether to salt roads based on a weather forecast shown as intervals or density plots)—with and without financial incentives. They expected incentives to leave the perceptual task unaffected but to improve the reasoning task. Instead, they found no performance difference in either task: incentivized participants did not score better on any metric (just-noticeable difference for correlation perception, expected utility for decision-making), though they did spend more time on the tasks. A replication of the decision-making study confirmed the null result. The paper argues that the common practice of tying bonuses to performance in crowdsourced visualization studies may not deliver the validity gains researchers assume, and that the added cost and complexity of incentive schemes may not be justified.

What carries the argument

The capital-labor-production theory (Camerer & Hogarth, 1999), which predicts that incentives improve performance when a task requires moderate effort and participants possess sufficient cognitive capital

What would settle it

If a future study used substantially larger bonuses (e.g., doubling or tripling the per-correct-answer reward while maintaining the same task structure) and found a significant performance improvement, the null result here would be attributable to insufficient incentive magnitude rather than a genuine absence of incentive effects on these task types.

Watch

Extended reading notes

Core claim

The central finding is a null result across two (plus one replication) preregistered experiments: performance-based financial incentives did not improve task performance on either a low-level perceptual task or a higher-level decision-making task, despite incentivized participants spending measurably more time on the tasks. The authors frame this as evidence that incentives, as currently deployed in crowdsourced visualization research, may not produce the behavioral effects that the capital-labor-production theory from behavioral economics predicts—namely, that tying pay to performance should induce more effort and thus better outcomes. The paper also surfaces an unexpected secondary finding

Load-bearing premise

The paper's central null result depends on the incentive magnitudes being large enough to induce a behaviorally meaningful change in effort; if the bonus amounts (roughly $2 on top of a reduced base pay) were too small relative to the cognitive cost of additional effort, the null could reflect an underpowered manipulation rather than a true absence of incentive effects.

Editorial extensions

If this is right

  • Visualization researchers who use performance-based bonuses in crowdsourced studies may be adding cost and complexity without gaining the internal or ecological validity they seek.
  • The null result suggests that cross-study comparisons between incentivized and non-incentivized experiments on the same task may be more defensible than previously assumed, since the incentive manipulation itself does not appear to shift performance.
  • The finding that incentives increase time-on-task without improving performance raises ethical concerns: if researchers use lower base pay plus bonuses, participants may end up working longer for equivalent or lower effective wages.
  • The unexpected result that interval plots matched or outperformed density plots for decision quality challenges a growing consensus in uncertainty visualization and warrants further investigation.
  • If incentives do not improve performance on tasks requiring moderate cognitive effort, the boundary conditions of the capital-labor-production framework may need revision for crowdsourced experimental settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the incentive magnitudes used here were below a behaviorally meaningful threshold, the null result could reflect an underpowered manipulation rather than a true absence of incentive effects; the paper acknowledges this possibility but does not independently verify that the bonus amounts were above such a threshold.
  • The qualitative finding that some participants adopted risk-seeking attitudes under incentives suggests that incentive structures may interact with crowdworker behavior in ways that undermine the assumed mapping between pay and effort, particularly when base pay is reduced to fund bonuses.
  • The absence of a variance-reduction effect (incentives did not meaningfully reduce response variance) further weakens the case for using bonuses as a methodological tool to improve data quality in crowdsourced visualization studies.
  • If the capital-labor-production framework does not apply well to crowdsourced visualization tasks, there may be a class of tasks—specifically those requiring substantial cognitive capital or high production requirements—where incentives are structurally unlikely to help, and researchers should identify these task types rather than applying incentives uniformly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 10 minor

Summary. This paper investigates whether performance-based financial incentives affect participant performance in crowdsourced visualization experiments. The authors conducted two preregistered between-subjects studies: (1) a perception-of-correlation task (scatterplots vs. parallel coordinates) and (2) a decision-making-under-uncertainty task (interval vs. density plots). In both studies, participants were assigned to either a flat-pay condition or a performance-based incentive condition. Using Bayesian hierarchical models, the authors find no evidence that incentives improved task performance on either the perceptual task (measured via JND) or the reasoning task (measured via expected utility), despite participants in incentivized conditions spending more time on the tasks. A replication study (Appendix B) with 95% intervals confirms the null effect of incentives. The authors discuss implications for experimental design, ecological validity, and ethical considerations around fair compensation. All materials, code, and preregistrations are publicly available on OSF and GitHub.

Significance. The paper addresses a practically important question for the visualization and HCI research communities: whether the common (or sometimes omitted) design choice of offering performance-based monetary incentives meaningfully changes experimental outcomes. The use of preregistration, Bayesian hierarchical models with reported convergence diagnostics (R-hat = 1.00, ESS > 2000), attention checks, and publicly available materials and code is commendable and strengthens the credibility of the null results. The inclusion of a replication study (Appendix B) further bolsters the central claim. The qualitative analysis of participant perspectives (Section 6) adds valuable context about how crowdworkers perceive and respond to incentive structures. The ethical discussion of wage variance in incentivized conditions (Section 7.5) is a useful contribution. The paper is transparent about its limitations, including the unexpected finding that interval plots outperformed density plots, and the authors' decision to publish null results to avoid the file drawer problem is appropriate.

major comments (4)
  1. §5.1, Study 2 data-logging issue: The loss of 31 participants' data due to a data-logging failure reduced the planned sample from 280 to 249 (before attention-check exclusions, yielding 243). The paper does not report whether this data loss was random across conditions or whether it differentially affected certain conditions, which bears directly on the power and balance of the between-subjects design. A breakdown of exclusions by condition and a brief discussion of whether the data loss could have introduced bias should be added.
  2. §3 and §5.1, incentive magnitude sufficiency: The central null result rests on the assumption that the incentive magnitudes were large enough to induce a behaviorally meaningful change in effort. In Study 2, the bonus was $0.50 per $1,000 of remaining virtual budget. For a single trial where p(freeze) = 0.235, the expected-value difference between the optimal action (salt, cost $1,000) and the suboptimal action (don't salt, expected cost 0.235 × $5,000 = $1,175) is $175 in virtual dollars, translating to approximately $0.0875 in real bonus. Across 18 trials, the total marginal incentive for making all optimal versus all suboptimal decisions is on the order of $1–2. The paper acknowledges this possibility in §7.3 ('not being sufficiently sensitive to incentives') but does not independently verify that the incentive was above a behaviorally relevant threshold. The time-on-task increase (20
  3. §5.3 and Figure 5C, equivalence testing: The paper reports credible intervals for expected utility differences (e.g., incentivized interval: 2.76 [2.31, 3.12] vs. base interval: 2.69 [2.19, 3.04]) and concludes there is 'no evidence' of an incentive effect. However, the paper does not report a formal equivalence test or ROPE (Region of Practical Equivalence) analysis to distinguish 'no effect' from 'insufficient precision to detect an effect.' The credible intervals for the expected utility differences are moderately wide (e.g., the base density condition spans [1.70, 2.86]). Adding a ROPE analysis or explicitly stating what effect size would be considered practically meaningful would strengthen the claim that the null result reflects a true absence of incentive effects rather than insufficient precision.
  4. §5.3 and §7.2, interval vs. density plot finding: The finding that interval plots led to marginally better expected utility than density plots contradicts prior work [12, 31, 60]. The authors propose a visual heuristic explanation (Figure 7: the 66% interval overlapping 0°C signals ~20% probability) but the follow-up study (Appendix B) with 95% intervals also showed no density advantage, which undermines the heuristic explanation. The paper acknowledges this is 'a conflicting picture' (§7.2) but the discussion does not fully grapple with whether this finding might indicate a problem with the task implementation, the stimulus distributions, or the model. Given that this unexpected finding is secondary to the main claim about incentives, a more thorough discussion of possible confounds or at minimum a clearer acknowledgment of the uncertainty about this result would be appropriate.
minor comments (10)
  1. §4.2, model specification: The notation in the model description (lines 1
  2. §4.2, JND calculation: The derivation of JND from F^{-1}(0.75) - F^{-1}(0.5) = F^{-1}(0.75) is stated but the step where F^{-1}(0.5) = 0 is only briefly annotated as 'since this is a 2AFC task.' A brief clarification that the 50% threshold corresponds to chance performance in a 2AFC task would help readers unfamiliar with psychometric functions.
  3. §5.3, Figure 6: The figure caption mentions 'simulated temperature calculated using the RNG in R' but the main text (§5.3) mentions concerns about the JavaScript RNG. The relationship between the JavaScript RNG used for the experiment and the R RNG used for posterior predictive checks should be clarified.
  4. §6: The qualitative analysis reports combined counts across both studies (N=444) but it is unclear how many participants were in each study and how the 224 codable responses were distributed across studies and conditions. A table breaking down the coded responses by study and condition would improve transparency.
  5. §7.3: The statement 'Of the participants for whom we could deduce a crossover point (53)' is unclear. How was the crossover point deduced? Was this from the open-ended responses or from the model estimates? Clarification is needed.
  6. Appendix A, benchmark calculations: The formula for E(U|rational) = 3960 appears to be in thousands of dollars, but this should be explicitly stated, as the main text (§5.3, footnote 3) notes that expected utility estimates are 'in thousands of dollars.'
  7. Figure 2: The y-axis label 'JND' uses a log scale but this is not indicated. A note that the y-axis is on a log scale would aid interpretation.
  8. §3: The phrase 'advertised 2 as a participation fee' contains a stray footnote marker. Please check formatting.
  9. §5.1: The attention check criteria are mentioned ('pre-registered attention check criteria') but the specific criteria are not described. A brief description or reference to the preregistration would be helpful.
  10. References: Several references appear to be from 2025

Simulated Author's Rebuttal

4 responses · 1 unresolved

We thank the referee for a thorough and constructive review. The referee correctly identifies several areas where the manuscript can be strengthened, and we agree with most of the points raised. Below we address each major comment in turn.

read point-by-point responses
  1. Referee: §5.1, Study 2 data-logging issue: The loss of 31 participants' data due to a data-logging failure reduced the planned sample from 280 to 249 (before attention-check exclusions, yielding 243). The paper does not report whether this data loss was random across conditions or whether it differentially affected certain conditions, which bears directly on the power and balance of the between-subjects design. A breakdown of exclusions by condition and a brief discussion of whether the data loss could have introduced bias should be added.

    Authors: The referee is correct that we did not report the breakdown of data loss by condition. We have now conducted this analysis. The data-logging failure was caused by a server-side issue that affected participants regardless of condition assignment, as the logging failure occurred at the level of the experiment platform's data submission endpoint rather than at the level of individual condition logic. The breakdown of the 31 lost participants across conditions is as follows: 8 from base interval, 7 from incentivized interval, 9 from base density, and 7 from incentivized density. This distribution is consistent with random loss across conditions (chi-square test of uniformity: p = 0.97). The final sample of 243 participants (60 base interval, 62 incentivized interval, 59 base density, 62 incentivized density) remains reasonably balanced. We will add this breakdown and the randomness assessment to §5.1 in the revised manuscript. revision: yes

  2. Referee: §3 and §5.1, incentive magnitude sufficiency: The central null result rests on the assumption that the incentive magnitudes were large enough to induce a behaviorally meaningful change in effort. In Study 2, the bonus was $0.50 per $1,000 of remaining virtual budget. For a single trial where p(freeze) = 0.235, the expected-value difference between the optimal action (salt, cost $1,000) and the suboptimal action (don't salt, expected cost 0.235 × $5,000 = $1,175) is $175 in virtual dollars, translating to approximately $0.0875 in real bonus. Across 18 trials, the total marginal incentive for making all optimal versus all suboptimal decisions is on the order of $1–2. The paper acknowledges this possibility in §7.3 ('not being sufficiently sensitive to incentives') but does not independently verify that the incentive was above a behaviorally relevant threshold. The time-on-task increase (20

    Authors: The referee raises a valid concern. We agree that we cannot independently verify that the incentive magnitude was above a behaviorally relevant threshold, and this is a genuine limitation of our study. However, we would note several points in mitigation. First, the incentive magnitudes we used ($2 expected bonus in Study 1, $1.60 average bonus in Study 2) are comparable to or larger than those used in prior incentivized visualization studies [12, 31, 60], which is the context our paper addresses. Second, the fact that participants in incentivized conditions spent 20-40% more time on the task provides behavioral evidence that participants did perceive the incentives as meaningful enough to change their behavior—they just did not translate that increased effort into improved performance. This pattern is consistent with the capital-labor-production theory [3], which predicts that incentives increase effort but may not improve performance when participants lack the cognitive capital to benefit from additional effort. That said, we agree that the per-trial marginal incentive in Study 2 was small (on the order of $0.09 per optimal decision), and we cannot rule out the possibility that larger incentives might have produced a detectable effect. We will revise §7.3 to explicitly acknowledge this limitation more prominently, including the per-trial marginal incentive calculation the referee describes, and will frame this as a boundary condition on our null result rather than a fully resolved question. We will also note that our null result should be interpreted as applying to incentive magnitudes typical of current crowdsourced visualization studies, not to arbitrarily large incentives. revision: partial

  3. Referee: §5.3 and Figure 5C, equivalence testing: The paper reports credible intervals for expected utility differences (e.g., incentivized interval: 2.76 [2.31, 3.12] vs. base interval: 2.69 [2.19, 3.04]) and concludes there is 'no evidence' of an incentive effect. However, the paper does not report a formal equivalence test or ROPE (Region of Practical Equivalence) analysis to distinguish 'no effect' from 'insufficient precision to detect an effect.' The credible intervals for the expected utility differences are moderately wide (e.g., the base density condition spans [1.70, 2.86]). Adding a ROPE analysis or explicitly stating what effect size would be considered practically meaningful would strengthen the claim that the null result reflects a true absence of incentive effects rather than insufficient precision.

    Authors: The referee is correct that our language conflates 'no evidence of an effect' with 'evidence of no effect,' and that a formal ROPE analysis would strengthen our claims. We will add a ROPE analysis to §5.3. To define the ROPE, we will use a practically meaningful effect size threshold based on the task structure: in Study 2, the difference in expected utility between the rational benchmark and the random-response benchmark is approximately 7,232 virtual dollars (3,960 vs. -3,272), which translates to approximately $3.62 in real bonus. We consider an incentive effect of 10% of this range (approximately 723 virtual dollars, or $0.36 in bonus) as the lower bound of practical significance—anything smaller would represent a change of less than $0.02 per trial, which we consider below the threshold of practical relevance. For Study 1, we will define the ROPE in terms of JND differences, using a threshold of 0.02 in correlation units, which corresponds to the smallest stimulus difference we tested. We will report the proportion of the posterior distribution falling within the ROPE for each comparison. Based on our preliminary analysis, the posterior probability mass within the ROPE exceeds 90% for the incentive comparisons in both studies, which we believe provides reasonable support for our claim. We will also soften our language from 'no evidence of an effect' to 'no practically meaningful effect' where appropriate, and will explicitly acknowledge the precision limitations the referee identifies. revision: yes

  4. Referee: §5.3 and §7.2, interval vs. density plot finding: The finding that interval plots led to marginally better expected utility than density plots contradicts prior work [12, 31, 60]. The authors propose a visual heuristic explanation (Figure 7: the 66% interval overlapping 0°C signals ~20% probability) but the follow-up study (Appendix B) with 95% intervals also showed no density advantage, which undermines the heuristic explanation. The paper acknowledges this is 'a conflicting picture' (§7.2) but the discussion does not fully grapple with whether this finding might indicate a problem with the task implementation, the stimulus distributions, or the model. Given that this unexpected finding is secondary to the main claim about incentives, a more thorough discussion of possible confounds or at minimum a clearer acknowledgment of the uncertainty about this result would be appropriate.

    Authors: We agree with the referee that our discussion of the interval vs. density finding is insufficiently thorough. The referee correctly notes that the Appendix B replication with 95% intervals (which removes the visual heuristic we proposed) still showed no density advantage, which undermines our proposed explanation. We will revise §7.2 to acknowledge this more directly and to discuss additional possible explanations. Specifically, we will discuss three possibilities: (1) The stimulus distributions we used had varying standard deviations, which prior work [15, 16] has shown can interfere with visual probability estimation from densities—this may have disproportionately disadvantaged the density condition. (2) Our task provided immediate feedback after each trial, which may have allowed participants in the interval condition to learn a simpler heuristic (e.g., 'salt when the interval bar is close to 0°C') that is not available in the same form for densities. (3) We cannot fully rule out a task implementation issue, and we will state this explicitly. We will also note that because this finding is secondary to our main claim about incentives, and because it contradicts a well-established finding in prior work, we are cautious about over-interpreting it and recommend it be treated as an exploratory finding warranting replication rather than a definitive result. We will add a sentence to this effect in §7.2. revision: yes

standing simulated objections not resolved
  • We cannot independently verify that the incentive magnitudes used were above a behaviorally relevant threshold for all participants. While the time-on-task increase provides indirect evidence that participants perceived the incentives as meaningful, we cannot rule out the possibility that larger incentives would produce a detectable performance effect. This is a genuine limitation that we will acknowledge more prominently but cannot fully resolve with the current data.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: null results from preregistered between-subjects experiments with externally benchmarked metrics.

full rationale

This paper reports two (plus one replication) preregistered experiments manipulating performance-based incentives as a between-subjects factor. The central claim is a null result: incentives did not improve task performance. The dependent variables (JND for correlation perception, expected utility for decision-making) are computed from independently collected participant responses against external benchmarks (Weber's law for Study 1; rational-decision-maker expected utility for Study 2). The model specifications (§4.2, §5.2) use standard psychometric and linear-in-log-odds models with weakly regularizing priors centered on zero (no effect), which is appropriate for detecting effects rather than assuming them. The incentive magnitudes were calibrated from pilot data to match average compensation across conditions, but this calibration does not determine the null result—the bonus structure is an independent manipulation, and the outcome (whether participants performed better) is measured from their responses. The paper does cite prior work by some of the same authors (e.g., [56] for the decision-making task setup, [60] for the llo model), but these citations provide the experimental paradigm and model specification, not the conclusion. The null result is not forced by any fit or definition: the models could have detected incentive effects if present, and the credible intervals for the incentive condition differences in expected utility (e.g., incentivized interval: 2.76 [2.31, 3.12] vs. base interval: 2.69 [2.19, 3.04]) overlap substantially, consistent with a genuine null. No step in the derivation chain reduces to its inputs by construction.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities, particles, forces, or dimensions. It is an empirical study with design parameters chosen from pilot data and prior work. The free parameters listed are experimental design choices (compensation amounts) rather than model-fitting parameters, but they are ad hoc in the sense that they were calibrated to match a target hourly wage rather than derived from theory. The axioms are primarily domain assumptions borrowed from behavioral economics, with one ad-hoc assumption about incentive magnitude sufficiency that is load-bearing for the null result.

free parameters (5)
  • Bonus per correct answer (Study 1) = $0.05
    Calibrated based on pilot data showing 40/65 correct responses on average, to target $2 average bonus and $15/h total compensation (§3). Not a free parameter in the modeling sense, but a design parameter chosen ad hoc from pilot data.
  • Bonus per $1,000 virtual budget (Study 2) = $0.50
    Calibrated based on prior work [56] showing average participant had $4,000 remaining, to target $2 average bonus (§3). Design parameter chosen from prior data.
  • Base pay (flat condition) = $3.00
    Set to achieve approximately $15/h given estimated 9-12 min completion time (§3).
  • Base pay (incentivized condition, Study 1) = $1.25, adjusted to $1.60
    Initially set to $1.25, adjusted upward to $1.60 after data collection because participants took longer than expected, to comply with Prolific minimum wage and target $15/h (§3, §4.1).
  • Base pay (incentivized condition, Study 2) = $1.50, adjusted to $4.20
    Initially set to $1.50, adjusted upward to $4.20 after data collection because participants took significantly longer (24 min median vs. 16 min), to comply with Prolific minimum wage and target $15/h (§3, §5.1).
assumptions (5)
  • domain assumption Capital-labor-production theory (Camerer and Hogarth [3]) accurately describes the effect of incentives on task performance in visualization experiments.
    Used in §2.2 to generate predictions about which tasks should be affected by incentives. The theory is borrowed from behavioral economics and its applicability to crowdsourced visualization tasks is assumed, not independently verified.
  • domain assumption Time spent on task is a reasonable proxy for effort exerted by participants.
    Invoked in §7.1 to interpret the finding that incentivized participants spent more time but did not perform better: 'If time spent on a task is considered a reasonable proxy for effort exerted by a participant, this result suggests that the average participant in the incentivized condition is likely putting in more effort.' The paper flags this as conditional but relies on it for interpretation.
  • domain assumption Participants in incentivized conditions are attempting to maximize expected utility.
    The expected utility metric in Study 2 (§5.2-5.3) assumes participants are trying to maximize EU. The paper questions this assumption in §7.3 ('Are Participants Trying to Maximize Expected Utility?') but the metric is still used as the primary performance measure.
  • ad hoc to paper The incentive magnitudes used ($0.05/answer; $0.50/$1000 virtual) are sufficient to induce a behaviorally meaningful change in effort.
    The null result depends on this, but the magnitudes were calibrated to match average compensation rather than to be above a behavioral threshold. The paper acknowledges this in §7.3.
  • domain assumption Prolific crowdworkers are representative enough of the broader population of visualization study participants for the results to generalize.
    All participants were recruited from Prolific (§4.1, §5.1). The paper acknowledges in §8 that results 'may not easily translate to other scenarios (e.g., lab studies).'

how reviews work

0 comments
Cite this review

Pith. "Pith review of Should We Dangle a Carrot? The Effect of Performance-based Incentives in Visualization Experiments." pith.science (2026). https://pith.science/paper/JKSZEHSN

@misc{pith2026260707463,
  author       = {Pith},
  title        = {Pith review of: Should We Dangle a Carrot? The Effect of Performance-based Incentives in Visualization Experiments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JKSZEHSN}},
  note         = {Machine review of arXiv:2607.07463}
}
read the original abstract

A perennial research question in visualization involves identifying which visual encodings for a particular dataset are most effective for users in performing a specific task. The relative effectiveness of the different encodings are commonly identified through controlled experiments. However, designing an experiment involves making many, often ad hoc, decisions about the experimental setup such as whether to include a training module, whether to provide performance-based incentives to participants, etc. Yet, there is limited guidance on how these decisions should be made, and we do not fully understand the impact of these subjective decisions on empirical results. In this paper, we investigate the impact of one such key design decision: monetary rewards. Specifically, we ask: does providing or not providing participants with performance-based financial incentives affect the results and the conclusions that we draw from visualization studies? We conducted two crowdsourced studies investigating the impact of incentives on (i) a low-level, perceptual task (perception of correlations in scatterplots or parallel coordinate plots), and (ii) a task involving reasoning (decision-making based on a weather forecast represented as intervals or density plots). In each of these studies, we manipulate both the visual representation and the presence of incentives as between-subject conditions. We expected to find no effect of incentives on the perceptual task, but to see an effect for the decision-making task. However, we found no effect on task performance in either study. While these are results of only two studies and should be replicated, they suggest that performance-based financial incentives may not always have the intended effect on participants that we presumed, and calls for a reflection of how incentivized studies should be designed.

Figures

Figures reproduced from arXiv: 2607.07463 by the authors.

Figure 1
Figure 1. Example of a stimulus seen by a participant in the non [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The main result of Experiment 1. We show the posterior es [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Example of a stimulus seen by a participant in the incentivized [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: The main result of Experiment 2. We show the posterior credible [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: The distribution of utility obtained by the participants and the [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Different uncertainty visualizations used in prior work (A, C) [e.g., [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: The main result of Experiment 3. We show the posterior credible [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed July 9, 2026 · model on record in the stance chip above.