REVIEW 3 major objections 4 minor 3 cited by
Paying for Failure in Expert Advice
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Reputational pressure makes expert advice more conservative at the top, and a success bonus can tune the risk-taking rate.
desk verdict Neat framework, but the common-cutoff lemma is unproven and generically false, and Theorem 2 is close to restating its assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The risky–safe advantage Δθ(s;π) = ϕ + E[V(π′)|a=1,θ,s] − E[V(π′)|a=0,θ,s] is the central object. Under MLRP it is strictly increasing in the private signal s, so each type's optimal advice is a threshold s*(π). The relative-diagnosticity condition—failures are weakly more revealing than successes at high standing—turns the threshold into a rising function of reputation. A success bonus b adds αb to the risky branch, shifting the threshold downward and generating a one-to-one bonus-to-experimentation mapping.
What would settle it
In the Gaussian benchmark with σL > σH, solve Δ_H(c;π)=0 and Δ_L(c;π)=0 separately over a grid of π. If the two roots differ for any reputation, Lemma 1's common-cutoff premise is false. In real data, estimate the switching signal threshold separately for high- and low-ability experts and test whether they coincide.
Extended reading notes
Core claim
The paper's central claim is that a one-shot advisory relationship with belief-based reputation is completely described by a single cutoff in the expert's private signal, and that this cutoff moves in a disciplined way. Under MLRP, each type's risky-safe advantage is strictly increasing, so an expert recommends the risky action if and only if her signal clears a threshold s*(π). When a relative-diagnosticity condition holds—failures are weakly more revealing than successes at high standing—the threshold rises with reputation: high-standing experts are conservative, recommending risk less often but with a higher success probability at the margin. The same cutoff logic yields a one-to-one map
Load-bearing premise
Both ability types are assumed to switch from safe to risky at exactly the same signal value, even though the high-ability type's signal is more informative; if their optimal thresholds differ, the single-cutoff characterization and the design results collapse.
Editorial extensions
If this is right
- At high reputation, risky recommendations become scarcer but more accurate; markets should read a risky recommendation from a top expert as a stronger signal of private confidence.
- A principal can implement any target experimentation rate in the implementable range by choosing a unique success bonus; no dynamic contract is required.
- If failure penalties are feasible, all interior experimentation rates are implementable, and the implementing contracts form an affine line in the bonus and penalty.
- Lower implementation probability (stricter gatekeeping) raises the cutoff and lowers experimentation; if gatekeeping loosens as reputation rises, conservatism is amplified.
- More informative signals or a higher prior success probability lower the cutoff, while stronger career concerns raise it, so the same transfer scheme works across environments once these primitives are known.
Reading between the lines
- The abstract's 'least-cost' claim is an interpretation of margin-selection: the body's theorems prove implementability and monotonicity, not a formal cost comparison. Turning it into a theorem requires a principal objective and a budget constraint.
- The single-cutoff assumption is load-bearing: if high- and low-ability types optimally use different thresholds, the conservatism theorem and the one-to-one bonus mapping must be re-derived. A Gaussian numerical check would reveal whether the difference is zero or small.
- The relative-diagnosticity condition leaves an observable signature: at high reputation, a failed risky recommendation should move the market's posterior more than a success does. That asymmetry is estimable from data on recommendations and outcomes.
- The committee extension yields a testable organizational prediction: raising the approval threshold should make the same experts recommend risky actions less often, visible in recommendation rates before and after a rule change.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a one-shot expert-advice model in which a privately informed expert (ability type H or L) recommends a risky or safe action and is evaluated through a belief-based reputational payoff that also allows outcome-contingent transfers. The main claims are: (i) equilibrium advice is a cutoff rule with a single cutoff common to both ability types (Lemma 1, Theorem 1); (ii) under a 'relative-diagnosticity' condition the cutoff is weakly increasing in prior reputation, generating reputational conservatism (Theorem 2); and (iii) a success-only bonus implements any target experimentation rate via a one-to-one mapping (Propositions 2, 4, 6). Extensions cover implementation/gatekeeping, outcome noise, and Gaussian closed forms.
Significance. If correct, the framework would offer a clean, distribution-free characterization of reputational bias in expert advice and a simple design lever through success bonuses. The Gaussian benchmark and the computational details in the appendices are valuable and make the model easy to calibrate. However, the central common-cutoff lemma is not proven and is generically false as stated, and the reputational-conservatism theorem rests on a condition defined through equilibrium objects rather than primitives. These issues undermine the foundation of the characterization and the design results.
major comments (3)
- [Lemma 1 / Appendix A, Lemma 2] Lemma 1 asserts a single cutoff s*(π) for both types, but Appendix A, Lemma 2 proves strict single-crossing only for ∆_H. Under (A2), p_H(s) and p_L(s) are different functions, so the two types' indifference conditions generically have different roots. For a common cutoff c one would need p_H(c)=p_L(c), which is a measure-zero coincidence unless a fixed-point argument is supplied. No such argument is given. Theorem 1 and all subsequent results that track a single cutoff therefore lose their foundation.
- [Definition 1 / Theorem 2 / Appendix B] Theorem 2's 'relative-diagnosticity' condition is a sign restriction on the derivative of the equilibrium reputational return R(π)=p_c(V(π_1,1)-V(π_0,0))+(1-p_c)(V(π_1,0)-V(π_0,0)). Along the equilibrium path, ∆_H(s*(π);π)=0 implies R(π)=-ϕ identically, so the total derivative dR/dπ is zero. The proof of Theorem 4 instead uses the partial derivative ∂_π∆_H, and the paper does not define which derivative enters RD. Thus RD is an endogenous object, not a primitive condition, and the theorem is close to a tautology. A primitive condition on signal distributions or V is needed to establish reputational conservatism.
- [Section 5.2 / Propositions 4 and 6] The experimentation rate ρ(π;β1) is defined as the H-type's risky frequency only, while the equilibrium requires both H and L cutoffs and the posteriors depend on both. If the common-cutoff claim fails, L's behavior changes the posterior mapping, so the claimed one-to-one bonus-to-experimentation relationship is not identified. Even if the common cutoff were restored, the design results must account for both types' behavior to be an equilibrium statement.
minor comments (4)
- [Title and Abstract] The full-text title is 'Risky Advice and Reputational Bias', while the arXiv metadata and the provided abstract correspond to 'Paying for Failure in Expert Advice'. The abstract also emphasizes failure protection, which is not developed in the main text. Please align title, abstract, and content.
- [Appendix A, Lemma 2 proof] The proof says 'V is increasing ((A2))', but (A2) is the informativeness ordering; V is increasing and convex under (A4). Please correct the reference.
- [Table 2 / Section 5.3] Table 2 reports negative β1 for ρ* = 0.65 and 0.80. Under the stated limited-liability assumption β1≥0, such targets are not implementable. The text notes β1 may become negative but does not reconcile this with the limited-liability design result.
- [Notation after Lemma 1] The notation p_c(π) ≡ p_θ(s*(π)) is said to be 'well-defined by Lemma 1'. Since the common-cutoff lemma is unproven, the definition is ambiguous until this point is resolved.
Circularity Check
No significant circularity: the conservatism theorem is a conditional sufficient-statistic result, not a definitional tautology; Lemma 1's common-cutoff issue is a correctness gap, not circularity.
full rationale
After walking the derivation chain, I find no circular step. Theorem 2 is explicitly conditional on the relative-diagnosticity (RD) condition, which is a stated assumption (Definition 1); the proof (Appendix B, Theorem 4) uses the implicit function theorem to show that the assumption's sign constraint on ∂π∆H directly implies s*'(π)≥0. This is a transparent sufficient-condition theorem, not a disguised identity: RD is a condition on the derivative of the reputational payoff difference, while the conclusion concerns the cutoff's response, and the link is a mathematical theorem rather than a definitional equivalence. The paper does not fit parameters or rename an empirical pattern. The self-citations in Section 2 (Lukyanov et al.) are merely related-literature notes and are not load-bearing. The main validity concern is Lemma 1's common-cutoff claim: the appendix proves only that ∆H(s;π) is strictly increasing, not that ∆L has the same root as ∆H, so the assertion that both types use one cutoff s*(π) is unsupported and generically suspect under (A2). That is a proof gap / correctness risk, not circularity. No step reduces by construction to its inputs.
Assumptions & free parameters
assumptions (6)
- domain assumption Monotone likelihood ratio property (MLRP): fθ(s|1)/fθ(s|0) strictly increasing in s for each θ.
- domain assumption Informativeness ordering: H is (weakly) more informative than L in the Blackwell sense (e.g., σ_L^2 > σ_H^2 in Gaussian).
- domain assumption Public observability: recommendations and, when a=1, outcomes are publicly observed; implementation is one in the baseline.
- domain assumption Career concerns: V(π) is increasing and convex; transfers are affine when used.
- ad hoc to paper Relative-diagnosticity (RD): the derivative of the expected reputational payoff at the cutoff with respect to π is non-positive for π above some threshold.
- ad hoc to paper Common cutoff: both types H and L use the same signal threshold s*(π).
Cite this review
Pith. "Pith review of Paying for Failure in Expert Advice." pith.science (2026). https://pith.science/paper/SCCCQJDY
@misc{pith2026250819707,
author = {Pith},
title = {Pith review of: Paying for Failure in Expert Advice},
year = {2026},
howpublished = {\url{https://pith.science/paper/SCCCQJDY}},
note = {Machine review of arXiv:2508.19707}
}
read the original abstract
A failed recommendation is visible; an unproposed project is not. How should an organization pay an adviser whose private confidence determines which projects reach the margin? We show that the least-cost instrument is protection after failure rather than a success bonus. Because risky advice selects the upper tail of confidence, success pay leaks to recommendations that would occur anyway, while failure protection is concentrated at the margin. We establish unique nonpooling implementation, show that full correction is never optimal, and, while advice still needs encouragement, find confidential internal review substitutes for explicit career insurance more effectively than transparent review.
Forward citations
Cited by 3 Pith papers
-
Endogenous Vindication: Reputation and Effort in Expert Advice
A dynamic expert-advice model where client effort responds to reputation produces a reputation-dependent advice threshold and reputational conservatism.
-
Designing Silence: Peer Feedback under Reputational Concerns
The abstract's concealment-ray theorem and 71.08% optimal revelation threshold are not present in the submitted full text, which is an unrelated paper.
-
Contrarian Incentives and Costly Social Learning
In a Gaussian social-learning model, contrarian preferences expand the set of public beliefs where agents invest in private information, as long as the no-signal action is the observed majority.
Reference graph
Works this paper leans on
-
[1]
Conflicts of Interest, Information Provision, and Competition in the Financial Services Industry,
Bolton, P., X. Freixas, and J. Shapiro (2007): “Conflicts of Interest, Information Provision, and Competition in the Financial Services Industry,” Journal of Financial Economics , 85, 297–330
work page 2007
-
[2]
Bolton, P. and C. Harris (1999): “Strategic Experimentation,” Econometrica, 67, 349–374
work page 1999
-
[3]
Strategic Information Transmission,
Crawford, V. P. and J. Sobel (1982): “Strategic Information Transmission,” Econometrica, 50, 1431–1451
work page 1982
-
[4]
Strictly Proper Scoring Rules, Prediction, and Estimation,
Gneiting, T. and A. E. Raftery (2007): “Strictly Proper Scoring Rules, Prediction, and Estimation,” Journal of the American Statistical Association , 102, 359–378. Holmstr¨om, B. (1999): “Managerial Incentive Problems: A Dynamic Perspective,” The Review of Economic Studies, 66, 169–182
work page 2007
-
[5]
Inderst, R. and M. Ottaviani (2012): “Financial Advice,” Journal of Economic Literature , 50, 494–512
work page 2012
-
[6]
Kamenica, E. and M. Gentzkow (2011): “Bayesian Persuasion,” American Economic Review, 101, 2590–2615
work page 2011
-
[7]
Strategic Experimentation with Poisson Bandits,
Keller, G. and S. Rady (2010): “Strategic Experimentation with Poisson Bandits,” Theoretical Economics, 5, 275–311
work page 2010
-
[8]
Eliciting Properties of Probability Distributions,
Lambert, N. S., D. M. Pennock, and Y. Shoham (2008): “Eliciting Properties of Probability Distributions,” in Proceedings of the 9th ACM Conference on Electronic Commerce (EC), 129–138
work page 2008
Show all 12 references
-
[9]
Herding Prices: Social Learning and Dynamic Competition in Duopoly,
Lukyanov, G. and A. Azova (2025): “Herding Prices: Social Learning and Dynamic Competition in Duopoly,” arXiv preprint arXiv:2509.01263
2025 arXiv
-
[10]
False Cascades and the Cost of Truth,
Lukyanov, G. and D. Cheredina (2025): “False Cascades and the Cost of Truth,” arXiv preprint, arXiv:2508.20538
2025 arXiv
-
[11]
Contrarian Motives in Social Learning: Information Cascades with Nonconformist Preferences,
Lukyanov, G. and V. Ivanik (2025): “Contrarian Motives in Social Learning: Information Cascades with Nonconformist Preferences,” arXiv preprint, arXiv:2508.21446
2025 arXiv
-
[12]
Dynamic Delegation with Reputation Feedback,
Lukyanov, G. and A. Vlasova (2025): “Dynamic Delegation with Reputation Feedback,” arXiv preprint arXiv:2508.19676, submitted. 21
2025 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.