REVIEW 2 major objections 5 minor 1 cited by
Contrarian Incentives and Costly Social Learning
T0 review · 2 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read When the default choice follows the crowd, contrarian tastes make agents more willing to pay for private information; extreme contrarianism then erodes accuracy.
desk verdict The model and threshold results are solid, but the paper's central comparative static (Prop. 4.4) is not proven as written, and the abstract promises results that the body doesn't deliver. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The contrarian bonus b(p_t(a)) = k(1 − p_t(a)) is tied to observed predecessor popularity only, so the equilibrium avoids any fixed point in anticipated popularity and keeps Bayesian updating standard. The action rule is the posterior cutoff c_t = 1/2 + k(p_t − 1/2), a threshold that tilts against the observed majority. The central object is the net value of information Φ_t(µ_t, p_t; k) = sup_ρ {A_t(ρ) + k B_t(ρ) − (c/2)ρ²}, where A_t is the correctness gain and B_t the bonus change from buying a signal. The crux is the sign of B_t: under the majority-aligned baseline, B_t(ρ) ≥ 0 for all ρ, giving increasing differences in (ρ,k) and monotone Φ_t. Gaussian LLR ℓ(s;ρ) = ρ(s − 1/2) puts thresho
What would settle it
A numerical check settles it: compute the cross-partial of the agent's expected payoff with respect to (ρ,k) under Gaussian signals, and search over (µ_t, p_t, k, F) with the no-signal action equal to the majority for any case where Φ_t falls as k rises; a single counterexample refutes Proposition 4.4 and the investment-region expansion. A lab or field experiment raising the contrarian bonus while popularity stays observable should show information-purchase rates rising in discrete jumps as the bonus crosses the entry threshold; absence of jumps counts against the interval prediction.
Extended reading notes
Core claim
Optimal actions follow the posterior cutoff c_t = 1/2 + k(p_t − 1/2), tilting against the observed majority. The central claim (Prop. 4.4): whenever the no-signal action coincides with the observed majority, the net value of information Φ_t is weakly increasing in contrarian intensity k, so the investment region {k : Φ_t > F} is an interval [k0, ∞) — stronger contrarian motives expand the histories in which agents buy signals, and at the entry threshold actions jump from uninformative to informative (Prop. 6.3). The binary-signal benchmark (abstract) adds that contrarian incentives expand the restart region and improve terminal beliefs and action accuracy in discrete steps until an informati
Load-bearing premise
The central message rests on an unproven step (Appendix A.2): the proof of Proposition 4.4 asserts, without computing the cross-partial and while treating the cutoff c_t as fixed though it shifts with k, that the objective has increasing differences in (ρ,k) when the no-signal action follows the majority; the Gaussian-quadratic specification is assumed throughout, with only a claim that MLRP families behave the same, and the abstract's binary-signal results are not derived in
Editorial extensions
If this is right
- When the no-signal action follows the observed majority, stronger contrarian preferences weakly raise the net value of information, so the investment region is an interval [k0, ∞): more agents pay for signals as k grows.
- Higher k shifts the posterior and signal thresholds against the majority, reducing the probability of choosing the majority action in both states (Corollaries 5.1–5.2).
- Public action informativeness is zero below k0 and jumps to strictly positive at k0, then weakly increases with k (Lemma 6.2, Proposition 6.3).
- In the binary-signal benchmark, contrarian incentives improve terminal beliefs and action accuracy in discrete steps up to an information-cost frontier; beyond a second threshold, long-run accuracy declines toward one half.
- For any general experiment menu with a positive fixed fee, expected information purchases are uniformly bounded; with full-support signals, learning is incomplete.
Reading between the lines
- Testable extension: a platform or lab experiment that raises the salience of popularity (or the social reward for distinctiveness) should show information-seeking behavior rising in discrete jumps as the contrarian incentive crosses a history-dependent boundary, not as a smooth gradient.
- The two thresholds—the information-cost frontier and the accuracy-decline point—define a design window: visibility of popularity sustains experimentation inside the window and backfires outside it; an open question the paper leaves implicit is where real communities sit relative to these thresholds.
- The threshold-tilt comparative statics (Corollaries 5.1–5.2) rest only on MLRP, so the directional prediction that majority choices decline with contrarian intensity is portable beyond Gaussian signals; the value-of-information and precision results are the Gaussian-dependent parts.
- An immediate corollary the paper does not draw: with heterogeneous contrarian tastes in a population, aggregate acquisition rates should be increasing in the average taste but kinked wherever mass crosses k0, a micro-data fingerprint of this mechanism.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a sequential model of social learning with endogenous information acquisition and nonconformist preferences. Each agent observes the history of actions, chooses a Gaussian private signal precision ρ after paying a fixed cost F and convex cost (c/2)ρ², and then chooses a binary action. Utility is correctness plus k times the unpopularity of the chosen action among predecessors. The paper characterizes the optimal action as a posterior/LLR threshold tilting away from the majority, derives the FOC for precision, and proves existence/regularity of the precision choice. Its central result (Prop 4.4) claims that when the no-signal action coincides with the majority, the value of information Φ_t is weakly increasing in k, so the investment region expands with k. It also derives comparative statics for thresholds and action probabilities and a local welfare/informativeness analysis.
Significance. The model is transparent and fully parametric, yielding closed-form thresholds and a clean characterization of the information-acquisition margin. The central monotonicity insight—contrarian tastes expand experimentation when the default action is the majority—is interesting and economically plausible, and the paper's comparative statics (Corollaries 5.1–5.2) are correct. However, the main proof is currently incomplete, and the abstract advertises dynamic results (stopped random walk, restart region, long-run frequency) that do not appear in the body. Once the proof of Proposition 4.4 is completed and the abstract aligned with the actual content, the paper would make a meaningful contribution to the literature on social learning and information choice.
major comments (2)
- [§4.4 / Appendix A.2] The proof of Proposition 4.4 is incomplete. The decomposition Φ_t = sup_{ρ≥0} {A_t(ρ)+k B_t(ρ)} treats A_t and B_t as independent of k, but the action probabilities P_1(ρ) depend on k through the threshold c_t(k)=1/2+k(p_t−1/2) and τ_t(k). The displayed inequality (1−2p_t)(P_1(ρ)−1)≥0 only shows that B_t(ρ)≥0 for each fixed k; it does not establish that G_t(ρ,k) has increasing differences in (ρ,k), nor that the supremum is increasing in k. The interval property of the investment region and Proposition 6.3 depend on this monotonicity. A correct argument must compute the total derivative dG_t/dk at fixed ρ (including the threshold-shift term) or use an envelope-theorem argument for Φ_t. Without this, the central claim is not proven.
- [Abstract / Introduction] The abstract claims results that are not in the manuscript: a 'binary-signal benchmark' with a 'restart region' and 'public log odds at information dates form a stopped random walk,' as well as long-run frequency statements about the frequency of correct actions. The main text is entirely Gaussian and presents single-history, local results (Propositions 4.4, 6.1, 6.3). No definition of a restart region or analysis of the dynamic belief process appears. The abstract must be revised to match the actual contributions, or the missing dynamic analysis must be supplied.
minor comments (5)
- [Appendix A.2] In the expression for E[1−p_t(a_t)] − E[1−p_t(a_t)]|_{ρ=0}, the dependence of P_1(ρ) on k is not made explicit; this ambiguity obscures the proof gap.
- [Section 2 (Introduction)] There is a stray '1' in the sentence 'Our approach is different: 1 we add a taste-based bonus...' which should be removed.
- [References] The reference Lu, Meyer, and Rosenbaum (2019) contains an internal note 'Working paper; update authors, venue, and link if you have them' and is not a complete citation.
- [Section 3.1 (footnote 2)] The claim that the qualitative results extend to any MLRP signal family is asserted without proof. Since the proof of Prop 4.4 relies on the specific form of the threshold c_t(k), the paper should state explicitly which results are proven only for Gaussian signals and which carry over to general MLRP families.
- [Proposition 6.3] The statement that 'upper hemicontinuity of the argmax yields right-continuity of ρ⋆_t(k) away from the entry point' is not immediate from UHC alone; a selection argument or further regularity is needed.
Circularity Check
No circularity: all results are derived from explicit primitives; the contested Proposition 4.4 step is a mathematical proof gap, not a circular reduction.
full rationale
The paper does not fit parameters to data, does not invoke self-citations for load-bearing premises, and does not import a uniqueness theorem or ansatz from the authors' prior work. The model is self-contained: preferences (correctness plus contrarian bonus) are stated as primitives, the Gaussian signal and quadratic cost are explicit assumptions, and the threshold rule (Lemma 4.1) follows directly from expected-utility maximization. Proposition 4.4's conclusion that the investment region is [k0,∞) rests on an asserted increasing-differences property; the appendix shows only B_t(ρ)≥0, which does not by itself establish monotonicity of the objective in (ρ,k). That is a validity gap in the proof, not a case where a prediction is equivalent to its input by construction. Under the hard rule that circularity must be exhibited as a specific reduction (e.g., Eq. X = Eq. Y by definition, or a fitted parameter renamed as prediction), no such reduction appears. The paper contains no self-citations at all, and its central comparative static is presented as a derived result rather than an assumption. Accordingly the circularity score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption Gaussian signals with precision ρ (A1)
- domain assumption Fixed entry cost F plus quadratic precision cost (A2)
- domain assumption Linear contrarian bonus b(p)=k(1-p) tied to observed predecessor popularity (A3)
- standard math Common prior 1/2 and PBE with Bayesian updating (Sections 3.4-3.5)
- domain assumption Unproved claim that results extend to any MLRP signal family (footnote 2)
- ad hoc to paper Increasing-differences property in the proof of Prop 4.4 (Appendix A.2)
Cite this review
Pith. "Pith review of Contrarian Incentives and Costly Social Learning." pith.science (2026). https://pith.science/paper/AOY73ZKL
@misc{pith2026250821446,
author = {Pith},
title = {Pith review of: Contrarian Incentives and Costly Social Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/AOY73ZKL}},
note = {Machine review of arXiv:2508.21446}
}
read the original abstract
We study sequential social learning when agents pay a fixed cost for private information and prefer less popular actions. Actions taken without new information leave beliefs unchanged but alter popularity and subsequent decision cutoffs, potentially restarting acquisition. In a binary-signal benchmark, we characterize the restart region and show that public log odds at information dates form a stopped random walk. Contrarian incentives initially expand this region and weakly improve terminal beliefs and action accuracy in discrete steps. After the region reaches an intrinsic information-cost frontier, beliefs stop improving; beyond a second threshold, the long-run frequency of correct actions declines toward one half. For general experiment menus, any positive fixed fee uniformly bounds expected purchases and, with full-support signals, implies incomplete learning. The restart mechanism extends to recency-weighted popularity indices and to endogenous Gaussian precision.
Forward citations
Cited by 1 Pith paper
-
Paying for Failure in Expert Advice
Higher reputation can make experts recommend risky actions less often, but the paper's key condition is assumed rather than derived, and its single-cutoff characterization is not proven for different ability types.
Reference graph
Works this paper leans on
-
[1]
Acemoglu, D., Dahleh, M., Lobel, I., and Ozdaglar, A. (2011). Bayesian learning in social networks. The Review of Economic Studies , 78(4):1201--1236
work page 2011
-
[2]
Akerlof, G. A. and Kranton, R. E. (2000). Economics and identity. Quarterly Journal of Economics , 115(3):715--753
work page 2000
-
[3]
Ali, S. N. and Kartik, N. (2012). Herding with collective preferences. Economic Theory , 51(3):601--626
work page 2012
-
[4]
B \'e nabou, R. and Tirole, J. (2011). Identity, morals, and taboos: Beliefs as assets. American Economic Review , 101(3):322--325
work page 2011
-
[5]
Bernheim, B. D. (1994). A theory of conformity. Journal of Political Economy , 102(5):841--877
1994
-
[6]
Bikhchandani, S., Hirshleifer, D., Tamuz, O., and Welch, I. (2024). Information cascades and social learning. Journal of Economic Literature , 62(3):1040--1093
2024
-
[7]
Bikhchandani, S., Hirshleifer, D., and Welch, I. (1992). A theory of fads, fashion, custom, and cultural change as informational cascades. Journal of Political Economy , 100(5):992--1026
1992
-
[8]
Bohren, A., Imas, A., and Rosenberg, M. (2019). The dynamics of discrimination: Theory and evidence. Quarterly Journal of Economics , 134(4):1567--1616
work page 2019
Show all 30 references
-
[9]
Bohren, J. A. (2016). Informational herding with model misspecification. Journal of Economic Theory , 163:222--247
2016
-
[10]
and Gale, D
Chamley, C. and Gale, D. (1994). Information revelation and strategic delay in a model of investment. Econometrica , 62(5):1065--1085
1994
-
[11]
and He, K
Dasaratha, K. and He, K. (2022). Network structure and naive sequential learning. Review of Economic Studies , 89(5):2469--2504
2022
-
[12]
Dasgupta, A. et al. (2000). Social learning with payoff complementarities. London School of Economics , 25
2000
-
[13]
Eyster, E., Rabin, M., and Weizs \"a cker, G. (2014). An experiment on social learning. Journal of the European Economic Association , 12(4):1143--1172
2014
-
[14]
Frick, M., Iijima, R., and Ishii, Y. (2020). Misinterpreting others and the fragility of social learning. Econometrica , 88(6):2281--2328
2020
-
[15]
Golman, R., Jain, A., and Saraf, S. (2021). Hipsters and the cool: A game theoretic analysis of identity expression, trends, and fads. Psychological Review . Earlier arXiv:1910.13385
2021 arXiv
-
[16]
and Van Weelden, R
Kartik, N. and Van Weelden, R. (2020). Informational herding with model uncertainty. American Economic Review , 110(12):3859--3893
2020
-
[17]
Kuhn, T. S. (1962). The Structure of Scientific Revolutions . University of Chicago Press, Chicago, 2nd edition
1962
-
[18]
Lu, E., Malmendier, U., and Zhang, D. (2021). Choosing the wrong pond: Social learning and peer effects in information acquisition. Quarterly Journal of Economics , 136(4):2185--2244
2021
-
[19]
Lukyanov, G. (2025). Social learning from experts with uncertain precision. Working paper, Toulouse School of Economics
2025
-
[20]
and Azova, A
Lukyanov, G. and Azova, A. (2025). Herding prices: Social learning and dynamic competition in duopoly. Working paper
2025
-
[21]
and Cheredina, D
Lukyanov, G. and Cheredina, D. (2025). False cascades and the cost of truth. arXiv preprint arXiv:2508.20538
2025 arXiv
-
[22]
Lukyanov, G., Popov, K., and Lashkery, S. (2025a). Self-employment as a signal: Career concerns with hidden firm performance. Working paper
-
[23]
and Vlasova, A
Lukyanov, G. and Vlasova, A. (2025). Dynamic delegation with reputation feedback. arXiv:2508.19676
2025 arXiv
-
[24]
Lukyanov, G., Vlasova, A., and Ziskelevich, M. (2025b). Risky advice and reputational bias. arXiv:2508.19707
-
[25]
Planck, M. (1949). Scientific Autobiography and Other Papers . Philosophical Library, New York
1949
-
[26]
and S rensen, P
Smith, L. and S rensen, P. (2000). Pathological outcomes of observational learning. Econometrica , 68(2):371--398
2000
-
[27]
and S rensen, P
Smith, L. and S rensen, P. (2011). Observational learning. In The New Palgrave Dictionary of Economics . Palgrave Macmillan
2011
-
[28]
Touboul, J. (2014). The hipster effect: When anticonformists all look the same. arXiv:1410.8001
2014 arXiv
-
[29]
Veldkamp, L. L. (2011). Information Choice in Macroeconomics and Finance . Princeton University Press
2011
-
[30]
Vives, X. (2008). Information and Learning in Markets: The Impact of Market Microstructure . Princeton University Press
2008
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.