{"id":"6aaaaeb7-210c-458e-a349-2b496c4f4ec7","arxiv_id":"2508.19676","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A dynamic expert-advice model where client effort responds to reputation produces a reputation-dependent advice threshold and reputational conservatism.","lead":"This paper models a long-lived expert whose recommendations are followed by short-lived clients whose effort depends on the expert's reputation. It shows that trusted experts recommend risky actions less often because failures become more informative when clients try harder.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 7 is not established: Appendix B's Lemma 30 is false for increasing convex V, and the diagnosticity Condition 6 is assumed rather than derived.","rationale":"The reader's conditional verdict is appropriate, and the stress-test confirms it. The central result Theorem 7 is not rigorously established: the proof's key monotone-comparative-statics lemma is false as stated for increasing convex V, and the diagnosticity Condition 6 is assumed at the equilibrium cutoff rather than derived from primitives. This is the single most load-bearing concern because the reputation-dynamics results, empirical predictions, and policy conclusions all rely on Theorem 7. The economic mechanism may still be salvageable, but the paper needs a repaired proof and an explicit primitive condition that delivers the needed failure-more-diagnostic-than-success asymmetry. The reader already assigned CONDITIONAL, so no verdict adjustment is needed.","tokens_in":54707,"tokens_out":13349,"duration_ms":151179,"concrete_test":"Independently compute the cross partial F(x,z)=∂²/∂x∂z V(σ(x+z)) for V(t)=e^t and the logistic σ at x+z=logit(0.8). The value is e^{0.8}(σ'^2+σ'') = e^{0.8}(0.0256−0.096)<0, directly contradicting the decreasing-differences claim in Lemma 30. This single analytic check pins the failure of the Appendix B proof; if a revised proof avoids Lemma 30, the theorem may survive, but as written the derivation collapses at this step.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline comparative static (Theorem 7) is the hinge for the reputation dynamics and empirical predictions, but its proof is invalid. In Appendix B, Lemma 30 asserts that Γ0(s,π)=V(σ(logit π+J0(s))) has decreasing differences because σ is increasing and concave and V increasing preserves the property. The logistic σ is not globally concave, and an increasing convex V does not preserve decreasing differences. With x=logit π, z=J0(s) (decreasing in s), decreasing differences of g(s,π) require F(x,z):=V''(σ)σ'^2 + V'(σ)σ'' ≥0. For V(t)=e^t at x+z=logit(0.8), σ=0.8, σ'=0.16, σ''=-0.096, so F=e^{0.8}(0.0256-0.096)<0: the required inequality fails and Γ0 has increasing differences at that point. Lemma 32's 'Jensen's inequality' step is similarly asserted rather than proved. Separately, Condition 6 (log L+≤-log L-) is imposed at the equilibrium cutoff and is never derived from A1-A3 or the quadratic benchmark; it is essentially the reputational-conservatism inequality and is the only route to the convexity of V (Theorem 4/OA.1). Thus Theorem 7, as stated, either silently assumes a condition close to its conclusion or lacks a valid proof.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a dynamic expert-advice model in which a long-lived expert of unknown ability advises a sequence of short-lived implementers. Implementer effort responds to the expert's current reputation, making outcomes more informative about ability when reputation is high. The expert anticipates this and chooses recommendations accordingly. The authors characterize a recursive belief-based equilibrium in which the high type uses a signal cutoff, show that the cutoff is increasing in reputation ('reputational conservatism'), derive comparative statics in precision, state prior, and patience, characterize reputation dynamics as a submartingale/supermartingale, and provide a Gaussian–quadratic benchmark, a surgery application, committee and monitoring extensions, and a success-bonus policy design. The headline result is Theorem 7, which asserts that the high type's risky-signal cutoff is weakly increasing in the public reputation under A1–A3 and an increasing convex value function.","tokens_in":55010,"tokens_out":10773,"duration_ms":121115,"significance":"The mechanism studied—reputation increases implementation effort, which in turn makes failures more diagnostic—is interesting and potentially important for advice markets, medical decision-making, finance, and organizational design. If Theorem 7 and the dynamic results were established, the paper would provide a clean, testable theory with explicit comparative statics, a tractable Gaussian benchmark, and a calibration recipe. The paper also contains many useful extensions (committees, monitoring, endogenous exit, continuous-time approximation) and a detailed empirical measurement appendix. These are real strengths. However, the central comparative static is supported by proofs that, as written, contain a false composition lemma and rely on an assumption (Condition 6) that is close to the conclusion. The martingale/boundary theorem is also misstated. The contribution is therefore currently not established to the standard required for publication.","major_comments":[{"comment":"There is a fundamental inconsistency in the proof of Theorem 7. In the model, the signal s is private: the implementer's effort e*(1,π) in (3) depends only on the public recommendation and reputation, and the public post-outcome reputations π+(π), π-(π) in (6) do not condition on the expert's private signal. Appendix B, however, repeatedly writes π±(π,s), J±(π,s), and e*(1,π;s), and Lemma 31 states that J- is decreasing in e*(1,π;s) with e* increasing in s. This treats s as publicly observed. Once the s-dependence is removed, the 'effort amplification' channel in Lemma 32(i) and the LLR-asymmetry channel used to obtain decreasing differences of Γ1 collapse. The proof of Theorem 7 as written therefore does not apply to the model that is actually specified.","section":"Appendix B / Section 3.5"},{"comment":"Lemma 30 is false as stated. It claims that Γ0(s,π)=V(σ(logit π+J0(s))) has decreasing differences because σ is increasing and concave and V increasing preserves the property. The logistic σ is not globally concave, and an increasing convex V does not preserve decreasing differences under a non-concave transformation. For the logistic link, the cross derivative is V''(σ)(σ')^2+V'(σ)σ''. Taking V(t)=e^{10t} and σ=0.4 (so x+z=logit(0.4)) makes this quantity strictly positive, so Γ0 has increasing differences at that point. Thus the claimed preservation property is not valid without additional restrictions. Since Lemma 30 is a load-bearing step, Theorem 7 is not established by the argument given.","section":"Appendix B, Lemma 30"},{"comment":"Condition 6—failures at least as diagnostic as successes at the equilibrium cutoff—is assumed rather than derived from A1–A3 or the Gaussian primitives. It is essentially the mechanism producing reputational conservatism. Moreover, Theorem 7 states its hypotheses as A1–A3 plus increasing convex V, but the convexity of V is itself obtained in Theorem 4 only under Condition 6 (see OA.1, Lemma 64). As stated, Theorem 7 either silently assumes a condition very close to its conclusion or has insufficient hypotheses. The binary-signal reduction of Condition 6 to q_H(1-q_H)≤q_L(1-q_L) illustrates its restrictiveness, but no analogous characterization is supplied for the continuous-signal case. This needs to be made explicit and, ideally, Condition 6 replaced by a primitive condition.","section":"Condition 6 / Theorem 7"},{"comment":"Theorem 12 contains a misstatement and an unsupported claim. Part (1) says that under θ=H the process is a submartingale and writes E[π_{t+1}|H_t]=π_t. The equality is the martingale property under the public prior with θ random; conditional on θ=H the correct inequality is >, as Lemma 33 itself proves. Part (3) asserts that the process hits a high-trust or low-trust region with probability one. The proof sketch says that if informative periods cease, the process is 'eventually constant and trivially hits a boundary region,' but an interior constant path does not hit a boundary region. The boundary-hitting claim therefore needs additional assumptions (e.g., experimentation occurs infinitely often on the relevant event) or a different formulation.","section":"Theorem 12"}],"minor_comments":[{"comment":"The manuscript is submitted under the title 'Endogenous Vindication: Reputation and Effort in Expert Advice,' but the full text uses 'Dynamic Delegation with Reputation Feedback.' This should be reconciled.","section":"Title"},{"comment":"Theorem 4 refers to 'Condition (6)' before Condition 6 is defined in Section 4.3. The numbering and order need adjustment.","section":"Section 4.2"},{"comment":"Propositions 9–11 and 14–16 are duplicate statements with identical proofs. This appears to be a manuscript assembly error.","section":"Sections 4.5 and 4.7"},{"comment":"The caption of Figure 1 reports baseline parameters (0,1,1,1.7,0.5,0.9), while OA.3.3 reports (0,1,0.8,1.6,0.5,0.95). The numerical values should be consistent.","section":"Figure 1 / OA.3"},{"comment":"The notation e*(1,π) and e*(1,π;s) is used interchangeably; given the model's information structure, e* should not depend on the private signal s. This ambiguity should be resolved in the main text and appendices.","section":"Notation throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is ambitious and likely a good fit for the journal if the central theorem can be repaired. The main concern is not the novelty of the mechanism but the correctness of the proof of the headline comparative static. I would recommend asking the authors to either prove Theorem 7 under a transparent primitive condition or explicitly state Condition 6 as an assumption and provide a continuous-signal characterization. The s-dependence inconsistency in the appendices is serious and should be addressed before any further consideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth your attention if you care about reputational cheap talk or delegated experimentation. The paper's core idea is genuinely new: a long-lived expert's reputation raises the effort that short-lived clients put into a risky project, and that effort makes failures more diagnostic of ability without making successes more diagnostic. That asymmetry is the microfoundation for reputational conservatism, and the recursive formulation is clean.\n\nThe paper does several useful things. The binary-signal example is transparent about the diagnosticity condition. The reputation dynamics—submartingale for competent types, supermartingale for incompetent, with boundary hitting—are plausible and the proof sketch is standard. The measurement appendix for surgery (how to proxy reputation, effort, and test the predictions) is practical and could be useful for empiricists.\n\nThe soft spots are real. The central theorem (Theorem 7) is not established. Condition 6, which says failures are at least as diagnostic as successes at the equilibrium cutoff, is essentially the mechanism itself; it is assumed rather than derived from primitives. More seriously, the stress-test is right that Lemma 30 in Appendix B is false for increasing convex V: the logistic map is not globally concave, and composing with an increasing convex function does not preserve decreasing differences. The given counterexample with V(t)=e^t works. So the proof of Theorem 7 is invalid as written.\n\nThere are also editorial problems: Theorem 12 misstates the martingale equality (it writes E[πt+1|Ht]=πt and then claims a submartingale, which is inconsistent—the equality should be ≥), Sections 4.5 and 4.7 are duplicates, and the abstract says 'diminishing career returns' (concave value) while the body uses convex V for the main result. These are fixable but contribute to the impression of a draft that isn't fully polished.\n\nI don't think the main idea is dead. The mechanism is economically meaningful, and the binary benchmark shows there is a primitive condition (qH(1−qH) ≤ qL(1−qL)) that delivers the asymmetry. A revised version that replaces Condition 6 with a derived primitive condition and repairs Appendix B could make the contribution solid.\n\nBottom line: this deserves a serious referee, not a desk reject. If the authors fix the proof and the internal inconsistencies, it could be a good contribution to the reputation and advice literature. I'd be willing to read a revision.\n\nBest,","headline":"Genuinely new reputation-effort feedback mechanism and a useful framework, but the central conservatism theorem is not proven: Appendix B has a false lemma and Condition 6 is assumed close to the conclusion.","tokens_in":55503,"tokens_out":2881,"would_cite":true,"duration_ms":30366,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an expert's advice follows a reputation-dependent cutoff: the higher the public reputation, the higher the private-signal threshold for recommending a risky action, because higher reputation makes implementers work ha","keywords":["dynamic delegation","expert advice","reputation feedback","moral hazard","experimentation","reputational conservatism","career concerns","belief-based equilibrium"],"falsifier":"Estimate the log-likelihood jumps after success and failure from panel data on recommendations and outcomes at each reputation level; if at the equilibrium cutoff log L+ > -log L-, Condition 6 fails and the predicted rising cutoff can reverse. Equivalently, estimate the risky-signal threshold as a function of reputation: a range in which the threshold falls with reputation would contradict Theorem 7.","tokens_in":54552,"feed_emoji":"🧠","tokens_out":9958,"duration_ms":102414,"temperature":0.7,"pith_summary":"An expert's advice is evaluated only when a client acts, and the quality of that evaluation depends on how hard the client works. The paper models a long-lived expert who repeatedly advises short-lived clients, with effort responding to the expert's reputation: trusted experts are heeded and their recommendations are executed more diligently. The paper's central claim is that, under a diagnosticity condition, the competent expert's optimal policy is a cutoff in her private signal that rises with reputation—reputational conservatism—and that reputation itself becomes a submartingale for competent types and a supermartingale for less competent types, hitting trust or distrust regions almost surely. If true, this explains why highly reputed surgeons or analysts may recommend risky actions less often, why early successes can suppress later experimentation, and why more informative tests need not produce more learning overall.","feed_headline":"Reputation breeds caution in expert advice","feed_subtitle":"In a new model, trusted experts recommend risky actions less often—and succeed more when they do.","key_machinery":"The engine is the effort response e*(1,π): because effort cost is convex, the implementer chooses effort equal to the posterior probability that the state is good, which rises with the expert's reputation. Higher effort raises the success probability and makes a failure much more informative about the expert's type, while leaving the informativeness of success unchanged. With a convex continuation value V(π), this asymmetry makes the risky-minus-safe payoff difference ΔH(s;π) have decreasing differences in (s,π), so the crossing point s*(π) increases with π. The formal load-bearing condition is Condition 6, log L+(π) ≤ -log L-(π), comparing the log-likelihood jumps of success and failure at","core_discovery":"At any public reputation π, the High-type expert recommends the risky action exactly when her private signal s is at least a cutoff s*(π) (Theorem 5). Under Condition 6—failures are at least as diagnostic as successes at the cutoff—this cutoff is weakly increasing in π (Theorem 7): more highly reputed experts are more conservative. Reputation dynamics follow a martingale law: posterior beliefs drift up under a competent expert, down under a less competent one, and hit boundary regions with probability one (Theorem 12). Comparative statics are transparent: better private information or a higher good-state prior lowers the cutoff and raises experimentation, whereas more patience raises conserv","pith_inferences":["If Condition 6 fails—successes more diagnostic than failures—the model's logic implies the reputation-advice slope could reverse, so estimating the sign of that slope from data is a direct test of the diagnosticity asymmetry.","The paper's monitoring result implies a concrete cross-setting prediction: environments with visible implementation effort (checklists, adherence logs) should show a flatter reputation-conservatism relationship than opaque environments, because visible effort reduces the reputational downside of failure.","The continuous-time limit suggests a reduced-form estimation strategy: recover the log-odds drift and jump sizes from observed recommendation frequencies and outcome jumps, without solving for the full value function.","Boundary absorption means long panels should show reputation distributions that are bimodal conditional on true ability—near full trust or near permanent distrust—rather than mean-reverting; this is a sharp, testable fingerprint of the model."],"forward_implications":["Higher reputation predicts fewer risky recommendations and, conditional on a risky recommendation, higher success rates, because implementation effort is higher.","Failures are more damaging to reputation at high reputation than at low reputation; transitory good news reduces experimentation frequency while making subsequent failures more revealing.","Improvements in the expert's signal precision or in the prior probability that the risky state is good lower the cutoff; greater patience raises it.","Competent experts' reputations drift up and can end in full vindication, while less competent experts drift down; with positive probability even a competent expert ends permanently distrusted, and boundary absorption is almost sure.","Individually more informative tests can coexist with less learning overall, because reputational incentives reduce the number of tests the expert is willing to run."],"supporting_citations":[{"why":"Supplies the career-concerns framework in which forward-looking experts distort current behavior to affect future assessments.","marker":"Holmström (1982)"},{"why":"Provides the conservatism-versus-gambling benchmark under career concerns that the paper's reputational conservatism extends.","marker":"Prendergast and Stole (1996)"},{"why":"Source of the impermanent-reputation and boundary-absorption dynamics that Theorem 12 builds on.","marker":"Cripps et al. (2004)"},{"why":"Baseline reputational cheap-talk model with continuous signals whose coarse-disclosure logic the cutoff equilibrium refines.","marker":"Ottaviani and Sørensen (2006)"},{"why":"Empirical evidence that clinician trust raises patient adherence, grounding the reputation-to-effort link.","marker":"Haskard Zolnierek and DiMatteo (2009)"},{"why":"Empirical evidence that surgical skill and volume affect complication rates, grounding the effort-to-outcome link.","marker":"Birkmeyer et al. (2013)"},{"why":"Empirical evidence that influential analyst recommendations move implementation, grounding the advice-to-implementation link.","marker":"Loh and Stulz (2011)"}],"fun_headline_variants":["Trusted experts play it safe, model shows","Reputation pushes experts to withhold advice","When experts are well-regarded, they hesitate","The reputational cost of bold advice","Higher standing makes experts more conservative"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Reputational conservatism stands on Condition 6—that, at the equilibrium cutoff, failure is at least as diagnostic of the expert's type as success—which the paper assumes rather than derives from primitives, and the proof also relies on an unproved composition claim about convex value functions and logistic posteriors.","fun_headline_variants_meta":{"raw":{"variants":["Trusted experts play it safe, model shows","Reputation pushes experts to withhold advice","When experts are well-regarded, they hesitate","The reputational cost of bold advice","Higher standing makes experts more conservative"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000361,"raw_usage":{"total_tokens":1768,"prompt_tokens":705,"completion_tokens":1063,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":999}},"tokens_in":449,"tokens_out":1063,"duration_ms":11945,"temperature":1.0,"reasoning_tokens":999,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:34:20.880997+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate the log-likelihood jumps after success and failure from panel data on recommendations and outcomes at each reputation level; if at the equilibrium cutoff log L+ > -log L-, Condition 6 fails and the predicted rising cutoff can reverse. Equivalently, estimate the risky-signal threshold as a function of reputation: a range in which the threshold falls with reputation would contradict Theorem 7.","supporting_citations":[{"cited_title":"Impetuous Youngsters and Jaded Old-Timers: Acquiring a Reputation for Learning,","cited_arxiv_id":null,"evidence_quote":"Provides the conservatism-versus-gambling benchmark under career concerns that the paper's reputational conservatism extends."},{"cited_title":"Imperfect Monitoring and Impermanent Reputation,","cited_arxiv_id":null,"evidence_quote":"Source of the impermanent-reputation and boundary-absorption dynamics that Theorem 12 builds on."},{"cited_title":"Reputational Cheap Talk,","cited_arxiv_id":null,"evidence_quote":"Baseline reputational cheap-talk model with continuous signals whose coarse-disclosure logic the cutoff equilibrium refines."},{"cited_title":"Surgical Skill and Complication Rates after Bariatric Surgery,","cited_arxiv_id":null,"evidence_quote":"Empirical evidence that surgical skill and volume affect complication rates, grounding the effort-to-outcome link."},{"cited_title":"When Are Analyst Recommendation Changes Influential?","cited_arxiv_id":null,"evidence_quote":"Empirical evidence that influential analyst recommendations move implementation, grounding the advice-to-implementation link."}],"review_version":1}