{"id":"55f4e589-86b0-4773-809d-5ba615fb7f02","arxiv_id":"2502.10500","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a finite-time search model, the optimal satisfaction threshold is finite and above the mean reward, and being overambitious is more costly than being underambitious.","lead":"This paper builds a mathematical search model showing that the best goal to set is slightly above average, not the absolute best. It argues the same logic explains moderation in entrepreneurship, dating, and politics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Smooth-landscape extension uses the reciprocal AR(1) variance, so the general T* > μ claim is unproven for φ > 0 and needs direct simulation.","rationale":"The reader identified the variance formula as the weakest assumption; my independent check of Eq. (1) confirms it. The stationary variance of the AR(1) in Eq. (1) is (1−φ)/(1+φ), not (1+φ)/(1−φ). This is not a minor typo: it reverses the direction of the threshold scaling and undermines the only analytical bridge from the proven φ=0 case to smooth landscapes. Since the abstract and title assert the result generally, the missing proof for φ>0 is the load-bearing concern. The proposed simulation is decisive. I do not recommend changing the reader's CONDITIONAL verdict, because the failure is correctable: the authors can either prove the autocorrelated case, restrict the theorem to φ=0, or replace the analytical scaling with direct simulations. If the decisive simulation shows T* ≤ 0 for some φ, the verdict should move to REJECT or the claims should be severely restricted.","tokens_in":16405,"tokens_out":10760,"duration_ms":102701,"concrete_test":"Run the Appendix B simulation with φ ∈ {0, 0.5, 0.9, 0.99}, t_max ∈ {100, 1000}, N(0,1) innovations, N = 10^4 agents, and a fine grid of thresholds T in [−1, 6] (in units of σ=1), computing mean total reward for each T; identify the maximizing T*. Also recompute the smooth-landscape curves using var = (1−φ)/(1+φ). If T* ≤ 0 for any φ ∈ [0,1) with t_max ≥ 2, the abstract's general claim fails; if T* > 0 for all tested φ, the claim may survive but the published scaling derivation is still wrong and must be corrected or replaced by simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section C, the paper states that for smoothness φ, 'the variance of the associated AR(1) process is var[X_t] = (1+φ)/(1−φ)'. But Eq. (1), X_t = φX_{t−1} + (1−φ)ε_t with ε_t i.i.d. N(0,1), has stationary variance (1−φ)^2/(1−φ^2) = (1−φ)/(1+φ), not its reciprocal. The paper then scales the threshold T by 1/√var[X_t], i.e., by √((1−φ)/(1+φ)) < 1, while the correct normalization would multiply by √((1+φ)/(1−φ)) > 1. The text itself says 'autocorrelation reduces the subsequent variance on smooth landscapes', which directly contradicts the printed formula. Because Appendix A's proof of T* > μ is explicitly for the φ=0 case, and the smooth-landscape curves in Fig. 3a are generated with this incorrect scaling, the abstract's general claim that the optimal threshold is strictly larger than the mean of available rewards is not demonstrated for autocorrelated reward landscapes. If one uses the correct variance or, better, directly simulates Eq. (1) for φ close to 1, the search process is nearly unable to change rewards, so the optimal threshold may fail to be uniquely defined or may not exceed μ. This is a proof gap and an internal inconsistency in the central claim's generality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalizes a finite-horizon search model in which an agent repeatedly samples rewards from an autoregressive process and stops searching once the current reward reaches a satisfaction threshold T. The central claims are that the optimal T is finite and strictly above the mean reward, that overshooting T is costlier than undershooting it, that longer time horizons and left-skewed or rugged reward landscapes raise optimal ambition, that upward social comparison harms performance, and that search costs do not overturn the main result unless they make all search unprofitable. An analytic expression for expected reward is derived for the Gaussian, maximally rugged (φ=0) case, and the other results are obtained by analytic scaling or simulation. The paper closes with qualitative applications to dating, college admissions, economic policy, wealth, and elections.","tokens_in":16693,"tokens_out":5890,"duration_ms":62370,"significance":"If the central theorem is established, the paper provides a simple and intuitive formalization of 'finite but above-average ambition' with a clean asymmetry result and several falsifiable qualitative predictions. The analytic expression in Eq. (2), the openly available simulation code, and the transparent numerical experiments are strengths. The paper is also careful to present the empirical examples as illustrations rather than as fitted validations. However, the significance is currently limited by two load-bearing gaps: the proof for the Gaussian rugged case relies on an unproven unimodality assertion, and the extension to smooth (φ>0) landscapes is based on an internally inconsistent variance formula. These gaps affect the abstract's general claim that the optimal threshold is strictly larger than the mean reward across reward landscapes. The manuscript is therefore promising but requires substantive revision before the central claims can be accepted.","major_comments":[{"comment":"The statement that the AR(1) process has variance var[X_t]=(1+φ)/(1−φ) is incorrect and self-contradictory. For Eq. (1), X_t=φX_{t−1}+(1−φ)ε_t with ε_t i.i.d. N(0,1), the stationary variance is (1−φ)^2/(1−φ^2)=(1−φ)/(1+φ). The very next sentence says that 'autocorrelation reduces the subsequent variance on smooth landscapes,' which is consistent with the correct formula and not with the printed one. Because the analytic smooth-landscape curves in Fig. 3a are generated by scaling the threshold by 1/sqrt(var[X_t]) using this erroneous variance, the reported dependence of optimal T on φ is not supported. The manuscript should either correct the variance calculation and redo the scaling, or preferably simulate Eq. (1) directly for φ>0 and report whether T*>μ continues to hold.","section":"Section C (page 5) and Fig. 3a"},{"comment":"The proof that the optimal threshold is strictly greater than the mean for φ=0 is incomplete. The key step asserts that every summand R(T,t_x)P(T,t_x) is unimodal, but this is not proven. The product of a reward component that grows roughly linearly in T and a probability component that is unimodal need not be unimodal, and the fact that the derivative of each summand at T=0 is positive does not by itself establish that the maximum of the total expected reward lies at T>0. The manuscript needs either a rigorous verification of unimodality (with explicit conditions) or a numerical proof/argument covering the relevant range of t_max. Without this, the main theorem for the rugged Gaussian case is not fully established.","section":"Appendix A.4.d"},{"comment":"The paper claims that the optimal threshold increases with the search time t_max, but Appendix A.2 provides only a verbal intuition rather than a proof. Since this monotonicity is presented as one of the main results and is used in the interpretation of Fig. 2a, a formal argument or a direct numerical demonstration should be supplied. In addition, the abstract states that the optimal threshold is 'strictly larger than the mean of available rewards' as a general theorem, yet the proof is restricted to the φ=0 Gaussian case and the smooth-landscape extension is invalidated by the variance error in Section C. The claims should be restricted to the cases actually proven, or the missing cases should be established by correct analysis or simulation.","section":"Appendix A.2 and abstract"}],"minor_comments":[{"comment":"There is a typo: 'aurocorrelation' should be 'autocorrelation.'","section":"Section C, paragraph beginning 'Rugged landscapes create...'"},{"comment":"The text appears to contain a long corrupted string of '/uni...' tokens. If this is present in the source file rather than an artifact of extraction, it should be removed or replaced with the intended figure caption or text.","section":"Appendix C, after Fig. 11"},{"comment":"Reference [44] is incomplete: it gives authors, title, and year but no journal, volume, or preprint identifier. Please provide full bibliographic information.","section":"Reference [44]"},{"comment":"The derivation of Eq. (2) assumes that the reward distribution is standard normal with μ=0 and σ=1; this should be stated explicitly before the equation, along with the affine rescaling that recovers general μ and σ.","section":"Equation (2)"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know the core of this paper is a restatement of standard reservation-wage search theory, with the main theorem being a special case of Kohn–Shavell. What's new is the packaging: satisfaction thresholds in standard-deviation units, plus extensions to autocorrelation, skewness, and social comparison. The paper is clearly written, honest about its simplifications, and ships code on GitHub; the empirical examples are properly labeled as illustrations rather than tests. That part is fine.\n\nThe soft spots are real. Section C asserts that the AR(1) process in Eq. (1) has stationary variance $(1+\\varphi)/(1-\\varphi)$. Direct computation from $X_t=\\varphi X_{t-1}+(1-\\varphi)\\epsilon_t$ with $\\epsilon_t\\sim N(0,1)$ gives $(1-\\varphi)^2/(1-\\varphi^2)=(1-\\varphi)/(1+\\varphi)$, the reciprocal. The text even says autocorrelation reduces variance on smooth landscapes, which contradicts the printed formula. This error drives the scaling used to generate Fig. 3a and to claim that the optimal threshold exceeds the mean for all $\\varphi>0$. The proof in Appendix A covers only $\\varphi=0$, and even there it assumes without proof that the product of the reward component and the probability component is unimodal in $T$. So the general claim in the abstract is not established for autocorrelated landscapes. That said, the error is fixable: replace the variance formula or directly simulate Eq. (1) for various $\\varphi$.\n\nThe skewness and social-comparison results are simulation-based and have plausible intuition — the ambition-versus-risk-taking distinction is a nice observation. But they are not the main theorem.\n\nWho benefits? A reader wanting an accessible, well-illustrated map of how optimal search connects to everyday ambition will get value from the first few sections. A mathematical reader will find the gap in the autocorrelated case. I would send this to peer review, because the conceptual contribution and the simulations are substantial enough that a good referee can push the authors to fix the variance issue, restrict the claims, and tighten the proof. I would not cite it in its current form, but I would revisit after revision.","headline":"A readable but overreaching formalization of folk ambition: the i.i.d. core is sound enough, yet the smooth-landscape extension rests on a wrong variance formula and the proof of $T^*>\\mu$ has an unproven unimodality step.","tokens_in":17204,"tokens_out":3022,"would_cite":false,"duration_ms":28994,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that in a finite search, the optimal satisfaction threshold is finite and strictly above the mean reward, and that overshooting that threshold costs more than undershooting it.","keywords":["optimal stopping","satisfaction threshold","explore-exploit","search theory","reward landscape","social comparison","skewness","ambition"],"falsifier":"Simulate equation (1) with φ close to 1 and a Gaussian reward distribution, compute the expected cumulative reward over a fixed horizon for a fine grid of thresholds T, and check whether the argmax exceeds the mean μ; if the optimum reaches or falls below μ for any φ>0, the general claim fails.","tokens_in":16126,"feed_emoji":"🎯","tokens_out":6975,"duration_ms":64759,"temperature":0.7,"pith_summary":"This paper formalizes a folk intuition about ambition: in a time-limited search for strategies with unknown rewards, the best policy is to hold out for a result that is better than average but not the best possible. The model gives each searcher a satisfaction threshold T, measured in standard deviations from the mean reward, and asks how high that threshold should be to maximize total reward over finitely many time steps. The main result is that the optimal threshold is always finite and strictly above the mean, matching the folk wisdom, and this survives search costs as long as any search is worthwhile. The paper also finds that overshooting the optimal threshold is more costly than undershooting it, and that longer time horizons, rugged reward landscapes, and left-skewed reward distributions all raise the optimal threshold. A final result is that judging one's own rewards through upward social comparison substantially lowers expected performance.","feed_headline":"Aim slightly above average, not at the moon","feed_subtitle":"A search model shows the best satisfaction threshold is finite and above the mean; overshooting costs more than undershooting.","key_machinery":"The central object is the satisfaction threshold T, expressed as a number of standard deviations above or below the mean reward μ, which controls an AR(1) explore-exploit process: while rewards sit below T the searcher samples a new reward, and once a reward reaches T the searcher keeps it. The key identity is equation (2), which expresses the expected cumulative reward as a sum over the exploration length t_x of (μ_explore(T) t_x + μ_exploit(T)(t_max - t_x))(1 - Φ(T)) Φ(T)^{t_x}, with μ_explore(T) = -φ(T)/Φ(T) and μ_exploit(T) = φ(T)/(1-Φ(T)), where φ(·) and Φ(·) are the standard normal density and CDF. This identity makes the expected reward a unimodal hump-shaped function of T, and the proof that each summand is maximized at a positive threshold is what carries the claim that T* > μ.","core_discovery":"The authors claim to prove that in a finite-horizon search, the optimal satisfaction threshold T* satisfies μ < T* < ∞, where μ is the mean reward across strategies. The proof is worked out for the maximally rugged Gaussian case, where successive rewards are independent, by writing the expected reward as a sum over the possible length of the exploration phase and showing that every summand peaks at a threshold above zero. The authors then argue, from numerical simulation and by scaling the threshold with the variance of the autoregressive process, that the same conclusion holds on smooth autocorrelated landscapes and under constant search costs, provided the costs do not eliminate the incentive to search altogether. They also derive the asymmetry that overambition is costlier than caution, and they connect the threshold logic to empirical patterns in online dating and college applications.","pith_inferences":["For practical settings, the model suggests a measurable rule: set aspiration levels from the median or mean of the observable outcome distribution, then shift only modestly upward, rather than anchoring on the top performers.","The difference between ambition and risk-taking means that policy advice should separate the two margins: left-skewed environments call for cautious actions but ambitious targets relative to the mean.","A laboratory or natural experiment that lengthens the decision horizon should produce higher observed aspiration thresholds; this prediction is directly testable."],"forward_implications":["A finite search horizon has a sweet spot: the always-settle and never-settle strategies both earn the mean on average, while a threshold strictly above the mean earns more.","Longer searches justify higher ambition; as the horizon grows, the optimal threshold rises, and only an infinite horizon permits unbounded ambition.","Reward landscapes matter: rugged, weakly autocorrelated landscapes and left-skewed reward distributions call for higher thresholds, while right-skewed distributions call for thresholds closer to the mean.","Search costs lower the optimal threshold and the value of search, but the above-mean result holds as long as searching is profitable at all.","Upward social comparison is doubly harmful: it raises the perceived mean and lowers the threshold, reducing both satisfaction and earned reward."],"supporting_citations":[{"why":"Supplies the inverse Mills ratio formulas for the mean rewards in the explore and exploit phases, which feed the exact expected-reward expression.","marker":"[33]"},{"why":"Provides asymptotic limits of the inverse Mills ratios used to show the reward component grows near-linearly in the threshold.","marker":"[59]"},{"why":"Provides the online-dating messaging data used to illustrate that people search most intensely just above their own desirability.","marker":"[31]"},{"why":"Provides the college-application data used to illustrate search behavior around the mean and the role of constraints.","marker":"[32]"}],"fun_headline_variants":["Aim above average, but not too high: math says so","Overshooting ambition is costlier than undershooting","Search longer? Aim higher, but still finite","Social comparison ruins your optimal ambition","Perfect isn't the enemy; too high ambition is"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analytic proof that the optimal threshold exceeds the mean is carried out only for the maximally rugged Gaussian case (φ=0), and the extension to smooth, autocorrelated landscapes rests on a variance scaling that is not derived from the model's own recurrence equation.","fun_headline_variants_meta":{"raw":{"variants":["Aim above average, but not too high: math says so","Overshooting ambition is costlier than undershooting","Search longer? Aim higher, but still finite","Social comparison ruins your optimal ambition","Perfect isn't the enemy; too high ambition is"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000843,"raw_usage":{"total_tokens":3664,"prompt_tokens":931,"completion_tokens":2733,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":2659}},"tokens_in":547,"tokens_out":2733,"duration_ms":16424,"temperature":1.0,"reasoning_tokens":2659,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T18:19:48.056944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate equation (1) with φ close to 1 and a Gaussian reward distribution, compute the expected cumulative reward over a fixed horizon for a fine grid of thresholds T, and check whether the argmax exceeds the mean μ; if the optimum reaches or falls below μ for any φ>0, the general claim fails.","supporting_citations":[{"cited_title":"Muller and M.-P","cited_arxiv_id":null,"evidence_quote":"Supplies the inverse Mills ratio formulas for the mean rewards in the explore and exploit phases, which feed the exact expected-reward expression."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the online-dating messaging data used to illustrate that people search most intensely just above their own desirability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the college-application data used to illustrate search behavior around the mean and the role of constraints."}],"review_version":1}