{"id":"c48a545b-e37f-471b-bebc-2ae222f34d96","arxiv_id":"2507.01817","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"On eToro, traders mirror popular investors rather than profitable ones, and the simulation used to argue for performance-based signals depends on an autocorrelation in returns that the data do not show.","lead":"Using trading and mirroring records from the eToro platform, this study shows that investors choose whom to copy mostly by follower count rather than by past returns, and that this popularity bias is linked to worse average performance. The authors simulate the platform to argue that shifting the displayed signals toward performance would improve outcomes, but the simulation relies on an assumption about return persistence that their own data contradict.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Fig. 3A counterfactual depends on performance persistence (OU θ=0.65 ⇒ daily autocorr ≈0.52) that the paper's own Fig. 10A contradicts; if daily returns are unpredictable, the headline design principle is unsubstantiated.","rationale":"The abstract's central prescriptive claim is that prioritizing performance over popularity would dramatically improve outcomes. For that claim to hold, performance must be at least somewhat predictable from past performance. The paper's own empirical section (SI §5.2, Fig. 10A) states that the autocorrelation of Rit matches a null model with shuffled daily returns, i.e., it is indistinguishable from the overlap artifact of a 30-day rolling window. But the simulation model uses Eq. 7 with θ=0.65, which yields AR(1)-type daily autocorrelation exp(-0.65)≈0.52. This injects the exact predictability the data deny, and it is the mechanism by which performance-weighted mirroring in Fig. 3A improves outcomes. The paper validates the model against the cross-correlation of popularity and performance (Fig. 11) and between-trader correlations (Fig. 9), but not against Fig. 10A; the omission is telling because the model's R autocorrelation under θ=0.65 should visibly exceed the null band. One caveat: independent daily returns are not logically incompatible with performance-based mirroring if traders differ in persistent mean skill. But the model has a common µ for all agents, so it cannot exploit such skill differences, and the paper provides no empirical test of cross-sectional persistence to support the counterfactual. Therefore the reader's identified weakest assumption is the load-bearing one. My recommendation is unchanged: the prescriptive claim should be rejected as currently supported, while the descriptive findings may be salvageable.","tokens_in":18000,"tokens_out":9077,"duration_ms":96296,"concrete_test":"Rerun the Fig. 3A simulations with εit drawn independently across days (θ=0 in Eq. 7, equivalently the per-trader daily-return shuffle used in Fig. 10A), re-calibrating µ and σ to keep the −10.98 bps mean; then recompute the gain from a 10% reduction in βpop/βper and overlay the model's own R autocorrelation on Fig. 10A. If the 6.6% gain collapses to ~0 while the model's autocorrelation under the original θ=0.65 lies outside the null band, the counterfactual is an artifact of the contradicted persistence assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The prescriptive claim in Fig. 3A—that shifting from βpop to βper raises platform ROI by ~6.6%—requires that past performance contain information about future returns. In the model this information is injected by the OU process in Eq. 7/SI §4.1: with θ=0.65, daily independent returns have autocorrelation exp(-0.65)≈0.52. Yet SI §5.2/Fig. 10A reports that the empirical autocorrelation of the 30-day rolling performance Rit matches a null model built by shuffling each trader's daily returns, so past daily performance is not predictive. The model is validated in SI §5.1 and §5.3 but not against Fig. 10A, and its own Rit autocorrelation under θ=0.65 should lie well above the null band. A performance-based copying strategy can only look good in Fig. 3A because the simulator makes yesterday's winner likely to be today's winner; if the eToro data show no such persistence, the headline 'prioritizing performance dramatically improves outcomes' is an artifact of an unvalidated and internally contradicted calibration choice. The descriptive findings (popularity bias; explorers vs keepers) are not equally affected. A persistent-trader-means alternative could rescue the counterfactual even with zero daily autocorrelation, but the paper provides no evidence of it and its model lacks it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies social learning on the eToro social trading platform, using trade and mirroring data from a seven-month stable period. It reports that mirroring decisions are driven far more by popularity than by past performance (logistic regression log-odds 15.68 vs. 0.34), that most traders maintain few and slowly changing mirrors, and that traders who revise their mirrors more frequently (explorers) perform better. The authors build a data-informed agent-based model of mirroring dynamics with an Ornstein-Uhlenbeck process for independent returns, and use it to argue that shifting weight from popularity to performance in the mirroring choice would improve platform performance by about 6.6% (Fig. 3A). The paper concludes that platform design should prioritize performance signals.","tokens_in":18274,"tokens_out":6873,"duration_ms":80142,"significance":"If the model-based counterfactual were sound, the paper would provide a valuable design principle for social trading platforms and a quantitative illustration of how popularity bias degrades collective outcomes. The descriptive findings—popularity bias in mirroring, the explorers vs. keepers performance gap, and the weak correlation between popularity and performance—are interesting and are supported by careful empirical analyses, including fixed-effects logistic regression, attention to missing data, and null-model comparisons. The authors also provide code and data. However, the central prescriptive claim depends on an assumption about return persistence that the paper's own empirical analysis contradicts, so the headline result is currently unsubstantiated.","major_comments":[{"comment":"The Ornstein-Uhlenbeck process in Eq. (7) is calibrated with θ = 0.65, which gives a one-day autocorrelation of independent returns of roughly exp(-0.65) ≈ 0.52. This is the mechanism that makes past performance predictive of future performance in the simulator. Yet SI §5.2, Fig. 10A shows that the autocorrelation of the 30-day rolling performance Rit in the eToro data is indistinguishable from a null model constructed by shuffling daily returns, and the main text explicitly states that past performance is not predictive of future performance (Results, 'Informational limitations'). The model is never validated against this key autocorrelation statistic. Consequently, the counterfactual in Fig. 3A — that shifting mirroring weight from popularity to performance raises platform ROI by 6.6% — is a direct consequence of an assumed persistence that the data do not exhibit. Without evidence of performance persistence, or an alternative mechanism (e.g., persistent latent trader skill) that the data support, the claim that prioritizing performance dramatically improves outcomes is not established.","section":"SI §4.1, Eq. (7); SI §5.2, Fig. 10A; Fig. 3A"},{"comment":"The model draws each trader's mirror capacity κi from a Poisson distribution with mean 10, while the empirical average reported in the main text is ⟨κit⟩ = 2.26 ± 0.01 simultaneous mirrors. This is a factor-of-four mismatch in a quantity that is central to the model's representation of cognitive limits and to the definition of γi = ηi/κi. The paper claims the model mimics the empirical distributions and the absence of correlation between η and κ, but the mean capacity is not calibrated to the data. This discrepancy may materially affect the quantitative conclusions in Fig. 3A and the comparison of explorer vs. keeper performance in Fig. 3C, and it should be corrected or explicitly justified.","section":"SI §4, model setup; main text, 'Dynamic strategy limitations'"}],"minor_comments":[{"comment":"The text cites SI Appendix Text 5.3 for the claim that past performance is not predictive of future performance, but the relevant analysis is the autocorrelation study in SI §5.2, Fig. 10A; SI §5.3 concerns cross-correlation between performance and popularity.","section":"Main text, 'Informational limitations'"},{"comment":"The legend contains a typo: 'βperf' should be 'βper'.","section":"SI Fig. 11"},{"comment":"The phrase 'reducing the ratio of popularity to performance used by users in their mirroring decisions by 10%' is ambiguous; the authors should clarify whether βpop is reduced by 10%, or whether the ratio βpop/βper is reduced by 10%, and report the resulting change in the metric.","section":"Main text, Fig. 3A"}],"recommendation":"major_revision","confidential_remarks":"The descriptive findings are likely salvageable, but the model-based counterfactual is the paper's main advertised contribution. If the authors cannot recalibrate the model to respect the zero performance autocorrelation shown in their own data (e.g., by modeling persistent trader quality rather than short-run return autocorrelation), the paper should be revised to remove the prescriptive claim and reframed as a purely descriptive study. The mismatch between the model's κ mean of 10 and the empirical mean of 2.26 also needs attention."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the empirical part is worth your time, the model counterfactual is not. The paper documents that on eToro, mirroring decisions are dominated by popularity (log-odds 15.68 vs 0.34 for performance), that mirror traders underperform, and that traders who churn their mirrors ('explorers') do better than those who hold steady ('keepers'). That last pattern is the genuinely new bit—earlier work by Krafft et al. and Pan et al. already established the popularity bias itself. The data handling is careful: they discuss missing days, impute missing parent trades, and use fixed-effects regressions. Code and data are on GitHub.\n\nThe soft spot is load-bearing. The simulation counterfactual in Fig 3A—that shifting weight from popularity to performance raises platform ROI by ~6.6%—depends on the Ornstein-Uhlenbeck process used for independent returns. They set θ=0.65, which gives daily return autocorrelation around 0.52. But their own SI 5.2 (Fig 10A) shows that the empirical autocorrelation of the 30-day rolling performance matches a null model built by shuffling daily returns. In plain terms: on eToro, past performance does not predict future performance, and their sim assumes it strongly does. The 'performance beats popularity' result is therefore written into the model, not discovered from the data. The authors never validate the model's return autocorrelation against that null model. That's an internal contradiction between the calibration and their own empirical finding.\n\nThe explorer advantage, by contrast, is observational and might survive. But it also lacks controls for selection—maybe high-churn traders differ in risk appetite or attention—so treat it as suggestive rather than established.\n\nSo: who is this for? Someone working on social learning or recommendation design will find the descriptive results useful and the methodological failure instructive. It deserves a serious referee, because the empirical work is substantial and the flaw is identifiable and fixable. But in current form the central design principle should not be cited.\n\nRecommendation: send to peer review, with the expectation of major revision. If the authors either calibrate the model to the observed non-persistence or reframe Fig 3A as an illustration of what would happen if returns were persistent, the paper could be salvageable. I wouldn't desk-reject it, but I'd be surprised if any referee lets the counterfactual through as is.","headline":"A genuinely interesting descriptive study of popularity bias on eToro, undercut by a model counterfactual that assumes exactly the return persistence the data say is absent.","tokens_in":18828,"tokens_out":5296,"would_cite":false,"duration_ms":48114,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["89.65.Gh","89.75.Fb"],"model":"deepseek-v4-flash","headline":"Investors on a large social trading platform copy others based on popularity, not performance, and a calibrated model shows that shifting the balance toward performance would improve returns by 6.6%.","keywords":["social learning","popularity bias","social trading","mirror trading","temporal networks","behavioral finance","eToro","Ornstein-Uhlenbeck process"],"falsifier":"Re-estimate the model's independent-return component directly from the platform's daily closed trades and rerun the simulation with that measured persistence, or with returns made effectively independent day to day; if the 6.6% platform-ROI gain from performance weighting disappears, the design prescription rests on assumed return persistence rather than on the measured mirroring behavior.","tokens_in":17736,"feed_emoji":"📉","tokens_out":12128,"duration_ms":138532,"temperature":0.7,"pith_summary":"Using seven months of trading and mirroring records from a large social trading platform, this paper argues that investors overwhelmingly choose whom to copy by social popularity, not by the 30-day performance numbers that are also on screen. Popularity is only weakly correlated with returns, and because popularity changes slowly while performance is volatile, following popularity steers capital toward traders who are visible rather than traders who are good. Mirroring traders lose on average 10.98 basis points (hundredths of a percent) while non-mirroring traders gain 1.15; mirrored trades themselves lose 61.24 basis points. A temporal-network model calibrated to the platform reproduces these losses and shows that reweighting mirroring decisions toward performance, a 10% reduction in the popularity-to-performance weight ratio, improves platform-wide returns by 6.6%. The paper also finds that users who frequently revise their mirrors outperform those who keep fixed connections, suggesting that adaptive revision partially compensates for the misleading signal.","feed_headline":"Popularity, not performance, drives whom traders copy","feed_subtitle":"A model calibrated to eToro shows that weighting mirrors toward returns would lift platform performance by 6.6%.","key_machinery":"The load-bearing machinery is a temporal-network model of mirroring combined with the logistic choice rule that governs which trader gets copied. Each trader $i$ has a fixed mirror capacity $\\kappa_i$ and a daily revision rate $\\eta^+_i$; when revising, $i$ chooses target $j$ with probability $\\logit^{-1}[\\beta_{\\mathrm{per}} R_{jt} + \\beta_{\\mathrm{pop}} P_{jt}]$, where $R_{jt}$ is the 30-day rolling return and $P_{jt}$ is popularity, then drops the lowest-ranked existing mirror to keep capacity constant. Returns on day $t$ are the average of the previous day's mirrored returns plus an independent component modeled with an Ornstein-Uhlenbeck process, a mean-reverting stochastic process with short-term correlation, calibrated to individual return dynamics (mean reversion $\\theta=0.65$). Changing $\\beta_{\\mathrm{pop}}$ and $\\beta_{\\mathrm{per}}$ in this model shifts what gets copied, and the Ornstein-Uhlenbeck process determines whether past performance carries enough information for performance-weighted copying to pay off.","core_discovery":"The paper's central claim is that social learning on the platform is driven by popularity rather than performance, and that this bias is why copying others is costly. In a logistic regression with day and trader fixed effects predicting who gets mirrored, popularity enters with log-odds $\\beta_{\\mathrm{pop}}=15.68$ versus $\\beta_{\\mathrm{per}}=0.34$ for 30-day performance. Popularity and performance are nearly unrelated (mean correlation $0.11$), popularity autocorrelates strongly across months, while performance autocorrelation matches a shuffled-returns null model, and users terminate mirrors using the same popularity-weighted ranking they used to create them. A temporal-network model using these estimated weights reproduces the platform's aggregate loss, weak popularity-performance correlation, and the advantage of frequently revising mirrors. In that model, reducing the popularity-to-performance weight ratio by 10% improves platform-wide returns by 6.6%.","pith_inferences":["Beyond the paper, the same decoupling between a visible popularity signal and an objective quality signal may apply to influencer markets, news recommendation, and content platforms, where follower counts are easy to process and quality is noisy; the design lesson would be to reorder or reweight the displayed signal rather than asking users to ignore popularity.","A testable extension would randomize the ordering of traders shown to new users; if performance-sorted lists increase mirroring of high-ROI traders, the causal role of display bias is confirmed.","The model's equal-weighting and homogeneous-risk assumptions mean the 6.6% performance gain may not transfer to heterogeneous risk preferences; a version with risk-averse utility could show smaller or larger gains.","The explorers' advantage suggests that even a modest weight on performance can be amplified by frequent revision, so platforms might nudge users to re-evaluate their mirrors periodically rather than only improving the ranking."],"forward_implications":["If platforms reweight the signals behind mirroring toward rolling performance, aggregate investor ROI rises; a 10% cut in the popularity-to-performance ratio is estimated to improve platform performance by 6.6%.","Users who frequently revise their mirroring relationships (trading explorers) outperform users who hold fixed mirrors, and the gap grows when performance signals are emphasized; in simulations the most adaptive traders can achieve up to 3 times the performance of the least adaptive.","Because popularity and performance are only weakly correlated, and popularity autocorrelates strongly while performance does not, platform features that rank traders by follower count will keep steering capital toward strategies with no persistent edge.","The calibrated model reproduces the observed negative average return of social learners, the weak popularity-performance correlation, and the explorer advantage, so the same mechanism can be used to evaluate design changes before deployment.","If designers reduce the influence of popularity-based signals, performance becomes a leading indicator of future popularity, aligning social influence with realized returns."],"supporting_citations":[{"why":"Supplies the preprocessed eToro trade and mirroring dataset that the empirical analysis and model calibration use.","marker":"21"},{"why":"Provides the Ornstein-Uhlenbeck process used to generate persistent, mean-reverting independent returns in the simulation.","marker":"35"},{"why":"Introduces the explorer-versus-keeper strategy taxonomy from limited communication capacity that the paper adapts to mirror revision rates.","marker":"34"},{"why":"Documents weak correlation between popularity and quality in a cultural market, the comparison case for the popularity-performance misalignment.","marker":"31"},{"why":"Provides the preferential-attachment mechanism invoked to explain why popularity evolves slowly and independently of performance.","marker":"29"},{"why":"Frames adaptive updating of strategies in markets, connecting weak performance weighting to bounded rationality and motivating the explorer analysis.","marker":"10"}],"fun_headline_variants":["Copying traders pick popularity over performance","Popularity, not returns, drives social trading mirrors","Social trading bias: popularity outweighs performance","Reweighting social signals could lift platform returns","Traders who revise mirrors outperform static copiers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The counterfactual improvement from performance-based mirroring assumes that a trader's independent returns are persistent enough from one day to the next that past performance carries usable information; if daily returns are actually independent, as the paper's own shuffled-return comparison suggests, the promised gain would shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["Copying traders pick popularity over performance","Popularity, not returns, drives social trading mirrors","Social trading bias: popularity outweighs performance","Reweighting social signals could lift platform returns","Traders who revise mirrors outperform static copiers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1390,"prompt_tokens":907,"completion_tokens":483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":523,"tokens_out":483,"duration_ms":5451,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:42:39.683856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the model's independent-return component directly from the platform's daily closed trades and rerun the simulation with that measured persistence, or with returns made effectively independent day to day; if the 6.6% platform-ROI gain from performance weighting disappears, the design prescription rests on assumed return persistence rather than on the measured mirroring behavior.","supporting_citations":[{"cited_title":"Popularity and Performance: A Large-Scale Study","cited_arxiv_id":"1406.7729","evidence_quote":"Supplies the preprocessed eToro trade and mirroring dataset that the empirical analysis and model calibration use."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Ornstein-Uhlenbeck process used to generate persistent, mean-reverting independent returns in the simulation."},{"cited_title":", author Lara, R","cited_arxiv_id":null,"evidence_quote":"Introduces the explorer-versus-keeper strategy taxonomy from limited communication capacity that the paper adapts to mirror revision rates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames adaptive updating of strategies in markets, connecting weak performance weighting to bounded rationality and motivating the explorer analysis."}],"review_version":1}