{"id":"610e238a-5644-4b3b-9c22-bb581f98dc93","arxiv_id":"2505.01921","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MLP models with two hidden layers outperform deeper networks and traditional linear benchmarks for pricing large-cap US stocks with portfolio factors.","lead":"This paper tests multilayer perceptron neural networks with a dynamic pyramid structure on factor pricing for 420 large-cap US stocks, using 182 portfolio factors. It finds that models with two or three hidden layers fit better out of sample than deeper networks, and that MLP-based factor investing mainly helps control downside risk rather than maximize returns.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section 3 universe filter requires full 2013-2021 survival, an ex post condition that biases the Tables 6-7 backtests and leaves the paper's downside-risk conclusion (Section 6) untested on a real-time tradable universe.","rationale":"I agree with the reader that the Section 3 survivorship filter is the most load-bearing weakness, and I would defend that choice over two other serious defects. First, the filter is provably ex post: requiring no missing data in 2013-2021 uses the future to select the universe, and the paper's own justification is internally inconsistent because Lehman Brothers failed in September 2008, inside the validation window, so the example does not illustrate the test-period filter, and the timing of the market-cap sort is unspecified. Second, the bias has a known direction: it removes the left tail of the return distribution, exactly the outcomes that determine maximum drawdown and downside deviation, so the paper's novel claim that MLP factor investing is about downside-risk control rather than absolute returns (Section 6) is tested only on a censored sample. The predictive pillar is more resilient: Table 3 compares all models on the same survivor universe, so the broad ranking that a shallow dynamic MLP beats linear benchmarks and deeper fixed-shape GKX2020 models on this universe is internally valid, which is why I keep the verdict at CONDITIONAL rather than moving to REJECT. The author's own Section 6 limitation statement concedes that 'stock selection' highly influences model and portfolio performance, which flags the same sensitivity. Two secondary concerns reinforce the conditional status but do not displace the survivorship issue. First, architecture selection and multiple comparisons: fw2 is identified as best on the same out-of-sample data used for the Diebold-Mariano tests in Table 5, and the best model is not stable across criteria (fw2 for equal-weighted R2 and MDD, fw3 for value-weighted MDD and Sortino, fw5 for pre-COVID Sharpe), so the 2-hidden-layers headline is partly an ex post selection. Second, Equation (43) defines Jensen's alpha as E(r) - E(r-hat), which is the mean out-of-sample prediction error, not the risk-adjusted intercept of Jensen (1968); the paper's observation that alpha is identical for equal- and value-weighted portfolios (Section 5.3) confirms it is not a portfolio-level risk-adjusted measure, so the significant t-statistics in Tables 6-7 support only the trivial claim that predictions are biased low. Both issues, like the survivorship filter, are addressable by re-analysis, so the reader's CONDITIONAL verdict stands (hence UNCHANGED) with a point-in-time re-run as the required condition.","tokens_in":27082,"tokens_out":16811,"duration_ms":166259,"concrete_test":"Re-run the Section 5 experiments on a point-in-time universe. At each annual rebalance from January 2013 to January 2021, select the top-15% market-cap NASDAQ/NYSE stocks with complete data available up to that rebalance date only, then add stocks that later delist, splicing in CRSP delisting returns (actual delisting return or the standard -30%/-55% convention), and re-estimate fw2, fw3, OLS, and the buy-and-hold benchmark. Recompute the Table 6 and Table 7 metrics (annual return, Sharpe, Sortino, MDD). If the full-testing-period fw2 MDD advantage (39.03% versus 53.10%) or its Sortino ranking shrinks or reverses, the downside-risk conclusion is an artifact of the ex post survival filter; the report should also state how many 2013 top-15% stocks delisted by 2021 and their market-cap weight.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 defines the 420-stock universe with the condition of 'no missing data in the testing period' (1/2013-12/2021), an ex post survival filter: an investor at the first rebalance in January 2013 cannot know which stocks survive to December 2021, so the evaluated universe is not the tradable opportunity set. The paper's justification that this 'going concern' filter is more realistic for practitioners is internally inconsistent: the Lehman Brothers example is temporally misplaced (Lehman failed in September 2008, inside the validation window 2/2003-12/2012, not the 2013-2021 test window), and the date at which the top-15% market-cap sort is applied is unspecified, so it too may be full-sample. The bias runs in a known direction for the paper's novel claim: the filter removes stocks that crashed or delisted during the test period, mechanically lowering maximum drawdown and downside deviation for every strategy in Tables 6-7. The headline result that MLP factor investing is more meaningful for downside risk control than for absolute annual returns (Section 6) is therefore evaluated on a censored return distribution: the reported fw2 advantage of 39.03% MDD versus 53.10% for buy-and-hold, and the Sortino rankings, are not established for a point-in-time universe. The R2 ranking in Table 3 is less affected because all models share the same survivor universe, so the broad shallow-versus-deep comparison is internally valid, but the backtest metrics and the buy-and-hold benchmark do not transfer to a real-time strategy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies multilayer perceptron (MLP) models with a dynamic pyramid structure to 182 firm-characteristic-sorted portfolio factors to forecast excess returns of 420 large-cap US stocks over 2013-2021, extending the GKX (2020) framework. The main empirical claims are (i) a two-hidden-layer MLP achieves the highest out-of-sample R2 (3.66%) and outperforms deeper MLPs, OLS, PLS, PCR, and fixed-shape GKX models; (ii) the COVID-19 period improves MLP OOS fit; and (iii) MLP factor investing is more valuable for downside-risk control than for absolute returns, based on backtests of a long-only signal strategy. The paper also reports variable importance and compares equal- and value-weighted portfolios.","tokens_in":27373,"tokens_out":6272,"duration_ms":58448,"significance":"If the results held, the paper would provide a useful practical data point for ML asset pricing: it shows that with a limited cross-section and a relatively short OOS window, shallow dynamic MLPs can be more effective than deeper fixed architectures, and that factor-sorted portfolio characteristics can substitute for a much larger predictor set. The within-universe OOS R2 comparison and DM tests are appropriate tools, and the paper is transparent about its data sources and provides pseudocode for its optimization procedures. However, the backtest and downside-risk conclusions rest on a survivorship-biased universe, the alpha measure is mislabeled and does not measure risk-adjusted performance, and the architecture is selected on the same test set used for reporting, so the headline claims require substantial revision before they can be accepted as stated.","major_comments":[{"comment":"The stock selection criterion 'having no missing data in the testing period' is an ex post survivorship filter: an investor at the start of the test period cannot know which stocks will survive through 2021. This biases the backtest metrics (Sharpe, Sortino, MDD, alpha) in Tables 6 and 7, because stocks that crashed or delisted during the test window are mechanically excluded. The paper's justification using Lehman Brothers is temporally misplaced: Lehman failed in 2008, which falls in the validation window (2/2003-12/2012), not the 2013-2021 test window. Consequently, the Section 6 conclusion that MLP factor investing is mainly useful for downside-risk control is not established for a point-in-time tradable universe. The R2 ranking in Table 3 is less affected because all models share the same survivor universe, but the backtest conclusions need to be re-run on an investable universe or explicitly reframed as conditional on survival.","section":"Section 3, Tables 6-7"},{"comment":"The best architecture (fw2, two hidden layers) is selected on the basis of the highest OOS R2 computed on the same 2013-2021 test period used to report results and to run DM tests. This creates a selection-on-the-test-set bias: the reported R2 gap for fw2 and the DM test significances against other models are inflated because the same data were used to choose the architecture. The paper should either use a separate validation period for architecture selection, or honestly report that the R2 values are conditional on in-sample selection and adjust the inference accordingly.","section":"Section 5.2, Table 3"},{"comment":"Equation (43) defines alpha as the difference between the expected out-of-sample excess return and the expected predicted excess return. This is not Jensen's alpha, which is the intercept from a time-series regression of portfolio excess returns on factor exposures. As a result, the alpha values and t-statistics in Tables 6 and 7 do not measure risk-adjusted performance or factor profitability; they merely say that the average predicted return is lower than the average realized return, which is not an economically meaningful performance metric. The interpretation that 'all models have significant positive alphas, which indicates the extra gain from factors' is therefore unsupported.","section":"Section 5.3, Equation (43)"},{"comment":"The empirical results are not reproducible because the hyperparameter values are not reported. The paper mentions L1 regularization (Equation (34)), early stopping, batch normalization, and Adam, and Appendix A gives pseudocode, but the actual values used (learning rate, batch size, maximum epochs, early-stopping patience, regularization strength lambda, and any hyperparameter tuning procedure) are missing. Given that the central claim is a comparative empirical evaluation, the absence of these details prevents verification and makes the results sensitive to unspecified choices.","section":"Sections 4.1-4.4"}],"minor_comments":[{"comment":"The dynamic pyramid formula appears to contain a typo: 'O(l0)' is not clearly defined, and the neuron counts in Table 2 (e.g., 36 and 6 for two hidden layers) do not obviously follow from the formula as printed. Please correct the notation and verify the computation.","section":"Equation (21)"},{"comment":"The description of the trading signal is ambiguous: 'if the sign of the actual excess return and predicted excess return are both positive, then long the stock from the next value.' Please clarify whether the actual return used is the realized return at time t (a momentum filter) or the contemporaneous return at time t+1 (which would introduce look-ahead bias).","section":"Section 5.1, trading signal rule"},{"comment":"The statement that 'portfolio weighting has no impact on extra gain alpha' and that the value-weighted alpha equals the equal-weighted alpha is incorrect for a weighted average of stock-level alphas; the value-weighted alpha is a weighted average, not generally equal to the equal-weighted alpha. This should be corrected or clarified.","section":"Section 5.3, value-weighted portfolio"},{"comment":"The reported average R2 and average MSE are non-monotonic across models (e.g., fw2 and pcr share the same average MSE but differ greatly in average R2). The text explains this with the multiple-labels caveat, but it would aid the reader to also report the cross-sectional dispersion of the underlying variances or per-stock R2 distributions.","section":"Table 3"},{"comment":"The date at which the top-15% market-capitalization sort is applied is not specified; please state whether the sort is point-in-time or uses full-sample information, because this affects the interpretation of the selected universe.","section":"Section 3"},{"comment":"There are numerous typos and formatting issues, including 'Univiersity', 'gamil.com', 'Rresearch', 'Shape Ratio' for 'Sharpe Ratio', and inconsistent author formatting for the Chen-Zimmermann reference. A careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is written more like a thesis chapter and contains several internal inconsistencies (e.g., the alpha definition, the survivorship justification, the statement about value-weighted alphas). The core idea is worth pursuing, but the survivorship bias and model-selection issue are central to the paper's practical claims, and the alpha mislabel undermines the factor-investing interpretation. I would encourage the editor to request a revision rather than reject, since the OOS R2 comparison within the selected universe is a legitimate and partially informative exercise. The authors should also be asked to place their contribution more clearly relative to the existing ML asset-pricing literature, particularly regarding the use of characteristic-sorted factors and the choice of the dynamic pyramid structure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nThe paper is a credible but flawed empirical study. The core finding — that a two-layer dynamic MLP beats fixed deep architectures and linear benchmarks on 420 large-cap US stocks with 182 characteristic-sorted portfolio factors — is plausible, and the model-comparison machinery is set up more honestly than most ML-in-finance papers. The portfolio and downside-risk conclusions, however, rest on a survivor-biased universe, and the thing called \"Jensen's alpha\" is not Jensen's alpha. I'd send it to referees, but with the expectation of major revision.\n\nWhat is actually new: the dynamic pyramid structure from Coqueret-Guida is applied to a new dataset, and the paper carefully compares OOS R2, DM tests, variable importance, and backtests across two periods (with and without COVID). The internal model ranking — shallow beats deep — is credible because all models share the same survivor universe, so the comparison is apples to apples. The DM tests broadly support the R2 ranking. That's real work, and the writing is clear.\n\nThe soft spots are significant. The Section 3 universe requires no missing data over the full 2013-2021 test window. That is an ex post survival filter: an investor at the first rebalance cannot know which stocks will survive to 2021. The Lehman Brothers example given as justification is temporally misplaced (Lehman failed in 2008, inside the validation window), and the bias direction is unambiguous: removing crash-and-delist stocks mechanically lowers maximum drawdown and downside deviation for every strategy, including the buy-and-hold benchmark. The claim that MLP factor investing \"is more meaningful for downside risk control\" is therefore not established on a point-in-time tradable universe. This is the load-bearing flaw.\n\nSecond, the fw2 headline is selected by evaluating OOS R2 on the same test set used for reporting. That is test-set selection; it inflates the apparent advantage. The paper reports all architectures, which helps, but it should use a validation-based choice or at least discuss the multiple testing.\n\nThird, Equation (43) defines alpha as the mean prediction error, E(r) - E(r-hat), not as Jensen's alpha from a factor regression. The t-tests on that quantity in Tables 6-7 are not tests of risk-adjusted performance. That's a mislabel that needs fixing.\n\nNo code or data is provided, which is frustrating but secondary.\n\nWho is this for? Researchers and practitioners in ML asset pricing, especially those interested in architecture choice and drawdown behavior. It deserves a serious referee: the empirical question is legitimate and the core model comparison is well constructed. My verdict on the portfolio claims is skeptical until re-run on a point-in-time universe with proper architecture selection and corrected alpha. I'd engage with a revised version.","headline":"Plausible model comparison, but survivorship bias and a mislabeled alpha undermine the portfolio conclusions; worth refereeing with major revision.","tokens_in":27945,"tokens_out":3700,"would_cite":false,"duration_ms":36500,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a two-hidden-layer dynamic pyramid MLP prices large-cap US stocks best out of sample, and that MLP factor investing matters mainly for downside risk control.","keywords":["asset pricing","multilayer perceptron","neural network","factor investing","out-of-sample R-squared","downside risk","large-cap US stocks","COVID-19"],"falsifier":"Re-run the backtest on a point-in-time universe that includes stocks that later delisted or stopped reporting; if the two-layer MLP's Sharpe ratio, Sortino ratio, and maximum drawdown advantage over buy-and-hold shrinks or flips, the downside-risk claim would be refuted.","tokens_in":26799,"feed_emoji":"📉","tokens_out":10993,"duration_ms":97362,"temperature":0.7,"pith_summary":"This paper asks whether multilayer perceptron networks can price individual large-cap US stocks from 182 characteristic-sorted portfolio factors, and how their structure should be chosen. Its central claim is that a two-hidden-layer MLP with dynamically sized layers (36 and 6 neurons) achieves the best out-of-sample fit, with an average R-squared of 3.66%, beating deeper networks, the fixed-shape benchmark network from the 2020 study, and linear methods such as OLS, PLS, and PCR. The second claim is that MLP factor investing does not beat buy-and-hold in absolute annual returns but is valuable for downside risk control: the winning model posts the lowest maximum drawdown and the highest risk-adjusted ratios in the full-period equal-weighted test. If the claims hold, practitioners should stop assuming 'deeper is better' for factor pricing and should evaluate neural factor strategies primarily by drawdown and downside-risk metrics.","feed_headline":"A two-layer MLP outperforms deeper networks in pricing large-cap stocks","feed_subtitle":"On 420 US large caps it posts 3.66% out-of-sample R2 and the lowest drawdown, beating OLS, PLS, PCR and deeper MLPs.","key_machinery":"The carrying mechanism is the dynamic pyramid MLP: a fully connected network whose hidden-layer widths shrink geometrically from the 182-factor input toward the single output, with the width schedule set by a formula that depends on the total number of hidden layers. For two hidden layers this produces 36 and 6 neurons, the configuration with the best out-of-sample fit. Estimation uses MSE loss, ReLU activation, adaptive-moment gradient optimization, early stopping, batch normalization, and L1 regularization; predictions are scored by out-of-sample R-squared and a pairwise forecast-error test, then turned into long-only sign-based equal- and value-weighted portfolios.","core_discovery":"The paper reports that, on 420 large-cap US stocks with 182 characteristic-sorted portfolio factors, a two-hidden-layer MLP whose widths are set by a dynamic pyramid rule achieves an average out-of-sample R-squared of 3.66% over 2013–2021, and 2.16% when the COVID-19 months are removed. This beats every alternative tested, including one- and three-layer dynamic networks, five-layer networks that fall to −1.04%, the fixed-shape benchmark network from the 2020 study, and OLS, PLS, and PCR. The author interprets the pattern as evidence that deeper MLPs overfit at this data size. In the long-only backtest, all models earn positive and significant alphas, yet none beats buy-and-hold annual returns; the two-layer model instead posts the lowest maximum drawdown (39.03%) and the highest Sharpe and Sortino ratios in the equal-weighted full-period test, supporting the paper's conclusion that MLP factor investing is mainly a downside-risk management tool.","pith_inferences":["Inference: a point-in-time tradable universe that includes delisted stocks would likely reduce the reported Sharpe and Sortino ratios, so the downside-risk advantage is best read as an upper bound.","Inference: the dynamic pyramid width rule should be portable to other markets; the depth result implies the best architecture depends on the data-to-parameter ratio, so smaller samples should favor even shallower networks.","Inference: adding a signal filter that requires a minimum predicted return before opening a long position may recover some of the shortfall against buy-and-hold in strong uptrends, where the paper finds unfiltered sign signals fail.","Inference: because Announcement Return, Earnings Forecast Disparity, and Size dominate variable importance across models, an ablation study using fewer than 182 factors could test whether the remaining factor zoo adds predictive value."],"forward_implications":["The optimal depth for MLP factor models at this data scale is two or three hidden layers; networks with five layers go to negative out-of-sample R-squared.","Practitioners should evaluate MLP factor strategies by Sharpe ratio, Sortino ratio, and maximum drawdown rather than annual return, since none of the tested models beat buy-and-hold in absolute return.","Including the COVID-19 months in the test window improves the proposed models' out-of-sample fit, indicating the dynamic pyramid networks stay usable in extreme market moves.","Value-weighting the portfolio lowers maximum drawdown further, with the three-layer model reaching 33.04% in the full testing period."],"supporting_citations":[{"why":"Supplies the benchmark MLP architecture, loss function, and evaluation protocol this paper extends; it is the main comparison target.","marker":"[12]"},{"why":"Provides the dynamic pyramid structure used to size the hidden layers, yielding the winning two-layer network.","marker":"[19]"},{"why":"Source of the 182 firm characteristic-sorted portfolio factors used as predictors.","marker":"[45]"},{"why":"Provides the firm characteristic-sorted return factors adopted as model inputs.","marker":"[18]"},{"why":"Earlier evidence that ML factor investing moderates downside risk, the conclusion this paper verifies.","marker":"[37]"},{"why":"The forecast-error comparison test used to establish that model performance differences are statistically significant.","marker":"[53]"}],"fun_headline_variants":["Two-layer MLP beats deeper nets on large-cap pricing","Shallow MLP wins: 3.66% out-of-sample R² on US large caps","MLP depth overfits: 2 hidden layers outperform 5 on US stocks","For large-cap pricing, a 2-layer MLP trims drawdown best","MLP factor investing shines for downside risk, not returns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 420 stocks with no missing data through the full test period are the same universe an investor could have traded from 2013, even though firms that delisted or stopped reporting were excluded.","fun_headline_variants_meta":{"raw":{"variants":["Two-layer MLP beats deeper nets on large-cap pricing","Shallow MLP wins: 3.66% out-of-sample R² on US large caps","MLP depth overfits: 2 hidden layers outperform 5 on US stocks","For large-cap pricing, a 2-layer MLP trims drawdown best","MLP factor investing shines for downside risk, not returns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1528,"prompt_tokens":934,"completion_tokens":594,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":550,"tokens_out":594,"duration_ms":5547,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:07:44.938522+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the backtest on a point-in-time universe that includes stocks that later delisted or stopped reporting; if the two-layer MLP's Sharpe ratio, Sortino ratio, and maximum drawdown advantage over buy-and-hold shrinks or flips, the downside-risk claim would be refuted.","supporting_citations":[{"cited_title":"CRC Press, New York, USA (2020)","cited_arxiv_id":null,"evidence_quote":"Provides the dynamic pyramid structure used to size the hidden layers, yielding the winning two-layer network."},{"cited_title":"Journal of Business and Economic Statistics 13(3), 253–263 (1995)","cited_arxiv_id":null,"evidence_quote":"The forecast-error comparison test used to establish that model performance differences are statistically significant."}],"review_version":1}