{"id":"96496770-2310-4502-b109-d0d93ccc6173","arxiv_id":"2511.09237","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Enrollment in Beijing's MaaS carbon-incentive program was associated with a 20.3 percentage-point increase in monthly low-carbon travel share and a model-estimated 1.8% citywide drop in gasoline-car trips.","lead":"This study analyzed 4.82 billion trips from roughly three million Beijing users to evaluate a carbon-incentive program built into a Mobility-as-a-Service app. It reports a 20.3 percentage-point rise in low-carbon travel share among enrolled users and a model-estimated annual CO2 reduction of about 94,000 tonnes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Random-forest counterfactual uses post-treatment trip duration/distance as inputs, making the 94,353 t CO2 and 1.8% car-trip figures potentially circular; re-estimate with exogenous features.","rationale":"The reader's weakest assumption identifies the same random-forest counterfactual concern, though the reader frames it more generally. My stress-test sharpens it: the inclusion of travel duration and distance as inputs creates a direct feedback loop, because these features are caused by the mode choice being predicted. This is not mere unobserved confounding; it is a structural flaw in the counterfactual estimator. The paper's explicit caveat that the carbon estimate is model-dependent is appropriate, but the caveat alone does not establish that the estimate is unbiased in the reported direction. The proposed concrete test—re-estimating with exogenous features and running a pre-period placebo—would settle whether the 94,353 t figure is an artifact. If the test shows large changes, the paper's central quantitative contribution would need revision, while the DiD behavioral effect may still stand. Thus the appropriate verdict is unchanged from the reader's CONDITIONAL: the paper should be accepted only with this re-analysis or a clear justification of why the duration/distance features do not bias the counterfactual. I do not see a basis for rejecting outright, since the paper is transparent about the model-dependence, but the concern is more concrete than mere 'selection bias.'","tokens_in":13838,"tokens_out":7372,"duration_ms":80139,"concrete_test":"Re-estimate the random-forest counterfactual excluding actual trip duration and distance features, using only predetermined trip attributes (origin, destination, departure time, workday, and exogenous network travel times from OSM). Recompute mode-shift ratios, daily car-trip reduction, and annual CO2. Additionally, run a placebo: apply the model (trained on pre-enrollment data) to pre-enrollment months with a fake enrollment date; nonzero 'shifts' indicate residual bias. If the recomputed CO2 differs from 94,353 t by more than 10%, the published carbon figure is not a reliable counterfactual estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline carbon figure (94,353 t CO2/yr; 1.8% citywide car-trip decline) is computed from a random-forest model that predicts each participant's counterfactual 'preferred mode' for every post-enrollment trip. Per Methods ('Estimation of the impacts of incentives on carbon emission reductions'), the model's inputs include travel duration and distance—both of which are themselves outcomes of the mode chosen. A trip that shifted from car to subway is observed with subway duration and distance, so the model, trained on pre-enrollment trips, will tend to predict 'subway' for that trip, not 'car.' The very mode shift the paper aims to count becomes largely invisible to the estimator. Conversely, trips whose post-enrollment duration/distance happen to resemble car trips may be misclassified as car, inflating apparent shifts. The reported accuracy (0.88) is measured on pre-enrollment data and does not validate the counterfactual use of the model. The DiD behavioral estimate (20.3%) is separate, but the paper's 'decarbonisation at scale' claim and the 94,353 t figure are direct linear functions of these predicted mode shifts. This is a specific, fixable—but load-bearing—correctness risk.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper evaluates a carbon-incentive program embedded in the Gaode MaaS platform in Beijing. Using 13 months of passively collected trip data from roughly 2.96 million users and a matched control group, the authors estimate a difference-in-differences effect of program enrollment on the monthly low-carbon trip share (reported as 20.3%), with event-study and placebo tests supporting the parallel-trends assumption. They then use a random-forest model trained on pre-enrollment trips to infer counterfactual modes for post-enrollment trips, from which they compute an estimated citywide 1.8% decline in gasoline-car trips and annual CO2 reductions of 94,353 tonnes (about 5.7% of Beijing's 2023 certified reductions). The paper also reports heterogeneity by gender, age, income, trip frequency, and urban zone characteristics. The abstract and discussion explicitly acknowledge that the emission estimate is model-dependent and that unobserved selection is a threat to causal identification.","tokens_in":14179,"tokens_out":9598,"duration_ms":101151,"significance":"The study's data scale (4.82 billion trip records, millions of users) and the use of passively collected records are major strengths, as are the matched DiD design, the event-study visualization, and the placebo test for the behavioral outcome. If the behavioral effect survives scrutiny, this would be one of the first city-scale demonstrations that MaaS-embedded carbon incentives can shift mode share. The authors also deserve credit for explicitly labeling the emission number as model-dependent rather than as a directly observed program effect. However, the headline emission reduction is not yet supported: the random-forest counterfactual is constructed from post-treatment features, and the citywide scaling uses an aggregate trip share rather than mode-specific shares. The paper is therefore best viewed as a credible behavioral evaluation with an as-yet-unverified carbon-accounting layer.","major_comments":[{"comment":"The random-forest counterfactual feeding Eqs. (8)–(10) uses travel duration as a feature, and enrollment itself increases trip duration by 7.4 min (Fig. 2b; Tables S13–S16). A car trip that shifts to subway is observed with subway duration, so a model trained on pre-enrollment trips will tend to predict 'subway' rather than 'car,' mechanically hiding the very shift the analysis aims to count. The reported pre-enrollment accuracy (0.88) does not validate predictions under this post-treatment distribution shift. Please re-estimate the counterfactual using only exogenous features (e.g., origin, destination, time of day, day type, individual covariates) and report how the 14.1% car-trip decline, the 1.8% citywide decline, and the 94,353 t CO2 figure vary. The current estimate is not merely uncertain; it is biased toward the program's intended mode shift.","section":"Methods: 'Estimation of the impacts of incentives on carbon emission reductions'; Results: 'Noticeable carbon emissions"},{"comment":"The 1.8% citywide decline in gasoline-car trips is derived from the participant-level decline (14.1%) and the statement that participants accounted for 12.6% of all daily trips. A mode-specific percentage change should be weighted by participants' share of gasoline-car trips, not by their share of all trips. If participants are less car-dependent than non-participants, 12.6% overstates the relevant weight; if they are more car-dependent, it understates it. Please report mode-specific trip shares and use them for the scaling calculation.","section":"Results: 'Noticeable carbon emissions reduction benefits' (citywide scaling)"},{"comment":"The abstract reports a '20.3 percentage-point increase in the monthly low-carbon travel share,' while the Results report a '20.3% increase relative to the matched control group' and a first-month increase of 20.4%. These are different quantities unless the baseline share is exactly 100%. Please state the baseline mean and report the coefficient in consistent units (percentage points vs. percent relative change) with confidence intervals.","section":"Abstract and Results: 'The MaaS incentive program increased low-carbon travel'"},{"comment":"The manuscript reports 3,983,027 registered program participants in Methods, 2,958,841 users in the final trip dataset, and a matched panel of 632,376 participants. It is unclear which sample underlies each analysis (behavioral DiD, carbon accounting, citywide scaling). The Discussion acknowledges that propensity-score matching cannot remove unobserved self-selection; given the voluntary enrollment, please state the analysis sample for each result and, if feasible, provide a sensitivity analysis (e.g., Oster's delta or bounds) quantifying how large unobserved selection would need to be to overturn the headline behavioral effect.","section":"Methods and Discussion: sample-size reconciliation and selection sensitivity"}],"minor_comments":[{"comment":"The paper uses 'tons' and 'tonnes' interchangeably (e.g., '94,000 tons' vs. '94,353 tonnes'). Use metric tonnes consistently.","section":"Throughout"},{"comment":"The statement 'subway network density exceeds 0.7 km²' should presumably be a density expressed in km/km²; please correct the units.","section":"Results: 'Urban characteristics associated with stronger behavioral responses'"},{"comment":"The abbreviation 'RCT' for the reduction in car travel ratio is confusing because RCT usually denotes randomized controlled trial; please rename, e.g., 'CRT' or spell out in full.","section":"Methods: Eq. (9)"},{"comment":"The definition of D_P_it,k is difficult to follow ('equals 1 for any period k that occurs after or during t-ST' reads backwards relative to the reference period). Please clarify the indexing so readers can map leads/lags to calendar months.","section":"Methods: Eq. (2)"},{"comment":"Some reference entries are incomplete or poorly formatted (e.g., reference 32 begins with 'Su Song, Miaoqing Zhong, D. T.'). Please check the reference list against journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The load-bearing issue with the random-forest counterfactual is fixable within the manuscript's scope: re-estimate with exogenous features and correct the mode-specific scaling. The behavioral DiD with event-study and placebo is a genuine strength. I saw no indication of selective reporting; the explicit 'model-dependent' caveat in the abstract is a positive sign. If the carbon accounting is revised appropriately, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the good news: this is genuinely new. 4.8 billion passively collected trips, a real carbon-market-funded incentive in a MaaS app, and a matched DiD with event-study and placebo tests. The 20.3% relative increase in low-carbon trip frequency is an observed behavioral estimate, and the parallel-trends and placebo checks give it some credibility. That alone is a useful, citable result for the MaaS and behavioral transport literature.\n\nThe carbon accounting is the problem. The random forest model that produces the 94,353 t CO2 figure uses travel duration (and, per the methods, travel conditions that include distance) as inputs. Those are outcomes of the mode actually chosen. A trip that shifted from car to subway is observed with subway duration, so the model trained on pre-enrollment trips will likely predict subway, not car. The mode shift the paper wants to count is largely invisible to the estimator. Accuracy of 0.88 on pre-enrollment data does not validate this counterfactual use. The stress-test note is right, and the paper's own disclaimer that the estimate is 'model-dependent' does not fix the circularity. The 1.8% car-trip decline and the 94,353 t figure are direct functions of those predicted shifts, so they should not be treated as evidence of decarbonisation at scale.\n\nOther soft spots are minor. The abstract says 'percentage-point increase' while the body says 20.3% relative; that matters for interpretation. Income is inferred from phone value, unobserved selection after matching is still possible, and the data are proprietary with only aggregated summaries on GitHub. Those are addressable, not fatal.\n\nThis is a conditional accept. The behavioral DiD is a solid contribution; the carbon accounting needs re-estimation with a counterfactual that excludes post-treatment outcomes, or a clear argument that the model actually identifies the underlying mode preference. If I were the editor, I'd send it to referees with this concern highlighted.","headline":"First city-scale evidence that MaaS carbon incentives shift behavior, but the headline CO2 figure rests on a circular counterfactual and should not be taken at face value.","tokens_in":14631,"tokens_out":2339,"would_cite":true,"duration_ms":21706,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a carbon-market-financed MaaS incentive in Beijing raised low-carbon trips by 20.3% and reduced annual CO2 by about 94,000 tonnes.","keywords":["carbon incentives","Mobility-as-a-Service","MaaS","travel behavior","difference-in-differences","propensity score matching","random forest","carbon accounting"],"falsifier":"Take the trained random forest and apply it to the non-participant sample in the post-enrollment months; if its predicted mode shares diverge significantly from the observed mode shares of non-participants (who were not exposed to the program), the counterfactual foundation for the 1.8% car-trip reduction and 94,353 t CO2 figure is invalid.","tokens_in":13781,"feed_emoji":"🚇","tokens_out":7111,"duration_ms":62560,"temperature":0.7,"pith_summary":"The paper set out to answer whether a carbon-incentive program embedded in a citywide Mobility-as-a-Service (MaaS) platform can shift travel behavior and cut CO2 at scale. Using 4.82 billion passively collected trip records from about 2.96 million people in Beijing over 13 months, it finds that enrolled participants increased their monthly low-carbon trip share by 20.3% relative to a matched control group, with a 12.8% increase still present eight months later. A random-forest counterfactual implies a 1.8% citywide decline in gasoline-car trips and an annual reduction of 94,353 tonnes CO2, about 5.7% of certified reductions traded in Beijing's carbon market in 2023. The authors frame these results as the first large-scale real-world evidence for market-financed digital incentives in MaaS, while explicitly noting that the carbon figure is model-dependent. If correct, the finding matters because it suggests a financially self-sustaining, behaviorally effective demand-side tool for urban transport decarbonization.","feed_headline":"MaaS carbon rewards lifted Beijing low-carbon trips by 20.3%","feed_subtitle":"Carbon-market-financed rewards also cut gasoline car trips 1.8% and saved ~94,000 t CO2 annually.","key_machinery":"The central object is the carbon-incentive program itself: verified emission reductions from subway, bus, and cycling trips are aggregated and sold in Beijing's carbon market, and the revenue funds rewards for users within the MaaS platform. The analysis rests on three pieces: (1) a propensity-score-matched difference-in-differences model for the causal behavioral effect; (2) a random-forest counterfactual, trained on pre-enrollment trips with features such as departure time, origin, destination, duration, workday, and date, that predicts counterfactual mode choice to compute mode-shift ratios and emission reductions; and (3) a graph-convolutional network with zone-level embeddings plus line","core_discovery":"The central discovery is that a MaaS-embedded carbon incentive can measurably shift mode choice at unprecedented scale. The behavioral effect is identified through propensity-score matching and a time-varying difference-in-differences design on a panel of 632,376 participants and 1,264,753 controls: participation raised monthly low-carbon trip share by 20.3 percentage points, with pre-enrollment trends near zero and no placebo effect. The carbon effect is derived from a random forest trained on pre-enrollment trips that predicts each participant's most likely mode absent the program; comparing actual modes to predicted ones yields a 1.8% reduction in citywide gasoline-car trips and 94,353 t","pith_inferences":["Editorial inference: The carbon-reduction estimate inherits any bias in the random forest's counterfactual; because trip duration and distance are among its inputs, and the program itself changes both, the 1.8% car-trip decline could be overstated or understated depending on how the model extrapolates.","Editorial inference: A direct test of the counterfactual would be to apply the trained random forest to non-participants and compare its predicted mode distribution to their actual modes in the same period; if accuracy degrades materially, the CO2 accounting needs rethinking.","Editorial inference: The authors do not disentangle the behavioral mechanism—reward value, environmental framing, or gamification; distinguishing these would help predict whether the effect transfers to cities with different carbon price levels or platform designs.","Editorial inference: The observed lower response among low-income users raises an equity question; future pilots could test whether adjusting reward amounts or payment timing changes uptake among lower-income groups."],"forward_implications":["If the behavioral effect is real, carbon-market-financed digital incentives can serve as a scalable, cost-recovering complement to infrastructure and technology measures for urban transport decarbonization.","The persistence of the effect (still 12.8% after eight months) suggests that repeated micro-incentives can create lasting habit change, not just short-term spikes.","The concentration of effects in zones with dense subway networks indicates that a program like this works best where low-carbon alternatives already exist, so infrastructure and incentives are complementary rather than substitutes.","A 1.8% reduction in gasoline-car trips at the city level, if replicated elsewhere, would put a MaaS incentive program on the same order as many conventional transport demand-management policies.","The heterogeneous responses across demographic and income groups point to room for adaptive reward design and targeted recruitment to raise overall effectiveness and address equity."],"fun_headline_variants":["Carbon rewards on Beijing MaaS lifted low-carbon trips 20.3 pp","Beijing MaaS carbon incentives: 20.3 pp more low-carbon trips","94k t CO2 saved via Beijing MaaS carbon rewards","20.3 pp low-carbon trip rise from Beijing MaaS carbon incentives","Beijing MaaS carbon rewards: +20.3 pp low-carbon, ~94k t CO2 cut"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the random-forest model trained on pre-enrollment trip attributes correctly predicts the travel modes participants would have used without the program—even though those same attributes (duration, distance) are altered by the program itself.","fun_headline_variants_meta":{"raw":{"variants":["Carbon rewards on Beijing MaaS lifted low-carbon trips 20.3 pp","Beijing MaaS carbon incentives: 20.3 pp more low-carbon trips","94k t CO2 saved via Beijing MaaS carbon rewards","20.3 pp low-carbon trip rise from Beijing MaaS carbon incentives","Beijing MaaS carbon rewards: +20.3 pp low-carbon, ~94k t CO2 cut"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001239,"raw_usage":{"total_tokens":4914,"prompt_tokens":727,"completion_tokens":4187,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":4084}},"tokens_in":471,"tokens_out":4187,"duration_ms":31474,"temperature":1.0,"reasoning_tokens":4084,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:38:59.931793+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained random forest and apply it to the non-participant sample in the post-enrollment months; if its predicted mode shares diverge significantly from the observed mode shares of non-participants (who were not exposed to the program), the counterfactual foundation for the 1.8% car-trip reduction and 94,353 t CO2 figure is invalid.","supporting_citations":[],"review_version":1}