{"id":"d45563d5-fd15-4b00-ad7e-1e21011c69a3","arxiv_id":"2505.09624","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"DBS-Gym is a configurable Kuramoto-based simulation environment that unifies 15 spatial, temporal, and bandwidth features for benchmark testing of adaptive DBS controllers, with RL and classical baselines evaluated at three complexity levels.","lead":"This paper introduces DBS-Gym, a simulation benchmark for adaptive deep brain stimulation in Parkinson's disease, built on a Kuramoto oscillator network with spatial, temporal, and frequency-band features. It provides a standardized testbed for reinforcement learning controllers to suppress pathological beta oscillations while saving stimulation energy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Parameter sensitivity is untested and key coupling constant K is unreported; reported algorithm rankings may be artifacts of a single hand-set configuration.","rationale":"The reader's weakest_assumption is the external validity of the Kuramoto-based environment: realism is asserted, not demonstrated. I agree that the only quantitative external check is the visual beta-burst comparison in Figure 1D. However, I see a more pressing internal-validity gap that is fully within the authors' control: the environment's distinguishing feature is configurability across 15 attributes, and the headline algorithm comparison is reported at a single, incompletely specified configuration. The coupling constant K, which directly controls synchronization and burst dynamics, is absent from Table A1, and the beta-locus size is stated inconsistently (0.55 in Table A1 versus about 25% in Appendix A.5.2). With no sensitivity analysis, the conclusion that SAC is the best online agent is a point estimate; a standard benchmark should provide evidence that rankings are robust to the very parameters that users are expected to configure. This concern is concrete and testable with the released code, and it strengthens rather than replaces the reader's conditional verdict. I marked agreement as partial because the reader's emphasis is on external clinical fidelity, while my concern is that even granting the model, the comparison lacks demonstrated stability.","tokens_in":30858,"tokens_out":4425,"duration_ms":48382,"concrete_test":"With the released code, run a systematic sensitivity sweep on Env2: vary coupling K over ±20% around the (currently unreported) training value, beta-locus size over 0.25–0.75, electrode-to-locus distance over 1–4 grid cells, and encapsulation/neural drift rates over 1–5% per event. Re-evaluate the trained SAC, IQL, and DDPG policies under each perturbed environment using the Table 1 protocol. If the best algorithm under the beta-power/energy trade-off changes across the sweep, the benchmark's comparison claim requires a stated default parameter regime and a robustness analysis; if the ranking is stable, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DBS-Gym is a benchmark for comparing aDBS algorithms, with a headline result that SAC outperforms other algorithms across environment levels (Section 4.2, Discussion). To support that claim, the comparison must be a stable property of the environment rather than of one hand-picked parameter point. Section 3.1 defines the dynamics through coupling constant K (Eq. 1), and Sections 3.2.1 and 3.2.3 state that beta bursting, locus size, and drift are controlled by K and associated coefficients; yet Table A1, the only parameter table, does not list K at all. Table A1 also gives beta-locus size as 0.55 (55%) while Appendix A.5.2 states it was set to about 25% of neurons, so even the reported configuration is internally inconsistent. No sensitivity sweep over K, beta-locus geometry, electrode distance, or drift rates is presented. If the ranking of SAC, IQL, and DDPG changes under plausible variation of these parameters, then the environment cannot serve as a standardized benchmark for comparing aDBS algorithms, independent of the separate question of whether the Kuramoto proxy is clinically faithful. The realism gap identified by the reader is real, but the more immediate internal weakness is that the reported comparison is not demonstrated to be robust to the arbitrary choice of unstated parameters.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DBS-Gym, a configurable simulation environment for adaptive deep brain stimulation (aDBS) based on a spatially embedded Kuramoto oscillator network. The environment implements three groups of features—bandwidth (beta sub-bands, nonstationary PSD), spatial (beta locus geometry, partial observability, directional/multi-contact electrodes), and temporal (neural and electrode drift, encapsulation, beta-burst modulation)—organized into three environment levels of increasing complexity. The authors benchmark classic HF-DBS, random stimulation, PI/PID controllers, and five RL algorithms (PPO, SAC, DDPG, IQL, CQL-SAC) under three reward functions, reporting beta-power suppression and energy consumption. The main empirical findings are that SAC performs consistently well, IQL performs comparably under the first reward, DDPG degrades under temporal drift, and CQL-SAC fails to learn an effective policy. The paper claims to be the first neurophysiologically realistic benchmark for comparing aDBS algorithms.","tokens_in":31136,"tokens_out":6671,"duration_ms":62700,"significance":"The environment is a potentially valuable open-source contribution: it is computationally efficient (JAX, 512 oscillators), integrates with Gymnasium/Stable-Baselines3, and covers a wider set of features than previous synthetic models (Figure 1F). The algorithm ranking is emergent rather than fitted to a desired outcome, and the paper ships code and uses standard reproducible RL implementations. If the configuration is made fully explicit and the ranking is shown to be robust to parameter variation, DBS-Gym could serve as a useful standard testbed for aDBS controller development. However, the 'neurophysiologically realistic' claim currently rests on a single visual burst-duration comparison, and the benchmark's usefulness for algorithm comparison hinges on the parameter sensitivity analysis that is not yet provided.","major_comments":[{"comment":"The environment's core parameters are not fully specified, and the reported configuration is internally inconsistent. The coupling constant K in Eq. (1) is stated to control beta bursting (§3.2.3) but is not listed in Table A1 nor given in Appendix A.5.2. 'Beta locus size, %' is 0.55 in Table A1, while Appendix A.5.2 says the locus was about 25% of neurons. In addition, §3.4 states the HF-DBS baseline uses a 90-microsecond pulse, but Table A1 sets 'Electrode stimulation duration' to 0.0015 s (1.5 ms), a 16-fold difference that changes the energy metric and the stimulation effect. These discrepancies make the exact training and evaluation setup unreproducible from the paper; please report K, reconcile the locus size, and correct or justify the pulse duration.","section":"§3.1, §3.4, Table A1, Appendix A.5.2"},{"comment":"The headline ranking (SAC > IQL/DDPG, CQL-SAC failure) is demonstrated at a single hand-set parameter point. No sensitivity analysis over K, beta-locus geometry, electrode distance, drift rates, or reward weights is presented. Because the environment's dynamics depend strongly on these choices (§3.2, Eq. (1)), the reported ranking may be an artifact of one configuration. Please provide a sensitivity sweep over at least K, beta-locus size and position, and electrode drift/encapsulation rates, and report how the algorithm ordering changes or, if it does not, the range over which it is stable.","section":"§4.2, Tables 1 and A2"},{"comment":"The central claim of 'neurophysiologically realistic' is supported only by a visual comparison of simulated beta burst durations with one patient cohort (Figure 1D); no quantitative goodness-of-fit is reported, and the other modeled attributes are justified by references rather than validated against experimental recordings. The limitations section (Appendix A.2) itself concedes that tissue volume activation, electrode capacitance, and other factors are not fully incorporated. Please add quantitative agreement measures (e.g., burst-duration distribution statistics, PSD peak location and width) and, if such validation is not yet available, soften the realism claim to 'mechanistically motivated' in the abstract and §1.","section":"§1, §3.2, Figure 1D"},{"comment":"The evaluation protocol trains and evaluates each algorithm on the same environment parameter distributions, so the reported performance differences reflect interpolation within one parameter regime rather than generalization across plausible clinical conditions. Even though Env2 includes drift events during evaluation, the underlying parameters remain in the training regime. For a benchmark whose purpose is to compare aDBS controllers for deployment, please include at least one held-out parameter configuration or an explicit out-of-distribution test to substantiate the claim that the ranking is robust and would transfer to unseen patient conditions.","section":"§4.1.2, §4.2"}],"minor_comments":[{"comment":"Units and labels are inconsistent: Table A2 mixes '%' and 'mV^2' for beta power across reward columns, and Figure 3 axis labels contain 'mV/two.numerator' and 'mV/two.numerator/Hz', evidently LaTeX artifacts. Please unify units and fix the axis labels.","section":"Table A2, Figure 3"},{"comment":"The discrete-time PID update uses t-1 for the integral and derivative terms; please clarify the discretization and report the tuned gains Kp, Ki, Kd for the PI/PID controllers, which are not listed despite being tuned with Optuna.","section":"Eq. (7), Appendix A.6.4"},{"comment":"The row 'Directed stimulation TURN ON - - -' is ambiguous: the main text says non-directional single-contact stimulation was used, so please replace the entries with explicit True/False values for each environment level.","section":"Table A1"},{"comment":"The observation window is described as a user-defined hyperparameter with a 1.2-second recommendation in §3.3.1, while Table A1 lists 'Observation window duration' as 1.17 s; please align these values.","section":"§3.3.1, Table A1"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim of 'first neurophysiologically realistic benchmark' is stronger than the evidence supports: the realism is asserted by construction with only one visual comparison for validation, and the internal parameter inconsistencies (missing K, locus size mismatch, pulse duration discrepancy) prevent reproduction. That said, the environment is a useful contribution and the algorithmic comparison is not circular. The requested sensitivity analysis and parameter reporting are within the scope of a revision and would materially decide whether the benchmark claim holds. Recommend major revision rather than rejection, contingent on the authors providing the missing parameter values, resolving the listed inconsistencies, and adding a sensitivity analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a solid engineering contribution, not a breakthrough. The DBS-Gym environment bundles 15 spatial, temporal, and bandwidth features into a configurable Kuramoto model and benchmarks eight controllers (PPO, SAC, DDPG, IQL, CQL-SAC, PI, PID, random) across three complexity levels. That unified testbed genuinely fills a gap: prior RL frameworks for aDBS are more limited, and the gymnasium-style API plus JAX implementation makes it usable. The benchmark numbers are internally coherent, the qualitative conclusions (SAC is the best online agent, CQL-SAC fails, temporal drift breaks several algorithms) are credible, and the limitations section is honest.\n\nThe soft spots are real but mostly fixable. The stress-test note lands: the main coupling constant K is not listed in Table A1, and beta locus size is given as 55% in Table A1 but ~25% in Appendix A.5.2. That internal inconsistency makes it hard to reproduce the exact configuration. More important, there is no sensitivity sweep over K, locus geometry, electrode distance, or drift rates. If the SAC/IQL ranking flips under plausible parameter changes, the 'benchmark' claim is just a single-point evaluation. As it stands, the comparison is a demonstration, not yet a benchmark.\n\nThe bigger marketing problem is the phrase 'neurophysiologically realistic.' The only external anchor is a visual comparison of beta burst durations to one patient cohort (Fig 1D), no quantitative fit or error bars. The frequency distribution, electrode kernels, and drift schedules are hand-set. That doesn't make the tool useless, but it means the realism claim is asserted, not demonstrated. Same for the code: abstract gives a GitHub URL, appendix says code will be open on publication. Sort that out.\n\nWho is this for? Researchers working on RL-based aDBS or neural control who want a common offline testbed for pre-training and controller comparison. They will get value even with the caveats. I'd send it to peer review, with a request for a reproducibility and sensitivity pass. It deserves referee time, and the issues are fixable.","headline":"A genuinely useful, well-built aDBS simulation testbed whose 'realistic' label is over-sold and whose headline algorithm ranking is not yet tested for robustness.","tokens_in":31681,"tokens_out":2678,"would_cite":true,"duration_ms":24887,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to provide the first neurophysiologically realistic benchmark for adaptive deep brain stimulation, a Kuramoto-oscillator environment with 15 previously ignored physiological features, and uses it to compare RL…","keywords":["deep brain stimulation","Parkinson's disease","reinforcement learning","Kuramoto model","beta oscillations","local field potentials","simulation environment","adaptive DBS"],"falsifier":"Record simultaneous subthalamic LFP from a cohort of Parkinson's patients under continuous and adaptive DBS, feed the same recorded signals or their statistics into DBS-Gym with the same electrode geometry, and check whether the simulated LFP responses, beta-burst statistics, and PSD changes match the patient data; a systematic mismatch would show that the benchmark's algorithm rankings need not transfer to clinical reality.","tokens_in":30625,"feed_emoji":"🧠","tokens_out":4593,"duration_ms":42785,"temperature":0.7,"pith_summary":"The paper aims to give the adaptive deep brain stimulation (aDBS) community a standardized, computationally efficient simulation in which different control algorithms, especially reinforcement learning agents, can be trained and compared before clinical testing. It argues that existing synthetic Parkinson's disease models omit key real-world complications and introduces DBS-Gym, a configurable Kuramoto-oscillator environment that recreates beta-band LFP features across bandwidth, spatial, and temporal domains. As a demonstration, the authors benchmark high-frequency DBS, PI/PID controllers, and five reinforcement learning algorithms across three difficulty levels, using the trade-off between beta power suppression and stimulation energy as the evaluation metric. A sympathetic reader would care because the field lacks a common training ground; if the environment's realism holds, RL agents pre-trained here could be fine-tuned on patient-specific data.","feed_headline":"A brain simulator benchmarks adaptive DBS on 15 real-world features","feed_subtitle":"Kuramoto-based testbed shows soft actor-critic best balances beta suppression and stimulation energy.","key_machinery":"The central object is a spatial Kuramoto model of N phase oscillators (default 512 on an 8×8×8 grid) with phase dynamics dθ_n/dt = ω_n + (K/N) Σ_m W_mn sin(θ_m − θ_n) + V(θ_n) A, where W_mn = cos(α_mn) encodes distance-dependent coupling, V(θ_n) = G(α_n,el) · PRC(θ_n) couples electrode stimulation, and the LFP is a conductance-weighted average of cos(θ_n). Around this core, the environment layers three feature groups—bandwidth (natural frequency distribution producing low and high beta), spatial (beta locus, partial observability, directional multi-contact electrodes), and temporal (STDP-based neural drift, electrode drift, encapsulation, burst modulation)—and exposes them as configurable parameters in a Gymnasium-style RL environment with sliding observation windows, a multi-contact action space, and reward functions trading beta power against energy.","core_discovery":"On its own terms, the paper introduces DBS-Gym as the first neurophysiologically realistic benchmark for adaptive deep brain stimulation, integrating 15 previously dismissed physiological attributes across spatial, temporal, and bandwidth feature groups, all modeled through Kuramoto phase oscillators on a 3D grid with distance-dependent coupling and electrode kernels. The environment reproduces PD-relevant LFP phenomena including beta bursting, partial observability, electrode drift, neural drift, and electrode encapsulation, and it supports configurable complexity levels (Env0, Env1, Env2) that isolate each feature group. Using this environment, the authors evaluate PI/PID and RL algorithms and report that Soft Actor-Critic (SAC) maintains the best balance of beta suppression and energy across all levels, while offline methods and DDPG degrade sharply under temporal drift.","pith_inferences":["If the environment's realism transfers to the clinic, this benchmark could make RL-aDBS studies comparable across labs; a natural extension would be a public leaderboard with fixed seeds and standardized hyperparameters.","The only external validation—matching beta burst durations from one patient cohort—is narrow; a stronger test would compare simulated PSD changes under real cDBS recordings, which the authors do not perform.","Since all parameters are hand-set, the benchmark's difficulty is somewhat arbitrary; one could invert the framework to tune environment parameters against real patient LFP datasets and thereby calibrate 'realism' quantitatively.","The paper's suppression-stability task introduces a new evaluation axis—control resilience to progressive degradation—that could generalize to other implantable closed-loop brain-computer interfaces."],"forward_implications":["Reinforcement learning agents can be pre-trained in the simulation and later fine-tuned on patient-specific data, addressing the scarcity of invasive aDBS training data.","The environment provides a standardized benchmark for comparing future aDBS algorithms on a common beta-suppression-versus-energy trade-off, which the authors argue no previous synthetic model unified.","The three-tiered complexity (Env0, Env1, Env2) allows researchers to isolate which feature domains (bandwidth, spatial, temporal) break which controllers, guiding algorithm design.","Stochastic-policy algorithms like SAC appear more robust to electrode drift and encapsulation than deterministic or offline policies, suggesting a design guideline for clinically deployed controllers.","Because the features are configurable, the testbed can be adapted to other stimulation pulse shapes, frequency bands, or neurological disorders without rebuilding the environment."],"supporting_citations":[{"why":"Supplies the Kuramoto oscillator model that forms the mathematical core of the environment.","marker":"[1]"},{"why":"Provides the patient beta-burst dynamics used for the only external comparison of the simulated LFP.","marker":"[114]"},{"why":"The prior reinforcement learning framework for DBS that this paper extends into a full benchmark environment.","marker":"[64]"},{"why":"Source for the clinical features (electrode drift, encapsulation, directional stimulation) that the benchmark includes.","marker":"[62]"},{"why":"Provides the offline RL training methodology and one of the reward functions used for CQL-SAC and IQL baselines.","marker":"[83]"},{"why":"Defines the Soft Actor-Critic algorithm that the paper reports as the best-performing controller.","marker":"[49]"},{"why":"Defines the PPO algorithm used as an online RL baseline in the comparison.","marker":"[101]"}],"fun_headline_variants":["DBS-Gym: realistic benchmark ranks adaptive DBS algorithms","Simulator adds 15 real-world quirks to test adaptive DBS","Adaptive DBS gets a realistic training ground with 15 features","Soft actor-critic wins in realistic adaptive DBS benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's realism rests on the unverified assumption that the manually chosen Kuramoto-oscillator dynamics and drift schedules reproduce the patient-relevant phenomena that actually determine how well an adaptive DBS algorithm performs in a living brain.","fun_headline_variants_meta":{"raw":{"variants":["DBS-Gym: realistic benchmark ranks adaptive DBS algorithms","Simulator adds 15 real-world quirks to test adaptive DBS","Adaptive DBS gets a realistic training ground with 15 features","Soft actor-critic wins in realistic adaptive DBS benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2809,"prompt_tokens":903,"completion_tokens":1906,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":1833}},"tokens_in":519,"tokens_out":1906,"duration_ms":13005,"temperature":1.0,"reasoning_tokens":1833,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:07:29.702108+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record simultaneous subthalamic LFP from a cohort of Parkinson's patients under continuous and adaptive DBS, feed the same recorded signals or their statistics into DBS-Gym with the same electrode geometry, and check whether the simulated LFP responses, beta-burst statistics, and PSD changes match the patient data; a systematic mismatch would show that the benchmark's algorithm rankings need not transfer to clinical reality.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the patient beta-burst dynamics used for the only external comparison of the simulated LFP."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source for the clinical features (electrode drift, encapsulation, directional stimulation) that the benchmark includes."}],"review_version":1}