REVIEW 4 major objections 5 minor 35 references
Uncertainty-guided active learning for surrogate prediction of stream-finishing wear fields
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper shows that an uncertainty-guided surrogate can map stream-finishing wear over all 696 feasible orientations from just 92 DEM simulations, with calibrated uncertainty that predicts where the reconstruction can be trusted.
desk verdict Useful surrogate for stream finishing wear, but the active-learning benefit is not actually demonstrated without a random-sampling baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by combining a deep ensemble with the Finnie wear relation. Rather than regressing wear rates directly, the surrogate predicts three smooth per-triangle fields from an 11-dimensional physical feature vector: rotated centroid position, surface normal, baseline coordinates, radial distance from the container axis, and alignment of the normal with the azimuthal flow. Wear is then reconstructed analytically through the Finnie formula, which confines the sharp angular and quadratic nonlinearity to an equation. The deep ensemble's disagreement provides the epistemic uncertainty that classifies orientations into accuracy bands and selects the next DEM simulations. A farthest
What would settle it
Run DEM on 30 orientations that the final surrogate labels low-uncertainty but were never simulated, then compare predicted and DEM wear fields. The central claim fails if a substantial fraction of these orientations show Spearman rank correlations below about 0.90, or if the predicted uncertainty no longer orders the orientations by realized error.
Extended reading notes
Core claim
The central claim is that the wear-rate field of a workpiece in stream finishing is a learnable function of geometry alone, and that its smooth constituents—per-facet normal impact velocity, tangential impact velocity, and particle impact flux—can be learned from a small fraction of all orientations. The paper shows that a deep ensemble seeded with 52 farthest-point orientations and refined through five rounds of ten uncertainty-selected DEM runs predicts these fields with held-out Spearman rank correlations of 0.93, 0.89, and 0.93. The ensemble's between-member standard deviation is a truthful pre-DEM estimate of realized error, with rank correlations of -0.77 and -0.82 against R² across he
Load-bearing premise
The load-bearing premise is that each surface facet's impact velocity and flux are smooth, deterministic functions of the 11 geometric features; if unresolved many-body collision noise dominates the variation, the surrogate's sparse-sample accuracy and its uncertainty calibration will not transfer to new geometries or operating conditions.
Editorial extensions
If this is right
- A new workpiece's full orientation-dependent wear map can be obtained from roughly a tenth of the exhaustive DEM programme, making orientation-sequence design feasible in practice.
- The input features are computable from any triangulated mesh, so the same pipeline transfers to arbitrary complex part shapes without new input engineering.
- Low-uncertainty orientations can be trusted without running DEM; only high-uncertainty orientations need additional simulation, and the uncertainty itself tells the user which ones.
- Rank-based agreement remains high even where absolute wear magnitudes degrade, so decisions based on relative wear across orientations remain usable at medium uncertainty.
- The low-uncertainty population is still growing at iteration 5, so a modest number of further batches is expected to bring nearly the whole feasible set into confident coverage.
Reading between the lines
- A natural but untested extension: the uncertainty thresholds are fit to the flat-wall workpiece, so applying the same pipeline to a new geometry would require recalibration of the σ_e-to-accuracy mapping with a small validation set.
- The 11 features are local and static; for parts with strong mutual shielding or internal passages, a facet's impact statistics may depend on the surrounding geometry, which could shrink the learnable smooth part and require more active-learning iterations.
- Because the DEM model uses a reduced Young's modulus, the predicted wear magnitudes are relative; converting surrogate outputs to absolute wear would still require experimental calibration.
- Choosing the next simulations by mean field-level uncertainty may be less efficient than targeting uncertainty in the reconstructed wear field or in the rank ordering of hotspots; that is a testable modification the paper does not explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a deep-ensemble surrogate to predict per-triangle normal impact velocity, tangential impact velocity, and particle impact flux for a cuboidal workpiece in stream finishing, using 11 geometric features per triangle. A farthest-point seeding scheme selects 52 initial DEM orientations, followed by five active-learning iterations that add the 10 orientations with the largest ensemble disagreement. The surrogate is evaluated on 14 held-out orientations and on uncertainty-band growth over the 696 feasible orientations. The authors report final-iteration Spearman rank correlations of 0.93, 0.89, and 0.93 for the three fields, claim the uncertainty is well calibrated, and reconstruct wear fields via the Finnie model matching DEM with Spearman correlation up to 0.97 at low uncertainty, using only 13% of the feasible orientations. The core idea is plausible, but the evidence presented does not yet establish the central claims about active-learning benefit and calibration.
Significance. The paper has several genuine strengths. The decision to predict smooth primary impact fields and to confine the nonlinear Finnie wear model to an analytical reconstruction is physically sensible and likely improves learnability. The use of an applicability-domain condition to avoid extrapolation is a thoughtful safeguard. The problem is practically relevant: orientation-dependent stream-finishing wear maps are expensive to obtain with DEM, and a reliable surrogate could substantially reduce cost. If the active-learning gains were supported by an appropriate control, the framework would be a useful template for geometry-to-wear surrogates. However, the manuscript currently stops short of demonstrating the central contribution: that uncertainty-guided acquisition outperforms simpler sampling strategies under an equal simulation budget.
major comments (4)
- [§2.2.6, §3.2] No baseline against random or diversity-only sampling is provided. The active-learning loop in §2.2.6 adds the 10 orientations with the largest σ_e and the abstract and §4 attribute the 13%-budget result to this uncertainty-guided strategy. But because the seed set already covers the feasible set by the d* criterion (§2.2.5), and because the training pool grows by 10 orientations per iteration, any reasonable rule that adds orientations within the covered region could plausibly produce similar gains. The reported monotone growth of the low-uncertainty band and the held-out accuracy is therefore not evidence for the value of the uncertainty signal. A control arm with the same seed set, batch size, retraining schedule, and either random or farthest-point additions is required to support the paper's central claim.
- [§3.1] The uncertainty calibration and its validation use the same 14 held-out orientations. Figure 5 reports Spearman correlations between σ_e and R²/ρ_s on those 14 orientations, and the σ_e thresholds for the low/medium/high bands are then fitted to the same 14-point σ_e–ρ_s relationship. This conflates calibration with validation, so the 'truthful uncertainty' claim and the band-growth figures are not independent evidence. A separate calibration split, or nested cross-validation, should be used, and the threshold values should be reported with uncertainty. With N=14, the rank correlations also have wide confidence intervals; point estimates of −0.77 and −0.82 should be accompanied by confidence intervals. In addition, §2.2.6 does not describe how the 14 validation orientations were selected; if hand-picked, this introduces another potential bias.
- [§3.2, Table 2.2, Figs. 6-7] No error bars or repeated realizations are reported for any main result. The deep ensemble uses random initializations and bootstrap subsets (§2.2.2), so the selected batches and final metrics are stochastic. Table 2.2 and Figures 6-7 present point estimates of R², ρ_s, and band fractions; the claims of monotone improvement and saturation at iteration 3 could reflect one realization. Re-running the loop with different ensemble seeds, or at least reporting variance over repeated training runs, is needed to support the quantitative 13%-data claim.
- [§3.1] The band thresholds fitted at iteration 5 are applied to earlier iterations ('the thresholds are fixed and applied unchanged to every iteration'), yet the σ_e distribution changes substantially across iterations, as shown in Table 2.2. The claim that a fixed threshold denotes the same accuracy level is reasonable only if the σ_e–ρ_s relationship is stable across iterations; this is asserted but not demonstrated. Reporting the σ_e–ρ_s relationship at each iteration, or fitting thresholds on a separate calibration set that is not also used for the Fig. 5 truthfulness test, would address this concern.
minor comments (5)
- [Abstract/Introduction] Typos and copyediting: 'uncertainity' in the Introduction, 'presenrlty' in the Summary, and the formatting of 'Ansys Rocky 2024R2' should be corrected.
- [§2.2.5] The relationship between d*, defined as the 95th percentile of nearest-neighbor distances within F, and its later use as a covering-radius stopping criterion needs clarification. Please state explicitly whether farthest-point sampling continued until the maximum distance to the nearest seed was below d*, and confirm that no feasible orientation lies beyond d* from a seed.
- [§3.1] The paper uses orientation-averaged R² and ρ_s, but notes that R² is unbounded below and one bad orientation can dominate an average. Reporting medians or per-orientation distributions instead of only averaged values would give a more complete picture, especially with only 14 validation orientations.
- [Global] No code or data availability statement is provided. Given the reproducibility value of the DEM surrogate pipeline, the authors should state whether the trained models, generated data, and scripts will be released.
- [§2.2.3] The 'irreducible limit' from unresolved stochastic collisions is acknowledged, but its size is never estimated. A comparison between ensemble error and the run-to-run variability of repeated DEM simulations at the same orientation would help contextualize the achievable accuracy of the surrogate.
Circularity Check
No significant circularity: surrogate and wear-field predictions are trained/evaluated against independent DEM data, with no load-bearing self-citation or by-construction reduction.
full rationale
The core derivation is self-contained. DEM provides independent per-triangle ground-truth velocities and flux; the deep ensemble is trained on a subset (52→92 orientations) and evaluated on a fixed held-out set of 14 orientations excluded from training. The wear field is reconstructed by applying the analytical Finnie relation (Eq. 2.1) to the predicted primary fields, so the wear output is not used as a training target and does not reduce to the surrogate's own inputs. Epistemic uncertainty is computed from ensemble disagreement before any DEM is run for a candidate orientation, and its correlation with realized error on held-out orientations is an empirical claim that could have failed, not an identity. The only caveats are validation-design issues: the σe-to-ρs thresholds for the uncertainty bands are fitted on the same 14 held-out orientations used for the truthfulness plot, and the active-learning benefit is not compared with a random-sampling control. These are limitations in evaluation rather than circular derivations: no equation or fitted parameter is renamed as an independent prediction. References are all external; no load-bearing self-citation appears.
Assumptions & free parameters
free parameters (5)
- sigma_e thresholds for uncertainty bands =
(4,7)e-3 normal; (10,15)e-3 tangential
- ensemble size M =
12
- active learning iterations =
5
- seed set size =
52
- validation set size =
14
assumptions (6)
- domain assumption DEM with Young's modulus reduced by 10^4 preserves relative bulk flow and wear patterns (Lommen et al. [26]).
- domain assumption Per-triangle expected impact velocities and flux are smooth, deterministic functions of the 11 geometric features.
- domain assumption Epistemic uncertainty from deep ensemble disagreement estimates model uncertainty (Lakshminarayanan et al. [22]).
- domain assumption The 14 held-out validation orientations are representative of the 696 feasible orientations for calibrating the sigma_e-to-rho_s relation.
- ad hoc to paper The kNN applicability domain with k=1 and the 95th percentile threshold (d*=1.21) ensures that all feasible orientations are interpolations.
- domain assumption Finnie model applies to this workpiece material and impact regime.
Cite this review
Pith. "Pith review of Uncertainty-guided active learning for surrogate prediction of stream-finishing wear fields." pith.science (2026). https://pith.science/paper/Y3NYGPBZ
@misc{pith2026260800593,
author = {Pith},
title = {Pith review of: Uncertainty-guided active learning for surrogate prediction of stream-finishing wear fields},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y3NYGPBZ}},
note = {Machine review of arXiv:2608.00593}
}
abstract
In stream finishing, the wear experienced by a workpiece depends strongly on its orientation within the rotating abrasive media. Determining suitable orientations to achieve uniform wear requires evaluating the wear-rate field over all feasible orientations. Although the discrete element method (DEM) accurately resolves particle interactions, simulating hundreds of feasible orientations for a new geometry is computationally expensive. We present an uncertainty-guided surrogate framework that predicts, directly from geometry, the three fields governing erosion: per-triangle normal impact velocity, tangential impact velocity, and particle impact flux. These fields are combined through the Finnie wear model to reconstruct the wear-rate distribution. The surrogate employs a deep ensemble whose disagreement estimates epistemic uncertainty, enabling an active-learning strategy that selectively performs DEM simulations for the most uncertain orientations. Trained using only $13\%$ of the $696$ feasible orientations, the surrogate achieves Spearman rank correlations of $0.93$, $0.89$, and $0.93$ for the normal impact velocity, tangential impact velocity, and particle impact flux, respectively. Moreover, the predicted uncertainty is well calibrated, reliably anticipating prediction error and the fidelity of the reconstructed wear field, which matches DEM with a Spearman rank correlation of up to $0.97$ for low-uncertainty orientations and degrades in a controlled manner as uncertainty increases.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Hashimoto.Abrasive Finishing Processes
F. Hashimoto.Abrasive Finishing Processes. Elsevier, 2006
work page 2006
-
[2]
LK Gillespie.Mass finishing handbook. Mass Finishing, Inc., 2007
work page 2007
-
[3]
Shengwei Ma, Keni Chih-Hua Wu, Stephen Wan, Cary Turangan, Kai Liang Tan, Wei Shin Cheng, Jun Ming Tan, and Bud Fox. Numerical simulation and experimental study of normal 25 force and particle speed in the robotic stream finishing process.Journal of Manufacturing Processes, 98:1–18, 2023
work page 2023
-
[4]
P. A. Cundall and O. D. L. Strack. A discrete numerical model for granular assemblies. G´ eotechnique, 29(1):47–65, 1979
work page 1979
-
[5]
I. Finnie. Erosion of surfaces by solid particles.Wear, 3(2):87–103, 1960
work page 1960
-
[6]
H. P. Zhu, Z. Y. Zhou, R. Y. Yang, and A. B. Yu.Discrete Particle Simulation of Particulate Systems: Theoretical Developments, volume 63. Chemical Engineering Science, 2008
work page 2008
-
[7]
Liming Zhao, Brooke Filanoski, Ruohong Chen, David Erickson, and Jingjie Yeo. Agent-based discrete element modeling of microbial-induced carbonate precipitation.Advanced Theory and Simulations, 9(3):e02233, 2026
work page 2026
-
[8]
P. W. Cleary. Large scale industrial dem modelling.Engineering Computations, 21(2/3/4):169– 204, 2004
work page 2004
Show all 35 references
-
[9]
Calibrated and validated wear prediction for bulk material handling equipment using dem simulations
Thomas Roessler and Andre Katterfeld. Calibrated and validated wear prediction for bulk material handling equipment using dem simulations. InICBMH2023: 14th International Conference on Bulk Materials Storage, Handling and Transportation, pages 284–295. The Institution of Engin...
2023
-
[10]
Simulation of solid particle erosion wear using discrete element method: Comparison of experimental and analysis results.Particuology, 2025
Mehmet Esat Aydin, Veysel Firat, and Mehmet Bagci. Simulation of solid particle erosion wear using discrete element method: Comparison of experimental and analysis results.Particuology, 2025
2025
-
[11]
Optimization of the stream finishing process for mechanical surface treatment by numerical and experimental process analysis.CIRP Annals, 68(1):373–376, 2019
Frederik Zanger, Andreas Kacaras, Patrick Neuenfeldt, and Volker Schulze. Optimization of the stream finishing process for mechanical surface treatment by numerical and experimental process analysis.CIRP Annals, 68(1):373–376, 2019
2019
-
[12]
Grinding surface roughness measurement combined with simulation data and transfer learning.Advanced Theory and Simulations, 7(5):2301100, 2024
Huaian Yi, Ziqiang Feng, and Aihua Shu. Grinding surface roughness measurement combined with simulation data and transfer learning.Advanced Theory and Simulations, 7(5):2301100, 2024. 26
2024
-
[13]
Efficient global optimization of expensive black-box functions.Journal of Global Optimization, 13(4):455–492, 1998
Donald R Jones, Matthias Schonlau, and William J Welch. Efficient global optimization of expensive black-box functions.Journal of Global Optimization, 13(4):455–492, 1998
1998
-
[14]
Liu, and X
Kaifeng Yang, S. Liu, and X. Chen. Bayesian optimization with gaussian process surrogates for high-dimensional design problems.Applied Mathematical Modelling, 104:215–231, 2022
2022
-
[15]
Kriging-based design optimization of an injection molding process for a light guide plate.Journal of Mechanical Science and Technology, 24(1):97–100, 2010
Soo-Hoon Chae, Hyo-Jin Kim, and Myung-Woo Cho. Kriging-based design optimization of an injection molding process for a light guide plate.Journal of Mechanical Science and Technology, 24(1):97–100, 2010
2010
-
[16]
Efficient global optimization applied to aerodynamic design of flyback booster.Journal of Spacecraft and Rockets, 44(5):1022–1030, 2007
Masahiro Kanazaki, Shinkyu Jeong, and Shigeru Obayashi. Efficient global optimization applied to aerodynamic design of flyback booster.Journal of Spacecraft and Rockets, 44(5):1022–1030, 2007
2007
-
[17]
Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. Practical bayesian optimization of machine learning algorithms. InAdvances in Neural Information Processing Systems, volume 25, pages 2951–2959, 2012
2012
-
[18]
Bayesian methods in political science
Jong Hee Park and Sooahn Shin. Bayesian methods in political science. In Luigi Curini and Robert Franzese, editors,The SAGE Handbook of Research Methods in Political Science & International Relations, pages 895–909. SAGE Publications, 2020
2020
-
[19]
Kriging-surrogate-based optimiza- tion considering expected hypervolume improvement in non-constrained many-objective test problems
Koji Shimoyama, Shinkyu Jeong, and Shigeru Obayashi. Kriging-surrogate-based optimiza- tion considering expected hypervolume improvement in non-constrained many-objective test problems. In2013 IEEE Congress on Evolutionary Computation, pages 2262–2269. IEEE, 2013
2013
-
[20]
Weight uncer- tainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. Weight uncer- tainty in neural networks. InProceedings of the 32nd International Conference on Machine Learning (ICML), volume 37, pages 1613–1622, 2015
2015
-
[21]
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. InProceedings of the 33rd International Conference on Machine Learning (ICML), volume 48, pages 1050–1059, 2016. 27
2016
-
[22]
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. InAdvances in Neural Information Processing Systems, volume 30, 2017
2017
-
[23]
Sculley, Sebastian Nowozin, Joshua V
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. InAdvances in Neural Information Proces...
2019
-
[24]
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. Deep ensembles: A loss landscape perspective. arxiv 2019.arXiv preprint arXiv:1912.02757, 2019
2019 arXiv
-
[25]
Mindlin and Herbert Deresiewicz
Raymond D. Mindlin and Herbert Deresiewicz. Elastic spheres in contact under varying oblique forces.Journal of Applied Mechanics, 20(3):327–344, 1953
1953
-
[26]
Dem speedup: Stiffness effects on behavior of bulk material.Particuology, 12:107–112, 2014
Stef Lommen, Dingena Schott, and Gabriel Lodewijks. Dem speedup: Stiffness effects on behavior of bulk material.Particuology, 12:107–112, 2014
2014
-
[27]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In3rd International Conference on Learning Representations (ICLR), 2015
2015
-
[28]
Active learning literature survey
Burr Settles. Active learning literature survey. Technical Report 1648, University of Wisconsin– Madison, 2009
2009
-
[29]
Physics-guided architecture (pga) of neural networks for quantifying uncertainty in lake temperature modeling
Arka Daw, R Quinn Thomas, Cayelan C Carey, Jordan S Read, Alison P Appling, and Anuj Karpatne. Physics-guided architecture (pga) of neural networks for quantifying uncertainty in lake temperature modeling. InProceedings of the 2020 siam international conference on data mining,...
2020
-
[30]
Defining a novel k-nearest neighbours approach to assess the applicability domain of a qsar model for reliable predictions.Journal of cheminformatics, 5(1):27, 2013
Faizan Sahigara, Davide Ballabio, Roberto Todeschini, and Viviana Consonni. Defining a novel k-nearest neighbours approach to assess the applicability domain of a qsar model for reliable predictions.Journal of cheminformatics, 5(1):27, 2013. 28
2013
-
[31]
Out-of-distribution detection with deep nearest neighbors
Yiyou Sun, Yifei Ming, Xiaojin Zhu, and Yixuan Li. Out-of-distribution detection with deep nearest neighbors. InInternational conference on machine learning, pages 20827–20840. PMLR, 2022
2022
-
[32]
Gonzalez
Teofilo F. Gonzalez. Clustering to minimize the maximum intercluster distance.Theoretical Computer Science, 38:293–306, 1985
1985
-
[33]
Active learning for convolutional neural networks: A core-set approach.arXiv preprint arXiv:1708.00489, 2017
Ozan Sener and Silvio Savarese. Active learning for convolutional neural networks: A core-set approach.arXiv preprint arXiv:1708.00489, 2017
2017 arXiv
-
[34]
Spearman
C. Spearman. The proof and measurement of association between two things.American Journal of Psychology, 15(1):72–101, 1904
1904
-
[35]
Fatigue life prediction of glare composites using regression tree ensemble-based machine learning model.Advanced Theory and Simulations, 3(6):2000048, 2020
Wei Sai, Gin Boay Chai, and Narasimalu Srikanth. Fatigue life prediction of glare composites using regression tree ensemble-based machine learning model.Advanced Theory and Simulations, 3(6):2000048, 2020. 29
2020
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.