GPT-4o and Claude Sonnet 4 show similar susceptibility to bias on GSM8K (1.3% vs 1.2%) but differ sharply in acknowledgment rates (13% vs 75%) under a rubric-defined metric.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
years
2026 3verdicts
UNVERDICTED 3representative citing papers
SW-DRSO optimizes a tractable surrogate of worst-case expected loss over plausible inference-time corruptions using a barycentric adversary approximated via simplex weights.
Physician-guided feature refinement in an interactive ML framework improves delirium detection discrimination and temporal robustness over automated baselines on 3862 hospital admissions.
citing papers explorer
-
Can Physician Expertise Improve Machine Learning Identification of Delirium?
Physician-guided feature refinement in an interactive ML framework improves delirium detection discrimination and temporal robustness over automated baselines on 3862 hospital admissions.