Randomized calibrated expert aggregation that Blackwell-refines a target is polynomial-time solvable; deterministic aggregation is NP-hard and has no multiplicative PTAS for proper losses.
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.
Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.
Quantile-based trading strategies for battery arbitrage fail to incentivize honest probabilistic forecasts and ignore price dependence, while stochastic programs using full distributions better connect forecast accuracy to economic value.
A reference-free proxy scoring framework combined with GIRB calibration produces better-aligned evaluation metrics for summarization and outperforms baselines across seven datasets.
citing papers explorer
-
Algorithmic Expert Aggregation
Randomized calibrated expert aggregation that Blackwell-refines a target is polynomial-time solvable; deterministic aggregation is NP-hard and has no multiplicative PTAS for proper losses.
-
Robust Aggregation of Calibrated Forecasts
Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.
-
Eigenvalue Calibration for Semantic Embeddings of Large Language Models
Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.
-
Probabilistic Forecasting for Day-ahead Electricity Prices, Battery Trading Strategies and the Economic Evaluation of Predictive Accuracy
Quantile-based trading strategies for battery arbitrage fail to incentivize honest probabilistic forecasts and ignore price dependence, while stochastic programs using full distributions better connect forecast accuracy to economic value.
-
Calibrating Model-Based Evaluation Metrics for Summarization
A reference-free proxy scoring framework combined with GIRB calibration produces better-aligned evaluation metrics for summarization and outperforms baselines across seven datasets.