Pith. sign in

REVIEW 4 major objections 6 minor 51 references

English-language news context systematically biases LLM territorial predictions toward Russian capture, and these biased pushes are wrong 64–72% of the time, a distortion the paper argues originates in the sources, not the models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

English news context systematically biases LLM predictions on Ukraine territorial markets toward Russian capture, and the bias originates in the text, not the model.

T0 review reviewed 2026-08-02 challenge →

load-bearing objection A clever measurement idea that overreaches in its headline: the push-accuracy claim ignores base rates, but the bias-shift and MAE results are worth a careful look. the 4 major comments →

arxiv 2607.20441 v1 pith:QT6M2A4J submitted 2026-05-12 cs.CL cs.LG

Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets

classification cs.CL cs.LG
keywords LLM beliefsprediction marketsinformation biasframing costterritorial predictionsablation studybelief elicitationUkraine conflict
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish a calibrated way to measure how far the beliefs an information ecosystem induces deviate from an external reference. It treats an LLM as an instrument that reads out, in probability form, the beliefs implied by the text it is given, and uses prediction-market prices — anchored at resolution by real outcomes — as the calibration scale. Applied to 111 Ukraine-related markets and roughly 93,000 predictions from four models, the paper finds that adding English-language news context systematically shifts territorial predictions toward Russian capture, and those shifts are wrong 64–72% of the time. A contaminated model that already knows the actual outcomes shows the same error rate, which the paper takes as evidence that the bias lives in the text, not in the models. Supplementing the context with Ukrainian military-analytical sources reduces the bias across all clean models, though absolute-error gains are partial and model-dependent.

Core claim

The paper's central claim is that the English-language information ecosystem about the war in Ukraine carries a measurable pro-capture bias, and that this bias propagates through any model that processes it. The evidence is a push-error rate: when the addition of English news context moves an LLM's probability toward Russian territorial capture, that movement is later contradicted by the outcome-anchored market reference in 64–72% of cases across all four models, with binomial p values below 10^-6. The contaminated-model control — a model whose training data includes the outcomes — shows the same failure rate, which the authors argue isolates the text corpus as the source of the distortion r

What carries the argument

The measuring instrument is an ablation ladder: the same model predicts under progressively richer information contexts — starting from bare market data, then adding a price chart, then English news articles, then the full English-language ecosystem, and finally augmented with Ukrainian military-analytical sources. The difference between conditions, expressed in percentage points relative to the prediction-market price trajectory, is the 'framing cost' of each text source. The decisive diagnostic is the push-error rate: when a context shift moves a prediction toward capture, does the later price path confirm it? The use of a contaminated model that already knows the outcomes serves as a cont

Load-bearing premise

The load-bearing premise is that the probability an LLM outputs is a direct, faithful measure of the belief the input text induces — an assumption adopted from prior work but not validated against an external belief measure such as human judgment.

What would settle it

A direct calibration experiment: have a panel of financially incentivized human forecasters read the same English news blocks and give probability updates under the same conditions; if human updates do not reproduce the 64–72% wrong-push rate relative to the market reference, then LLM output probabilities are not faithful readouts of text-induced beliefs, and the measured framing cost would be an artifact of the model rather than a property of the information ecosystem.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • English-language LLM forecasting of territorial conflicts will systematically overstate the attacking side's success unless the information diet is broadened.
  • Adding sources from the affected side (e.g., its military-analytical ecosystem) can reduce the directional bias, but the benefit varies by model and is not guaranteed to improve absolute error.
  • The bias is a property of the corpus, so any downstream system that consumes the same English news text will inherit it.
  • Model selection becomes a strategic lever: conservative reasoning may be preferable to deep reasoning when the text itself carries a directional bias.
  • The method offers a general, probability-calibrated measure of 'framing cost' that can be applied to other conflicts or contested topics.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One testable extension would be to apply the same push-error metric directly to news outlets, ranking them by how often their inclusion shifts model forecasts toward capture and then proves wrong; the paper's data hint that the most-cited analytical source correlates with worse directional accuracy.
  • The contaminated-model result does not separate two possible mechanisms — offense-dominant framing within the text versus exclusion of mitigating sources; a cleaner ablation would compare English-only context with Ukrainian-only context (without the English mix) to quantify the relative contribution of each.
  • If the instrument-fidelity assumption holds, the method could be used as a general 'belief calibration' audit for any text corpus on any question with a prediction market, not just war outcomes.
  • The paper's ethical discussion suggests the bias may influence human policy and public opinion; the method could theoretically be extended to measure human belief update rates on the same texts, though that goes beyond the current study.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a method for quantifying information-ecosystem bias in LLM predictions, using Polymarket price trajectories as an external calibration reference. On 111 Ukraine-related markets (~93,000 predictions), five information conditions (A: blind; B: +chart; C: +English news; D: full English context; DUA: D + Ukrainian military sources) are run on three clean models plus a contaminated model whose training data include realized outcomes. The main claims are that English-language context systematically shifts territorial predictions toward Russian capture, that such pro-capture pushes are wrong 64–72% of the time (binomial p<10^-6), that a contaminated model shows the same push-error rate and therefore the bias originates in the text, and that adding Ukrainian military-analytical sources reduces the directional bias while MAE gains are partial and model-dependent.

Significance. If the central claims hold, the paper offers a novel, externally anchored way to quantify the cost of framing in LLM world models, with practical implications for multilingual RAG and forecasting. The design has genuine strengths: market-level clustering for the bias-shift tests, Bonferroni correction, permutation tests, a diplomatic-market placebo, per-horizon analyses, a contaminated-model control, and an honest limitations section. The released dataset of predictions and reasoning traces is a useful resource. However, the headline push-accuracy result currently rests on an inappropriate null hypothesis, and the attribution to 'English news text' is clouded by the composition of condition D. These issues are load-bearing and require reanalysis before the main quantitative claim is accepted.

major comments (4)
  1. [§4.2 / Table 1] The headline 'wrong 64–72%' compares upward/pro-capture push accuracy against a 50% binomial null. The appropriate baseline is the unconditional rate of upward price moves in the same market/horizon subset. With 62/65 territorial markets resolving NO and Polymarket carrying a +3.5 pp pro-capture bias, the base rate of sign(p_horizon − p_current) > 0 may be well below 50%. If it is ~28–30%, the observed 27.9–36.3% accuracies are close to what any upward-predicting rule would produce. The binomial p<10^-6 only shows that upward pushes are often wrong, not that English text creates the error. The contaminated-model comparison does not fix this: a model that knows final outcomes can still have low short-horizon direction accuracy because 7-day price direction is not the same as final binary resolution. The paper should report (a) the base rate of upward price moves in the same instances, (b)
  2. [Appendix B / Table 1] The paper states that 'all tests use market-level aggregates with cluster-robust inference,' but the push-accuracy p-values are binomial tests across individual predictions, despite ~44% overlap between adjacent cutoffs and intra-cluster correlations up to 0.141. This violates the stated unit-of-independence assumption and makes the p<10^-6 values unreliable. These tests are also explicitly excluded from the Bonferroni family. Please report a market-clustered test (e.g., cluster bootstrap or market-level accuracy means) and either include push accuracy in the multiple-testing family or justify the exclusion.
  3. [§3.1 / §4.2] The main claim is that 'English news' systematically biases predictions, but the push-accuracy table appears to be for condition D, which bundles English news with a price chart, a war map, and Polymarket trader comments. The introduction defines C as A + English news blocks, yet no C push accuracy is reported in Table 1. The conclusion that 'the bias originates primarily in the text' therefore conflates the full English-language ecosystem with news text. If the claim is about news, report condition C; if it is about the full ecosystem, revise the abstract and discussion. The contaminated model is run only in conditions A–D, so the source attribution cannot separate text from the other components of D.
  4. [Section 1 / §4.2] The measurement instrument assumes that 'the output probability is the induced belief,' justified by two LLM theory papers. This is load-bearing: if output probabilities partly reflect prompt formatting, instruction-following, or reasoning heuristics rather than text-induced beliefs, the measured 'framing cost' is not what is claimed. The contaminated model rules out model ignorance but not instrument infidelity. Add a validation study — for example, compare LLM probabilities under C/D against a human belief-elicitation benchmark, or show that prompt/order perturbations do not change the measured push error rates.
minor comments (6)
  1. [Abstract / Table 1] The abstract reports 'wrong 64–72%' while Table 1 reports accuracies of 27.9–36.3%. Stating the accuracy and its complement explicitly would avoid confusion.
  2. [Appendix L, Table 9] The entry for 'advance into' reads '770∞'; this appears to be a formatting error for '77 / 0' and should be corrected.
  3. [§3.1] The bullet list has a typo: 'Aprovides' should be 'A provides'.
  4. [§3.2 / Appendix J] The contaminated model is called 'Gemini 3.1 Pro Preview' in the text and 'Pro 3.1*' in tables; define this abbreviation at first use.
  5. [Appendix G] 'Bias 2 accounts for only 2–9%' should use a proper superscript or spell out 'bias squared' for clarity.
  6. [Limitations] The limitations section notes that DUA vs D MAE improvement is not significant and that MAE effects are underpowered; this should be reflected in the abstract's claims about 'accuracy gains'.

Circularity Check

0 steps flagged

No significant circularity: the core measurements are anchored to external market trajectories and independent controls, with no self-citation chain or fit-renamed-as-prediction.

full rationale

The paper's derivation chain is not circular. The central quantities—bias shifts (A→D, D→DUA) and push accuracy—are computed from LLM outputs against Polymarket price trajectories and realized outcomes, which are external to the model and not fitted to the model's outputs. The contaminated-model control (Pro 3.1*) is a genuine out-of-sample manipulation: its outcome knowledge is inferred from behavior, and its same push-error rate is offered as a control, not as a definitional consequence. The formal framework in Appendix A uses latent λ weights and β shifts only to state a conditional mechanism; these parameters are not estimated from data and do not enter the empirical estimates, so Proposition 1 is a toy implication of the weighted-average definition rather than a fitted prediction. The citations for the core instrument assumption (Xie et al. 2022; von Oswald et al. 2023) are external, not self-citations, and are used as supporting theory rather than as the sole evidence for the empirical result. The main validity concern—that push accuracy is tested against a 50% binomial null rather than the base rate of upward price moves in mostly-NO territorial markets—is a statistical design issue, not a circularity, because it does not reduce the claimed result to its inputs by construction. Overall, the quantitative claims stand on external benchmarks and independent controls, so no circular step is identified.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The central claim rests on two domain assumptions: (1) LLM output probabilities faithfully reflect text-induced beliefs, and (2) prediction-market trajectories are a valid calibration reference. Both are plausible, cited, but not independently validated. No free parameters are fitted to produce the empirical results; the Appendix A λ weights are conceptual effective parameters. No invented entities.

axioms (5)
  • domain assumption LLM output probabilities are the induced belief from the input text (Section 1: 'The output probability is the induced belief.')
    The entire measurement instrument treats the model's probability as a faithful readout of text-induced belief, justified by citations (Xie et al. 2022; von Oswald et al. 2023) but not directly validated against an external belief measure.
  • domain assumption Prediction market prices are financially incentivized, eventually outcome-anchored probability estimates usable as an external calibration reference (Section 1, citing Wolfers & Zitzewitz)
    The bias metric compares model probabilities to market price trajectories, assuming the market reference is meaningful even though the paper shows Polymarket itself is biased (+3.5 pp) versus realized outcomes.
  • domain assumption The market price trajectory at each horizon (6h–7d) is a suitable continuous reference for model comparisons
    Section 3.3/4.1: continuous trajectory used as calibration reference; binary resolutions too sparse (3/65 YES).
  • domain assumption The GDELT subset and the DUA corpus are representative of the English and Ukrainian military-analytical information ecosystems respectively
    Section 3.1 and Appendix F: corpus composition is described, but representativeness is asserted, not measured.
  • domain assumption Polymarket resolutions for territorial markets are a valid realization of the underlying events (accepting ISW/DeepState expert assessments)
    Section 4.1 acknowledges resolution relies on ISW/DeepState frontline updates, expert assessments rather than direct observation.

reviewed 2026-08-02 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets." pith.science (2026). https://pith.science/paper/QT6M2A4J

@misc{pith2026260720441,
  author       = {Pith},
  title        = {Pith review of: Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QT6M2A4J}},
  note         = {Machine review of arXiv:2607.20441}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Every information ecosystem produces beliefs that shape strategic decisions. Both human analysts and AI systems inherit the blind spots of their information sources. We show that LLMs, combined with prediction markets, function as a calibrated instrument for measuring how far ecosystem-induced beliefs deviate from an external reference: LLMs extract the beliefs a text corpus implies, and prediction market price trajectories, anchored at resolution by realised outcomes, provide the calibration reference against which to quantify the deviation. We isolate the bias contribution of specific text through ablation: varying information context while holding the model fixed, with a contaminated model that knows actual outcomes as control. Applied to 111 Ukraine-related prediction markets, comprising approximately 93,000 predictions across four models, we find that English news context systematically biases territorial predictions, wrong 64 to 72 percent of the time when it pushes predictions toward territorial capture. A contaminated model that knows actual outcomes shows the same error rate, indicating that the bias originates primarily in the text. Supplementing with Ukrainian military-analytical sources reduces the bias for all clean models, while absolute-error gains are partial and model-dependent. We show that the distortion originates primarily in the sources, not the models. Consistent across four architectures, it will persist in any system that processes them and propagate into downstream decisions.

Figures

Figures reproduced from arXiv: 2607.20441 by Anton Polishko, Artur Kiulian, Dmytro Zamriy, Kostiantyn Kozlov, Mykola Khandoga, Yevhen Kostiuk, Yurii Filipchuk.

Figure 1
Figure 1. Figure 1: One evaluation instance. The model receives [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: MAE vs no-change baseline across all in [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Bias vs accuracy for each model under condi [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Left: pro-capture bias grows with prediction [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Bias decomposition. Flash and Pro share training data but differ 5× in context-induced bias. Ukrainian sources provide a similar absolute correc￾tion for all models, but it is swamped by Pro’s context damage. A B C D DUA Flash −5.3 −4.9 −1.6 +1.0 −1.8 Pro +3.5 +2.3 +5.9 +15.5 +12.3 GPT +12.9 +5.6 +16.9 +19.4 +15.5 3.1* −10.4 −8.2 −9.6 −9.4 – [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Per-market MAE under D (English) vs DUA (supplemented with Ukrainian sources). Each dot is one territorial market. Points below the diagonal indicate DUA outperforms D. DUA wins the majority of markets for Flash (37/65) and Pro (32/65). Territorial markets, 7-day horizon. Flash Pro GPT 10 5 0 5 10 15 20 MAE vs no-change (%) better worse -5.3% +3.5% +12.9% +1.0% +15.5% +19.4% -1.8% +12.3% Context hurts +15.… view at source ↗
Figure 7
Figure 7. Figure 7: Domain specificity: English context dam [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: MAE by time remaining to market resolution. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 1 linked inside Pith

  1. [1]

    Mohammed Al-Harbi and 1 others. 2026. An evaluation of LLMs for political bias in Western media: Israel-Hamas and Ukraine-Russia wars. arXiv preprint arXiv:2601.06132

  2. [2]

    Mohammad Ali and Naeemul Hassan. 2022. A survey of computational framing analysis approaches. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9335--9348

  3. [3]

    Rohan Alur, Bradly C Stadie, Daniel Kang, Ryan Chen, Matt McManus, Michael Rickert, Tyler Lee, Michael Federici, and 1 others. 2025. AIA forecaster: Technical report. arXiv preprint arXiv:2511.07678

  4. [4]

    Ramy Baly, Giovanni Da San Martino, James Glass, and Preslav Nakov. 2020. We can detect your bias: Predicting the political ideology of news articles. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 4982--4991

  5. [5]

    Dallas Card, Amber E Boydstun, Justin H Gross, Philip Resnik, and Noah A Smith. 2015. The media frames corpus: Annotations of frames across issues. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics, pages 438--444

  6. [6]

    Robert M Entman. 2004. Projections of Power: Framing News, Public Opinion, and U.S. Foreign Policy . University of Chicago Press

  7. [7]

    Shangbin Feng, Chan Young Park, Yohan Liu, and Yulia Tsvetkov. 2023. From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics

  8. [8]

    Isabel O Gallegos, Ryan A Rossi, Joe Barber, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Bui, Cheonbok Kim, Besmira Nushi, Duen Horng Yu, and 1 others. 2024. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3)

  9. [10]

    Ezra Karger, Houtan Bastani, Chen Yueh-Han, Zachary Jacobs, Danny Halawi, Fred Zhang, and Philip E Tetlock. 2025. ForecastBench : A dynamic benchmark of AI forecasting capabilities. In International Conference on Learning Representations

  10. [11]

    Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In International Conference on Learning Representations

  11. [12]

    Preslav Nakov, Jisun An, Haewoon Kwak, Muhammad Arslan Mansurov, and Momin Mansurov. 2024. A survey on predicting the factuality and the bias of news media. In Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics

  12. [13]

    Markus Ojala, Mervi Pantti, and Jenni Kangas. 2024. Framing the war in Ukraine : A comparative study of news coverage. Journalism Studies

  13. [14]

    Yulia Otmakhova, Shima Khanehzar, and Lea Frermann. 2024. Media framing: A typology and survey of computational approaches across disciplines. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

  14. [15]

    Grzegorz Ptaszek, Bogdan Yuskiv, and Serhii Khomych. 2024. War on frames: Text mining of conflict in Russian and Ukrainian news agency coverage on Telegram . Media, War & Conflict, 17(1):41--61

  15. [16]

    Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Jo \ a o Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2023. Transformers learn in-context by gradient descent. In International Conference on Machine Learning, pages 35151--35174

  16. [17]

    Justin Wolfers and Eric Zitzewitz. 2004. Prediction markets. Journal of Economic Perspectives, 18(2):107--126

  17. [18]

    Justin Wolfers and Eric Zitzewitz. 2006. Interpreting prediction market prices as probabilities. NBER Working Paper 12200, National Bureau of Economic Research

  18. [19]

    Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit Bayesian inference. In International Conference on Learning Representations

  19. [20]

    Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

    A Survey of Computational Framing Analysis Approaches , author=. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

  20. [21]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year=

    Media Framing: A Typology and Survey of Computational Approaches Across Disciplines , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year=

  21. [22]

    Science , volume=

    The Promise of Prediction Markets , author=. Science , volume=

  22. [23]

    Journal of Communication , volume=

    Convergent News? A Longitudinal Study of Similarity and Dissimilarity in the Domestic and Global News Agendas , author=. Journal of Communication , volume=

  23. [24]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

    Unsupervised Cross-lingual Representation Learning at Scale , author=. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

  24. [25]

    Projections of Power: Framing News, Public Opinion, and

    Entman, Robert M , year=. Projections of Power: Framing News, Public Opinion, and

  25. [26]

    From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair

    Feng, Shangbin and Park, Chan Young and Liu, Yohan and Tsvetkov, Yulia , booktitle=. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair

  26. [27]

    Computational Linguistics , volume=

    Bias and Fairness in Large Language Models: A Survey , author=. Computational Linguistics , volume=

  27. [28]

    arXiv preprint arXiv:2402.18563 , year=

    Approaching Human-Level Forecasting with Language Models , author=. arXiv preprint arXiv:2402.18563 , year=

  28. [29]

    Framing the War in

    Ojala, Markus and Pantti, Mervi and Kangas, Jenni , journal=. Framing the War in

  29. [30]

    How Multilingual is Multilingual

    Pires, Telmo and Schlinger, Eva and Garrette, Dan , booktitle=. How Multilingual is Multilingual

  30. [31]

    Ukrainian

    Romanyshyn, Mariana , booktitle=. Ukrainian

  31. [32]

    Schoenegger, Philipp and Park, Sami and Karger, Ezra and Tetlock, Philip E , journal=

  32. [33]

    Journal of Political Economy , volume=

    Explaining the Favorite--Long Shot Bias: Is it Risk-Love or Misperceptions? , author=. Journal of Political Economy , volume=

  33. [34]

    Journal of Economic Perspectives , volume=

    Prediction Markets , author=. Journal of Economic Perspectives , volume=

  34. [35]

    2006 , number=

    Interpreting Prediction Market Prices as Probabilities , author=. 2006 , number=

  35. [36]

    Findings of the Association for Computational Linguistics: ACL 2024 , year=

    A Survey on Predicting the Factuality and the Bias of News Media , author=. Findings of the Association for Computational Linguistics: ACL 2024 , year=

  36. [37]

    Advances in Neural Information Processing Systems , year=

    Forecasting Future World Events with Neural Networks , author=. Advances in Neural Information Processing Systems , year=

  37. [38]

    International Conference on Learning Representations , year=

    Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , author=. International Conference on Learning Representations , year=

  38. [39]

    An Evaluation of

    Al-Harbi, Mohammed and others , journal=. An Evaluation of

  39. [40]

    Alur, Rohan and Stadie, Bradly C and Kang, Daniel and Chen, Ryan and McManus, Matt and Rickert, Michael and Lee, Tyler and Federici, Michael and others , journal=

  40. [41]

    Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages=

    We Can Detect Your Bias: Predicting the Political Ideology of News Articles , author=. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , pages=

  41. [42]

    Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics , pages=

    The Media Frames Corpus: Annotations of Frames Across Issues , author=. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics , pages=

  42. [43]

    Introducing

    Chaplynskyi, Dmytro , booktitle=. Introducing

  43. [44]

    Framing and Agenda-Setting in

    Field, Anjalie and Kliger, Doron and Wintner, Shuly and Pan, Jennifer and Jurafsky, Dan and Tsvetkov, Yulia , booktitle=. Framing and Agenda-Setting in

  44. [45]

    A Contemporary News Corpus of

    Fischer, Stefan and Haidarzhyi, Kateryna and Knappen, J. A Contemporary News Corpus of. Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) , year=

  45. [46]

    Karger, Ezra and Bastani, Houtan and Yueh-Han, Chen and Jacobs, Zachary and Halawi, Danny and Zhang, Fred and Tetlock, Philip E , booktitle=

  46. [47]

    From Bytes to Borsch: Fine-tuning

    Kiulian, Artur and Polishko, Anton and Khandoga, Mykola and Chubych, Oryna and Connor, Jack and Ravishankar, Raghav and Shirawalmath, Adarsh , booktitle=. From Bytes to Borsch: Fine-tuning

  47. [48]

    Piskorski, Jakub and Stefanovitch, Nicolas and Da San Martino, Giovanni and Nakov, Preslav , booktitle=

  48. [49]

    War on Frames: Text Mining of Conflict in

    Ptaszek, Grzegorz and Yuskiv, Bogdan and Khomych, Serhii , journal=. War on Frames: Text Mining of Conflict in

  49. [50]

    An Explanation of In-context Learning as Implicit

    Xie, Sang Michael and Raghunathan, Aditi and Liang, Percy and Ma, Tengyu , booktitle=. An Explanation of In-context Learning as Implicit

  50. [51]

    International Conference on Machine Learning , pages=

    Transformers Learn In-Context by Gradient Descent , author=. International Conference on Machine Learning , pages=

  51. [52]

    Romanyshyn, Mariana and Syvokon, Oleksiy and Kyslyi, Roman , booktitle=. The

This paper was first reviewed by deepseek-v4-flash on August 2, 2026.