REVIEW 3 major objections 5 minor 47 references
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read State-of-the-art LLMs, evaluated on six real-world healthcare prediction tasks under three learning frameworks, often barely outperform random guessing and consistently predict less favorable outcomes for African American patients.
desk verdict Solid multi-task fairness benchmark in low-resource healthcare, but the headline 'persistent disparities' claim needs confidence intervals before it can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a benchmark suite of six healthcare tasks built from four datasets, paired with the standard fairness metrics Demographic Parity Difference (DPD) and Equal Opportunity Difference (EOD). DPD measures how much more or less often a group receives the favorable prediction compared with everyone else; EOD measures the same gap restricted to patients whose true outcome is favorable. The evaluation deliberately balances classes and demographic attributes so that each subgroup contributes roughly equally, then applies the same three frameworks—chain-of-thought in-context learning, LoRA fine-tuning, and a web-searching agent—across the tasks. This setup is what lets the authors attribute differences in accuracy and fairness to the models and frameworks rather than to skewed test distributions.
What would settle it
Bootstrap the reported DPD and EOD estimates from the same test sets: if the 95% confidence interval for the African-American-versus-White gap on, say, the 500-example MIMIC mortality test contains zero, then the paper's central claim of persistent racial disparity on that task would not survive.
Extended reading notes
Core claim
On its own terms, the paper claims that the general-domain success of large language models does not transfer to low-resource healthcare classification. In the authors' six benchmarks, implementations such as zero-shot and few-shot chain-of-thought, LoRA fine-tuning, and a ReAct-style agent frequently hover near the random-guess baseline, with the weakest cases in readmission and the two schizophrenia/bipolar diagnosis tasks. Fairness measurements using Demographic Parity Difference and Equal Opportunity Difference show a consistent direction: relative to White patients, African American patients receive favorable predictions less often and are correctly classified as favorable less often in most tasks, with racial gaps larger than gender gaps. The paper further claims that explicitly inserting demographic attributes into prompts does not reliably improve accuracy or fairness, and that the LLM-as-agent approach can retrieve current guidelines yet still reach wrong conclusions by misapplying them. Finally, the authors show that models can infer patients' race from dialogue and that the reasoning used is often stereotyped, which they take as evidence of a hidden pathway for biased predictions in conversational health tasks.
Load-bearing premise
The conclusions rest on the assumption that the small, demographically balanced test sets (60 to 500 examples per task) give stable estimates of how often each group receives favorable predictions and true positives, so sampling noise could change the size or even the sign of some reported disparities.
Editorial extensions
If this is right
- Claimed few-shot competence of LLMs on generic classification does not extend to these real-world healthcare tasks; accuracy can sit at or near random guessing even with chain-of-thought examples.
- Racial disparity, not gender disparity, is the dominant fairness pattern: African American patients receive less favorable predictions and lower true positive rates in most tasks and frameworks.
- Explicitly providing demographic information is not a dependable fairness intervention; its effects on accuracy, DPD, and EOD vary by task, model, and metric.
- Giving an LLM web access to current guidelines does not guarantee correct health predictions; retrieved facts can be relevant while the model's reasoning from them remains wrong.
- Race inference from conversational text is feasible and biased, so dialogue-based health tasks carry a risk of hidden demographic bias even when no demographics are supplied.
Reading between the lines
- The reported fairness gaps are point estimates from test sets of 60 to 500 examples; a bootstrap or confidence-interval analysis could reveal that several gaps are statistically indistinguishable from zero, so the 'persistent' pattern should be treated as provisional until quantified with uncertainty.
- A natural testable extension is to rerun the mortality and readmission evaluations on naturalistic (unbalanced) patient distributions; the balanced sampling used here may understate or overstate the real-world disparity depending on how base rates interact with model bias.
- The agent's failure mode—retrieving correct guidelines but misapplying them to spoken dialogue—suggests a concrete intervention: require the model to quote the specific guideline clause it used and justify its application, which could be evaluated as a direct follow-up.
- Because models infer race from dialect features like 'ain't' and 'gonna', a targeted stress test could use dialogues with matched content but systematically varied dialect markers to measure how much of the diagnosis disparity is driven by linguistic stereotyping.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates GPT-4, Claude-3, and LLaMA-3 across six healthcare tasks (MIMIC mortality, MIMIC readmission, health coaching goal completion, schizophrenia/bipolar neighbor and landlord scene classification, and MedQA) using three frameworks: in-context learning with chain-of-thought, LoRA fine-tuning, and an LLM-as-agent retrieval pipeline. Performance is measured by global accuracy and by two fairness metrics, Demographic Parity Difference (DPD) and Equal Opportunity Difference (EOD), reported for race and gender subgroups. The central claims are that LLMs struggle with real-world healthcare classification (some implementations barely exceeding random guessing), that they show persistent demographic disparities (especially less favorable predictions and lower equality of opportunity for African American patients), that explicit demographic prompting gives mixed results, and that access to up-to-date guidelines via agents does not guarantee accurate predictions.
Significance. If the quantitative claims are reliable, this is a useful and timely cautionary result for the deployment of LLMs in healthcare. The paper compares three model families and three learning paradigms on six tasks, including realistic dialogue-based and clinical-note tasks, which is broader than many single-benchmark evaluations. The qualitative examples and the sociolinguistic consultation on race inference are valuable and give concrete evidence of stereotyped reasoning. However, the headline fairness conclusions rest entirely on point estimates from small test sets with no confidence intervals or significance tests, and some accuracy comparisons are statistically indistinguishable from random baselines at the reported sample sizes. The paper's strength is therefore the breadth and plausibility of the patterns, not yet the statistical support for the word 'persistent.'
major comments (3)
- [Section 6.2, Table 3] The central fairness claim that LLMs 'consistently predict less favorable outcomes for African American patients' and show 'persistent' disparities is not supported by the reported statistics alone, because all DPD and EOD values are point estimates without confidence intervals, bootstrap estimates, or significance tests. With the test sizes in Table 1, the standard errors are large: for a balanced MIMIC split of 500 (roughly 250 per group), the SE of a DPD is about 4.5 percentage points, so a value of -8.2 is within a 95% interval that includes zero; EOD, being conditioned on Y=1, has a still larger SE. For Health Coaching (n=60), the subgroup sizes are about 20, giving a DPD SE of roughly 13.7 points, so the reported -4.2 and -12.5 values are statistically indistinguishable from zero. For Neighbor and Landlord (n=261), the SE is about 6.6 points, so most reported DPDs are within noise. Please report per-group counts, standard errors or bootstrap confidence intervals, and significance tests for the DPD/EOD estimates, or explicitly weaken the abstract and Section 6.2 wording from 'persistent' and 'consistently' to directional patterns. A sign test across the 13 rows would be a minimal additional check.
- [Section 6.2, Table 2 and Section 4] The accuracy claims that some implementations 'barely surpass random guessing' also need uncertainty quantification. For example, the ICL GPT-4 readmission accuracy of 55.3 on n=500 has an approximate 95% confidence interval of [50.9, 59.7] under the usual normal approximation, which includes the 50% random baseline; the same applies to several other values in the 52-55% range on MIMIC tasks. In addition, Section 4 says 'We report the best performance between Zero-Shot and N-Shot in this setting,' but the paper does not state which variant produced each number or how many shots were used, and selecting the best of several prompted configurations without multiple-comparison adjustment tends to inflate apparent performance. Please report the exact prompt settings per task and provide confidence intervals or error bars for the accuracy values that are near the random baseline.
- [Section 3 and Table 1] The paper states that each dataset was sampled so that classes, demographic attributes, and P(C=c|Z=z) are 'roughly balanced,' but it does not report the actual demographic counts in the test sets, which are necessary both for interpreting the fairness metrics and for computing their precision. Without exact subgroup sizes, a reader cannot verify the standard-error calculations or know whether the 'roughly balanced' condition holds for the small Health Coaching and Schizophrenia/Bipolar datasets. Please include a supplementary table with the per-subgroup test counts for each task and framework, or at least for each task, and use those counts in the uncertainty analysis requested above.
minor comments (5)
- [Table 2 caption] The MedQA column reports ICL results only inside parentheses (with explicit demographics), with the outside cell left as '-'. It would be clearer to state explicitly that MedQA was evaluated only with the demographic-prompt variant, and to give the sample size and the fact that these are percentages from n=175.
- [Section 5 and Table 3 caption] The text defines 'Demographic Parity Difference (DPD)' but the Table 3 caption and several places in Section 6.2 use 'PDP' instead of 'DPD.' Please make the acronym consistent throughout.
- [Section 6.3, Table 6] The race-inference results for Health Coaching (LLaMA-3 40.0 vs random 33.3) and the Claude-3 refusals are reported without any uncertainty or note on sample size. Given n=60, the 40.0% value is within one standard error of the random baseline; please add a caveat or a confidence interval.
- [Section 4, ICL baselines] The phrase 'four to eight-shot in-context examples' is vague; please specify the exact number of shots used for each task, and whether the same examples were used for all demographic-prompt conditions.
- [Limitations] The Limitations section does not mention the small-sample precision issue that affects the main fairness and accuracy comparisons. Adding a sentence acknowledging that the DPD/EOD estimates are noisy and that the conclusions are preliminary would improve the paper's self-assessment.
Circularity Check
No significant circularity: the paper is an empirical evaluation whose fairness and accuracy claims rest on external benchmarks and standard metrics, not on a derivation that reduces to its own inputs.
full rationale
This paper is an empirical evaluation rather than a derivation, so there is no equation-level chain in which a prediction is equivalent to an input by construction. The fairness metrics DPD and EOD are standard, externally defined quantities applied to model outputs and ground-truth labels; the reported disparities are computed from held-out test examples, not fitted from the same statistic that is later called a prediction. The health coaching dataset does originate in the authors' prior work (Gupta et al. 2020; Zhou et al. 2024b), and one in-context-learning reference is by a co-author (Zhou et al. 2024c), but these citations are data provenance and related-work context, not load-bearing justifications for the central claim that LLMs underperform and exhibit disparities. The accuracy and fairness conclusions are anchored in multiple external datasets (MIMIC, MedQA, and the Schizophrenia/Bipolar Interview corpus) and in direct model outputs. Concerns about small test-set sizes and the absence of confidence intervals are legitimate statistical precision issues, but they are not circularity: they do not show that the reported numbers were derived from the conclusions or that the metrics are self-definitional. The limitations section acknowledges the need for mitigation strategies and further study, but it does not reveal a circular step. Accordingly, no circular step is identified.
Assumptions & free parameters
free parameters (6)
- Inference temperature =
0.3
- LoRA rank =
8
- LoRA alpha =
8
- Learning rate =
1e-5
- Batch size =
8
- Dropout rate =
0.1
assumptions (4)
- domain assumption Demographic group labels in the datasets are accurate and complete.
- domain assumption The sampled test sets are representative of the clinical populations for each task.
- standard math Random guess baselines assume balanced classes across the one-vs-all comparisons.
- domain assumption The transcribed dialogues contain sufficient linguistic cues for valid diagnosis by either clinicians or models.
Cite this review
Pith. "Pith review of Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective." pith.science (2026). https://pith.science/paper/7YBRQWW5
@misc{pith2026241200554,
author = {Pith},
title = {Pith review of: Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/7YBRQWW5}},
note = {Machine review of arXiv:2412.00554}
}
read the original abstract
This paper studies the performance of large language models (LLMs), particularly regarding demographic fairness, in solving real-world healthcare tasks. We evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLMs to real-world healthcare tasks and persistent fairness issues across demographic groups. We also find that explicitly providing demographic information yields mixed results, while LLM's ability to infer such details raises concerns about biased health predictions. Utilizing LLMs as autonomous agents with access to up-to-date guidelines does not guarantee performance improvement. We believe these findings reveal the critical limitations of LLMs in healthcare fairness and the urgent need for specialized research in this area.
Figures
Reference graph
Works this paper leans on
-
[1]
Ankit Aich, Avery Quynh, Varsha Badal, Amy Pinkham, Philip Harvey, Colin Depp, and Natalie Parde. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.208 Towards intelligent clinically-informed language analyses of people with bipolar disorder and schizophrenia . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 2871--2887,...
-
[2]
AI@Meta. 2024. https://github.com/meta-llama/llama3?tab=readme-ov-file Llama 3 model card
2024
-
[3]
Anthropic . 2024. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf The claude 3 model family: Opus, sonnet, haiku
2024
-
[4]
Paula Braveman. 2006. Health disparities and health equity: concepts and measurement. Annual Review of Public Health, 27(1):167--194
work page 2006
-
[5]
Jan Clusmann, Fiona R Kolbinger, Hannah Sophie Muti, Zunamys I Carrero, Jan-Niklas Eckardt, Narmin Ghaffari Laleh, Chiara Maria Lavinia L \"o ffler, Sophie-Caroline Schwarzkopf, Michaela Unger, Gregory P Veldhuizen, et al. 2023. The future landscape of large language models in medicine. Communications medicine, 3(1):141
2023
-
[6]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. https://arxiv.org/abs/2305.14314 Qlora: Efficient finetuning of quantized llms . Preprint, arXiv:2305.14314
arXiv 2023
-
[7]
George T Doran. 1981. There's a SMART way to write management’s goals and objectives. Management review, 70(11):35--36
work page 1981
-
[8]
Carol Friedman, George Hripcsak, et al. 1999. Natural language processing and its future in medicine. Acad Med, 74(8):890--5
work page 1999
Show all 47 references
-
[9]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2023. https://arxiv.org/abs/2309.00770 Bias and fairness in large language models: A survey . Preprint, arXiv:2309.00770
2023 arXiv
-
[10]
Itika Gupta, Barbara Di Eugenio, Brian Ziebart, Aiswarya Baiju, Bing Liu, Ben Gerber, Lisa Sharp, Nadia Nabulsi, and Mary Smart. 2020. https://aclanthology.org/2020.sigdial-1.30 Human-human health coaching via text messages: Corpus, annotation, and analysis . In Proceedings of...
2020
-
[11]
Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria. 2024. https://arxiv.org/abs/2310.05694 A survey of large language models for healthcare: from data, technology, and applications to accountability and ethics . Preprint, arXiv:2310.05694
2024 arXiv
-
[12]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. https://arxiv.org/abs/2106.09685 Lora: Low-rank adaptation of large language models . Preprint, arXiv:2106.09685
2021 arXiv
-
[13]
Yan Hu, Qingyu Chen, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, et al. 2024. Improving large language models for clinical named entity recognition via prompt engineering. Journal of the American Medical Informa...
2024
-
[14]
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. arXiv preprint arXiv:2009.13081
2020 arXiv
-
[15]
Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. 2020. Mimic-iv. PhysioNet. Available online at: https://physionet. org/content/mimiciv/1.0/(accessed August 23, 2021)
2020
-
[16]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023. https://arxiv.org/abs/2205.11916 Large language models are zero-shot reasoners . Preprint, arXiv:2205.11916
2023 arXiv
-
[17]
William Labov. 1969. A study of non-standard english
1969
-
[18]
Haylee Lane, Mitchell Sarkies, Jennifer Martin, and Terry Haines. 2017. Equity in healthcare resource allocation decision making: a systematic review. Social science & medicine, 175:11--27
2017
-
[19]
Thomas A. LaVeist. 2005. Minority populations and health: An introduction to health disparities in the United States, volume 4. John Wiley & Sons
2005
-
[20]
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, et al. 2023. https://arxiv.org/abs/2211.09110 Holistic evaluation of language models . Preprint, arXiv:2211.09110
2023 arXiv
-
[21]
Yanchen Liu, Srishti Gautam, Jiaqi Ma, and Himabindu Lakkaraju. 2023. Investigating the fairness of large language models for predictions on tabular data. arXiv preprint arXiv:2310.14607
2023 arXiv
-
[22]
Masoud Monajatipoor, Jiaxin Yang, Joel Stremmel, Melika Emami, Fazlolah Mohaghegh, Mozhdeh Rouhsedaghat, and Kai-Wei Chang. 2024. https://arxiv.org/abs/2404.07376 Llms in biomedicine: A study on clinical named entity recognition . Preprint, arXiv:2404.07376
2024 arXiv
-
[23]
Nambi Ndugga and Samantha Artiga. 2021. Disparities in health and health care: 5 key questions and answers. Kaiser Family Foundation, 11
2021
-
[24]
OpenAI. 2023. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774
2023 arXiv
-
[25]
Joao Pereira. 1993. What does equity in health mean? Journal of Social Policy, 22(1):19--48
1993
-
[26]
Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, et al
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, et al. 2022. https://arxiv.org/abs/2112.11446 Scaling language models: Methods, analysis and insights from training gopher . Preprint, arXiv:2112.11446
2022 arXiv
-
[27]
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. https://doi.org/10.1162/tacl_a_00434 Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp . Transactions of the Association for Computational Linguistics, 9:1408--1424
2021 doi
-
[28]
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. https://arxiv.org/abs/2303.11366 Reflexion: Language agents with verbal reinforcement learning . Preprint, arXiv:2303.11366
2023 arXiv
-
[29]
Edward Shortliffe. 1976. Computer-based medical consultations: MYCIN. Elsevier
1976
-
[30]
Sara Mahdavi, Joelle Barral, Dale Webster, Greg S
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y ...
2023 arXiv
-
[31]
Brown, et al
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, et al. 2023. https://openreview.net/forum?id=uyTL5Bvosj Beyond the imitation game: Quantifying and extrapolating the capabilities of language models . Transactions on...
2023
-
[32]
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, ...
2024 arXiv
-
[33]
Shubo Tian, Qiao Jin, Lana Yeganova, Po-Ting Lai, Qingqing Zhu, Xiuying Chen, Yifan Yang, Qingyu Chen, Won Kim, Donald C Comeau, Rezarta Islamaj, Aadit Kapoor, Xin Gao, and Zhiyong Lu. 2023. https://doi.org/10.1093/bib/bbad493 Opportunities and challenges for chatgpt and large...
2023 doi
-
[34]
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. 2023. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. arXiv preprint arXiv:2306.11698
2023 arXiv
-
[35]
Guanchu Wang, Junhao Ran, Ruixiang Tang, Chia-Yuan Chang, Chia-Yuan Chang, Yu-Neng Chuang, Zirui Liu, Vladimir Braverman, Zhandong Liu, and Xia Hu. 2024 a . https://arxiv.org/abs/2408.08422 Assessing and enhancing large language models in rare disease question-answering . Prep...
2024
-
[36]
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024 b . https://doi.org/10.1007/s11704-024-40231-1 A survey on large language model based autonomous agents . Fron...
2024 doi
-
[37]
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024 c . A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345
2024
-
[38]
Hugh R Waters. 2000. Measuring equity in access to health care. Social science & medicine, 51(4):599--612
2000
-
[39]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824--24837
2022
-
[40]
Laura Weidinger, John F. J. Mellor, M. Rauh, C. Griffin, J. Uesato, Po-Sen Huang, M. Cheng, Mia Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William S. Isaac, Sean Legassick...
2021 arXiv
-
[41]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. https://arxiv.org/abs/2210.03629 React: Synergizing reasoning and acting in language models . Preprint, arXiv:2210.03629
2023 arXiv
-
[42]
Rich Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork. 2013. https://proceedings.mlr.press/v28/zemel13.html Learning fair representations . In Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Researc...
2013
-
[43]
Chen, Peilin Zhou, Junling Liu, Yining Hua, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, and David A
Hongjian Zhou, Fenglin Liu, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S. Chen, Peilin Zhou, Junling Liu, Yining Hua, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, and David A. Clifton. 2024 a . https://arxiv.org/abs/2311.05...
2024 arXiv
-
[44]
Yue Zhou, Barbara Di Eugenio, Brian Ziebart, Lisa Sharp, Bing Liu, and Nikolaos Agadakos. 2024 b . https://aclanthology.org/2024.lrec-main.1005 Modeling low-resource health coaching dialogues via neuro-symbolic goal summarization and text-units-text generation . In Proceedings...
2024
-
[45]
Yue Zhou, Yada Zhu, Diego Antognini, Yoon Kim, and Yang Zhang. 2024 c . https://doi.org/10.18653/v1/2024.naacl-long.153 Paraphrase and solve: Exploring and exploiting the impact of surface form on mathematical reasoning in large language models . In Proceedings of the 2024 Con...
2024 doi
-
[46]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[47]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.