REVIEW 3 major objections 2 minor 109 references
Spatial navigation in preclinical Alzheimer's disease: A review
T0 review · 3 major / 2 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Spatial navigation tasks can flag Alzheimer's risk years before memory symptoms, tracking early pathology biomarkers in people who still test as cognitively normal.
desk verdict Useful narrative synthesis on navigation as a preclinical AD marker, but I only have the abstract—and the attached full text is the wrong paper—so treat claims about p-tau correlations and screening utility as still unchecked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The alignment of spatial navigation computations (path integration and wayfinding) with the neural circuits that are the earliest sites of AD pathology—used to explain why navigation measures can detect risk before episodic memory decline.
What would settle it
A prospective study in cognitively unimpaired biomarker-positive people showing that baseline path-integration or wayfinding scores do not predict later cognitive decline or conversion once age, vascular burden, education, and non-AD pathology are controlled.
Extended reading notes
Core claim
In cognitively unimpaired individuals with AD biomarkers, performance on spatial navigation tasks—particularly path integration and wayfinding—correlates with plasma and CSF markers of AD pathology (notably p-tau), making navigation assessment a sensitive candidate for preclinical risk detection.
Load-bearing premise
That the observed links between navigation scores and AD biomarkers in still-unimpaired people truly reflect early AD circuit damage, not age, vascular disease, education, task quirks, or other pathology—and that those links will predict future clinical decline.
Editorial extensions
If this is right
- Path integration and wayfinding tests could be added to preclinical screening batteries alongside or ahead of standard memory tests.
- Plasma or CSF p-tau levels may be interpretable together with navigation scores as a joint early-risk signal.
- Scalable navigation assessments could help select asymptomatic at-risk participants for prevention trials.
- Interventions timed to navigation decline might slow progression before clinically significant impairment appears.
- Future work should prioritize longitudinal designs that test whether navigation change predicts clinical outcomes.
Reading between the lines
- If navigation truly tracks earliest circuit failure, digital or VR path-integration tasks could become remote, low-cost screening tools outside specialty clinics.
- Dissociating path integration from wayfinding may help separate entorhinal-centered from broader network contributions to early AD risk.
- Combining navigation metrics with plasma p-tau could tighten enrichment of prevention trials beyond biomarkers alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is submitted as a narrative review arguing that spatial navigation—especially path integration and wayfinding—is a particularly sensitive cognitive marker in preclinical Alzheimer’s disease. The abstract claims that navigation depends on circuits that are among the earliest sites of AD pathology, that performance on such tasks correlates with plasma and CSF AD biomarkers (notably p-tau) in cognitively unimpaired biomarker-positive individuals, and that navigation assessment could therefore provide a sensitive, scalable approach for early risk detection and for informing future interventions. The abstract further contrasts this putative early sensitivity with episodic memory decline, which is said to appear only after more substantial medial temporal damage.
Significance. If the synthesis is accurate and the cited associations are robust after appropriate confound control and prospective validation, the review would be clinically and scientifically useful: it would consolidate a circuit-to-cognition rationale for navigation as a preclinical marker, highlight task classes (path integration, wayfinding) that may outperform standard memory tests at the asymptomatic stage, and motivate scalable digital or behavioral screening. That contribution would be of clear interest to cognitive neuroscience and AD biomarker communities. However, the significance of the present submission cannot be assessed from the materials provided, because the body text supplied under this arXiv ID does not correspond to the titled review.
major comments (3)
- The full manuscript text attached under paper_id 2603.23082 / title “Spatial navigation in preclinical Alzheimer’s disease: A review” is not that review. The body is an unrelated methods paper on MedCausalX (adaptive causal reasoning for medical vision–language models; arXiv 2603.23085), with its own abstract, CRMed dataset, reflective tokens, DPO/GRPO pipeline, and medical VQA tables. None of the promised content—overview of spatial navigation computations and tasks, mapping to earliest AD pathology sites, or synthesis of cognitively unimpaired biomarker-positive cohorts—is present. Peer review of the stated central claim is therefore impossible on the supplied file.
- Even taking only the abstract of the intended review as the claim set, the load-bearing assertion that path integration and wayfinding performance “correlates with plasma and CSF biomarkers of AD pathology, notably p-tau” in cognitively unimpaired at-risk individuals, and that this “can represent a sensitive and scalable approach for early detection,” cannot be checked: there is no methods section defining inclusion criteria, no table of primary studies with effect sizes or quality appraisal, no discussion of confounds (age, vascular burden, education, task design, non-AD pathology), and no longitudinal outcome data supporting prospective risk prediction. Those elements are required for a review whose clinical recommendation rests on that correlation.
- The abstract’s mechanistic premise—that navigation relies on neural circuits corresponding to the earliest sites of AD pathology and is therefore more sensitive than episodic memory—is a central organizing claim. Without the actual review body (circuit mapping, staging of pathology, and head-to-head comparison with memory measures in the same biomarker-defined cohorts), this premise remains an unexamined assertion rather than a documented synthesis.
minor comments (2)
- Abstract wording is generally clear, but “individuals at-risk of AD” should be defined more precisely (e.g., amyloid/tau biomarker criteria, genetic risk, or both) once the correct full text is available.
- The abstract asserts clinical utility (“will inform future interventions”) without qualifying the current evidence level (cross-sectional association vs. prospective prediction). Softening that language until longitudinal data are reviewed would improve accuracy.
Circularity Check
Narrative review of external navigation–biomarker correlations; no derivation-by-construction or load-bearing circular step.
full rationale
Paper 2603.23082 is a narrative review whose strongest claim is that path-integration and wayfinding performance correlates with plasma/CSF AD biomarkers (notably p-tau) in cognitively unimpaired at-risk individuals, and may aid preclinical detection. That claim is framed as a synthesis of external empirical studies, not as a first-principles derivation, fitted parameter renamed as prediction, or uniqueness theorem imported from the authors. There is no equation chain in which navigation scores are defined from the same biomarkers they are said to predict, nor any self-definitional loop (X defined via Y then used to derive Y). Residual risks for a review of this type—selective citation, cross-sectional confounds, lack of new longitudinal outcome data—are correctness/generalization concerns, not circularity under the enumerated patterns. The supplied full-text block is an unrelated MedCausalX VLM manuscript and cannot be used to invent circular steps for 2603.23082. Honest finding: no significant circularity; score 0; steps empty.
Assumptions & free parameters
assumptions (4)
- domain assumption AD has a prolonged preclinical phase in which neuropathology accumulates before clinical cognitive symptoms.
- domain assumption Spatial navigation depends on neural circuits that are among the earliest sites of AD pathology (e.g., medial temporal / entorhinal-related systems).
- domain assumption Plasma and CSF p-tau (and related AD biomarkers) validly index AD pathology in cognitively unimpaired individuals.
- domain assumption Episodic memory decline typically appears only after substantial medial temporal damage, making it less sensitive preclinically than navigation.
Cite this review
Pith. "Pith review of Spatial navigation in preclinical Alzheimer's disease: A review." pith.science (2026). https://pith.science/paper/VYMPWEAT
@misc{pith2026260323082,
author = {Pith},
title = {Pith review of: Spatial navigation in preclinical Alzheimer's disease: A review},
year = {2026},
howpublished = {\url{https://pith.science/paper/VYMPWEAT}},
note = {Machine review of arXiv:2603.23082}
}
read the original abstract
Alzheimer's disease (AD) develops over a prolonged preclinical phase, during which neuropathological changes accumulate long before cognitive symptoms appear. Identifying cognitive functions affected at early stages is critical for the preclinical detection of asymptomatic individuals at-risk of AD. Early risk identification could enable timely interventions aimed at mitigating the development of significant future cognitive impairment. While episodic memory decline typically appears after substantial medial temporal lobe damage, spatial navigation has emerged as a particularly sensitive cognitive function in preclinical AD. In this review, we provide an overview of spatial navigation computations and the tasks used to assess them, highlighting how spatial navigation relies on neural circuits corresponding to the earliest sites of AD pathology. We synthesize evidence from cognitively unimpaired individuals with AD biomarkers, i.e. individuals at-risk of AD, and discuss future research directions. Overall, performance on spatial navigation tasks, particularly path integration and wayfinding, correlates with plasma and CSF biomarkers of AD pathology, notably p-tau. Spatial navigation assessment can represent a sensitive and scalable approach for early detection of individuals at-risk of AD in preclinical stages, and will inform future interventions to mitigate the progression toward clinically significant cognitive impairment.
Reference graph
Works this paper leans on
-
[1]
Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023. 1
arXiv 2023
-
[2]
Shuai Bai and et al. Qwen2.5-vl: Technical report. https : / / arxiv . org / abs / 2502 . 13923, 2025. arXiv preprint arXiv:2502.13923. 5, 6, 7, 9
arXiv 2025
-
[3]
Constitutional ai: Harmlessness from ai feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, and et al. Constitutional ai: Harmlessness from ai feedback, 2022. 3
2022
-
[4]
Coun- terfactual causal-effect intervention for interpretable medi- cal visual question answering.IEEE Transactions on Med- ical Imaging, 2024
Linqin Cai, Haodu Fang, Nuoying Xu, and Bo Ren. Coun- terfactual causal-effect intervention for interpretable medi- cal visual question answering.IEEE Transactions on Med- ical Imaging, 2024. 3
2024
-
[5]
Causality matters in medical imaging.Nature Communications, 11 (1):3673, 2020
Daniel C Castro, Ian Walker, and Ben Glocker. Causality matters in medical imaging.Nature Communications, 11 (1):3673, 2020. 2
2020
-
[6]
360+x: A panoptic multi- modal scene understanding dataset
Hao Chen, Yuqi Hou, Chenyuan Qu, Irene Testini, Xiao- han Hong, and Jianbo Jiao. 360+x: A panoptic multi- modal scene understanding dataset. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19373–19382, 2024. 3
2024
-
[7]
Junying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao, Shunian Chen, Guiming Hardy Chen, Xidong Wang, Ruifei Zhang, Zhenyang Cai, Ke Ji, et al. Huatuogpt-vision, to- wards injecting medical visual knowledge into multimodal llms at scale.arXiv preprint arXiv:2406.19280, 2024. 3
arXiv 2024
-
[8]
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Sch ¨arli, and Denny Zhou. Teaching large language models to self-debug. In ICLR, 2024. 2
2024
Show all 109 references
-
[9]
Bridging radiology and pathology foundation models via concept- based multimodal co-adaptation
Yihang Chen, Yanyan Huang, Fuying Wang, Maximus Ye- ung, Yuming Jiang, Shujun Wang, and Lequan Yu. Bridging radiology and pathology foundation models via concept- based multimodal co-adaptation. InThe Fourteenth Inter- national Conference on Learning Representations. 1
-
[10]
Expanding performance boundaries of open-source multimodal models with model, data, and test- time scaling.arXiv preprint arXiv:2412.05271, 2024
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhang- wei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test- time scaling.arXiv preprint arXiv:2412.05271, 2024. 5...
2024 arXiv
-
[11]
Internvl: Scaling up vision founda- tion models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision founda- tion models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE/CVF conference on computer vi...
2024
-
[12]
Deep reinforcement learn- ing from human preferences.Advances in neural informa- tion processing systems, 30, 2017
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learn- ing from human preferences.Advances in neural informa- tion processing systems, 30, 2017. 3
2017
-
[13]
Tinymig: Trans- ferring generalization from vision foundation models to single-domain medical imaging
LIU Chuang, Hongyan Xu, Yichao Cao, Xiu Su, Zhe Qu, Tianfa Li, Shan An, and Haogang Zhu. Tinymig: Trans- ferring generalization from vision foundation models to single-domain medical imaging. InForty-second Interna- tional Conference on Machine Learning, 2025. 1
2025
-
[14]
Gaussiandwm: 3d gaussian driving world model for unified scene understanding and multi-modal generation.arXiv preprint arXiv:2512.23180,
Tianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen, Yuyao Xu, Lijin Yang, Le Xu, Yu Zhang, Bo Zhang, Wuxiong Huang, and Hesheng Wang. Gaussiandwm: 3d gaussian driving world model for unified scene understanding and multi-modal generation.arXiv preprint arXiv:2512.23180,
-
[15]
What is the best 3d scene representation for robotics? from geometric to foundation models.arXiv preprint arXiv:2512.03422, 2025
Tianchen Deng, Yue Pan, Shenghai Yuan, Dong Li, Chen Wang, Mingrui Li, Long Chen, Lihua Xie, Danwei Wang, Jingchuan Wang, Javier Civera, Hesheng Wang, and Wei- dong Chen. What is the best 3d scene representation for robotics? from geometric to foundation models.arXiv preprint ...
2025
-
[16]
En- hancing large vision language models with self-training on image comprehension.Advances in Neural Information Processing Systems, 37:131369–131397, 2024
Yihe Deng, Pan Lu, Fan Yin, Ziniu Hu, Sheng Shen, Quan- quan Gu, James Zou, Kai-Wei Chang, and Wei Wang. En- hancing large vision language models with self-training on image comprehension.Advances in Neural Information Processing Systems, 37:131369–131397, 2024. 1
2024
-
[17]
Prpo: Aligning process reward with outcome reward in policy optimization
Ruiyi Ding, Yongxuan Lv, Xianhui Meng, Jiahe Song, Chao Wang, Chen Jiang, and Yuan Cheng. Prpo: Aligning process reward with outcome reward in policy optimization. arXiv preprint arXiv:2601.07182, 2026. 3
2026
-
[18]
Neurosymb-mrg: Differentiable abductive reasoning with active uncertainty minimization for radiology report generation.arXiv preprint arXiv:2603.01756, 2026
Rong Fu, Yiqing Lyu, Chunlei Meng, Muge Qi, Yabin Jin, Qi Zhao, Li Bao, Juntao Gao, Fuqian Shi, Nilanjan Dey, et al. Neurosymb-mrg: Differentiable abductive reasoning with active uncertainty minimization for radiology report generation.arXiv preprint arXiv:2603.01756, 2026. 3
2026 arXiv
-
[19]
Shortcut learning in deep neural net- works.Nature Machine Intelligence, 2(11):665–673, 2020
Robert Geirhos, J ¨orn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Fe- lix A Wichmann. Shortcut learning in deep neural net- works.Nature Machine Intelligence, 2(11):665–673, 2020. 1, 2
2020
-
[20]
Med-cmr: A fine-grained benchmark integrating visual evidence and clinical logic for medical complex multimodal reasoning.arXiv preprint arXiv:2512.00818, 2025
Haozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu, Xiaoxiao Yan, Jingjing Liu, Kai Wu, Jiazhen Pan, Bailiang Jian, Jiangning Zhang, et al. Med-cmr: A fine-grained benchmark integrating visual evidence and clinical logic for medical complex multimodal reasoning.arXiv preprint ...
2025
-
[21]
Deepseek-r1: Incentivizing reason- ing capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reason- ing capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025. 3
2025 arXiv
-
[22]
Meddr: Diagnosis-guided bootstrapping for large-scale medical vision-language learning.arXiv preprint arXiv:2404.15127, 2024
Sunan He, Yuxiang Nie, Zhixuan Chen, Zhiyuan Cai, Hongmei Wang, Shu Yang, and Hao Chen. Meddr: Diagnosis-guided bootstrapping for large-scale medical vision-language learning.arXiv preprint arXiv:2404.15127, 2024. 5, 6, 7
2024 arXiv
-
[23]
Pathological visual ques- tion answering.arXiv preprint arXiv:2010.12435, 2020
Xuehai He, Zhuo Cai, Wenlan Wei, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathological visual ques- tion answering.arXiv preprint arXiv:2010.12435, 2020. 5, 6, 9
2010 arXiv
-
[24]
Codev: Code with images for faithful visual reasoning via tool-aware policy optimization.arXiv preprint arXiv:2511.19661, 2025
Xinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li, Jia Liu, Todd C Hollon, and Bryan Wang. Codev: Code with images for faithful visual reasoning via tool-aware policy optimization.arXiv preprint arXiv:2511.19661, 2025. 2, 3
2025
-
[25]
Advancing medical imaging with language mod- els: A journey from n-grams to chatgpt.arXiv preprint arXiv:2304.04920, 2023
Mingzhe Hu, Shaoyan Pan, Yuheng Li, and Xiaofeng Yang. Advancing medical imaging with language mod- els: A journey from n-grams to chatgpt.arXiv preprint arXiv:2304.04920, 2023. 3
2023 arXiv
-
[26]
Large language models cannot self-correct reasoning yet
Jie Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. Large language models cannot self-correct reasoning yet. In Proceedings of the International Conference on Learn- ing Representations (ICLR) 2024, 2024. Paper ID: 8b4add8b0aa8749d80a34ca5...
2024
-
[27]
Medreflect: Teaching medical llms to self-improve via reflective correc- tion.arXiv preprint arXiv:2510.03687, 2025
Yue Huang, Yanyuan Chen, Dexuan Xu, Weihua Yue, Huamin Zhang, Meikang Qiu, and Yu Huang. Medreflect: Teaching medical llms to self-improve via reflective correc- tion.arXiv preprint arXiv:2510.03687, 2025. 1, 2
2025
-
[28]
Gpt-4o system card.arXiv preprint arXiv:2410.21276, 2024
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card.arXiv preprint arXiv:2410.21276, 2024. 5, 12
2024 arXiv
-
[29]
Med-moe: Mixture of domain-specific experts for lightweight medical vision-language models
Songtao Jiang, Tuo Zheng, Yan Zhang, Yeying Jin, Li Yuan, and Zuozhu Liu. Med-moe: Mixture of domain-specific experts for lightweight medical vision-language models. In Findings of the association for computational linguistics: EMNLP 2024, pages 3843–3860, 2024. 3
2024
-
[30]
Hulu-med: A transparent general- ist model towards holistic medical vision-language under- standing.arXiv preprint arXiv:2510.08668, 2025
Songtao Jiang, Yuan Wang, Sibo Song, Tianxiang Hu, Chenyi Zhou, Bin Pu, Yan Zhang, Zhibo Yang, Yang Feng, Joey Tianyi Zhou, et al. Hulu-med: A transparent general- ist model towards holistic medical vision-language under- standing.arXiv preprint arXiv:2510.08668, 2025. 3
2025
-
[31]
Why adam can beat sgd: Second-moment normalization yields sharper tails.arXiv preprint arXiv:2603.03099, 2026
Ruinan Jin, Yingbin Liang, and Shaofeng Zou. Why adam can beat sgd: Second-moment normalization yields sharper tails.arXiv preprint arXiv:2603.03099, 2026. 6
2026 arXiv
-
[32]
Tic-grpo: Provable and efficient optimiza- tion for reinforcement learning from human feedback
Ruinan Jin et al. Tic-grpo: Provable and efficient optimiza- tion for reinforcement learning from human feedback. 3
-
[33]
Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019
Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019. 5, 9
2019
-
[34]
Large language models are zero-shot reasoners.Advances in Neural Information Pro- cessing Systems, 35:22199–22213, 2022
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners.Advances in Neural Information Pro- cessing Systems, 35:22199–22213, 2022. 2
2022
-
[35]
Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McK- inney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, and Aleksandra Faust. ...
2024
-
[36]
Med-r1: Re- inforcement learning for generalizable medical reasoning in vision-language models.IEEE Transactions on Medical Imaging, 2026
Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, Yuheng Li, Konstantinos Psounis, and Xiaofeng Yang. Med-r1: Re- inforcement learning for generalizable medical reasoning in vision-language models.IEEE Transactions on Medical Imaging, 2026. 1, 2, 3, 5, 6, 11, 12
2026
-
[37]
A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018. 5, 6, 9
2018
-
[38]
Exploring efficient open- vocabulary segmentation in the remote sensing.arXiv preprint arXiv:2509.12040, 2025
Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, and Xuelong Li. Exploring efficient open- vocabulary segmentation in the remote sensing.arXiv preprint arXiv:2509.12040, 2025. 3
2025
-
[39]
Maris: Marine open-vocabulary in- stance segmentation with geometric enhancement and se- mantic alignment.arXiv preprint arXiv:2510.15398, 2025
Bingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao, Junyu Gao, and Xuelong Li. Maris: Marine open-vocabulary in- stance segmentation with geometric enhancement and se- mantic alignment.arXiv preprint arXiv:2510.15398, 2025
2025
-
[40]
Stitchfusion: Weaving any visual modalities to en- hance multimodal semantic segmentation
Bingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao, and Xue- long Li. Stitchfusion: Weaving any visual modalities to en- hance multimodal semantic segmentation. InProceedings of the 33rd ACM International Conference on Multimedia, pages 1308–1317, 2025. 3
2025
-
[41]
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36: 28541–2...
2023
-
[42]
The choice of divergence: A neglected key to mitigating diversity collapse in reinforcement learning with verifiable reward.arXiv preprint arXiv:2509.07430,
Long Li, Jiaran Hao, Jason Klein Liu, Zhijian Zhou, Yant- ing Miao, Wei Pang, Xiaoyu Tan, Wei Chu, Zhe Wang, Shirui Pan, et al. The choice of divergence: A neglected key to mitigating diversity collapse in reinforcement learning with verifiable reward.arXiv preprint arXiv:2509.07430,
-
[43]
M²IV: Towards efficient and fine-grained multimodal in-context learning via representation engineering
Yanshu Li, Yi Cao, Hongyang He, Qisen Cheng, Xiang Fu, Xi Xiao, Tianyang Wang, and Ruixiang Tang. M²IV: Towards efficient and fine-grained multimodal in-context learning via representation engineering. InSecond Con- ference on Language Modeling, 2025. 3
2025
-
[44]
Mmt-ard: Multimodal multi-teacher adversarial dis- tillation for robust vision-language models.arXiv preprint arXiv:2511.17448, 2025
Yuqi Li, Junhao Dong, Chuanguang Yang, Shiping Wen, Piotr Koniusz, Tingwen Huang, Yingli Tian, and Yew-Soon Ong. Mmt-ard: Multimodal multi-teacher adversarial dis- tillation for robust vision-language models.arXiv preprint arXiv:2511.17448, 2025. 3
2025
-
[45]
Catp: Contextually adap- tive token pruning for efficient and enhanced multimodal in-context learning.arXiv preprint arXiv:2508.07871,
Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, and Ruixiang Tang. Catp: Contextually adap- tive token pruning for efficient and enhanced multimodal in-context learning.arXiv preprint arXiv:2508.07871,
-
[46]
Taco: Enhancing multimodal in-context learning via task mapping-guided sequence con- figuration
Yanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng, Jinfa Huang, and Ruixiang Tang. Taco: Enhancing multimodal in-context learning via task mapping-guided sequence con- figuration. InProceedings of the 2025 Conference on Em- pirical Methods in Natural Language Processing, pages...
2025
-
[47]
Efficient Medical Image Segmentation via Reinforcement Learning-Driven K-Space Sampling.IEEE Transactions on Emerging Topics in Com- putational Intelligence, 2025
Yuqi Li, Hansheng Zeng, Fuyan Zhang, Chuanguang Yang, Yanli Li, and Weiping Ding. Efficient Medical Image Segmentation via Reinforcement Learning-Driven K-Space Sampling.IEEE Transactions on Emerging Topics in Com- putational Intelligence, 2025. 3
2025
-
[48]
Manxi Lin, Nina Weng, Kamil Mikolaj, Zahra Bashir, Morten B. S. Svendsen, Martin G. Tolsgaard, Anders Ny- mark Christensen, and Aasa Feragen. Shortcut learning in medical image segmentation. InMedical Image Computing and Computer Assisted Intervention – MICCAI 2024, pages 623–...
2024
-
[49]
Surgical post-training: Cutting er- rors, keeping knowledge.arXiv preprint arXiv:2603.01683,
Wenye Lin and Kai Han. Surgical post-training: Cutting er- rors, keeping knowledge.arXiv preprint arXiv:2603.01683,
-
[50]
Gamebot: Transparent assessment of llm reasoning in games
Wenye Lin, Jonathan Roberts, Yunhan Yang, Samuel Al- banie, Zongqing Lu, and Kai Han. Gamebot: Transparent assessment of llm reasoning in games. InProceedings of the 63rd Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), pages 7656– 768...
2025
-
[51]
Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1650–1654. IEEE, 2021. 5, 6, 9
2021
-
[52]
Perturbating, tuning, and collaborating: Har- nessing vision foundation models for single domain gen- eralization on medical imaging
Chuang Liu, Yichao Cao, YingYing Zhang, Xiu Su, and Haogang Zhu. Perturbating, tuning, and collaborating: Har- nessing vision foundation models for single domain gen- eralization on medical imaging. InProceedings of the AAAI Conference on Artificial Intelligence, pages 5370– 5...
2025
-
[53]
Cdrrm: Contrast-driven rubric gener- ation for reliable and interpretable reward modeling.arXiv preprint arXiv:2603.08035, 2026
Dengcan Liu, Fengkai Yang, Xiaohan Wang, Shurui Yan, Jiajun Chai, Jiahao Li, Yikun Ban, Zhendong Mao, Wei Lin, and Guojun Yin. Cdrrm: Contrast-driven rubric gener- ation for reliable and interpretable reward modeling.arXiv preprint arXiv:2603.08035, 2026. 3
2026
-
[54]
On the intrinsic self-correction capability of llms: Uncertainty and latent concept.arXiv preprint arXiv:2406.02378, 2024
Guangliang Liu, Haitao Mao, Bochuan Cao, Zhiyu Xue, Xitong Zhang, Rongrong Wang, Jiliang Tang, and Kris- ten Johnson. On the intrinsic self-correction capability of llms: Uncertainty and latent concept.arXiv preprint arXiv:2406.02378, 2024. 1, 2
2024 arXiv
-
[55]
Improved baselines with visual instruc- tion tuning.arXiv preprint arXiv:2310.03744, 2024
Haotian Liu et al. Improved baselines with visual instruc- tion tuning.arXiv preprint arXiv:2310.03744, 2024. 11
2024 arXiv
-
[56]
Medcot: Medical chain of thought via hierarchical expert
Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Tianyi Zhou, and Zuozhu Liu. Medcot: Medical chain of thought via hierarchical expert. InProceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Process- ing (EMNLP 2024), pages 17371–17389, 2024. 1, 3, 6
2024
-
[57]
Automated opti- mization modeling via a localizable error-driven perspec- tive.arXiv preprint arXiv:2602.11164, 2026
Weiting Liu, Han Wu, Yufei Kuang, Xiongwei Han, Tao Zhong, Jianfeng Feng, and Wenlian Lu. Automated opti- mization modeling via a localizable error-driven perspec- tive.arXiv preprint arXiv:2602.11164, 2026. 3
2026
-
[58]
Scanext: Enhancing 3d medical image segmentation with dual attention network and depth-wise convolution.He- liyon, 10(5), 2024
Yajun Liu, Zenghui Zhang, Jiang Yue, and Weiwei Guo. Scanext: Enhancing 3d medical image segmentation with dual attention network and depth-wise convolution.He- liyon, 10(5), 2024. 3
2024
-
[59]
M 3 hl: Mutual mask mix with high-low level feature consistency for semi-supervised medical im- age segmentation
Yajun Liu, Zenghui Zhang, Jiang Yue, Weiwei Guo, and Dongying Li. M 3 hl: Mutual mask mix with high-low level feature consistency for semi-supervised medical im- age segmentation. InInternational Conference on Medi- cal Image Computing and Computer-Assisted Intervention, pages...
2025
-
[60]
Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017. 6
2017 arXiv
-
[61]
Gastric-x: A multimodal multi-phase bench- mark dataset for advancing vision-language models in gas- tric cancer analysis, 2026
Sheng Lu, Hao Chen, Rui Yin, Juyan Ba, Yu Zhang, and Yuanzhe Li. Gastric-x: A multimodal multi-phase bench- mark dataset for advancing vision-language models in gas- tric cancer analysis, 2026. 3
2026
-
[62]
Thinking with blueprints: Assisting vision-language models in spatial reasoning via structured object representation, 2026
Weijian Ma, Shizhao Sun, Tianyu Yu, Ruiyu Wang, Tat- Seng Chua, and Jiang Bian. Thinking with blueprints: Assisting vision-language models in spatial reasoning via structured object representation, 2026. 2
2026
-
[63]
Self- refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hal- linan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self- refine: Itera...
2023
-
[64]
Med-flamingo: a multimodal medical few-shot learner
Michael Moor, Qian Huang, Shirley Wu, Michihiro Ya- sunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Ed- uardo Pontes Reis, and Pranav Rajpurkar. Med-flamingo: a multimodal medical few-shot learner. InMachine Learn- ing for Health (ML4H), pages 353–367. PMLR, 2023. 5, 6, 7
2023
-
[65]
Self-taught self-correction for small language mod- els.arXiv preprint arXiv:2503.08681, 2025
Viktor Moskvoretskii, Chris Biemann, and Irina Nik- ishina. Self-taught self-correction for small language mod- els.arXiv preprint arXiv:2503.08681, 2025. 2
2025 arXiv
-
[66]
Training lan- guage models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Car- roll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training lan- guage models to follow instructions with human feedback. Advances in neural information processing systems, 35:...
2022
-
[67]
Medvlm-r1: Incentivizing medical reasoning capability of vision-language models (vlms) via reinforce- ment learning
Jiazhen Pan, Che Liu, Junde Wu, Fenglin Liu, Jiayuan Zhu, Hongwei Bran Li, Chen Chen, Cheng Ouyang, and Daniel Rueckert. Medvlm-r1: Incentivizing medical reasoning capability of vision-language models (vlms) via reinforce- ment learning. InInternational Conference on Medical I...
2025
-
[68]
Toward understanding why adam converges faster than sgd for transformers.arXiv preprint arXiv:2306.00204, 2023
Yan Pan and Yuanzhi Li. Toward understanding why adam converges faster than sgd for transformers.arXiv preprint arXiv:2306.00204, 2023. 6
2023 arXiv
-
[69]
On the theory and practice of grpo: A trajectory-corrected approach with fast conver- gence.arXiv preprint arXiv:2508.02833, 2025
Lei Pang and Ruinan Jin. On the theory and practice of grpo: A trajectory-corrected approach with fast conver- gence.arXiv preprint arXiv:2508.02833, 2025. 3
2025
-
[70]
Elements of causal inference: foundations and learning al- gorithms
Jonas Peters, Dominik Janzing, and Bernhard Scholkopf. Elements of causal inference: foundations and learning al- gorithms. MIT press, 2017. 2
2017
-
[71]
Causal inference and counterfac- tual prediction in machine learning for actionable health- care.Nature Machine Intelligence, 2(7):369–375, 2020
Mattia Prosperi, Yi Guo, Matt Sperrin, James S Koopman, Jae S Min, Xing He, Shannan Rich, Mo Wang, Iain E Buchan, and Jiang Bian. Causal inference and counterfac- tual prediction in machine learning for actionable health- care.Nature Machine Intelligence, 2(7):369–375, 2020. 2
2020
-
[72]
Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christo- pher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023. 3, 9
2023
-
[73]
Multimodal generative ai for medical image in- terpretation.Nature, 639(8056):888–896, 2025
Vishwanatha M Rao, Michael Hla, Michael Moor, Subathra Adithan, Stephen Kwak, Eric J Topol, and Pranav Ra- jpurkar. Multimodal generative ai for medical image in- terpretation.Nature, 639(8056):888–896, 2025. 1
2025
-
[74]
Im- proving the accuracy of medical diagnosis with causal ma- chine learning.Nature communications, 11(1):3923, 2020
Jonathan G Richens, Ciar ´an M Lee, and Saurabh Johri. Im- proving the accuracy of medical diagnosis with causal ma- chine learning.Nature communications, 11(1):3923, 2020. 2
2020
-
[75]
Proximal policy optimization algo- rithms.arXiv preprint arXiv:1707.06347, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Rad- ford, and Oleg Klimov. Proximal policy optimization algo- rithms.arXiv preprint arXiv:1707.06347, 2017. 3
2017 arXiv
-
[76]
Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junx- iao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024. 2, 3, 9
2024 arXiv
-
[77]
Reflexion: Language agents with verbal reinforcement learning.Advances in Neural In- formation Processing Systems, 36:8634–8652, 2023
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning.Advances in Neural In- formation Processing Systems, 36:8634–8652, 2023. 2
2023
-
[78]
Hume: Introducing system-2 thinking in visual- language-action model.arXiv preprint arXiv:2505.21432,
Haoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen, Qi Lv, Yiwen Tang, Modi Shi, Guanghui Ren, Maoqing Yao, Bin Zhao, et al. Hume: Introducing system-2 thinking in visual- language-action model.arXiv preprint arXiv:2505.21432,
-
[79]
Bottom-up policy optimization: Your language model policy secretly contains internal policies.arXiv preprint arXiv:2512.19673, 2025
Yuqiao Tan, Minzheng Wang, Shizhu He, Huanxuan Liao, Chengfeng Zhao, Qiunan Lu, Tian Liang, Jun Zhao, and Kang Liu. Bottom-up policy optimization: Your language model policy secretly contains internal policies.arXiv preprint arXiv:2512.19673, 2025. 3
2025 arXiv
-
[80]
Collaboration between clinicians and vision–language models in radiology report generation.Nature Medicine, 31 (2):599–608, 2025
Ryutaro Tanno, David GT Barrett, Andrew Sellergren, Sumedh Ghaisas, Sumanth Dathathri, Abigail See, Jo- hannes Welbl, Charles Lau, Tao Tu, Shekoofeh Azizi, et al. Collaboration between clinicians and vision–language models in radiology report generation.Nature Medicine, 31 (2)...
2025
-
[81]
Causal reasoning in medical imaging
Athanasios Vlontzos, Christian M ¨uller, and Bernhard Kainz. Causal reasoning in medical imaging. InTrustwor- thy AI in Medical Imaging, pages 367–381. Elsevier, 2025. 2
2025
-
[82]
Interpretable bilingual multimodal large language model for diverse biomedical tasks.arXiv preprint arXiv:2410.18387, 2024
Lehan Wang, Haonan Wang, Honglong Yang, Jiaji Mao, Zehong Yang, Jun Shen, and Xiaomeng Li. Interpretable bilingual multimodal large language model for diverse biomedical tasks.arXiv preprint arXiv:2410.18387, 2024. 5, 6, 7, 11
2024 arXiv
-
[83]
Chain-of-thought reason- ing without prompting.arXiv preprint arXiv:2402.10200,
Xuezhi Wang and Denny Zhou. Chain-of-thought reason- ing without prompting.arXiv preprint arXiv:2402.10200,
-
[84]
Self-consistency improves chain of thought reason- ing in language models.arXiv preprint arXiv:2203.11171,
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reason- ing in language models.arXiv preprint arXiv:2203.11171,
-
[85]
Care what fails: Contrastive anchored-reflection for verifiable multimodal.arXiv preprint arXiv:2512.19554,
Yongxin Wang, Zhicheng Yang, Meng Cao, Mingfei Han, Haokun Lin, Yingying Zhu, Xiaojun Chang, and Xiaodan Liang. Care what fails: Contrastive anchored-reflection for verifiable multimodal.arXiv preprint arXiv:2512.19554,
-
[86]
Spatialclip: Learning 3d- aware image representations from spatially discriminative language
Zehan Wang, Sashuai Zhou, Shaoxuan He, Haifeng Huang, Lihe Yang, Ziang Zhang, Xize Cheng, Shengpeng Ji, Tao Jin, Hengshuang Zhao, et al. Spatialclip: Learning 3d- aware image representations from spatially discriminative language. InProceedings of the Computer Vision and Pat- ...
2025
-
[87]
Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022. 2
2022
-
[88]
Kicgpt: Large language model with knowledge in context for knowledge graph completion
Yanbin Wei, Qiushi Huang, Yu Zhang, and James Kwok. Kicgpt: Large language model with knowledge in context for knowledge graph completion. InFindings of the asso- ciation for computational linguistics: EMNLP 2023, pages 8667–8683, 2023. 2
2023
-
[89]
Dynamicgtr: Leveraging graph topology representation preferences to boost vlm ca- pabilities on graph qas.arXiv preprint arXiv:2602.21864,
Yanbin Wei, Jiangyue Yan, Chun Kang, Yang Chen, Hua Liu, James Kwok, and Yu Zhang. Dynamicgtr: Leveraging graph topology representation preferences to boost vlm ca- pabilities on graph qas.arXiv preprint arXiv:2602.21864,
-
[90]
Towards generalist foundation model for radiology.arXiv preprint arXiv:2308.02463, 2023
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Towards generalist foundation model for radiology.arXiv preprint arXiv:2308.02463, 2023. 5, 6
2023 arXiv
-
[91]
Yuechen Xie, Jie Song, Huiqiong Wang, and Mingli Song. Training data provenance verification: Did your model use synthetic data from my generative model for training? In Proceedings of the Computer Vision and Pattern Recogni- tion Conference, pages 23817–23827, 2025. 1
2025
-
[92]
Spatialqa: A benchmark for evaluating spatial log- ical reasoning in vision-language models.arXiv preprint arXiv:2602.20901, 2026
Yuechen Xie, Xiaoyan Zhang, Yicheng Shan, Hao Zhu, Rui Tang, Rong Wei, Mingli Song, Yuanyu Wan, and Jie Song. Spatialqa: A benchmark for evaluating spatial log- ical reasoning in vision-language models.arXiv preprint arXiv:2602.20901, 2026. 2
2026
-
[93]
Structure causal models and llms integration in medi- cal visual question answering.IEEE Transactions on Med- ical Imaging, 2025
Zibo Xu, Qiang Li, Weizhi Nie, Weijie Wang, and Anan Liu. Structure causal models and llms integration in medi- cal visual question answering.IEEE Transactions on Med- ical Imaging, 2025. 3
2025
-
[94]
Your group-relative advantage is biased.arXiv preprint arXiv:2601.08521, 2026
Fengkai Yang, Zherui Chen, Xiaohan Wang, Xiaodong Lu, Jiajun Chai, Guojun Yin, Wei Lin, Shuai Ma, Fuzhen Zhuang, Deqing Wang, et al. Your group-relative advantage is biased.arXiv preprint arXiv:2601.08521, 2026. 3
2026
-
[95]
Tooltree: Efficient llm tool planning via dual- feedback monte carlo tree search and bidirectional pruning
Shuo Yang, Caren Han, Yihao Ding, Shuhe Wang, and Ed- uard Hovy. Tooltree: Efficient llm tool planning via dual- feedback monte carlo tree search and bidirectional pruning. InThe Fourteenth International Conference on Learning Representations. 3
-
[96]
Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026
Shuo Yang, Soyeon Caren Han, Xueqi Ma, Yan Li, Moham- mad Reza Ghasemi Madani, and Eduard Hovy. Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026. 3
2026
-
[97]
Visual grounding with multi-modal conditional adaptation
Ruilin Yao, Shengwu Xiong, Yichen Zhao, and Yi Rong. Visual grounding with multi-modal conditional adaptation. InProceedings of the 32nd ACM International Conference on Multimedia, pages 3877–3886, 2024. 3
2024
-
[98]
Sa-med2d-20m dataset: Segment anything in 2d medical imaging with 20 million masks
Jin Ye, Junlong Cheng, Jianpin Chen, Zhongying Deng, Tianbin Li, Haoyu Wang, Yanzhou Su, Ziyan Huang, Jilong Chen, Lei Jiang, et al. Sa-med2d-20m dataset: Segment anything in 2d medical imaging with 20 million masks. arXiv preprint arXiv:2311.11969, 2023. 5, 9
2023 arXiv
-
[99]
Futuresight- drive: Thinking visually with spatio-temporal cot for au- tonomous driving.arXiv preprint arXiv:2505.17685, 2025
Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, and Xing Wei. Futuresight- drive: Thinking visually with spatio-temporal cot for au- tonomous driving.arXiv preprint arXiv:2505.17685, 2025. 3
2025 arXiv
-
[100]
Janusvln: Decoupling semantics and spatiality with dual implicit memory for vision-language navigation.arXiv preprint arXiv:2509.22548, 2025
Shuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong, Shichao Xie, Xiaolong Wu, Shiyi Liang, Mu Xu, and Xing Wei. Janusvln: Decoupling semantics and spatiality with dual implicit memory for vision-language navigation.arXiv preprint arXiv:2509.22548, 2025. 3
2025
-
[101]
A generalist vision–language foundation model for diverse biomedical tasks.Nature Medicine, 30(11):3129–3141, 2024
Kai Zhang, Rong Zhou, Eashan Adhikarla, Zhiling Yan, Yixin Liu, Jun Yu, Zhengliang Liu, Xun Chen, Brian D Davison, Hui Ren, et al. A generalist vision–language foundation model for diverse biomedical tasks.Nature Medicine, 30(11):3129–3141, 2024. 5
2024
-
[102]
Improve vision language model chain-of-thought reasoning.arXiv preprint arXiv:2410.16198, 2024
Ruohong Zhang et al. Improve vision language model chain-of-thought reasoning.arXiv preprint arXiv:2410.16198, 2024. 2
2024 arXiv
-
[103]
Pmc-vqa: Visual instruction tuning for medical visual question answering
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415, 2023. 6
2023 arXiv
-
[104]
Non-intrusive graph-based bot detection for e- commerce using inductive graph neural networks.arXiv preprint arXiv:2601.22579, 2026
Sichen Zhao, Zhiming Xue, Yalun Qi, Xianling Zeng, and Zihan Yu. Non-intrusive graph-based bot detection for e- commerce using inductive graph neural networks.arXiv preprint arXiv:2601.22579, 2026. 2
2026
-
[105]
Explainable differential diagnosis with dual- inference large language models.npj Health Systems, 2(1): 12, 2025
Shuang Zhou, Mingquan Lin, Sirui Ding, Jiashuo Wang, Canyu Chen, Genevieve B Melton, James Zou, and Rui Zhang. Explainable differential diagnosis with dual- inference large language models.npj Health Systems, 2(1): 12, 2025. 4
2025
-
[106]
Unified thinker: A general reasoning modular core for image generation, 2026
Sashuai Zhou, Qiang Zhou, Jijin Hu, Hanqing Yang, Yue Cao, Junpeng Ma, Yinchao Ma, Jun Song, Tiezheng Ge, Cheng Yu, Bo Zheng, and Zhou Zhao. Unified thinker: A general reasoning modular core for image generation, 2026. 3
2026
-
[107]
Drop- ping experts, recombining neurons: Retraining-free prun- ing for sparse mixture-of-experts llms
Yixiao Zhou, Ziyu Zhao, Dongzhou Cheng, Zhiliang Wu, Jie Gui, Yi Yang, Fei Wu, Yu Cheng, and Hehe Fan. Drop- ping experts, recombining neurons: Retraining-free prun- ing for sparse mixture-of-experts llms. InFindings of the Association for Computational Linguistics: EMNLP 2025...
2025
-
[108]
Look inward to explore outward: Learning tem- perature policy from llm internal states via hierarchical rl
Yixiao Zhou, Yang Li, Dongzhou Cheng, Hehe Fan, and Yu Cheng. Look inward to explore outward: Learning tem- perature policy from llm internal states via hierarchical rl. arXiv preprint arXiv:2602.13035, 2026. 3
2026
-
[109]
Fine-tuning language models from human preferences.arXiv preprint arXiv:1909.08593, 2019
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences.arXiv preprint arXiv:1909.08593, 2019. 3
1909 arXiv
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.