REVIEW 4 major objections 4 minor 205 references
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read General-purpose multimodal models fall short in specialized fields, so this survey organizes the domain-specific benchmarks that measure that gap into a single eight-discipline taxonomy.
desk verdict A workmanlike survey that will be handy as a pointer resource, but its central coverage claim is not checkable as written and several table entries are off-scope. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the eight-branch domain hierarchy (Figure 1): disciplines → domains → sub-domains → application areas. Each branch is paired with a consolidation table that records the benchmark's scale, task type, input modality, models evaluated, and performance. That hierarchy carries the survey's argument: it converts scattered benchmark papers into a structured picture of where the MLLM evaluation ecosystem is dense, where it is thin, and where model failure is systematic.
What would settle it
A systematic replication that runs the paper's own search terms against the same databases and screens for domain-specific MLLM benchmarks could settle the coverage claim: finding a substantial discipline or an established benchmark family that fits none of the eight taxonomy branches would refute comprehensiveness. A cheaper check is row-level: every benchmark named in the domain tables should be traceable to a paper that actually introduces or evaluates it, and any miscategorized row would weaken the resource's reliability.
Extended reading notes
Core claim
The central claim is that the field has reached the point where general benchmarks no longer tell the full story: MLLMs need domain-specific benchmarks to expose and guide their specialized capabilities. The paper's positive contribution is a taxonomy, announced as seven disciplines in the abstract but presented as eight in the methodology and Figure 1, with per-domain tables consolidating each benchmark's scale, task type, input modality, models evaluated, and reported performance. Across the domains the assembled evidence shows a recurring pattern: frontier models handle high-level reasoning and familiar text well, but stumble on fine-grained perception, specialized data formats, and domain-specific reasoning—low accuracy on financial question answering, weak geospatial localization, pathology understanding far below human experts, and poor materials property prediction. The survey frames these failures not as isolated results but as a systematic gap that domain-specific benchmarking is meant to close.
Load-bearing premise
The load-bearing premise is that the search strategy described in Section 2 captured the full landscape of domain-specific MLLM benchmarks, so that the taxonomy and summary tables are representative and complete; the paper's own inconsistency between seven disciplines (abstract) and eight (methodology) shows those coverage boundaries are not clearly pinned down.
Editorial extensions
If this is right
- If the taxonomy is accurate, a researcher entering an unfamiliar domain can locate the relevant benchmarks and their reported baselines in one place, lowering the cost of designing a new evaluation.
- The cross-domain pattern of failures—text shortcuts in pathology, poor counting and localization in geospatial data, low accuracy on financial QA—implies that benchmark design should include controls that isolate each input modality.
- The survey's evidence supports making future benchmarks 'living' and multimodal, and scoring robustness, efficiency, and safety as well as accuracy.
- Domain-specific benchmark results argue for domain-adapted fine-tuning and hybrid pipelines (LLM plus retrieval, symbolic solvers, or verification tools) rather than relying on a generalist model alone.
- If the map is complete, it gives a concrete way to track whether MLLM progress toward generally capable AI is actually broadening across disciplines, the paper's stated long-term aim.
Reading between the lines
- Going beyond the paper: its own tables suggest a testable hypothesis that MLLM performance on a benchmark depends less on model size than on how closely the task's input format and vocabulary match the model's training distribution.
- The abstract's seven-discipline count versus the methodology's eight suggests the taxonomy's boundaries were still shifting; a natural extension is a living, community-maintained registry that updates the map as new benchmarks appear.
- The recurring 'text shortcut' failure—models answering from language cues instead of analyzing images—implies that future benchmarks should include diagnostic variants that remove one modality at a time, so scores cannot be gamed by linguistic priors.
- The survey's cited fine-tuning results imply that domain benchmarks can serve as training signal, not just evaluation tools; adversarially noisy or domain-specific examples measurably improve robustness and accuracy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of domain-specific benchmarks for multimodal large language models (MLLMs). It proposes a taxonomy of disciplines—Engineering, Science, Technology, Mathematics, Humanities, Finance, Healthcare, and Language Understanding—and, for each, provides summary tables with scale, task type, input modality, models, performance, and key focus, along with narrative discussion of trends and limitations. The stated contribution is a comprehensive, accessible resource that maps the domain-specific evaluation landscape and highlights where current MLLMs succeed or fail in specialized fields.
Significance. If the survey's scope and categorization were reliable, it would be a useful entry point for researchers seeking domain-specific evaluation resources. The paper's strength is its breadth: it organizes a large number of benchmarks and studies across eight disciplines and identifies recurring gaps (e.g., in medical image reasoning and in modality-dependent performance). However, the value of the survey depends entirely on the reproducibility of its selection process and the accuracy of its categorizations; the current inconsistencies substantially weaken the central claim of comprehensiveness. The paper does not provide machine-checked proofs or reproducible code; its contribution is the structured synthesis itself.
major comments (4)
- [§1 (Abstract) vs §2] The abstract states that the paper introduces 'a taxonomy of seven key disciplines,' while Section 2 explicitly lists eight disciplines and Figure 1 displays eight. This is a direct internal contradiction about the scope of the survey. Since the paper's central claim is that it provides a comprehensive taxonomy, the intended number of disciplines must be clarified; as written, a reader cannot tell whether one section is extraneous or the abstract is wrong.
- [§2, Search Strategy and Scope] The search strategy paragraph names databases and general keyword categories but provides no screening criteria, deduplication procedure, date range, number of records retrieved, number of records screened, or exclusion log. The paper claims to be a 'systematic review' and concludes that its resource is comprehensive, but without these details the coverage is not checkable or reproducible. The claim that the review 'examines eight key disciplines' and that the tables consolidate 'relevant benchmarks and survey papers' is therefore not falsifiable; the authors should add a principled inclusion/exclusion protocol and a flow diagram or equivalent transparency measure.
- [Tables 1, 2, and 8] Several entries in the summary tables are not domain-specific benchmarks, which contradicts the paper's stated scope. Table 1 places BIG-bench under Software Engineering/Knowledge Graphs & Semantic Systems, but BIG-bench is a general-purpose benchmark spanning 204 tasks across many areas. Table 2 includes WeatherBench 2, a weather forecasting benchmark that is not an LLM or MLLM evaluation benchmark. Table 8 lists ToT (a prompting method), LongLLaVA, LLaVA-OneVision, KOSMOS-1, KOSMOS-2, and ChatGLM (models or architectures) as if they were benchmarks. These inclusions make the consolidated resource unreliable and undermine the central claim that the paper catalogs domain-specific benchmarks for evaluating MLLMs.
- [§1.2, paragraph 3] The manuscript states: "OpenAI's GPT-4o (referred to as 'o3') at 69.1%." This is a factual error: GPT-4o and o3 are different models, and the parenthetical misattributes a reported score. In a survey whose purpose is to summarize model performance on benchmarks, such a mislabeling is a load-bearing accuracy problem because readers may rely on the reported comparisons for model selection. The sentence should be corrected and the source of the 69.1% figure should be cited.
minor comments (4)
- [Figure 1] The label 'Chain & Crypto' should read 'Blockchain & Cryptocurrency' to match the terminology used in Section 5.3 and Table 3.
- [Throughout] The model name 'LLaVA' is repeatedly typeset as 'LLaV A' with an erroneous space; please correct this globally.
- [§6.1 vs §8.2.1] The benchmark cited as [104] is called 'KnowledgeMath' in Table 4 and Section 6.3 but 'FinanceMath' in Table 6; the text in §6.3 notes the alternative name, but the tables should use one canonical name or explicitly cross-reference the alias.
- [Table 6] The row for FinanceBench lists 'GPT-4+retriever' with performance '19% correct'; the surrounding text (Section 8.2.3) says GPT-4 provided correct responses to only 19% of questions, but the table does not indicate whether this is the best result among the evaluated models or a representative one; please clarify.
Circularity Check
No circularity: the survey organizes external benchmarks and derives no predictions from fitted inputs or self-citation chains.
full rationale
This paper is a survey and taxonomy of domain-specific MLLM benchmarks. It contains no equations, fitted parameters, or derived quantitative predictions whose outputs could reduce to its inputs by construction. The central contribution is an organizing framework: the authors select eight disciplines, assign existing benchmark papers to sub-domains, and summarize their characteristics. The taxonomy is a descriptive categorization of external work, not a result derived from those same papers, so the self-definitional and fitted-input patterns do not apply. There is no load-bearing appeal to a prior uniqueness theorem by the same authors, and no ansatz is smuggled in via citation; the benchmark descriptions are taken from the cited papers themselves, which are independent external sources. The paper's coverage claim is supported by a stated search strategy over standard databases, and while that strategy is high-level, the absence of screening detail is a reproducibility or rigor concern, not a circularity concern. Internal inconsistencies such as the abstract stating 'seven key disciplines' while Section 2 and Figure 1 list eight, and table entries that are models or methods rather than benchmarks (e.g., ToT, LongLLaVA, KOSMOS-1, ChatGLM) or general-purpose rather than domain-specific (e.g., BIG-bench under Software Engineering), are scope and consistency defects. They do not make the survey's claims equivalent to its inputs by definition. Under the hard rule requiring a specific reduction to label circularity, no such reduction can be quoted from this paper. The honest finding is therefore no significant circularity, with a score of 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The eight discipline taxonomy is complete and non-overlapping.
- domain assumption The literature search and inclusion criteria capture all or most relevant domain-specific benchmarks.
Cite this review
Pith. "Pith review of Domain Specific Benchmarks for Evaluating Multimodal Large Language Models." pith.science (2026). https://pith.science/paper/HWYET6G2
@misc{pith2026250612958,
author = {Pith},
title = {Pith review of: Domain Specific Benchmarks for Evaluating Multimodal Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/HWYET6G2}},
note = {Machine review of arXiv:2506.12958}
}
read the original abstract
Large language models (LLMs) are increasingly being deployed across disciplines due to their advanced reasoning and problem solving capabilities. To measure their effectiveness, various benchmarks have been developed that measure aspects of LLM reasoning, comprehension, and problem-solving. While several surveys address LLM evaluation and benchmarks, a domain-specific analysis remains underexplored in the literature. This paper introduces a taxonomy of seven key disciplines, encompassing various domains and application areas where LLMs are extensively utilized. Additionally, we provide a comprehensive review of LLM benchmarks and survey papers within each domain, highlighting the unique capabilities of LLMs and the challenges faced in their application. Finally, we compile and categorize these benchmarks by domain to create an accessible resource for researchers, aiming to pave the way for advancements toward artificial general intelligence (AGI)
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Milli- can, et al., Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805 (2023)
arXiv 2023
-
[3]
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosse- lut, E. Brunskill, et al., On the opportunities and risks of foundation models, arXiv preprint arXiv:2108.07258 (2021)
arXiv 2021
-
[4]
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Ka- dian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al., The llama 3 herd of models, arXiv preprint arXiv:2407.21783 (2024)
arXiv 2024
- [5]
-
[6]
Islam, A
P. Islam, A. Kannappan, D. Kiela, R. Qian, N. Scherrer, B. Vidgen, Financebench: A new benchmark for financial question answering, arXiv preprint (2023). arXiv:2311. 11944
2023
- [7]
-
[8]
J. Li, Y . Zhu, Z. Xu, J. Gu, M. Zhu, X. Liu, N. Liu, Y . Peng, F. Feng, J. Tang, Mmro: Are multimodal llms eligible as the brain for in-home robotics?, arXiv preprint arXiv:2406.19693 (2024)
arXiv 2024
Show all 205 references
-
[9]
A. C. Doris, D. Grandi, R. Tomich, M. F. Alam, M. Ataei, H. Cheong, F. Ahmed, Designqa: A multimodal bench- mark for evaluating large language models’ understanding of engineering documentation, Journal of Computing and Information Science in Engineering 25 (2) (2024) 021009. ...
2024 doi
-
[10]
Kernan Freire, C
S. Kernan Freire, C. Wang, M. Foosherian, S. Wellsandt, S. Ruiz-Arenas, E. Niforatos, Knowledge sharing in man- ufacturing using llm-powered tools: user study and model benchmarking, Frontiers in Artificial intelligence 7 (2024) 1293084
2024
-
[11]
X. Liu, S. Yang, X. Dong, H. Rong, B. Fu, Manu-eval: A chinese language understanding benchmark for manu- facturing industry, in: China Conference on Knowledge Graph and Semantic Computing, Springer, 2024, pp. 309– 317
2024
-
[12]
Eslaminia, A
A. Eslaminia, A. Jackson, B. Tian, A. Stern, H. Gor- don, R. Malhotra, K. Nahrstedt, C. Shao, Fdm-bench: A comprehensive benchmark for evaluating large language models in additive manufacturing tasks, arXiv preprint arXiv:2412.09819 (2024)
2024 arXiv
-
[13]
Fakih, R
M. Fakih, R. Dharmaji, Y . Moghaddas, G. Quiros, O. Ogundare, M. A. Al Faruque, Llm4plc: Harnessing large language models for verifiable programming of plcs in industrial control systems, in: Proceedings of the 46th International Conference on Software Engineering: Software En...
2024
-
[14]
Tizaoui, R
T. Tizaoui, R. Tan, Towards a benchmark dataset for large language models in the context of process automation, Digital Chemical Engineering (2024) 100186
2024
-
[15]
Y . Xia, J. Zhang, N. Jazdi, M. Weyrich, Incorporating large language models into production systems for en- hanced task automation and flexibility, arXiv preprint arXiv:2407.08550 (2024)
2024 arXiv
-
[16]
Ogundare, S
O. Ogundare, S. Madasu, N. Wiggins, Industrial engi- neering with large language models: A case study of chatgpt’s performance on oil & gas problems, in: 2023 11th International Conference on Control, Mechatronics and Automation (ICCMA), IEEE, 2023, pp. 458–461
2023
-
[17]
S. A. Rahman, S. Chawla, M. Yaqot, B. Menezes, Lever- aging large language models for supply chain manage- ment optimization: A case study, in: International Confer- ence on Innovative Intelligent Industrial Production and Logistics, Springer, 2024, pp. 175–197
2024
-
[18]
B. Li, K. Mellou, B. Zhang, J. Pathuri, I. Menache, Large language models for supply chain optimization, arXiv preprint arXiv:2307.03875 (2023)
2023 arXiv
-
[19]
Raman, A
R. Raman, A. Sreenivasan, M. Suresh, A. Gunasekaran, P. Nedungadi, Ai-driven education: a comparative study on chatgpt and bard in supply chain management contexts, Cogent Business & Management 11 (1) (2024) 2412742. 32
2024
-
[20]
Meyer, J
L.-P. Meyer, J. Frey, K. Junghanns, F. Brei, K. Bulert, S. Gründer-Fahrer, M. Martin, Developing a scalable benchmark for assessing large language models in knowl- edge graph engineering (2023).arXiv:2308.16622. URLhttps://arxiv.org/abs/2308.16622
2023 arXiv
-
[21]
bench authors, Beyond the imitation game: Quantify- ing and extrapolating the capabilities of language models, Transactions on Machine Learning Research (2023)
B. bench authors, Beyond the imitation game: Quantify- ing and extrapolating the capabilities of language models, Transactions on Machine Learning Research (2023). URL https://openreview.net/forum?id= uyTL5Bvosj
2023
-
[22]
Azanza, B
M. Azanza, B. P. Lamancha, E. Pizarro, Tracking the moving target: A framework for continuous evaluation of llm test generation in industry (2025). arXiv:2504. 18985. URLhttps://arxiv.org/abs/2504.18985
2025 arXiv
-
[23]
N. Shah, Z. Genc, D. Araci, Stackeval: Benchmarking llms in coding assistance, Advances in Neural Informa- tion Processing Systems 37 (2024) 36976–36994
2024
-
[24]
Zhang, C
Q. Zhang, C. Fang, Y . Xie, Y . Zhang, Y . Yang, W. Sun, S. Yu, Z. Chen, A survey on large language models for software engineering (2024).arXiv:2312.15223. URLhttps://arxiv.org/abs/2312.15223
2024 arXiv
-
[25]
R. Bell, R. Longshore, R. Madachy, Introducing syseng- bench: A novel benchmark for assessing large language models in systems engineering, Tech. rep., Acquisition Research Program (2024)
2024
-
[26]
Y . Hu, Y . Goktas, D. D. Yellamati, C. De Tassigny, The use and misuse of pre-trained generative large language models in reliability engineering, in: 2024 Annual Reli- ability and Maintainability Symposium (RAMS), IEEE, 2024, pp. 1–7
2024
-
[27]
Vendrow, E
J. Vendrow, E. Vendrow, S. Beery, A. Madry, Do large lan- guage model benchmarks test reliability?, arXiv preprint arXiv:2502.03461 (2025)
2025 arXiv
-
[28]
Y . Liu, Y . Yao, J.-F. Ton, X. Zhang, R. Guo, H. Cheng, Y . Klochkov, M. F. Taufiq, H. Li, Trustworthy llms: a sur- vey and guideline for evaluating large language models’ alignment (2024).arXiv:2308.05374. URLhttps://arxiv.org/abs/2308.05374
2024 arXiv
-
[29]
J. A. Irvin, E. R. Liu, J. C. Chen, I. Dormoy, J. Kim, S. Khanna, Z. Zheng, S. Ermon, TEOChat: A Large Vision-Language Assistant for Temporal Earth Observa- tion Data, _eprint: 2410.06234 (2024). URLhttps://arxiv.org/abs/2410.06234
2024 arXiv
-
[30]
Xiong, F
Z. Xiong, F. Zhang, Y . Wang, Y . Shi, X. X. Zhu, Earth- Nets: Empowering artificial intelligence for Earth obser- vation, IEEE Geoscience and Remote Sensing Magazine (2024) 2–36doi:10.1109/MGRS.2024.3466998
2024
-
[31]
Zhang, S
C. Zhang, S. Wang, Good at captioning, bad at counting: Benchmarking gpt-4v on earth observation data, arXiv preprint arXiv:2401.17600 (2024)
2024 arXiv
-
[32]
W. Li, D. Yao, R. Zhao, W. Chen, Z. Xu, C. Luo, C. Gong, Q. Jing, H. Tan, J. Bi, STBench: Assessing the Ability of Large Language Models in Spatio-Temporal Analysis, arXiv preprint arXiv:2406.19065 (2024)
2024 arXiv
-
[33]
Sapkota, R
R. Sapkota, R. Qureshi, S. Z. Hassan, J. Shutske, M. Shoman, M. Sajjad, F. A. Dharejo, A. Paudel, J. Li, Z. Meng, others, Multi-modal LLMs in agriculture: A comprehensive review, Authorea PreprintsPublisher: Au- thorea (2024)
2024
-
[34]
M. S. Danish, M. A. Munir, S. R. A. Shah, K. Kuck- reja, F. S. Khan, P. Fraccaro, A. Lacoste, S. Khan, GEOBench-VLM: Benchmarking Vision-Language Mod- els for Geospatial Tasks, _eprint: 2411.19325 (2024). URLhttps://arxiv.org/abs/2411.19325
2024 arXiv
-
[35]
Roberts, T
J. Roberts, T. Lüddecke, R. Sheikh, K. Han, S. Albanie, Charting new territories: Exploring the geographic and geospatial capabilities of multimodal llms, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 554–563
2024
-
[36]
X. Liu, Z. Lian, RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mix- ture of Experts, _eprint: 2412.05679 (2024). URLhttps://arxiv.org/abs/2412.05679
2024 arXiv
-
[37]
C. Lin, H. Lyu, X. Xu, J. Luo, INS-MMBench: A Com- prehensive Benchmark for Evaluating LVLMs’ Perfor- mance in Insurance, arXiv preprint arXiv:2406.09105 (2024)
2024 arXiv
-
[38]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, J. Steinhardt, Measuring massive multitask lan- guage understanding, arXiv preprint arXiv:2009.03300 (2020)
2020 arXiv
-
[39]
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, A. Kalyan, Learn to explain: Multi- modal reasoning via thought chains for science question answering, Advances in Neural Information Processing Systems 35 (2022) 2507–2521
2022
-
[40]
D. Fu, R. Guo, G. Khalighinejad, O. Liu, B. Dhingra, D. Yogatama, R. Jia, W. Neiswanger, Isobench: Bench- marking multimodal foundation models on isomorphic representations, arXiv preprint arXiv:2404.01266 (2024)
2024 arXiv
-
[41]
Jiang, Z
Z. Jiang, Z. Yang, J. Chen, Z. Du, W. Wang, B. Xu, J. Tang, Visscience: An extensive benchmark for evalu- ating k12 educational multi-modal scientific reasoning, arXiv preprint arXiv:2409.13730 (2024)
2024 arXiv
-
[42]
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, S. R. Bowman, Gpqa: A graduate- level google-proof q&a benchmark, in: First Conference on Language Modeling, 2024. 33
2024
-
[43]
Anand, J
A. Anand, J. Kapuriya, A. Singh, J. Saraf, N. Lal, A. Verma, R. Gupta, R. Shah, Mm-phyqa: Multimodal physics question-answering with multi-image cot prompt- ing, in: Pacific-Asia Conference on Knowledge Discovery and Data Mining, Springer, 2024, pp. 53–64
2024
-
[44]
Mirza, N
A. Mirza, N. Alampara, S. Kunchapu, M. Ríos-García, B. Emoekabu, A. Krishnan, T. Gupta, M. Schilling- Wilhelmi, M. Okereke, A. Aneesh, et al., Are large language models superhuman chemists?, arXiv preprint arXiv:2404.01475 (2024)
2024 arXiv
-
[45]
S. Zhu, X. Liu, G. Khalighinejad, Chemqa: a multimodal question-and-answering dataset on chemistry reasoning, https://huggingface.co/datasets/shangzhu/ ChemQA(2024)
2024
-
[46]
T. Guo, B. Nan, Z. Liang, Z. Guo, N. Chawla, O. Wiest, X. Zhang, et al., What can large language models do in chemistry? a comprehensive benchmark on eight tasks, Advances in Neural Information Processing Systems 36 (2023) 59662–59688
2023
-
[47]
B. Yu, F. N. Baker, Z. Chen, X. Ning, H. Sun, Llasmol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tun- ing dataset, arXiv preprint arXiv:2402.09391 (2024)
2024 arXiv
-
[48]
M. Zaki, N. Krishnan, et al., Mascqa: A question answering dataset for investigating materials science knowledge of large language models, arXiv preprint arXiv:2308.09115 (2023)
2023 arXiv
-
[49]
A. N. Rubungo, K. Li, J. Hattrick-Simpers, A. B. Di- eng, Llm4mat-bench: benchmarking large language mod- els for materials property prediction, arXiv preprint arXiv:2411.00177 (2024)
2024 arXiv
-
[50]
H. Cao, Y . Shao, Z. Liu, Z. Liu, X. Tang, Y . Yao, Y . Li, Presto: progressive pretraining enhances synthetic chem- istry outcomes, arXiv preprint arXiv:2406.13193 (2024)
2024 arXiv
-
[51]
X. Liu, Y . Guo, H. Li, J. Liu, S. Huang, B. Ke, J. Lv, Drugllm: Open large language model for few-shot molecule generation, arXiv preprint arXiv:2405.06690 (2024)
2024 arXiv
-
[52]
S. Rasp, S. Hoyer, A. Merose, I. Langmore, P. Battaglia, T. Russell, A. Sanchez-Gonzalez, V . Yang, R. Carver, S. Agrawal, et al., Weatherbench 2: A benchmark for the next generation of data-driven global weather models, Journal of Advances in Modeling Earth Systems 16 (6) (20...
2024
-
[53]
J. Chen, P. Zhou, Y . Hua, D. Chong, M. Cao, Y . Li, Z. Yuan, B. Zhu, J. Liang, Vision-language models meet meteorology: Developing models for extreme weather events detection with heatmaps, arXiv preprint arXiv:2406.09838 (2024)
2024
-
[54]
C. Ma, Z. Hua, A. Anderson-Frey, V . Iyer, X. Liu, L. Qin, Weatherqa: Can multimodal language models reason about severe weather?, arXiv preprint arXiv:2406.11217 (2024)
2024 arXiv
-
[55]
H. Li, Z. Wang, J. Wang, A. K. H. Lau, H. Qu, Cllmate: A multimodal llm for weather and climate events fore- casting, arXiv preprint arXiv:2409.19058 (2024)
2024 arXiv
-
[56]
Y . Sun, C. Wang, Y . Peng, Unleashing the potential of large language model: Zero-shot vqa for flood disaster scenario, in: Proceedings of the 4th International Confer- ence on Artificial Intelligence and Computer Engineering, 2023, pp. 368–373
2023
-
[57]
Rawat, Disasterqa: A benchmark for assessing the performance of llms in disaster response, arXiv preprint arXiv:2410.20707 (2024)
R. Rawat, Disasterqa: A benchmark for assessing the performance of llms in disaster response, arXiv preprint arXiv:2410.20707 (2024)
2024 arXiv
-
[58]
Z. B. Patel, Y . Bachwana, N. Sharma, S. Gut- tikunda, N. Batra, Vayubuddy: an llm-powered chat- bot to democratize air quality insights, arXiv preprint arXiv:2411.12760 (2024)
2024 arXiv
-
[59]
Pafilis, S
E. Pafilis, S. P. Frankild, L. Fanini, S. Faulwetter, C. Pavloudi, A. Vasileiadou, C. Arvanitidis, L. J. Jensen, The species and organisms resources for fast and accurate identification of taxonomic names in text, PloS one 8 (6) (2013) e65390
2013
-
[60]
Abdelmageed, F
N. Abdelmageed, F. Löffler, L. Feddoul, A. Algergawy, S. Samuel, J. Gaikwad, A. Kazem, B. König-Ries, Bio- divnere: Gold standard corpora for named entity recog- nition and relation extraction in the biodiversity domain, Biodiversity Data Journal 10 (2022) e89481
2022
-
[61]
A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y . Zhao, X. Du, M. R. G. Madani, et al., Are we done with mmlu?, arXiv preprint arXiv:2406.04127 (2024)
2024 arXiv
-
[62]
A. M. Richard, R. Huang, S. Waidyanatha, P. Shinn, B. J. Collins, I. Thillainadarajah, C. M. Grulke, A. J. Williams, R. R. Lougee, R. S. Judson, et al., The tox21 10k com- pound library: collaborative chemistry advancing toxi- cology, Chemical Research in Toxicology 34 (2) (20...
2020
-
[63]
S. Kim, P. A. Thiessen, E. E. Bolton, J. Chen, G. Fu, A. Gindulyte, L. Han, J. He, S. He, B. A. Shoemaker, et al., Pubchem substance and compound databases, Nu- cleic acids research 44 (D1) (2016) D1202–D1213
2016
-
[64]
W. Jin, C. Coley, R. Barzilay, T. Jaakkola, Predicting or- ganic reaction outcomes with weisfeiler-lehman network, Advances in neural information processing systems 30 (2017)
2017
-
[65]
Edwards, C
C. Edwards, C. Zhai, H. Ji, Text2mol: Cross-modal molecule retrieval with natural language queries, in: Pro- ceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 595–607. 34
2021
-
[66]
J. J. Irwin, K. G. Tang, J. Young, C. Dandarchuluun, B. R. Wong, M. Khurelbaatar, Y . S. Moroz, J. Mayfield, R. A. Sayle, Zinc20—a free ultralarge-scale chemical database for ligand discovery, Journal of chemical information and modeling 60 (12) (2020) 6065–6073
2020
-
[67]
Davies, M
M. Davies, M. Nowotka, G. Papadatos, N. Dedman, A. Gaulton, F. Atkinson, L. Bellis, J. P. Overington, Chembl web services: streamlining access to drug dis- covery data and utilities, Nucleic acids research 43 (W1) (2015) W612–W620
2015
-
[68]
S. Rasp, P. D. Dueben, S. Scher, J. A. Weyn, S. Mouatadid, N. Thuerey, Weatherbench: a benchmark data set for data- driven weather forecasting, Journal of Advances in Mod- eling Earth Systems 12 (11) (2020) e2020MS002203
2020
-
[69]
S. Li, W. Yang, P. Zhang, X. Xiao, D. Cao, Y . Qin, X. Zhang, Y . Zhao, P. Bogdan, Climatellm: Efficient weather forecasting via frequency-aware large language models, arXiv preprint arXiv:2502.11059 (2025)
2025 arXiv
-
[70]
Sachdeva, N
E. Sachdeva, N. Agarwal, S. Chundi, S. Roelofs, J. Li, M. Kochenderfer, C. Choi, B. Dariush, Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2024, pp. 7513–7522
2024
-
[71]
X. Ding, J. Han, H. Xu, X. Liang, W. Zhang, X. Li, Holistic autonomous driving understanding by bird’s-eye- view injected multi-modal large models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13668–13677
2024
-
[72]
T. Qian, J. Chen, L. Zhuo, Y . Jiao, Y .-G. Jiang, NuScenes-QA: A Multi-Modal Visual Question An- swering Benchmark for Autonomous Driving Scenario, Proceedings of the AAAI Conference on Artificial Intelligence 38 (5) (2024) 4542–4550, number: 5. doi:10.1609/aaai.v38i5.28253. ...
2024 doi
-
[73]
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, others, Cambrian-1: A fully open, vision-centric exploration of multimodal llms, arXiv preprint arXiv:2406.16860 (2024)
2024 arXiv
-
[74]
Q. Zhou, S. Chen, Y . Wang, H. Xu, W. Du, H. Zhang, Y . Du, J. B. Tenenbaum, C. Gan, HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments, ArxivUniv Massachusetts Amherst Peking Univ MIT MIT (2024)
2024
-
[75]
Z. Yang, X. Jia, H. Li, J. Yan, LLM4Drive: A Survey of Large Language Models for Autonomous Driving, Arx- ivOpenDriveLab (2024)
2024
-
[76]
Q. Kong, Y . Kawana, R. Saini, A. Kumar, J. Pan, T. Gu, Y . Ozao, B. Opra, Y . Sato, N. Kobori, WTS: A Pedestrian- Centric Traffic Video Dataset for Fine-Grained Spatial- Temporal Understanding, V ol. 15134, 2025, pp. 1–18, woven Toyota. doi:10.1007/978-3-031-73116-7_ 1
2025 doi
-
[77]
Malla, C
S. Malla, C. Choi, I. Dwivedi, J. H. Choi, J. Li, Drama: Joint risk localization and captioning in driving, in: Pro- ceedings of the IEEE/CVF winter conference on applica- tions of computer vision, 2023, pp. 1043–1052
2023
-
[78]
J. Yang, S. Gao, Y . Qiu, L. Chen, T. Li, B. Dai, K. Chitta, P. Wu, J. Zeng, P. Luo, J. Zhang, A. Geiger, Y . Qiao, H. Li, Generalized Predictive Model for Autonomous Driving, 2024, pp. 14662–14672. URL https://openaccess.thecvf.com/content/ CVPR2024/html/Yang_Generalized_Pred...
2024
-
[79]
M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, L. Zhang, Reason2Drive: Towards Interpretable and Chain-Based Reasoning for Autonomous Driving, in: A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sat- tler, G. Varol (Eds.), Computer Vision – ECCV 2024, Springer Nature Swi...
2024 doi
-
[80]
C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, H. Li, DriveLM: Driving with Graph Visual Question Answering, in: A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sat- tler, G. Varol (Eds.), Computer Vision – ECCV 2024, Springer Nat...
2024 doi
- [81]
-
[82]
Z. Xiao, Q. Wang, H. Pearce, S. Chen, Logic meets magic: Llms cracking smart contract vulnerabilities (2025).arXiv:2501.07058. URLhttps://arxiv.org/abs/2501.07058
2025 arXiv
-
[83]
Z. Wei, J. Sun, Z. Zhang, X. Zhang, M. Li, Z. Hou, Llm- smartaudit: Advanced smart contract vulnerability detec- tion (2024).arXiv:2410.09381. URLhttps://arxiv.org/abs/2410.09381
2024 arXiv
-
[84]
Zhang, K
L. Zhang, K. Li, K. Sun, D. Wu, Y . Liu, H. Tian, Y . Liu, Acfix: Guiding llms with mined common rbac practices for context-aware repair of access control vulnerabilities in smart contracts (2024).arXiv:2403.06838. URLhttps://arxiv.org/abs/2403.06838 35
2024 arXiv
-
[85]
K. I. Roumeliotis, N. D. Tselikas, D. K. Nasiopoulos, Llms and nlp models in cryptocurrency sentiment anal- ysis: A comparative classification study, Big Data and Cognitive Computing 8 (6) (2024) 63. doi:10.3390/ bdcc8060063
2024
-
[86]
Makri, G
E. Makri, G. Palaiokrassas, S. Bouraga, A. Polychroni- adou, L. Tassiulas, Ethereum price prediction employing large language models for short-term and few-shot fore- casting (2025).arXiv:2503.23190. URLhttps://arxiv.org/abs/2503.23190
2025 arXiv
-
[87]
Q. Wang, Y . Gao, Z. Tang, B. Luo, N. Chen, B. He, Exploring llm cryptocurrency trading through fact- subjectivity aware reasoning (2025). arXiv:2410. 12464. URLhttps://arxiv.org/abs/2410.12464
2025 arXiv
-
[88]
Z. He, Z. Li, S. Yang, H. Ye, A. Qiao, X. Zhang, X. Luo, T. Chen, Large language models for blockchain secu- rity: A systematic literature review (2025). arXiv: 2403.14280. URLhttps://arxiv.org/abs/2403.14280
2025 arXiv
-
[89]
Trozze, T
A. Trozze, T. Davies, B. Kleinberg, Large language models in cryptocurrency securities cases: Can a gpt model meaningfully assist lawyers? (2024). arXiv: 2308.06032. URLhttps://arxiv.org/abs/2308.06032
2024 arXiv
-
[90]
C. Cui, Y . Ma, X. Cao, W. Ye, Y . Zhou, K. Liang, J. Chen, J. Lu, Z. Yang, K.-D. Liao, et al., A survey on multi- modal large language models for autonomous driving, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 958–979
2024
-
[91]
Y . Shi, K. Jiang, J. Li, Z. Qian, J. Wen, M. Yang, K. Wang, D. Yang, Grid-centric traffic scenario percep- tion for autonomous driving: A comprehensive review, IEEE Transactions on Neural Networks and Learning Systems (2024)
2024
-
[92]
P. Tong, E. Brown, P. Wu, S. Woo, A. J. V . IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang, et al., Cambrian-1: A fully open, vision-centric exploration of multimodal llms, Advances in Neural Information Pro- cessing Systems 37 (2024) 87310–87356
2024
-
[93]
J. Yang, S. Gao, Y . Qiu, L. Chen, T. Li, B. Dai, K. Chitta, P. Wu, J. Zeng, P. Luo, et al., Generalized predictive model for autonomous driving, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14662–14672
2024
-
[94]
R. S. Shah, K. Chawla, D. Eidnani, A. Shah, W. Du, S. Chava, N. Raman, C. Smiley, J. Chen, D. Yang, When flue meets flang: Benchmarks and large pre-trained lan- guage model for financial domain, arXiv preprint (2022). arXiv:2211.00083
2022 arXiv
-
[95]
W. Guan, J. Cao, S. Qian, J. Gao, C. Ouyang, Logllm: Log-based anomaly detection using large language mod- els (2025).arXiv:2411.08561. URLhttps://arxiv.org/abs/2411.08561
2025 arXiv
-
[96]
Sinha, A
R. Sinha, A. Elhafsi, C. Agia, M. Foutter, E. Schmerling, M. Pavone, Real-time anomaly detection and reactive planning with large language models (2024). arXiv: 2407.08735. URLhttps://arxiv.org/abs/2407.08735
2024 arXiv
-
[97]
Zhang, D
R. Zhang, D. Jiang, Y . Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K.-W. Chang, Y . Qiao, et al., Mathverse: Does your multi-modal llm truly see the diagrams in vi- sual math problems?, in: European Conference on Com- puter Vision, Springer, 2024, pp. 169–186
2024
-
[98]
J. Fan, S. Martinson, E. Y . Wang, K. Hausknecht, J. Brenner, D. Liu, N. Peng, C. Wang, M. Brenner, HARDMATH: A benchmark dataset for challenging problems in applied mathematics, in: The 4th Workshop on Mathematical Reasoning and AI at NeurIPS’24, 2024. URL https://openreview....
2024
-
[99]
H. Liu, Y . Zhang, Y . Luo, A. C.-C. Yao, Augmenting Math Word Problems via Iterative Question Composing, ArxivShanghai Qizhi Inst Beijing Univ Posts & Telecom- mun (2024)
2024
-
[100]
Liang, D
Z. Liang, D. Yu, W. Yu, W. Yao, Z. Zhang, X. Zhang, D. Yu, Mathchat: Benchmarking mathematical reasoning and instruction following in multi-turn interactions, arXiv preprint arXiv:2405.19444 (2024)
2024 arXiv
-
[101]
Zhang, L
Z. Zhang, L. Xu, Z. Jiang, H. Hao, R. Wang, Multiple- choice questions are efficient and robust llm evaluators. 2024d, URL https://doi. org/10.48550/arXiv 2405
-
[102]
Anantheswaran, H
U. Anantheswaran, H. Gupta, K. Scaria, S. Verma, C. Baral, S. Mishra, Cutting through the noise: Boosting llm performance on math word problems, arXiv preprint arXiv:2406.15444 (2024)
2024
-
[103]
K. Yang, A. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. J. Prenger, A. Anandkumar, Leandojo: Theo- rem proving with retrieval-augmented language models, Advances in Neural Information Processing Systems 36 (2023) 21573–21612
2023
-
[104]
Y . Zhao, H. Liu, Y . Long, R. Zhang, C. Zhao, A. Cohan, Financemath: Knowledge-intensive math reasoning in finance domains, in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 2024, pp. 12841–12858
2024
-
[105]
Pezeshkpour, E
P. Pezeshkpour, E. Hruschka, Large language models sensitivity to the order of options in multiple-choice questions, in: K. Duh, H. Gomez, S. Bethard (Eds.), Findings of the Association for Computational Linguis- tics: NAACL 2024, Association for Computational 36 Linguistics, ...
2024 doi
-
[106]
LaBelle, Monte carlo tree search applications to neural theorem proving, Ph.D
E. LaBelle, Monte carlo tree search applications to neural theorem proving, Ph.D. thesis, Massachusetts Institute of Technology (2024)
2024
-
[107]
Z. Yuan, K. Wang, S. Zhu, Y . Yuan, J. Zhou, Y . Zhu, W. Wei, Finllms: A framework for financial reasoning dataset generation with large language models, IEEE Transactions on Big Data (2024)
2024
-
[108]
Y . Jin, M. Choi, G. Verma, J. Wang, S. Kumar, Mm- soc: Benchmarking multimodal large language models in social media platforms, arXiv preprint arXiv:2402.14154 (2024)
2024 arXiv
-
[109]
Y . Chen, S. Yan, Q. Guo, J. Jia, Z. Li, Y . Xiao, Hotv- com: Generating buzzworthy comments for videos, in: L.-W. Ku, A. Martins, V . Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 2024, pp. 2198–2224. doi: 10.18653/v1...
2024 doi
-
[110]
Y . Chen, S. Yan, Z. Zhu, Z. Li, Y . Xiao, Xmecap: Meme caption generation with sub-image adaptability, in: Pro- ceedings of the 32nd ACM International Conference on Multimedia, MM ’24, Association for Computing Ma- chinery, New York, NY , USA, 2024, pp. 3352–3361. doi:10.1145...
2024
-
[111]
Shahriar, R
S. Shahriar, R. Dara, Priv-iq: A benchmark and com- parative evaluation of large multimodal models on pri- vacy competencies, AI 6 (2) (2025) 29. doi:10.3390/ ai6020029
2025
-
[112]
I. O. Gallegos, et al., Bias and fairness in large language models: A survey, Comput. Linguist. 50 (3) (2024) 1097– 1179.doi:10.1162/coli_a_00524
2024 doi
-
[113]
Liu, et al., Culturevlm: Characterizing and improving cultural understanding of vision-language models for over 100 countries, arXiv (Jan
S. Liu, et al., Culturevlm: Characterizing and improving cultural understanding of vision-language models for over 100 countries, arXiv (Jan. 2025). arXiv:arXiv:2501. 01282,doi:10.48550/arXiv.2501.01282
-
[114]
Ghaboura, et al., Time travel: A comprehensive bench- mark to evaluate lmms on historical and cultural arti- facts, arXiv (Feb
S. Ghaboura, et al., Time travel: A comprehensive bench- mark to evaluate lmms on historical and cultural arti- facts, arXiv (Feb. 2025). arXiv:arXiv:2502.14865, doi:10.48550/arXiv.2502.14865
- [115]
-
[116]
Y . Chen, S. Yan, S. Liu, Y . Li, Y . Xiao, Emotionqueen: A benchmark for evaluating empathy of large language mod- els, in: L.-W. Ku, A. Martins, V . Srikumar (Eds.), Find- ings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 2024, pp. 2149–2176...
2024 doi
-
[117]
Yang, et al., Editworld: Simulating world dynamics for instruction-following image editing, CoRRAccessed: Feb
L. Yang, et al., Editworld: Simulating world dynamics for instruction-following image editing, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= aLGe5a823O
2025
-
[118]
He, et al., Llms meet multimodal generation and editing: A survey, CoRRAccessed: Feb
Y . He, et al., Llms meet multimodal generation and editing: A survey, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= sYfBSrHed5
2025
-
[119]
J. Cao, Y . Liu, Y . Shi, K. Ding, L. Jin, Wenmind: A comprehensive benchmark for evaluating large language models in chinese classical literature and language arts, Adv. Neural Inf. Process. Syst. 37 (2025) 51358–51410
2025
-
[120]
Li, et al., The music maestro or the musically chal- lenged, a massive music evaluation benchmark for large language models, in: L.-W
J. Li, et al., The music maestro or the musically chal- lenged, a massive music evaluation benchmark for large language models, in: L.-W. Ku, A. Martins, V . Srikumar (Eds.), Findings of the Association for Computational Lin- guistics: ACL 2024, Bangkok, Thailand, 2024, pp. 32...
2024 doi
-
[121]
B. Weck, I. Manco, E. Benetos, E. Quinton, G. Fazekas, D. Bogdanov, Muchomusic: Evaluating music un- derstanding in multimodal audio-language models, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= ViQaFx6bmj
2025
-
[122]
Zhou, et al., Can llms ’reason’ in music? an evaluation of llms’ capability of music understanding and generation, CoRRAccessed: Feb
Z. Zhou, et al., Can llms ’reason’ in music? an evaluation of llms’ capability of music understanding and generation, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= 39KGGrrCoj
2025
-
[123]
Hachmeier, R
S. Hachmeier, R. Jäschke, A benchmark and robustness study of in-context-learning with large language models in music entity detection, in: O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schock- aert (Eds.), Proceedings of the 31st International Conferen...
2025
-
[124]
Z. Wang, et al., Muchin: a chinese colloquial descrip- tion benchmark for evaluating language models in the field of music, in: Proceedings of the Thirty-Third In- ternational Joint Conference on Artificial Intelligence, IJCAI ’24, Jeju, Korea, 2024, pp. 7771–7779. doi: 10.249...
2024 doi
-
[125]
Zhao, et al., Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space, arXiv (Feb
Y . Zhao, et al., Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space, arXiv (Feb. 2025). arXiv:arXiv:2502.12532, doi: 10.48550/arXiv.2502.12532. 37
-
[126]
Zhang, et al., Transportationgames: Benchmarking transportation knowledge of (multimodal) large language models, CoRRAccessed: Feb
X. Zhang, et al., Transportationgames: Benchmarking transportation knowledge of (multimodal) large language models, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= S2LbBjEs7x
2025
- [127]
-
[128]
Zhang, J
W. Zhang, J. Han, Z. Xu, H. Ni, H. Liu, H. Xiong, Urban foundation models: A survey, in: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, Association for Computing Machinery, New York, NY , USA, 2024, pp. 6633–6643. doi:10.1145/363...
2024
-
[129]
Zheng, et al., Urbanplanbench: A comprehensive assessment of urban planning abilities in large language modelsAccessed: Feb
Y . Zheng, et al., Urbanplanbench: A comprehensive assessment of urban planning abilities in large language modelsAccessed: Feb. 25, 2025 (Oct. 2024). URL https://openreview.net/forum?id= Dl5JaX7zoN
2025
-
[130]
Feng, et al., Citybench: Evaluating the capabilities of large language model as world model, CoRRAccessed: Feb
J. Feng, et al., Citybench: Evaluating the capabilities of large language model as world model, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= TCDS3eYNtD
2025
-
[131]
J. Ji, Y . Chen, M. Jin, W. Xu, W. Hua, Y . Zhang, Moralbench: Moral evaluation of llms, CoRRAccessed: Feb. 27, 2025 (Jan. 2024). URL https://openreview.net/forum?id= uZWqTso8oK
2025
- [132]
-
[133]
G. F. G. Marraffini, A. Cotton, N. F. Hsueh, A. Frid- man, J. Wisznia, L. D. Corro, The greatest good bench- mark: Measuring llms’ alignment with utilitarian moral dilemmas, in: Y . Al-Onaizan, M. Bansal, Y .-N. Chen (Eds.), Proceedings of the 2024 Conference on Empir- ical Me...
2024
-
[134]
Yao, et al., Value compass leaderboard: A platform for fundamental and validated evaluation of llms values, arXiv (Jan
J. Yao, et al., Value compass leaderboard: A platform for fundamental and validated evaluation of llms values, arXiv (Jan. 2025). arXiv:arXiv:2501.07071, doi: 10.48550/arXiv.2501.07071
- [135]
-
[136]
Y . Yang, Y . Xu, C. Huang, J. Jurgensen, H. Hu, Interideas: An llm and expert-enhanced dataset for philosophical intertextualityAccessed: Feb. 27, 2025 (Oct. 2024). URL https://openreview.net/forum?id= cA8iQJFioL
2025
-
[137]
Trepczy´nski, Religion, theology, and philosophical skills of llm–powered chatbots, Disput
M. Trepczy´nski, Religion, theology, and philosophical skills of llm–powered chatbots, Disput. Philos. Int. J. Phi- los. Relig. 25 (1) (2023) 19–36
2023
-
[138]
Deng, et al., Deconstructing the ethics of large lan- guage models from long-standing issues to new-emerging dilemmas, CoRRAccessed: Feb
C. Deng, et al., Deconstructing the ethics of large lan- guage models from long-standing issues to new-emerging dilemmas, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= fa1CtyfDrd
2025
-
[139]
Wang, et al., Piecing it all together: Verifying multi-hop multimodal claims, in: O
H. Wang, et al., Piecing it all together: Verifying multi-hop multimodal claims, in: O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schock- aert (Eds.), Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, UAE, 2025, ...
2025
-
[140]
Jin, et al., Agentreview: Exploring peer review dy- namics with llm agents, in: Y
Y . Jin, et al., Agentreview: Exploring peer review dy- namics with llm agents, in: Y . Al-Onaizan, M. Bansal, Y .-N. Chen (Eds.), Proceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Pro- cessing, Miami, Florida, USA, 2024, pp. 1208–1226. doi:10.18653...
2024 doi
-
[141]
Song, et al., Mosabench: Multi-object sentiment analy- sis benchmark for evaluating multimodal large language models understanding of complex image, arXiv (Nov
S. Song, et al., Mosabench: Multi-object sentiment analy- sis benchmark for evaluating multimodal large language models understanding of complex image, arXiv (Nov. 2024). arXiv:arXiv:2412.00060, doi:10.48550/ arXiv.2412.00060
2024 doi
-
[142]
M. F. Adilazuarda, et al., Towards measuring and mod- eling ’culture’ in llms: A survey, in: Y . Al-Onaizan, M. Bansal, Y .-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, 2024, pp. 15763– 15784.doi:1...
2024 doi
-
[143]
Y . Chen, Y . Xiao, Recent advancement of emotion cognition in large language models, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= BMOLRz7ko6
2025
-
[144]
Chakrabarty, P
T. Chakrabarty, P. Laban, D. Agarwal, S. Muresan, C.- S. Wu, Art or artifice? large language models and the false promise of creativity, in: Proceedings of the 2024 CHI Conference on Human Factors in Comput- ing Systems, CHI ’24, Association for Computing Ma- chinery, New York...
2024
- [145]
-
[146]
Bulla, S
L. Bulla, S. De Giorgis, M. Mongiovì, A. Gangemi, Large language models meet moral values: A comprehen- sive assessment of moral abilities, Comput. Hum. Behav. Rep. 17 (2025) 100609. doi:10.1016/j.chbr.2025. 100609
2025 doi
-
[147]
Q. Xie, W. Han, X. Zhang, Y . Lai, M. Peng, A. Lopez- Lira, J. Huang, Pixiu: A large language model, instruction data and evaluation benchmark for finance, arXiv preprint (2023).arXiv:2306.05443
2023 arXiv
-
[148]
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Lang- don, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, W. Y . Wang, Finqa: A dataset of numerical reason- ing over financial data, arXiv preprint (2021). arXiv: 2109.00122
2021 arXiv
-
[149]
Reddy, R
V . Reddy, R. Koncel-Kedziorski, V . D. Lai, M. Krumdick, C. Lovering, C. Tanner, Docfinqa: A long-context finan- cial reasoning dataset, arXiv preprint (2024). arXiv: 2401.06915
2024
-
[150]
Z. Chen, S. Li, C. Smiley, Z. Ma, S. Shah, W. Y . Wang, Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering, arXiv preprint (2022).arXiv:2210.03849
2022 arXiv
-
[151]
Webersinke, M
N. Webersinke, M. Kraus, J. A. Bingler, M. Leippold, Climatebert: A pretrained language model for climate- related text, arXiv preprint (2021). arXiv:2110.12010
2021 arXiv
-
[152]
Sharma, T
S. Sharma, T. Nayak, A. Bose, A. K. Meena, K. Dasgupta, N. Ganguly, P. Goyal, Finred: A dataset for relation ex- traction in financial domain, in: Companion Proceedings of the Web Conference 2022, 2022, pp. 595–597
2022
-
[153]
X. Wu, J. Liu, H. Su, Z. Lin, Y . Qi, C. Xu, J. Su, J. Zhong, F. Wang, S. Wang, F. Hua, Golden touchstone: A comprehensive bilingual benchmark for evaluating fi- nancial large language models, arXiv preprint (2024). arXiv:2411.06272
2024
-
[154]
Subrahmanyam, Behavioural finance: A review and synthesis, European Financial Management 14 (1) (2008) 12–29
A. Subrahmanyam, Behavioural finance: A review and synthesis, European Financial Management 14 (1) (2008) 12–29
2008
-
[155]
Rubbaniy, A
G. Rubbaniy, A. A. Khalid, K. Syriopoulos, E. Polyzos, Dynamic returns connectedness: Portfolio hedging impli- cations during the COVID-19 pandemic and the Russia– Ukraine war, Journal of Futures Markets 44 (10) (2024) 1613–1639
2024
-
[156]
S. Li, H. Hoque, J. Liu, Investor sentiment and firm capital structure, Journal of Corporate Finance 80 (2023) 102426
2023
-
[157]
Karadima, H
M. Karadima, H. Louri, Economic policy uncertainty and non-performing loans: The moderating role of bank con- centration, Finance Research Letters 38 (2021) 101458
2021
-
[158]
Hodbod, S
A. Hodbod, S. Huber, K. Vasilev, Sectoral risk-weights and macroprudential policy, Journal of Banking and Fi- nance 112 (2020) 105336
2020
-
[159]
Danielsson, K
J. Danielsson, K. R. James, M. Valenzuela, I. Zer, Model risk of risk models, Journal of Financial Stability 23 (2016) 79–91
2016
-
[160]
Oehler, M
A. Oehler, M. Horn, Does chatgpt provide better advice than robo-advisors?, Finance Research Letters 60 (2024) 104898
2024
-
[161]
Dowling, B
M. Dowling, B. Lucey, Chatgpt for (finance) research: The bananarama conjecture, Finance Research Letters 53 (2023) 103662
2023
-
[162]
Polyzos, A
E. Polyzos, A. Fotiadis, T.-C. Huan, The asymmetric im- pact of Twitter sentiment and emotions: Impulse response analysis on European tourism firms using micro-data, Tourism Management 104 (2024) 104909
2024
-
[163]
Kalamara, A
E. Kalamara, A. Turrell, C. Redl, G. Kapetanios, S. Ka- padia, Making text count: economic forecasting using newspaper text, Journal of Applied Econometrics 37 (5) (2022) 896–919
2022
-
[164]
G. P. Herrera, M. Constantino, J.-J. Su, A. Naranpanawa, Renewable energy stocks forecast using twitter investor sentiment and deep learning, Energy Economics 114 (2022) 106285
2022
-
[165]
Dicks, P
D. Dicks, P. Fulghieri, Uncertainty, investor sentiment, and innovation, The Review of Financial Studies 34 (3) (2021) 1236–1279
2021
-
[166]
X.-Y . Liu, G. Wang, H. Yang, D. Zha, Fingpt: Democ- ratizing internet-scale data for financial large language models, arXiv preprint arXiv:2307.10485 (2023)
2023 arXiv
-
[167]
S. Wu, O. Irsoy, S. Lu, V . Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, G. Mann, Bloomberggpt: A large language model for finance, arXiv preprint (2023).arXiv:2303.17564
2023 arXiv
-
[168]
Y . Liu, N. Bu, Z. Li, Y . Zhang, Z. Zhao, At-fingpt: Fi- nancial risk prediction via an audio-text large language model, Finance Research Letters (2025) 106967
2025
-
[169]
E. Polyzos, Inflation and the war in Ukraine: Evidence using impulse response functions on economic indicators and twitter sentiment, Research in International Business and Finance 66 (2023) 102044
2023
-
[170]
S. Long, B. Lucey, Y . Xie, L. Yarovaya, I just like the stock, The role of Reddit sentiment in the GameStop share rally. Financial Review 58 (1) (2023) 19–37. 39
2023
-
[171]
Lyócsa, E
Š. Lyócsa, E. Baumöhl, T. Výrost, Yolo trading: Rid- ing with the herd during the gamestop episode, Finance Research Letters 46 (2022) 102359
2022
-
[172]
Polyzos, A
E. Polyzos, A. Samitas, I. Kampouris, Quantifying mar- ket efficiency: Information dissemination through social media, Available at SSRN 4082899 (2022)
2022
-
[173]
Vasileiou, Does the short squeeze lead to market ab- normality and antileverage effect? evidence from the gamestop case, Journal of Economic Studies 49 (8) (2022) 1360–1373
E. Vasileiou, Does the short squeeze lead to market ab- normality and antileverage effect? evidence from the gamestop case, Journal of Economic Studies 49 (8) (2022) 1360–1373
2022
-
[174]
A. H. Huang, H. Wang, Y . Yang, Finbert: A large lan- guage model for extracting information from financial text, Contemporary Accounting Research 40 (2) (2023) 806–841
2023
-
[175]
Loughran, B
T. Loughran, B. McDonald, When is a liability not a liability? textual analysis, dictionaries, and 10-ks, The Journal of Finance 66 (1) (2011) 35–65
2011
-
[176]
Ni¸ toi, M
M. Ni¸ toi, M. M. Pochea, ¸ S. C. Radu, Unveiling the senti- ment behind central bank narratives: A novel deep learn- ing index, Journal of Behavioral and Experimental Fi- nance 38 (2023) 100809. doi:10.1016/j.jbef.2023. 100809
2023 doi
-
[177]
Schimanski, A
T. Schimanski, A. Reding, N. Reding, J. A. Bingler, M. Kraus, M. Leippold, Bridging the gap in esg mea- surement: Using nlp to quantify environmental, social, and governance communication, Finance Research Let- ters 61 (2024) 104979
2024
-
[178]
Polyzos, A
E. Polyzos, A. Samitas, M.-S. Katsaiti, Who is unhappy for Brexit? a machine-learning, agent-based study on financial instability, International Review of Financial Analysis 72 (2020) 101590
2020
-
[179]
Polyzos, K
E. Polyzos, K. Abdulrahman, A. Christopoulos, Good management or good finances? An agent-based study on the causes of bank failure, Banks & Bank Systems 13, Iss. 3 (2018) 95–105
2018
-
[180]
Stevens Institute of Technology, Applying large language models to financial decision- making, https://www.stevens.edu/news/ applying-large-language-models-to-financial-decision-making , [Accessed 3 March 2025] (2023)
2023
-
[181]
J. Ye, G. Wang, Y . Li, Z. Deng, W. Li, T. Li, H. Duan, Z. Huang, Y . Su, B. Wang, et al., Gmai-mmbench: A com- prehensive multimodal evaluation benchmark towards general medical ai, Advances in Neural Information Pro- cessing Systems 37 (2024) 94327–94427
2024
-
[182]
J. Liu, W. Wang, Y . Su, J. Huan, W. Chen, Y . Zhang, C.-Y . Li, K.-J. Chang, X. Xin, L. Shen, et al., A spectrum evalu- ation benchmark for medical multi-modal large language models, arXiv preprint arXiv:2402.11217 (2024)
2024 arXiv
-
[183]
K. Keat, R. Venkatesh, Y . Huang, R. Kumar, S. Tuteja, K. Sangkuhl, B. Li, L. Gong, M. Whirl-Carrillo, T. E. Klein, et al., Pgxqa: A resource for evaluating llm perfor- mance for pharmacogenomic qa tasks, in: Biocomputing 2025: Proceedings of the Pacific Symposium, World Sci- ...
2025
-
[184]
H. Liu, H. Wang, Genotex: A benchmark for eval- uating llm-based exploration of gene expression data in alignment with bioinformaticians, arXiv preprint arXiv:2406.15341 (2024)
2024 arXiv
-
[185]
F. Bai, Y . Du, T. Huang, M. Q.-H. Meng, B. Zhao, M3d: Advancing 3d medical image analysis with multi-modal large language models, arXiv preprint arXiv:2404.00578 (2024)
2024 arXiv
-
[186]
P. R. A. S. Bassi, M. C. Yavuz, K. Wang, X. Chen, W. Li, S. Decherchi, A. Cavalli, Y . Yang, A. Yuille, Z. Zhou, Radgpt: Constructing 3d image-text tumor datasets (2025).arXiv:2501.04678. URLhttps://arxiv.org/abs/2501.04678
2025 arXiv
-
[187]
S. Mo, P. P. Liang, MultiMed: Massively Multimodal and Multitask Medical Understanding, arXiv preprint arXiv:2408.12682 (2024)
2024 arXiv
-
[188]
M. S. Sepehri, Z. Fabian, M. Soltanolkotabi, M. Soltanolkotabi, MediConfusion: Can you trust your AI radiologist? Probing the reliability of mul- timodal medical foundation models, arXiv preprint arXiv:2409.15477 (2024)
2024 arXiv
-
[189]
Neehal, B
N. Neehal, B. Wang, S. Debopadhaya, S. Dan, K. Muruge- san, V . Anand, K. P. Bennett, Ctbench: A comprehensive benchmark for evaluating language model capabilities in clinical trial design, arXiv preprint arXiv:2406.17888 (2024)
2024 arXiv
-
[190]
Khandekar, Q
N. Khandekar, Q. Jin, G. Xiong, S. Dunn, S. Apple- baum, Z. Anwar, M. Sarfo-Gyamfi, C. Safranek, A. An- war, A. Zhang, et al., Medcalc-bench: Evaluating large language models for medical calculations, Advances in Neural Information Processing Systems 37 (2024) 84730– 84745
2024
-
[191]
Jiang, P
J. Jiang, P. Chen, J. Wang, D. He, Z. Wei, L. Hong, L. Zong, S. Wang, Q. Yu, Z. Ma, et al., Benchmarking large language models on multiple tasks in bioinformat- ics nlp with prompting, arXiv preprint arXiv:2503.04013 (2025)
2025 arXiv
-
[192]
L. Chen, X. Han, S. Lin, H. Mai, H. Ran, Trimedlm: Advancing three-dimensional medical image analysis with multi-modal llm, in: 2024 IEEE International Con- ference on Bioinformatics and Biomedicine (BIBM), 2024, pp. 4505–4512. doi:10.1109/BIBM62325.2024. 10822809. 40
2024
-
[193]
Lozano, J
A. Lozano, J. Nirschl, J. Burgess, S. R. Gupte, Y . Zhang, A. Unell, S. Yeung, Micro-bench: A microscopy bench- mark for vision-language understanding, Advances in Neural Information Processing Systems 37 (2024) 30670– 30685
2024
-
[194]
Y . Sun, H. Wu, C. Zhu, S. Zheng, Q. Chen, K. Zhang, Y . Zhang, D. Wan, X. Lan, M. Zheng, et al., Pathmmu: A massive multimodal expert-level benchmark for un- derstanding and reasoning in pathology, in: European Conference on Computer Vision, Springer, 2024, pp. 56– 73
2024
-
[195]
Burgess, J
J. Burgess, J. J. Nirschl, L. Bravo-Sánchez, A. Lozano, S. R. Gupte, J. G. Galaz-Montoya, Y . Zhang, Y . Su, D. Bhowmik, Z. Coman, S. M. Hasan, A. Johannesson, W. D. Leineweber, M. G. Nair, R. Yarlagadda, C. Zuraski, W. Chiu, S. Cohen, J. N. Hansen, M. D. Leonetti, C. Liu, E. ...
2025 arXiv
-
[196]
Y . Chen, G. Wang, Y . Ji, Y . Li, J. Ye, T. Li, M. Hu, R. Yu, Y . Qiao, J. He, Slidechat: A large vision-language assistant for whole-slide pathology image understanding (2025).arXiv:2410.11761. URLhttps://arxiv.org/abs/2410.11761
2025 arXiv
-
[197]
Lozano, J
A. Lozano, J. Nirschl, J. Burgess, S. R. Gupte, Y . Zhang, A. Unell, S. Yeung-Levy, µ-bench: A vision-language benchmark for microscopy understanding (2024). arXiv: 2407.01791. URLhttps://arxiv.org/abs/2407.01791
2024 arXiv
-
[198]
X. Wang, D. Song, S. Chen, C. Zhang, B. Wang, LongLLaV A: Scaling Multi-modal LLMs to 1000 Im- ages Efficiently via a Hybrid Architecture, arXiv preprint arXiv:2409.02889 (2024)
2024
-
[199]
B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liu, et al., Llava- onevision: Easy visual task transfer, arXiv preprint arXiv:2408.03326 (2024)
2024 arXiv
-
[200]
Huang, L
S. Huang, L. Dong, W. Wang, Y . Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, Q. Liu, K. Ag- garwal, Z. Chi, J. Bjorck, V . Chaudhary, S. Som, X. Song, F. Wei, Language is not all you need: Aligning perception with language models (2023).arXiv:2302.14045. UR...
2023 arXiv
-
[201]
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, F. Wei, Kosmos-2: Grounding multimodal large language models to the world (2023).arXiv:2306.14824. URLhttps://arxiv.org/abs/2306.14824
2023 arXiv
-
[202]
T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al., Chatglm: A family of large language models from glm-130b to glm-4 all tools, arXiv preprint arXiv:2406.12793 (2024)
2024 arXiv
-
[203]
Y . Wang, Y . Liu, F. Yu, C. Huang, K. Li, Z. Wan, W. Che, H. Chen, Cvlue: A new benchmark dataset for chinese vision-language understanding evaluation, in: Proceed- ings of the AAAI Conference on Artificial Intelligence, V ol. 39, 2025, pp. 8196–8204
2025
-
[204]
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y . Cao, K. Narasimhan, Tree of thoughts: Deliberate problem solving with large language models, Advances in Neural Information Processing Systems 36 (2024)
2024
-
[205]
M. A. Arshad, T. Z. Jubery, T. Roy, R. Nassiri, A. K. Singh, A. Singh, C. Hegde, B. Ganapathysubramanian, A. Balu, A. Krishnamurthy, et al., Leveraging vision lan- guage models for specialized agricultural tasks, in: 2025 IEEE/CVF Winter Conference on Applications of Com- pute...
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.