Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 100 inbound Pith citation observations for arXiv:2211.09110.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:54:33.121732Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
21 of 21 outbound references displayed
External citation measurements
16
pith, observed 2026-08-05T02:28:24.338817Z
Observation fd2b9ddb-6b51-4c64-b546-8f994c8bb60e · outbound
Holistic Evaluation of Language Models Language Models are Few-Shot Learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4fc7afca-6d7e-46e3-af1d-835c2d1bc741 · outbound
Holistic Evaluation of Language Models doi: 10.18653/v1/2021.acl-long.150
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d1b20050-8ea6-4d4b-ba5b-c3abdc75b098 · outbound
Holistic Evaluation of Language Models URLhttps://glottolog.org/accessed2021-08-08
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7b7f0655-19dc-4471-b9be-85b3bb247e18 · outbound
Holistic Evaluation of Language Models Measuring Coding Challenge Competence With APPS
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e60302d-56ca-4737-8cc1-67f5948ec5e3 · outbound
Holistic Evaluation of Language Models In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 333e3eee-7498-4e14-a12c-24d304c33015 · outbound
Holistic Evaluation of Language Models Cognition , year =
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 655b68ab-0010-4a0d-9fbf-3528748e4909 · outbound
Holistic Evaluation of Language Models The Natural Language Decathlon: Multitask Learning as Question Answering
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7150b94-318f-4e30-8f44-5d369f022534 · outbound
Holistic Evaluation of Language Models Red Teaming Language Models with Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7b09df6-a83c-4bda-affe-0aebbf696222 · outbound
Holistic Evaluation of Language Models Measuring and Narrowing the Compositionality Gap in Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 02724bfe-2e4b-42b2-beaf-c73880838e9c · outbound
Holistic Evaluation of Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bbe34219-0bfc-444e-9d46-7e8bfdb9d9c2 · outbound
Holistic Evaluation of Language Models Yes” or “No
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73b5c53d-ac5d-40d7-80f5-ae782b89d9e2 · outbound
Holistic Evaluation of Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99063e66-2f6c-4e86-9ca2-88dc4e348cd1 · outbound
Holistic Evaluation of Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d10ca844-e51b-418c-b639-d9e2b9f95339 · outbound
Holistic Evaluation of Language Models Grandfather
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 85ae1069-f0e0-4dac-9230-449b76a51acf · outbound
Holistic Evaluation of Language Models (2017), which derives its list form Greenwald et al
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation edd232ba-c256-498d-8d6d-8b31e9d535ce · outbound
Holistic Evaluation of Language Models (2018), which derives its list form Chalabi & Flowers (2017)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20e4603d-6d17-4152-bf89-71769e843b82 · outbound
Holistic Evaluation of Language Models It came from down here
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 740f62a9-1862-4a24-bf73-b9fb2a035933 · outbound
Holistic Evaluation of Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ce64fb07-b7be-4c50-8c71-8c8963c0a340 · outbound
Holistic Evaluation of Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37e2122b-5091-448f-890e-09361fe234a0 · outbound
Holistic Evaluation of Language Models The capital of France is __
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 95af2df3-d032-4c58-b63e-9a27e2ff1c79 · outbound
Holistic Evaluation of Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a2e31f5-aeec-4826-b78d-e5c493f93d40 · inbound
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Holistic Evaluation of Language Models
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 21df8cda-88d8-46fd-8b46-d991c83a6b65 · inbound
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Holistic Evaluation of Language Models
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b857d7ae-e7a1-4bad-b72b-3dc02cf82a60 · inbound
BloombergGPT: A Large Language Model for Finance Holistic Evaluation of Language Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · inbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 109acafb-4654-4b97-ade5-b7d1d0f50f30 · inbound
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting Holistic Evaluation of Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 284311c4-f8f7-4c9e-9863-d59a231203bf · inbound
StarCoder: may the source be with you! Holistic Evaluation of Language Models
Reference 278
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fac3c34-f5b5-44a0-95d9-1cfb280ef061 · inbound
Towards Expert-Level Medical Question Answering with Large Language Models Holistic Evaluation of Language Models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ccaf4683-caab-4354-85e0-1b4d6528f6c1 · inbound
PaLM 2 Technical Report Holistic Evaluation of Language Models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3114b01c-f1c7-4dac-b5ad-6c9c791af106 · inbound
QLoRA: Efficient Finetuning of Quantized LLMs Holistic Evaluation of Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c50b9ac0-713b-4088-a30b-8a22fa424d79 · inbound
Scaling Data-Constrained Language Models Holistic Evaluation of Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c4b2457-6247-403d-a0b9-40662455db54 · inbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Holistic Evaluation of Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7f06c6e-425d-4089-b4da-5e359e40c414 · inbound
H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Holistic Evaluation of Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4db573e6-6781-430a-b62c-0d42808378ed · inbound
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment Holistic Evaluation of Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2276c250-7782-4a1b-9da1-8fc07679cd7c · inbound
GAIA: a benchmark for General AI Assistants Holistic Evaluation of Language Models
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f3811a93-4e13-4ad1-945e-4db5cd664d97 · inbound
The Falcon Series of Open Language Models Holistic Evaluation of Language Models
Reference 212
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dacd4082-fada-42c9-bcd6-eccb45ff3db4 · inbound
TrustLLM: Trustworthiness in Large Language Models Holistic Evaluation of Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 24054ef2-9599-4b38-be4d-b42b0b32f8c9 · inbound
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Holistic Evaluation of Language Models
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a3635d8-5b20-4c17-9a12-330413744d0b · inbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Holistic Evaluation of Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ca8c94d0-bb92-4d0e-ba17-fd0a9cad37e2 · inbound
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Holistic Evaluation of Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4b8d1bf-352d-402b-b588-85accc1e953a · inbound
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models Holistic Evaluation of Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4e1c07a-6472-4e33-8bce-b595859eb664 · inbound
GPAI Evaluations Standards Taskforce: Towards Effective AI Governance Holistic Evaluation of Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e17965-2ffe-431c-bd3b-2ba97c179f1d · inbound
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Holistic Evaluation of Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · inbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38462fc9-d789-48ce-8a39-a724681605a7 · inbound
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Holistic Evaluation of Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b60683-06f7-44b2-9629-5734beb4a3c3 · inbound
Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks Holistic Evaluation of Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffce77d-4554-48c2-b7a0-00122bd02305 · inbound
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach Holistic Evaluation of Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7290a1-7f3f-4dad-81a5-8aaeb574c22d · inbound
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection Holistic Evaluation of Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c82884-8e57-4f5a-8d98-b09919060252 · inbound
Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks Holistic Evaluation of Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d99f40-47cc-4c61-aea0-6812515a4d73 · inbound
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective Holistic Evaluation of Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627e066d-cf78-46da-beef-dbc4afd38e81 · inbound
Best Practices for Large Language Models in Radiology Holistic Evaluation of Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd2df8d-b59f-4695-9cc1-eb5b565deece · inbound
Addressing Data Leakage in HumanEval Using Combinatorial Test Design Holistic Evaluation of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc408e55-2172-4b16-ae57-05e086bfd276 · inbound
CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks Holistic Evaluation of Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396228b2-8607-415e-a327-5bec2dd767af · inbound
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Holistic Evaluation of Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496bbc29-aa5d-4013-9cc9-e5810ee50300 · inbound
C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation Holistic Evaluation of Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29116bd-d715-461e-bc48-61ca2bbfff9a · inbound
Code LLMs: A Taxonomy-based Survey Holistic Evaluation of Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd859e7b-4303-43e9-85be-3d014470063d · inbound
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Holistic Evaluation of Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84169d9f-473d-4b0e-ae81-120693f2d837 · inbound
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks Holistic Evaluation of Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8c4ce8-bd46-455b-b93b-5700c83e709a · inbound
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws Holistic Evaluation of Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add45c71-8f27-4ea8-87d5-bd84b9256717 · inbound
PickLLM: Context-Aware RL-Assisted Large Language Model Routing Holistic Evaluation of Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdbf6ba-05ba-469f-aa78-396ca4958c90 · inbound
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation Holistic Evaluation of Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f689af3-fa83-48a2-b8ad-cb04e59ac950 · inbound
Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models Holistic Evaluation of Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf64c9c-8a63-44e6-abab-2941bf075d34 · inbound
INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models Holistic Evaluation of Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e3189a-48ed-4eda-8eb0-f730cbd65023 · inbound
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases Holistic Evaluation of Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d4414f-df42-44b7-8b4f-fe720994635c · inbound
A Generative AI-driven Metadata Modelling Approach Holistic Evaluation of Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11554f1f-f4e1-40dc-bc36-013fbacd976d · inbound
A Survey on Large Language Models with some Insights on their Capabilities and Limitations Holistic Evaluation of Language Models
Reference 190
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6439d2-e77b-4d61-a9ac-b786c99d9dec · inbound
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Holistic Evaluation of Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32469f6b-cb2d-49f3-b768-e031bb9cdc87 · inbound
Addressing the sustainable AI trilemma: a case study on LLM agents and RAG Holistic Evaluation of Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f47109e-10b6-477e-b1cf-8a0cc5bbacf4 · inbound
Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change Holistic Evaluation of Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7cc790-2d6c-4439-912f-fb955a6bec0b · inbound
RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models Holistic Evaluation of Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06999197-03dc-4598-b2eb-405ce3fdcf68 · inbound
Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences Holistic Evaluation of Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6887e6a-0877-46d1-8f8a-f439804b5877 · inbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Holistic Evaluation of Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706c8388-aaf6-4f52-ad7d-565da31aaef2 · inbound
Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences Holistic Evaluation of Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651ef3b7-6b2e-4988-ad39-5486161d1280 · inbound
MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers Holistic Evaluation of Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e73f1a-be3b-4574-869d-ebff3a8d7aaa · inbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Holistic Evaluation of Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaeada13-0cc4-4762-9209-69f34cdc99f6 · inbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Holistic Evaluation of Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1aac824-a3ea-4db0-b811-de8ebc76936e · inbound
Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey Holistic Evaluation of Language Models
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce47f047-a3fb-4c34-a689-8de137f31e16 · inbound
White Hat Search Engine Optimization using Large Language Models Holistic Evaluation of Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ccb9e5d-77dc-4605-97b2-3ff4ae39264b · inbound
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Holistic Evaluation of Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d6a1a3-2b5b-4296-9a88-b5697a9071e9 · inbound
Salamandra Technical Report Holistic Evaluation of Language Models
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08eb5c24-8535-4b84-8652-abfed5201ad3 · inbound
Measuring Diversity in Synthetic Datasets Holistic Evaluation of Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · inbound
The Science of Evaluating Foundation Models Holistic Evaluation of Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 298370d4-fe6e-481a-ae59-4f4ed00a8fd4 · inbound
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages Holistic Evaluation of Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53288b3d-354a-46b5-91bd-77643a3ce62d · inbound
Instruction-Based Fine-tuning of Open-Source LLMs for Predicting Customer Purchase Behaviors Holistic Evaluation of Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7004c954-9449-48ce-8b1e-ba9edb649fcd · inbound
Towards an AI co-scientist Holistic Evaluation of Language Models
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f0bcf5e2-1a4f-4abd-9852-fb73ae166fe1 · inbound
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation Holistic Evaluation of Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 17351bac-c664-4485-a0ae-67855d85dc35 · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives Holistic Evaluation of Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7185838f-adde-4da3-89d5-65093b10910a · inbound
Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers Holistic Evaluation of Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0f962621-e721-4d8a-8038-118360d4715b · inbound
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling Holistic Evaluation of Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0167f08b-7e13-4d3f-a7c7-7424af1f6053 · inbound
Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning Holistic Evaluation of Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a23f07-3b2a-4d2d-82e2-b3b24dd28cf7 · inbound
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance Holistic Evaluation of Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e817d53-75a9-4742-972f-1035cb756a88 · inbound
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Holistic Evaluation of Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e760b2af-b696-4799-8cb7-c3d8d1e693a6 · inbound
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs Holistic Evaluation of Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa802e0-ff55-4d3e-a7f4-23a95b019711 · inbound
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning Holistic Evaluation of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869f92eb-023a-433b-9a5d-87cc4520c761 · inbound
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Holistic Evaluation of Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed57de1c-10d0-4c87-bbb7-59015f1042b5 · inbound
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs Holistic Evaluation of Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651a6269-6e48-46a1-b3f1-48a56e9347e9 · inbound
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models Holistic Evaluation of Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb0f193-1d3f-4d51-bf99-f699c382f548 · inbound
Human-Centric Evaluation for Foundation Models Holistic Evaluation of Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ace072c-7f89-43fc-97c8-da9aef52b629 · inbound
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs Holistic Evaluation of Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734ea006-f8da-425d-8a43-001721b94a15 · inbound
AetherVision-Bench: An Open-Vocabulary RGB-Infrared Benchmark for Multi-Angle Segmentation across Aerial and Ground Perspectives Holistic Evaluation of Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e7e1dd-fb91-4302-abe1-0fe58aa54d11 · inbound
Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures Holistic Evaluation of Language Models
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e50d866-5b5a-462c-ba6a-267b6524b1a1 · inbound
Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs Holistic Evaluation of Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07129ff-b59a-4975-a41d-cda583b50758 · inbound
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Holistic Evaluation of Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e25f433-5a68-4b69-b0ec-849b35998983 · inbound
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation Holistic Evaluation of Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45d531a0-0a4f-4002-ba70-99ca07e15346 · inbound
Metritocracy: Representative Metrics for Lite Benchmarks Holistic Evaluation of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6027413f-2408-4b54-b3b3-a7bddedb0f2a · inbound
A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications Holistic Evaluation of Language Models
Reference 149
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9f4759-c3dc-40ae-b1c3-776290aa4053 · inbound
Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law Holistic Evaluation of Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4e5a4b-cb9c-4437-9fbb-da517ee04f3b · inbound
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models Holistic Evaluation of Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ec63c1-4f0b-4b96-94f9-a00835c174b3 · inbound
The NordDRG AI Benchmark for Large Language Models Holistic Evaluation of Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09e3ada-863a-4154-a90a-278abcdf75a1 · inbound
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs Holistic Evaluation of Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c7ccc0-ea19-4d32-8541-39d64b58b2c4 · inbound
Enterprise Large Language Model Evaluation Benchmark Holistic Evaluation of Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34dd9cdc-8b85-4a39-9792-2ecfad2c4309 · inbound
Potemkin Understanding in Large Language Models Holistic Evaluation of Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4365fa1-c857-42c8-b687-24f496d5e057 · inbound
From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models Holistic Evaluation of Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5c1a77-f935-4365-893d-707fbd1e56f6 · inbound
JointRank: Rank Large Set with Single Pass Holistic Evaluation of Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927ef670-9513-4afa-9905-0015758b138f · inbound
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments Holistic Evaluation of Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f7a3a2-f019-4284-862a-59d997839c37 · inbound
PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies? Holistic Evaluation of Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2d86e8-0460-49ad-9524-a01d06d5ddaa · inbound
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions Holistic Evaluation of Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4bf49c-402e-4618-9f6d-3ce0e18e69a8 · inbound
Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances Holistic Evaluation of Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c5ee0f-35ef-4ac2-a32a-516d5b1947a7 · inbound
Harnessing Pairwise Ranking Prompting Through Sample-Efficient Ranking Distillation Holistic Evaluation of Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a50514e0-674e-4837-929c-1719d1a01b66 · inbound
ASSURE: Metamorphic Testing for AI-powered Browser Extensions Holistic Evaluation of Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0697e7-fef1-419d-8dab-d1b460a130e7 · inbound
Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models Holistic Evaluation of Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.