Pith. sign in

Paper Citation Record · LEDGER

Holistic Evaluation of Language Models

As of 12 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 100 inbound Pith citation observations for arXiv:2211.09110.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.09110 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 100 of 346 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:54:33.121732Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch7

External citation measurements

16
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fd2b9ddb-6b51-4c64-b546-8f994c8bb60e · outbound

This paper cites Language Models are Few-Shot Learners.

Holistic Evaluation of Language Models Language Models are Few-Shot Learners

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.266423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2c69a3a5eeefb52e4fcd22669be482963798ee58e5cfda467f109cf088a5141d

Observation 4fc7afca-6d7e-46e3-af1d-835c2d1bc741 · outbound

This paper cites doi: 10.18653/v1/2021.acl-long.150.

Holistic Evaluation of Language Models doi: 10.18653/v1/2021.acl-long.150

Reference 2

Resolution
verified exact
doi, observed 2026-05-24T10:09:19.299324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:cfbd407a27ced653d6d8de65ca92bc5af05db6291558e3120939f9089f997482

Observation d1b20050-8ea6-4d4b-ba5b-c3abdc75b098 · outbound

This paper cites URLhttps://glottolog.org/accessed2021-08-08.

Holistic Evaluation of Language Models URLhttps://glottolog.org/accessed2021-08-08

Reference 3

Resolution
verified exact
doi, observed 2026-05-24T10:09:19.283766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:88b3055a94309b7c0b9d272b97c89e909a933bef2f56f62b03f139ab160ac1f2

Observation 7b7f0655-19dc-4471-b9be-85b3bb247e18 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Holistic Evaluation of Language Models Measuring Coding Challenge Competence With APPS

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.289641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:eef86e681db064a10244198df3d1bbc47e1877a9d5fc0096289f68038e393b74

Observation 8e60302d-56ca-4737-8cc1-67f5948ec5e3 · outbound

This paper cites In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp.

Holistic Evaluation of Language Models In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-05-24T10:09:19.254183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:7d01ad983088dcb2f47f8a877568471401dd0f22db07b09a454976e179e73ad5

Observation 333e3eee-7498-4e14-a12c-24d304c33015 · outbound

This paper cites Cognition , year =.

Holistic Evaluation of Language Models Cognition , year =

Reference 6

Resolution
metadata mismatch
doi, observed 2026-05-24T10:09:19.273599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:9165da2e640df5e252516acb6ab15c40c2144e560fe96f16528a19814c3118a4

Observation 655b68ab-0010-4a0d-9fbf-3528748e4909 · outbound

This paper cites The Natural Language Decathlon: Multitask Learning as Question Answering.

Holistic Evaluation of Language Models The Natural Language Decathlon: Multitask Learning as Question Answering

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.279889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:81f154bd8404f9723c4e08d3492c25d3aee9e018944a08f661f30b7ae932f436

Observation c7150b94-318f-4e30-8f44-5d369f022534 · outbound

This paper cites Red Teaming Language Models with Language Models.

Holistic Evaluation of Language Models Red Teaming Language Models with Language Models

Reference 8

Resolution
malformed identifier
local_arxiv, observed 2026-05-24T10:09:19.693997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:15ff901b8be6a1eb6e1172a9db624d6ebe6c5d1dc893677a05e05e913091cca5

Observation c7b09df6-a83c-4bda-affe-0aebbf696222 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Holistic Evaluation of Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.294997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:d803292a12922e3137e7358cd71bc9f70d3c0612b764604998e516a7240b8330

Observation 02724bfe-2e4b-42b2-beaf-c73880838e9c · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Holistic Evaluation of Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.247836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:c7f65d8385ff30f3ad198819ce9e5d53fd17814780afeb5db3bcea7d4a73c594

Observation bbe34219-0bfc-444e-9d46-7e8bfdb9d9c2 · outbound

This paper cites Yes” or “No.

Holistic Evaluation of Language Models Yes” or “No

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T10:09:19.260372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:ae020760b2cae5a7f8b802933b6e619d1015bc6ede32f806fb8500beed32458e

Observation 73b5c53d-ac5d-40d7-80f5-ae782b89d9e2 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.794025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:393eb3f9cbf08c5d6ffcb9d80f41494be3171dd9b51dda9b0af9d856e6e87528

Observation 99063e66-2f6c-4e86-9ca2-88dc4e348cd1 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.797403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:b351ed0f8520619b0a2f3adcb58c8a100797568fea77627bdaa1ecf9e45b588a

Observation d10ca844-e51b-418c-b639-d9e2b9f95339 · outbound

This paper cites Grandfather.

Holistic Evaluation of Language Models Grandfather

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.813945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:7c83df0d75a8f98fd6555181fd965f1acf202b4d79d0e85f85ea977e2167a089

Observation 85ae1069-f0e0-4dac-9230-449b76a51acf · outbound

This paper cites (2017), which derives its list form Greenwald et al.

Holistic Evaluation of Language Models (2017), which derives its list form Greenwald et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.805151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:00e671e052f4a2d91cd449bac9fc9127ca71529222173d98447de7c09417c00f

Observation edd232ba-c256-498d-8d6d-8b31e9d535ce · outbound

This paper cites (2018), which derives its list form Chalabi & Flowers (2017).

Holistic Evaluation of Language Models (2018), which derives its list form Chalabi & Flowers (2017)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.811250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:0fdde4698d69502add0037df224cc800ffdfa7408c60d3a6814c7d53e58d2127

Observation 20e4603d-6d17-4152-bf89-71769e843b82 · outbound

This paper cites It came from down here.

Holistic Evaluation of Language Models It came from down here

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.802573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:9eb953363f078878ad6983f06196d593a2d4e308de9c9278117b2ed3c5401423

Observation 740f62a9-1862-4a24-bf73-b9fb2a035933 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.800328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:de4844c28b66b1248454f8dd203569996499f11bfe69b95e6f3911e0b81b2bb6

Observation ce64fb07-b7be-4c50-8c71-8c8963c0a340 · outbound

This paper cites an unresolved cited work.

Holistic Evaluation of Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:19:19.807714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:a54c2df448211763c249bf68135ad706b26c19f951c5fbb850e3520dd1021290

Observation 37e2122b-5091-448f-890e-09361fe234a0 · outbound

This paper cites The capital of France is __.

Holistic Evaluation of Language Models The capital of France is __

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:19:19.791793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:07ba65a5faca5813ddf85fb9e0b41bc91708b767273a2aaf8c215bdbcac39e6b

Observation 95af2df3-d032-4c58-b63e-9a27e2ff1c79 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Holistic Evaluation of Language Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T10:09:19.667240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T10:06:30.552785Z digest=sha256:ccecc3c00e92717186bcade6ced662ffb4d76ef32c7ee5abef534059f3e19799

Pith citing papers

Observation 2a2e31f5-aeec-4826-b78d-e5c493f93d40 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Holistic Evaluation of Language Models

Reference 187

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.310415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:7fc77beb8570fc5b235f8f838e942fe10800619f37b0bb74aeae486a81146941

Observation 21df8cda-88d8-46fd-8b46-d991c83a6b65 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Holistic Evaluation of Language Models

Reference 266

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T00:51:11.617142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:27a4be33c2d6690fc264a9641543bbd995cd42b1972209a85d9f74aaa588638d

Observation b857d7ae-e7a1-4bad-b72b-3dc02cf82a60 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Holistic Evaluation of Language Models

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:19:46.406429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:27dbd6171a232cafde846bbb8083bc23a148c294cecc302577f1fc2bc7213a70

Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · inbound

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models cites this paper.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.393937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f4ff59b5b1e26a1ae7d32f4540028e5de43a735d2d23ccfe160d04ee5d389865

Observation 109acafb-4654-4b97-ade5-b7d1d0f50f30 · inbound

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting cites this paper.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting Holistic Evaluation of Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:01:19.833619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:91658abe4a4e8578252847be86101d69b6e3dc80a06c58fb1252a192e2fd1a9a

Observation 284311c4-f8f7-4c9e-9863-d59a231203bf · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Holistic Evaluation of Language Models

Reference 278

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:33:01.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:eeb7cc855befa51ea320d8ce2ddefa5c22198ffc418e6c44a8bd7fccef9313d4

Observation 2fac3c34-f5b5-44a0-95d9-1cfb280ef061 · inbound

Towards Expert-Level Medical Question Answering with Large Language Models cites this paper.

Towards Expert-Level Medical Question Answering with Large Language Models Holistic Evaluation of Language Models

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T04:32:33.596042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-24T04:32:33.271634Z digest=sha256:4908db48e2437062e6cb2ef3d3df6d6219a9cd0b6adecacc45ff19a6520b1e52

Observation ccaf4683-caab-4354-85e0-1b4d6528f6c1 · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Holistic Evaluation of Language Models

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:59:27.187433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:62c025ad0f9099eb4da7606f4d40e16d05f15ff2ff84d6843f673a293a188b70

Observation 3114b01c-f1c7-4dac-b5ad-6c9c791af106 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs Holistic Evaluation of Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:29:53.548060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:16099aceac7a895c26e14e9fad439478b67f64831ade194d487a73050aee0f29

Observation c50b9ac0-713b-4088-a30b-8a22fa424d79 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Holistic Evaluation of Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:35:21.244425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:40cf534f1d3861c60b10cecbabc9686a29eef603d02ae521b5d2c39ca1fd0fa7

Observation 4c4b2457-6247-403d-a0b9-40662455db54 · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Holistic Evaluation of Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.095035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:992642edb618cc5484dead757d50676b9ee88fee5837b48038e360b86cec79f0

Observation e7f06c6e-425d-4089-b4da-5e359e40c414 · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Holistic Evaluation of Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.427058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:d093461eb289fd7c15467690c7da67779e4d1178db5f83728de6f13a2abffd71

Observation 4db573e6-6781-430a-b62c-0d42808378ed · inbound

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment cites this paper.

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment Holistic Evaluation of Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:30:44.671076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T22:30:44.520703Z digest=sha256:75ee8902a624fa1d0920094f440d389bb74f6fd81e41dd7c7ba01e303d46868b

Observation 2276c250-7782-4a1b-9da1-8fc07679cd7c · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Holistic Evaluation of Language Models

Reference 115

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T15:46:03.400665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:262eedd281b275ecbaf269c62b359c7b1730816d3020030ac180ce925f02d44d

Observation f3811a93-4e13-4ad1-945e-4db5cd664d97 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Holistic Evaluation of Language Models

Reference 212

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:46:09.906965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:93d50dfb167b53757115bae24bef7659e250e0c1f3342f8f4f1a5da60e86e720

Observation dacd4082-fada-42c9-bcd6-eccb45ff3db4 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Holistic Evaluation of Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:17:08.428332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:e595cfa279ae21dd7338da93ba0c7fb00ee1841b72aa3f4fb851495ce7044df5

Observation 24054ef2-9599-4b38-be4d-b42b0b32f8c9 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Holistic Evaluation of Language Models

Reference 121

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:04:44.477071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:e7ed9f6b0281c1a19d75d0fd5cf94fb446664d164061ef116761f9927662c8a4

Observation 8a3635d8-5b20-4c17-9a12-330413744d0b · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Holistic Evaluation of Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.467354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:8c67b414e169814c149e3651d0cf82520231d6611cf4959840429653898dd006

Observation ca8c94d0-bb92-4d0e-ba17-fd0a9cad37e2 · inbound

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline cites this paper.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Holistic Evaluation of Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:03:38.982197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:66e2f63b6ac2a0b90c883e604990ee923b9d5d01d1badf8bbabe67d44aeebb08

Observation d4b8d1bf-352d-402b-b588-85accc1e953a · inbound

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models cites this paper.

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models Holistic Evaluation of Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:18:31.890032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T22:15:52.638622Z digest=sha256:deebcd57e8554d28baf46c95f989eeb9b4cfc29ec88292ee96c5fb4dd7399275

Observation c4e1c07a-6472-4e33-8bce-b595859eb664 · inbound

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance cites this paper.

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance Holistic Evaluation of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:54:33.121732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:54:33.121732Z digest=sha256:ecaa9ec3b3da713444b01ecba00cd182c648fe8e769f2c8131a16caefa97dad5

Observation 45e17965-2ffe-431c-bd3b-2ba97c179f1d · inbound

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models cites this paper.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.226325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.226325Z digest=sha256:c59c0a372cd025dd58977dfff257af121876928060c1164486ef9c067f8f1351

Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · inbound

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set cites this paper.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.413062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.413062Z digest=sha256:c266449363fe3d1bebbc06d11cb7bd04dbf20edfacba85000615428848f0fe06

Observation 38462fc9-d789-48ce-8a39-a724681605a7 · inbound

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity cites this paper.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Holistic Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.690101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.690101Z digest=sha256:a6e3442f5fe72fabee63e79a7d9878c50068294eaf368d6131d7cdd1f3f6270d

Observation e4b60683-06f7-44b2-9629-5734beb4a3c3 · inbound

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks cites this paper.

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks Holistic Evaluation of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:28:14.337637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:28:14.337637Z digest=sha256:1b5e2a9e50b4bc9ea7b42b500b985a3b4253ff9828d53bd9080862e659181844

Observation 7ffce77d-4554-48c2-b7a0-00122bd02305 · inbound

Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach cites this paper.

Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:21:33.895762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:21:33.895762Z digest=sha256:72e6b53bc88d3987c90eac270a0dffdbe4ccc141dec69fcff00211c9d4573ac3

Observation 4e7290a1-7f3f-4dad-81a5-8aaeb574c22d · inbound

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection cites this paper.

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection Holistic Evaluation of Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:56.136541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:56.136541Z digest=sha256:5a034a543915fd43f20da18c17fa1b23940ec8146213fcde65967c62cb17db11

Observation 63c82884-8e57-4f5a-8d98-b09919060252 · inbound

Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks cites this paper.

Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:55.117620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:55.117620Z digest=sha256:0fd3413c6a90509a81dc1f993be6925a1a431586d6e15848ed804cda7e9e263a

Observation 33d99f40-47cc-4c61-aea0-6812515a4d73 · inbound

Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective cites this paper.

Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:18:55.812296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:18:55.812296Z digest=sha256:234e4b247b7a8471e79855e045b2f714eea2efb39d6cb08e3cb6fd68a5fcb2d9

Observation 627e066d-cf78-46da-beef-dbc4afd38e81 · inbound

Best Practices for Large Language Models in Radiology cites this paper.

Best Practices for Large Language Models in Radiology Holistic Evaluation of Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:59.502032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:35:59.502032Z digest=sha256:603b3d09a1de57fa83479135035b5047448d0e3e977004510ce66e7dd7aced13

Observation ffd2df8d-b59f-4695-9cc1-eb5b565deece · inbound

Addressing Data Leakage in HumanEval Using Combinatorial Test Design cites this paper.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.483043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.483043Z digest=sha256:0fcc30d63960dafffc92249504f82893967ce89252fda6ae2ecafb6c6259ac13

Observation dc408e55-2172-4b16-ae57-05e086bfd276 · inbound

CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks cites this paper.

CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:02.813497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:02.813497Z digest=sha256:dba2610fed3e449eab6f926902fb0d0969bdbfe5630ea8ce861fe0af46a24fa6

Observation 396228b2-8607-415e-a327-5bec2dd767af · inbound

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization cites this paper.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.303089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.303089Z digest=sha256:6c6dfd10be285891a992b0df4532f0fd91752317765ad28323b5f2c2775a8ba8

Observation 496bbc29-aa5d-4013-9cc9-e5810ee50300 · inbound

C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation cites this paper.

C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation Holistic Evaluation of Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T21:10:35.151341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:10:35.151341Z digest=sha256:43e823930154cebe6ecbf5c63f702b7181398660871a2438b0397ed9294f73b3

Observation c29116bd-d715-461e-bc48-61ca2bbfff9a · inbound

Code LLMs: A Taxonomy-based Survey cites this paper.

Code LLMs: A Taxonomy-based Survey Holistic Evaluation of Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:57.173433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:57.173433Z digest=sha256:4851dee1f0e9bb5fa7b135e12d0a1b360e4a6233ffbcb847bd447f0988fbb1c2

Observation bd859e7b-4303-43e9-85be-3d014470063d · inbound

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model cites this paper.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.759469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.759469Z digest=sha256:8807b47adaee85f9385e2f29e163c3cd0ee4300f1d2914c8d06f3926fed49fdb

Observation 84169d9f-473d-4b0e-ae81-120693f2d837 · inbound

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks cites this paper.

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:21.846142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:19:21.846142Z digest=sha256:ae0f4369207758fc3dab4ff461011caa352c51b2cdb97faad6aa69870655fe26

Observation 3f8c4ce8-bd46-455b-b93b-5700c83e709a · inbound

Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws cites this paper.

Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws Holistic Evaluation of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:33.258973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:02:33.258973Z digest=sha256:fee32938ec1ca75745bb3fb5681458a6759a6f59b22a66b6d1022e5e13fb3698

Observation add45c71-8f27-4ea8-87d5-bd84b9256717 · inbound

PickLLM: Context-Aware RL-Assisted Large Language Model Routing cites this paper.

PickLLM: Context-Aware RL-Assisted Large Language Model Routing Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:29:02.773849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:29:02.773849Z digest=sha256:1120b3f55caeca5ba14b6854c07283e5d3732068ac4d9d0cb671c2a9136f9ed1

Observation bcdbf6ba-05ba-469f-aa78-396ca4958c90 · inbound

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation cites this paper.

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:26.892141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:26.892141Z digest=sha256:6710ea57093683671daa45b15feefae6502c779f62db74e58fb1fabbe329e077

Observation 9f689af3-fa83-48a2-b8ad-cb04e59ac950 · inbound

Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models cites this paper.

Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models Holistic Evaluation of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:09:29.199485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:09:29.199485Z digest=sha256:ba2bc4b376dcb7325587f594ef2857248f7ae402bb7037df93b1c2ec3f6375b9

Observation aaf64c9c-8a63-44e6-abab-2941bf075d34 · inbound

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models cites this paper.

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:48:02.339502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:48:02.339502Z digest=sha256:39a765e80205b9ee39cb0bc5e062715c2c05e751006636961ec90e4fdb9e2b3c

Observation f7e3189a-48ed-4eda-8eb0-f730cbd65023 · inbound

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases cites this paper.

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:29.802738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:56:29.802738Z digest=sha256:420af03d731e74fb7404954d1ef72d1413d9565a62d4465bf7bce5a489bfa481

Observation 24d4414f-df42-44b7-8b4f-fe720994635c · inbound

A Generative AI-driven Metadata Modelling Approach cites this paper.

A Generative AI-driven Metadata Modelling Approach Holistic Evaluation of Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:07.811499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:32:07.811499Z digest=sha256:64e71efb9ab654b90446af21273a92970b710f5e9b4ae31ca6575c9b91aa358e

Observation 11554f1f-f4e1-40dc-bc36-013fbacd976d · inbound

A Survey on Large Language Models with some Insights on their Capabilities and Limitations cites this paper.

A Survey on Large Language Models with some Insights on their Capabilities and Limitations Holistic Evaluation of Language Models

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:56.116010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:56.116010Z digest=sha256:049dabfc44742e807d92ba8ac28e7f3eeab7c75af4df480757df473dde16e56a

Observation 9d6439d2-e77b-4d61-a9ac-b786c99d9dec · inbound

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values cites this paper.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values Holistic Evaluation of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.736259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.736259Z digest=sha256:461e1d337f6b0981d7cb83d35131d65c5661df26f84f582084af5ab826fe5bda

Observation 32469f6b-cb2d-49f3-b768-e031bb9cdc87 · inbound

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG cites this paper.

Addressing the sustainable AI trilemma: a case study on LLM agents and RAG Holistic Evaluation of Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:34.841850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:34.841850Z digest=sha256:7a842a8b0797e34e19111f12cf08718c11a5c1ae46922894dfadff2e3b2e525f

Observation 9f47109e-10b6-477e-b1cf-8a0cc5bbacf4 · inbound

Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change cites this paper.

Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change Holistic Evaluation of Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:39:33.230512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:39:33.230512Z digest=sha256:5f0f4e4a1405308661762ed3466e978778ac913c38591558acb0ccc4af20f14f

Observation 5d7cc790-2d6c-4439-912f-fb955a6bec0b · inbound

RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models cites this paper.

RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models Holistic Evaluation of Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T04:37:31.933903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:36:32.897387Z digest=sha256:4572f01a02707c0d3fe02d88ae15fc12b035fba73717be932b10bf06ed832ecd

Observation 06999197-03dc-4598-b2eb-405ce3fdcf68 · inbound

Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences cites this paper.

Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences Holistic Evaluation of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T16:34:01.122641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:34:01.122641Z digest=sha256:d9ccef81cb1ec496b32b4a9fdb8baa17ac28b2813890753d60a394d64a4de20a

Observation e6887e6a-0877-46d1-8f8a-f439804b5877 · inbound

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing cites this paper.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.913059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.913059Z digest=sha256:bc7c0d467c7257461c5eef9ad9910890f2d83da4b55e5e9c555b1517a18352ed

Observation 706c8388-aaf6-4f52-ad7d-565da31aaef2 · inbound

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences cites this paper.

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences Holistic Evaluation of Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:52.861137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:51:52.861137Z digest=sha256:d40ac156b106313ed72af9fd45789a7fa9b7175168d741b1cefc2ac8e546ffc2

Observation 651ef3b7-6b2e-4988-ad39-5486161d1280 · inbound

MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers cites this paper.

MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T01:05:23.072022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:05:23.072022Z digest=sha256:9250bde4f0716a5657ef2024f93f20f2cb8d12044435a3be1f42fe6de8b06f51

Observation f7e73f1a-be3b-4574-869d-ebff3a8d7aaa · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Holistic Evaluation of Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.036681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.036681Z digest=sha256:71ec5cd84341b69015af9b79584442094e43de0204d189d17b703eb68b08e811

Observation eaeada13-0cc4-4762-9209-69f34cdc99f6 · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Holistic Evaluation of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.426227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.426227Z digest=sha256:a4868c43949b8beaae824f9c687c6346caed1b2fd96d97b2e69e60ca86852d3d

Observation e1aac824-a3ea-4db0-b811-de8ebc76936e · inbound

Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey cites this paper.

Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey Holistic Evaluation of Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:25.382826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:25.382826Z digest=sha256:a342cf8d0699f9a27ab662fbe611806302a5e942f2aa1fc5a03fb8f57861b8cd

Observation ce47f047-a3fb-4c34-a689-8de137f31e16 · inbound

White Hat Search Engine Optimization using Large Language Models cites this paper.

White Hat Search Engine Optimization using Large Language Models Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T13:10:56.377693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:10:56.377693Z digest=sha256:af82d35f80f58aceff63e45734e479792d0b461748826514ed15d8addae22bf5

Observation 6ccb9e5d-77dc-4605-97b2-3ff4ae39264b · inbound

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories cites this paper.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Holistic Evaluation of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.733056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.733056Z digest=sha256:1fb5111a71d34f4db603fdfcbf3800e93f387ae99dd8bf32f7101f1c9788d218

Observation 29d6a1a3-2b5b-4296-9a88-b5697a9071e9 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Holistic Evaluation of Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:33.096413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:33.096413Z digest=sha256:14335fd175d1ab3a06e63314ddd08ebf7587624a7e15f71bbc6468da607fe380

Observation 08eb5c24-8535-4b84-8652-abfed5201ad3 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.768611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.768611Z digest=sha256:2c0d0b8e30f6f5252fe90e9f6027372d30830cf2cceda30b6b03d26839e08219

Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.752311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.752311Z digest=sha256:09a08797f5570c0e942f1ac8800deef5583a74c48b818bf56026b490d43c4267

Observation 298370d4-fe6e-481a-ae59-4f4ed00a8fd4 · inbound

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages cites this paper.

Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages Holistic Evaluation of Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T19:19:38.204950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:19:38.204950Z digest=sha256:1241f84cd621763b23c5ca03f2e472123f16b9c1c18f87ae1e5dd22a63b5678c

Observation 53288b3d-354a-46b5-91bd-77643a3ce62d · inbound

Instruction-Based Fine-tuning of Open-Source LLMs for Predicting Customer Purchase Behaviors cites this paper.

Instruction-Based Fine-tuning of Open-Source LLMs for Predicting Customer Purchase Behaviors Holistic Evaluation of Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T10:07:33.771199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:07:33.771199Z digest=sha256:cf0b07feb1c3521f810885e38045e0598a0263b9c1a55066a3349f98e44e9a2c

Observation 7004c954-9449-48ce-8b1e-ba9edb649fcd · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Holistic Evaluation of Language Models

Reference 169

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:45.077542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:ecd571cd2ef20b8c13ed634e5f7499a6983d410fba81609ca5b1168d7be65626

Observation f0bcf5e2-1a4f-4abd-9852-fb73ae166fe1 · inbound

Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation cites this paper.

Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation Holistic Evaluation of Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T23:42:16.115559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T23:41:22.017848Z digest=sha256:53e8a3cd56d31cca5acce4c9d5abab0dce19d648de2870af9a77ff85088cd476

Observation 17351bac-c664-4485-a0ae-67855d85dc35 · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Holistic Evaluation of Language Models

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:36:58.792994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:10d5dfc5d82b5dff5766b1377034a38e17d8d03430693ac7432c3c87e3277e2a

Observation 7185838f-adde-4da3-89d5-65093b10910a · inbound

Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers cites this paper.

Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers Holistic Evaluation of Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:14:57.484055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T15:13:28.927880Z digest=sha256:37779bdfcb8be00885dba9e058ae903512cc90e8eac5f604b8b68da17c638ae4

Observation 0f962621-e721-4d8a-8038-118360d4715b · inbound

Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling cites this paper.

Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling Holistic Evaluation of Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:21:39.613433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T14:21:18.355360Z digest=sha256:96a1a4fda6a9a391989f9533dfdf2ea98d6c5af51ba71554397d22bfd4bc87a3

Observation 0167f08b-7e13-4d3f-a7c7-7424af1f6053 · inbound

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning cites this paper.

Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning Holistic Evaluation of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:13.198498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:12:13.198498Z digest=sha256:bdbdb542770dc9e9371f73caa2ba922918e8360c4d1652d6e93131a0e7a6a599

Observation 18a23f07-3b2a-4d2d-82e2-b3b24dd28cf7 · inbound

How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance cites this paper.

How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:00.526886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:00.526886Z digest=sha256:f1fdca4af0eac6b656af626b0a683889f877cd31b2748d7855add9dd4b989878

Observation 9e817d53-75a9-4742-972f-1035cb756a88 · inbound

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs cites this paper.

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs Holistic Evaluation of Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:22.022328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:22.022328Z digest=sha256:936b085bb00da4c853c039da17ca682d377f17c73a6ffcf3cee3342ee868aa96

Observation e760b2af-b696-4799-8cb7-c3d8d1e693a6 · inbound

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs cites this paper.

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs Holistic Evaluation of Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:34.517834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:34.517834Z digest=sha256:aa7f820e5b6a5d260360fcfca10f64dde4d6aede97f6990dd949bbba22b537ee

Observation 2fa802e0-ff55-4d3e-a7f4-23a95b019711 · inbound

REARANK: Reasoning Re-ranking Agent via Reinforcement Learning cites this paper.

REARANK: Reasoning Re-ranking Agent via Reinforcement Learning Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:12.880824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:07:12.880824Z digest=sha256:7d73fb5fc4f909807c37dd6b8495138b41747b2cb5a1c1631bcfc1830bb3098f

Observation 869f92eb-023a-433b-9a5d-87cc4520c761 · inbound

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts cites this paper.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Holistic Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.704962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.704962Z digest=sha256:dedab2fdcbd418469fde15aca101c3dc828e66bd8b6c9cfbbbc09f1384e29091

Observation ed57de1c-10d0-4c87-bbb7-59015f1042b5 · inbound

Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs cites this paper.

Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs Holistic Evaluation of Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.487420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:06.487420Z digest=sha256:cdb60a6f4e16bb40e00d879564b59b9e01a7864135ecec930f35b99505b3f1e5

Observation 651a6269-6e48-46a1-b3f1-48a56e9347e9 · inbound

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models cites this paper.

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models Holistic Evaluation of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:19.678444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:19.678444Z digest=sha256:9416d08ca16979b65ea5a8a46124099f0e52156265b4a7ff724ee42115482f66

Observation fcb0f193-1d3f-4d51-bf99-f699c382f548 · inbound

Human-Centric Evaluation for Foundation Models cites this paper.

Human-Centric Evaluation for Foundation Models Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.881359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.881359Z digest=sha256:17b1085725f8b4525756c2880410194bb1ebf52fce84b210461c3a4a94cbbb07

Observation 1ace072c-7f89-43fc-97c8-da9aef52b629 · inbound

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs cites this paper.

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs Holistic Evaluation of Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:46.171747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:46.171747Z digest=sha256:aa2b52cfc6a74dc6312ee3603fd3a4dd535c937b48f1297d9c1d888977d2e2c4

Observation 734ea006-f8da-425d-8a43-001721b94a15 · inbound

AetherVision-Bench: An Open-Vocabulary RGB-Infrared Benchmark for Multi-Angle Segmentation across Aerial and Ground Perspectives cites this paper.

AetherVision-Bench: An Open-Vocabulary RGB-Infrared Benchmark for Multi-Angle Segmentation across Aerial and Ground Perspectives Holistic Evaluation of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:54.811123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:54.811123Z digest=sha256:e1b2cbb05e564b8c7e71e984cc0db987a13a6d8dc28419c623276a6cac754132

Observation a6e7e1dd-fb91-4302-abe1-0fe58aa54d11 · inbound

Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures cites this paper.

Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures Holistic Evaluation of Language Models

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:34.893274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:54:34.893274Z digest=sha256:e92dd17e75e6aaf8ae27cc546e5cf47cd27ed4c2ca4ec6b78eb8a5929f216063

Observation 0e50d866-5b5a-462c-ba6a-267b6524b1a1 · inbound

Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs cites this paper.

Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs Holistic Evaluation of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:37.169704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:37.169704Z digest=sha256:d64553ff5e378498a495586e2489a1f57ec9b4c0fb671617b46289c9df4efcfd

Observation a07129ff-b59a-4975-a41d-cda583b50758 · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Holistic Evaluation of Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.034053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.034053Z digest=sha256:2dc4aa0c18ad064f2e1e5d70ed0030bdaa6e38aa237b305eb9b779de677ac0b7

Observation 4e25f433-5a68-4b69-b0ec-849b35998983 · inbound

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation cites this paper.

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation Holistic Evaluation of Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.736114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.736114Z digest=sha256:c05d17d90c099732650ab1be6453ace9c9c51b63a3427485e509a63756f097b2

Observation 45d531a0-0a4f-4002-ba70-99ca07e15346 · inbound

Metritocracy: Representative Metrics for Lite Benchmarks cites this paper.

Metritocracy: Representative Metrics for Lite Benchmarks Holistic Evaluation of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:34.163425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:50:34.163425Z digest=sha256:81ae9285d9565d6e1a6e0238a5b1f43fbdb34dde322bdce7bd6264fd81c6eb13

Observation 6027413f-2408-4b54-b3b3-a7bddedb0f2a · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications Holistic Evaluation of Language Models

Reference 149

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:20.489334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:20.489334Z digest=sha256:fe74bb4141910df9ed6b2e0c62c4b7ef21cc178b3852e7e6776c5b15e028817a

Observation 3a9f4759-c3dc-40ae-b1c3-776290aa4053 · inbound

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law cites this paper.

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law Holistic Evaluation of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.924943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.924943Z digest=sha256:8360cf2fd5762f65af7358a95405d4e92a460e4f08eae97ccf4c3e3fa7d4e7c6

Observation 9b4e5a4b-cb9c-4437-9fbb-da517ee04f3b · inbound

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models cites this paper.

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models Holistic Evaluation of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:55.961100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:31:55.961100Z digest=sha256:eccd705a8c77cbf74ce2b4313460b118cadeb758c75f5a9e2064f3094aca292e

Observation 51ec63c1-4f0b-4b96-94f9-a00835c174b3 · inbound

The NordDRG AI Benchmark for Large Language Models cites this paper.

The NordDRG AI Benchmark for Large Language Models Holistic Evaluation of Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:18.042981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:47:18.042981Z digest=sha256:f5693b5aad188066466597ad1cdc928820d3547d787a3874f2daf62e2b88e16a

Observation c09e3ada-863a-4154-a90a-278abcdf75a1 · inbound

Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs cites this paper.

Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs Holistic Evaluation of Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:55.366632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:55.366632Z digest=sha256:355797e729bd89103d676d947bfbeceeb8a27711132630fb8e61afcd27f67327

Observation d4c7ccc0-ea19-4d32-8541-39d64b58b2c4 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark Holistic Evaluation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:31.015321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:31.015321Z digest=sha256:135bb8f04040c5a2901e15977682f43fb26fd5592d6d20c461f41c84fd89fe1d

Observation 34dd9cdc-8b85-4a39-9792-2ecfad2c4309 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models Holistic Evaluation of Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.910951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.910951Z digest=sha256:b26a4134234a4958516e0544b1d3a7885b2eec4f1407656b6195fe5b1efc9016

Observation e4365fa1-c857-42c8-b687-24f496d5e057 · inbound

From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models cites this paper.

From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models Holistic Evaluation of Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:29:48.278185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:29:48.278185Z digest=sha256:797a1891cde58a0013cd8e0c9163a2dc2aca216f46648fe7d1e9a41677a6f866

Observation 4f5c1a77-f935-4365-893d-707fbd1e56f6 · inbound

JointRank: Rank Large Set with Single Pass cites this paper.

JointRank: Rank Large Set with Single Pass Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:08.918938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:08.918938Z digest=sha256:7a4aa538063aafaf16b0401e665ce2706a5201fc70abe98915587677b55717e2

Observation 927ef670-9513-4afa-9905-0015758b138f · inbound

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments cites this paper.

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments Holistic Evaluation of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:02.598843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:02.598843Z digest=sha256:a393d69c8eca2c4524c8ee322cee046cd3a8d5521cba25feaabab548836248b5

Observation 48f7a3a2-f019-4284-862a-59d997839c37 · inbound

PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies? cites this paper.

PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies? Holistic Evaluation of Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:51.600014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:51.600014Z digest=sha256:02a33aaedaeb956e9dad8405a9f32034b053281034e7e2283bc7278976d20879

Observation 2b2d86e8-0460-49ad-9524-a01d06d5ddaa · inbound

Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions cites this paper.

Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions Holistic Evaluation of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:43.336489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:43.336489Z digest=sha256:2bc8dbee96f1f27f4cd0735409fedce8d7b18262c3592871898f0e6b29ca2a97

Observation 4a4bf49c-402e-4618-9f6d-3ce0e18e69a8 · inbound

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances cites this paper.

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances Holistic Evaluation of Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:46.292328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:47:46.292328Z digest=sha256:9dd15f04487ecec11680ef5f71979c06b1bdd5ff83739b67785bdd0b041cfa4f

Observation 38c5ee0f-35ef-4ac2-a32a-516d5b1947a7 · inbound

Harnessing Pairwise Ranking Prompting Through Sample-Efficient Ranking Distillation cites this paper.

Harnessing Pairwise Ranking Prompting Through Sample-Efficient Ranking Distillation Holistic Evaluation of Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:57.117708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:57.117708Z digest=sha256:de8a4a739badae0444c79f1a257bc0af2311b62a680ae6aa6bd41a8665bd5462

Observation a50514e0-674e-4837-929c-1719d1a01b66 · inbound

ASSURE: Metamorphic Testing for AI-powered Browser Extensions cites this paper.

ASSURE: Metamorphic Testing for AI-powered Browser Extensions Holistic Evaluation of Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:28.417193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:28.417193Z digest=sha256:069b22ec0686e048eac051976065ad5b7585f4fef5ac090ef7a38f163f4abc22

Observation db0697e7-fef1-419d-8dab-d1b460a130e7 · inbound

Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models cites this paper.

Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models Holistic Evaluation of Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:26:58.398605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:26:58.398605Z digest=sha256:5d262d8387733b1888112bc40db30bdebfb78dfa2424516750fa0e31d47df688