Pith. sign in

Paper Citation Record · LEDGER

AHELM: A Holistic Evaluation of Audio-Language Models

As of 8 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 10 inbound Pith citation observations for arXiv:2508.21376.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21376 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:22:34.997009Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:52:45.801592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:52.874154Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9781f1e-d579-44da-aab7-1f362649c634 · outbound

This paper cites GPT-4 Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.697071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.697071Z digest=sha256:1f39e4dbbb190996cde2c100bee335bdaf2493cf0cca7c5eec63dd645c356072

Observation 51e0b0aa-3846-40c3-8ef3-af0dd4b33398 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models The Claude 3 model family: Opus, Sonnet, Haiku, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.739040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:33.778200Z digest=sha256:5a1291af155c2d3c64fcaa6ff670c7332dadca05be79fe2d008675811493dd17

Observation a2426604-3b6c-4168-a590-8cb736b79234 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

AHELM: A Holistic Evaluation of Audio-Language Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.923565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.923565Z digest=sha256:528f465c21c45bf5e3cf7f74ea9e87ff84a4b144e79ab072f0e2252053d46b1b

Observation 2602b30f-96dd-4597-8e93-c1a7b37ea665 · outbound

This paper cites Qwen Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.102234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.102234Z digest=sha256:b743de03b1f7f33e22be6b39e6c21691a04c2033abafad42f8c63144e3eeeb4d

Observation 6b69d948-705f-4460-9ef4-0a4967a3ff44 · outbound

This paper cites Towards multimodal sarcasm detection (an _obviously_ perfect paper).

AHELM: A Holistic Evaluation of Audio-Language Models Towards multimodal sarcasm detection (an _obviously_ perfect paper)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.731229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.243183Z digest=sha256:dab94674e0c360a687a416086b36f39ee9380e8b60f375d0ba6a0946c8f47e48

Observation 2b74c9ce-de68-4f7c-b1c5-70410f7cdd73 · outbound

This paper cites Qwen2-Audio Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.352067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.352067Z digest=sha256:5a211dc322d553b4c6287afcd73168387ca53b6d5092bf2b3694b06d1d3f946f

Observation 41cacf6c-12bb-4b8c-a526-e4abebb8b96f · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

AHELM: A Holistic Evaluation of Audio-Language Models VoxCeleb2: Deep Speaker Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.385152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.385152Z digest=sha256:d4a667063b910680983725ef66b1e6ccf2d7a43c4c443dcf07db8bbfd14ae65f

Observation e4269de4-646a-4529-86dc-f4865d5969a2 · outbound

This paper cites Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit.

AHELM: A Holistic Evaluation of Audio-Language Models Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.722887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.540549Z digest=sha256:ad16ff16ab2fa486ca956836db365380ebefa89d324226fbe08c02f8677064f5

Observation 38f7207c-cba3-4f5a-9765-493bc4c10f6b · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech.

AHELM: A Holistic Evaluation of Audio-Language Models FLEURS: Few-shot learning evaluation of universal representations of speech

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.714660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.635896Z digest=sha256:e5c4d51eade3973b44ea3bc41fc3cf4638593ad695cc734b3ffa1353a8b6cf9f

Observation c7aecf5c-9e8e-491b-8fa9-a0115a3c3be6 · outbound

This paper cites MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector.

AHELM: A Holistic Evaluation of Audio-Language Models MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:22:35.199235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.753589Z digest=sha256:3fcb856e46e12e16120b8c9d809953995f90ad77e7a3f8d8759353e2d1bdeab9

Observation d728dde4-d52a-4a2f-b473-9ccde920fbd1 · outbound

This paper cites Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition.

AHELM: A Holistic Evaluation of Audio-Language Models Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.706024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.775890Z digest=sha256:be14df8d6328bb9ee3ce2656a65bafd0b83b0e0328a7ff5d5b30bfd6fd8eb055

Observation fe123ce9-b696-4805-9234-726bc74cbbe1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

AHELM: A Holistic Evaluation of Audio-Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.779733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.779733Z digest=sha256:55c44005e23711bcaaf14f583df5c9a4ae983fa226d125612590357b98f4dd1b

Observation 9a8d9c12-d1d7-4781-8e23-3bd439c1ee47 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

AHELM: A Holistic Evaluation of Audio-Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.697626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.783925Z digest=sha256:8e6302cabb5ecfec5f0d5d81bf35edd12eda07c2a7eb06c1efdd0c67a41f8dd7

Observation 02a1fa96-c5ad-42d8-b7c0-252adacb5815 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

AHELM: A Holistic Evaluation of Audio-Language Models Clap learning audio concepts from natural language supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.689202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.787398Z digest=sha256:840b01fbaddefd1be90a805b2e052be4b47c36507e5255582d32d8d6f5a3867f

Observation 6b2cf18e-ff38-4505-a3d2-fc3396052351 · outbound

This paper cites Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images.

AHELM: A Holistic Evaluation of Audio-Language Models Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:22:35.170240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.790680Z digest=sha256:44d828a387d6af647eb9b0a058964561f1c3779e4f10d2b052d65f8aca11e748

Observation dd5a1352-ec15-4aed-98d6-51d73912d090 · outbound

This paper cites CSR-I (WSJ0) Complete.

AHELM: A Holistic Evaluation of Audio-Language Models CSR-I (WSJ0) Complete

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.680312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.794039Z digest=sha256:be2b1b8c03a71717d73d9bae5ca5b22bef70efb3adac6b74c787d708c2793fbd

Observation a2cebc79-52ca-4f7e-bb5c-7b9e40df3c96 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.797552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.797552Z digest=sha256:53a37bab7e400725f5165fd0395637e023d24a304d17d7ae0c944e5a131139d0

Observation 80c80121-1169-4858-a503-2b0fc6c46247 · outbound

This paper cites Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities.

AHELM: A Holistic Evaluation of Audio-Language Models Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.669917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.801817Z digest=sha256:af3cd7f29f93962387ab7c80b0df0d935f282206954159173abc21366f3b7ec2

Observation 0afda4f7-ceff-44f8-9dc6-e6014bebf7d0 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

AHELM: A Holistic Evaluation of Audio-Language Models V ocalsound: A dataset for improving human vocal sounds recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.661754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.805468Z digest=sha256:ce941b158f721250bb2e5bd410d5ca7d8dc0b5e49f1609e32b9e83ee2000a662

Observation 93ad2fea-62ad-447d-84c6-01898825071b · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

AHELM: A Holistic Evaluation of Audio-Language Models Sequence Transduction with Recurrent Neural Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.808889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.808889Z digest=sha256:bf4c450923936298f1b835da504a7512ca0156bb9e9c123d318d72d29a0783a8

Observation 71697c35-c355-43ba-b779-faeba1af1613 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

AHELM: A Holistic Evaluation of Audio-Language Models Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.653539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.812080Z digest=sha256:821b715c053d1a4076a127a1bcf0054ac2f485da75ddcee9e58ccbd756949b09

Observation 6641787a-fc0b-4f71-8245-0d829f09500b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AHELM: A Holistic Evaluation of Audio-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.815275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.815275Z digest=sha256:3c566ecc8883382e2bbf76255b8ac9cea417c8aca347540844a3272e50755b9b

Observation 9668280e-98e8-4b15-88ba-4a84c824c22c · outbound

This paper cites Design of a linguistic statistical decoder for the recognition of continuous speech.

AHELM: A Holistic Evaluation of Audio-Language Models Design of a linguistic statistical decoder for the recognition of continuous speech

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.645260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.818670Z digest=sha256:ffa49436d806b3bc36c9397dc6939dcf19722938026943cb2041a4c00d39029b

Observation 366e1f1f-8735-45f1-bd81-92a809687cb2 · outbound

This paper cites Gemini 2.5: Our most intelligent AI model.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini 2.5: Our most intelligent AI model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.636480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.821970Z digest=sha256:5b03fefd06bb786e94d2b1573b1c32503e0fef54529d1b22feb3102b9fa54b69

Observation 13b680d9-4fd5-4b52-95f4-db89ba74f8ea · outbound

This paper cites AudioCaps: Generat- ing captions for audios in the wild.

AHELM: A Holistic Evaluation of Audio-Language Models AudioCaps: Generat- ing captions for audios in the wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.628926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.825065Z digest=sha256:1891e562b0cac18c9d9be2a6801f6853244dc2179aadabf1b1ccdca6f9a61d8f

Observation 106c11c1-51a3-405e-9cb2-fed3cd7d70fc · outbound

This paper cites Prometheus-vision: Vision-language model as a judge for fine-grained evaluation.

AHELM: A Holistic Evaluation of Audio-Language Models Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.620586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.828327Z digest=sha256:c94024a70ede13a26bd6f5c509b0787a3ce78e0fa8cdfe1ea086a0cee4f1e845

Observation 7d73ef13-c850-4071-8564-9fe6da7cf9b3 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.

AHELM: A Holistic Evaluation of Audio-Language Models Vhelm: A holistic evaluation of vision language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.611540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.832092Z digest=sha256:dfdd29ba76eb36abe90a9a28f25ad5b5a5f8d166b3d54127d2300d664988c1d0

Observation e85f6957-3677-45d6-8bd4-bf9d69372f0f · outbound

This paper cites Holistic evaluation of text-to-image models.

AHELM: A Holistic Evaluation of Audio-Language Models Holistic evaluation of text-to-image models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.601221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.835355Z digest=sha256:ef01e76e05afa472b8ebe98dc09ac2228e2b592407ad88a059e31c24706579b3

Observation 779edca4-f6cc-4ca8-a1b6-cf978c979b3e · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.591876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.838505Z digest=sha256:cdf51d9b3e97df21d25fb5b6aa324dd323dfb33531a94306d581e5387eb663f6

Observation e0df2478-580a-4889-a915-2f06b3634273 · outbound

This paper cites The next chapter of the Gemini era for developers.

AHELM: A Holistic Evaluation of Audio-Language Models The next chapter of the Gemini era for developers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.582196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.842252Z digest=sha256:7d7bb1867978218b9a511b52d845c6d6b641745f1e527e01b31d872341c72a78

Observation 280baaad-fbf2-4893-88cf-f56237ef9faa · outbound

This paper cites Hello GPT-4o, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models Hello GPT-4o, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.572830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.845109Z digest=sha256:0bb0de219282dab6aa52252edfc9d0df07a711c66c7b4bac2a4a8c02825d799d

Observation 0ed65fc5-1a08-49c4-95dc-657b8e8cfabb · outbound

This paper cites Introducing our next-generation audio models, Mar 2025.

AHELM: A Holistic Evaluation of Audio-Language Models Introducing our next-generation audio models, Mar 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.563178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.848693Z digest=sha256:f0401ed9f22951e58e9e70e3fd1eb3cb72ae16d3edadb974274d3fc85fff7d31

Observation 1d37ba2b-6924-4f40-bb80-fbaad1e10bbf · outbound

This paper cites LibriSpeech: an ASR corpus based on public domain audio books.

AHELM: A Holistic Evaluation of Audio-Language Models LibriSpeech: an ASR corpus based on public domain audio books

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.553483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.851608Z digest=sha256:d78984d87c9c105fb5e6b07c40eeb9908038c2ec22f89b27f93dae2743bc0546

Observation 0a627daf-dfb1-40e9-82a3-149fe1217b18 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions.

AHELM: A Holistic Evaluation of Audio-Language Models MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.545031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.855344Z digest=sha256:742ae94d9a99e8676b57b361c4448ddad915b14455e640862e68036cf705ecbc

Observation 4a345e82-50e5-47fb-8f4e-f8f82df64bad · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

AHELM: A Holistic Evaluation of Audio-Language Models MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.858777Z digest=sha256:70e31b8224d8ec724376b4ea40d1174a5fc53c70c4441788dd64437b3ac6d3ac

Observation 0c9017b8-63ac-46df-81f9-2c5038f0525e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

AHELM: A Holistic Evaluation of Audio-Language Models Robust speech recognition via large-scale weak supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.862427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.862427Z digest=sha256:54b32b28f231db05cf0c38688b3feb17fbd8b55cf0a1bc4dbe71972902f13eaf

Observation 20219ee9-5565-425d-80c8-3895efb66f6e · outbound

This paper cites Speech Robust Bench: A Robustness Benchmark For Speech Recognition.

AHELM: A Holistic Evaluation of Audio-Language Models Speech Robust Bench: A Robustness Benchmark For Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.865850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.865850Z digest=sha256:dc126142c560686d1b46aab3805ea26e0de700646ef5e03de222e854fdfc4a26

Observation 90cb00ee-5387-4568-af70-1a4c4781ecf3 · outbound

This paper cites V oice jailbreak attacks against GPT-4o, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models V oice jailbreak attacks against GPT-4o, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.530681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.869568Z digest=sha256:5cd7370096f0e4c51a942319ed4fa456617b195b3cb23380be5661fd2eb35e37

Observation 09d16adf-5107-46f1-bb0e-e5f9d8bf2d42 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

AHELM: A Holistic Evaluation of Audio-Language Models Salmonn: Towards generic hearing abilities for large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.521493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.872720Z digest=sha256:d2a11cb923b6d21d93327b80d77e1b75cc2a59cf324eade7c68eeb801830a709

Observation 4e4f081b-e1f0-4e0c-887e-5c4cd491836a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.875777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.875777Z digest=sha256:b58a605c30aa4cd5e622489f314e326cfe1ebe5ed18053ed6583f4e73a1070a2

Observation 3abc1829-0a24-4cb6-bc49-48c29cb16690 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

AHELM: A Holistic Evaluation of Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.879225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.879225Z digest=sha256:f2bfca55144afabaf4a4df9e20a5ede0d1388d0fbd59b5dc30dc8ba6b690e1aa

Observation b57f1fcb-69fa-49dc-8cfe-fa202f93dd4e · outbound

This paper cites Qwen2.5-Omni Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen2.5-Omni Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.882641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.882641Z digest=sha256:3e71a6fa37a93037af641b616501f1114fc86a58d9c44c616a2b103675029783

Observation 81ee07bf-d9ee-44bb-b7c7-8ea7ded00799 · outbound

This paper cites Qwen3 Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.886194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.886194Z digest=sha256:28f2f82a9609c1af09e65d677607016983bf1579ceb77946d059abf993f160bb

Observation d1b45dbc-f6be-4405-871f-4538421e4457 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

AHELM: A Holistic Evaluation of Audio-Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.889693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.889693Z digest=sha256:ae01489f309f87596a06972f04110ebb47648408c4e3a020445ccbfa69a86c63

Observation 4d5afb4d-7645-48d4-86bc-306b24464242 · outbound

This paper cites Speechlm: Enhanced speech pre-training with unpaired textual data.

AHELM: A Holistic Evaluation of Audio-Language Models Speechlm: Enhanced speech pre-training with unpaired textual data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.512709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.893240Z digest=sha256:f4f0d062e93bdd0ac8afa042c139f2b43e11bc0f67bc5675f2db6752297cb241

Observation e4df7286-3747-4109-a699-fcaec0306b8a · outbound

This paper cites Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese.

AHELM: A Holistic Evaluation of Audio-Language Models Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.896569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.896569Z digest=sha256:78c33cb25808a904f04c4a49e2d56f9467d03e63eb639f64dac8e88de384b686

Observation f46aaaf8-49d0-4b7f-9471-8852c1804a6d · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.503933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.901477Z digest=sha256:00f31a225ce3ebf949a0f8715bb17cc8b55cfc84297c38112f619e478c2808d2

Observation 6b02af22-b3d2-4244-a430-85bf59d46073 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.495380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.905490Z digest=sha256:cbd2817649087ba38109138a6230095c26aa3d206bace95661f22aaa31de2689

Observation 94f55932-df50-4f9f-ae84-315f1317d399 · outbound

This paper cites humorous and imaginative.

AHELM: A Holistic Evaluation of Audio-Language Models humorous and imaginative

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.487336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.908871Z digest=sha256:c86e8b3f2a20b428e3e260c00410a7a1108f8e87ef9712837bdc423b871116c3

Observation 9122540e-17ca-44f8-9f7d-e72828c7eedb · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.478978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.912676Z digest=sha256:320b45e259ca3750086228b35b7a9ed5bbafc1d15bcde94994de9d3a9a8b3331

Observation 36bc4d4e-83e2-4d0e-bf9d-0e3c89ed6915 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.470623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.916173Z digest=sha256:5cfa5d3929a04d84c30f124bcbf7c040de936835a9518814069fd29e141c8824

Observation 746c16fa-8b67-4ab6-8ddd-f90736f57628 · outbound

This paper cites E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content.

AHELM: A Holistic Evaluation of Audio-Language Models E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.461880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.919399Z digest=sha256:fc5e0a30ba7a04065a1c059cf48bb0c502c7e811d961d9a4454eb30dd1ab65ab

Observation 1953059c-e1a3-4b3a-9680-cad11029ab8e · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.452859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.922630Z digest=sha256:c502003662e39ba487842104f961e0967c7380b3800f364e5c659207b863838d

Observation 7100bcb0-534f-44b1-853f-e0a1fb4429cb · outbound

This paper cites You should refer to the score rubric.

AHELM: A Holistic Evaluation of Audio-Language Models You should refer to the score rubric

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.445050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.926181Z digest=sha256:10d4fde3bb778856f27c82ae0eec2cf2e9494d776b839608ce527417df233685

Observation 91bc551c-76b9-45aa-bc4f-37b5f5d10b12 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.437420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.929628Z digest=sha256:a9ff98419a8a244fe01ece6e92404a586154401083db663143e9af212ad6b5b0

Observation fd9eec2a-1908-42c0-81cc-67730c0c68d4 · outbound

This paper cites haha”) or throat clearing (e.g., “ahem.

AHELM: A Holistic Evaluation of Audio-Language Models haha”) or throat clearing (e.g., “ahem

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.429327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.933794Z digest=sha256:eddb7eb053e2f8aea39665cf44ba7d6ff3a556403347cd75b598f3dafc87ed58

Observation cafd4911-0adb-4276-8e8d-1cd50d4a9aeb · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.421481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.937749Z digest=sha256:4a2ddaa877d43cb5a286a162100d512442b63925d1b1244f62dd03bcc4848bae

Observation d000f3ec-6c48-4317-99a0-4a1f55975eaf · outbound

This paper cites From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash.

AHELM: A Holistic Evaluation of Audio-Language Models From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.413407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.940870Z digest=sha256:7600c5dddcbc429dbc8796bc3d5019eec487c4ce49ab46028c7119a2603e05ec

Observation 0ad1c622-00ad-4d0f-be3a-b2dc963f88b7 · outbound

This paper cites When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack.

AHELM: A Holistic Evaluation of Audio-Language Models When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.404930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.944391Z digest=sha256:2db7904de18aed843dcf50676cf249b8b9482582d4f079413b4df38788769b06

Observation ce0143f0-edec-44b0-a2c5-f3104c8eec86 · outbound

This paper cites We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5.

AHELM: A Holistic Evaluation of Audio-Language Models We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.396940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.947659Z digest=sha256:0fb26c8010f015664798d5e96135d3fef34806fe8e68fd6ca12f8bc40395b422

Observation fe1d0241-7fb1-4404-878f-8d5d953d958a · outbound

This paper cites Limitations.

AHELM: A Holistic Evaluation of Audio-Language Models Limitations

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.388203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.950813Z digest=sha256:9ebd0396bb8c148458cf66d19e7382fc120967946e14db7a922fe477589cc0ed

Observation 94467b2d-8bef-4025-b527-944ac890cf07 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.379320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.954108Z digest=sha256:2b5f7346390de333cb9315146f7957aeeba74ff1b7f92fa7660bf444d9505b8b

Observation 865d1775-6ed3-449e-a803-01997f63b213 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.371248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.957811Z digest=sha256:c825dd35b09a79eb65cd3e8555f225006f017cc2e3951b6ed8adca7bcbabf136

Observation bed55278-f3a2-4c5b-942a-36e402193d8b · outbound

This paper cites com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1.

AHELM: A Holistic Evaluation of Audio-Language Models com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.362856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.961310Z digest=sha256:cd3adc3354dcac11279df6be9bd202de0cf1322290551c7e00f730f3eed9caae

Observation 31b3c5a6-d911-4336-a6d4-ab1b7ac171cd · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.354147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.964523Z digest=sha256:86707fb52e5f09ac6e6ca83cdae35a48b6995b10409c1933ae5e739d305954ac

Observation c9c61ba0-91ae-45a4-bcfd-6cd5a60f4db5 · outbound

This paper cites But we do not compute error bars for other scenarios.

AHELM: A Holistic Evaluation of Audio-Language Models But we do not compute error bars for other scenarios

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.344723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.967592Z digest=sha256:9d2c2f1477d7ffccf5d0e605ef1271eb8271cfb2c99ab8aacab258c5e9542056

Observation 2b1cd651-9153-45be-a311-d04b174b3f1d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.335836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.970937Z digest=sha256:5f03c49a2de633eca2ea5ba0cfbd2cf225e98dca792794bb0846fe6646967310

Observation 0ba18208-7e04-4b5c-8bd6-64e4f23efcec · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.327529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.974113Z digest=sha256:72b43081b92c3a582d1325cdf7a7f094add249b92815fb369babc390811e60ee

Observation 1aacc2fa-b354-4c1b-82a3-e84f0db5faf4 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.318849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.977488Z digest=sha256:a8cba473e6e10ad393e0818c54e835f2c177963526ae35682e7119a1aea6ff4c

Observation 462c6a8a-0ff0-46e1-ae74-e02ca8c24b93 · outbound

This paper cites Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata.

AHELM: A Holistic Evaluation of Audio-Language Models Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.309287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.980801Z digest=sha256:4998ee5539de15590acf2ddb527b9d3d497dda2c2c16079421373cb3c9e79417

Observation 979de6e4-0d9a-433a-8dfa-f1b0852bb2da · outbound

This paper cites We cite all the datasets and models used in our work.

AHELM: A Holistic Evaluation of Audio-Language Models We cite all the datasets and models used in our work

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.300092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.983943Z digest=sha256:8823a9ee48ea1a2f6b720d71234265546f634a8c2ff174e88a6577df2e49e07d

Observation e119bfdd-abc9-4fae-92af-603f17650782 · outbound

This paper cites PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio.

AHELM: A Holistic Evaluation of Audio-Language Models PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.289975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.987392Z digest=sha256:d2dbace1d2f1b4046fbdc8027221b595d748e409ea4a09a75764dae55768907c

Observation 85324c3b-c3d3-4ad6-904e-fa33b94ec416 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.279914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.990404Z digest=sha256:387649b2a1d4ce62218b3c32fe68990822669d682956d4e0d0a07ad140548e7f

Observation 48de1972-aa06-47ee-883c-a22d9a349457 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.269972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.993270Z digest=sha256:a8b58da2a62f325897f9f0e58031bb808fb8782f8ba011a50495b71478043921

Observation 913828ab-683a-4f62-91d4-fe2b0213da28 · outbound

This paper cites Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark.

AHELM: A Holistic Evaluation of Audio-Language Models Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.259930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:22:34.997009Z digest=sha256:fb886bc65dc2f02b410295c40ab126ebc7b52f4dc4eb734562512dc9d9d36359

Pith citing papers

Observation 438bb6eb-afa4-4118-9e01-f38ceee8c266 · inbound

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs cites this paper.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs AHELM: A Holistic Evaluation of Audio-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.185900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:f2c037ad45d6eece87e3a1c01da27bb729234761b01e53e73183e6a514bf79ce

Observation 420de55c-2f6c-49e9-9414-845e6f2c69af · inbound

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics cites this paper.

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics AHELM: A Holistic Evaluation of Audio-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:52:45.801592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:52:45.801592Z digest=sha256:a819acf161e3419843232a7ba05a709a4746e3d18ee3cf58db02ca919053f10c

Observation b118258d-1bab-4c12-ba6a-d91a2b208abd · inbound

PRiSM: Benchmarking Phone Realization in Speech Models cites this paper.

PRiSM: Benchmarking Phone Realization in Speech Models AHELM: A Holistic Evaluation of Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:31.131370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:31.131370Z digest=sha256:0f24d0330686e07a567da8e77c4c5997038ec9c418edfe0dee1b77cc78914e7e

Observation 3af163d3-1a39-4b7a-8f15-d33576402f1d · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where AHELM: A Holistic Evaluation of Audio-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:22.088783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:f545226e36abf472d9631568f18663da85915589d284b25793630aa3e7d48f42

Observation a2e3a273-2a73-4c70-9492-e80937080f18 · inbound

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech cites this paper.

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech AHELM: A Holistic Evaluation of Audio-Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.226416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:46:50.923340Z digest=sha256:82083d38627213967762eec26ce8d4937c200919b6687ccccff95b1606cfbbd3

Observation c7a02f6c-211f-4b94-ba6b-cd699f7bcfc5 · inbound

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI cites this paper.

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI AHELM: A Holistic Evaluation of Audio-Language Models

Reference 227

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:55:29.197360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:27:18.774649Z digest=sha256:66d968936a3d2e8454c08241897461f72d420916838980c2a0bdaa72f5f9bec3

Observation ec1072ce-d1cb-4a1e-97ae-f072a79f6845 · inbound

AudioMosaic: Contrastive Masked Audio Representation Learning cites this paper.

AudioMosaic: Contrastive Masked Audio Representation Learning AHELM: A Holistic Evaluation of Audio-Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:53:28.861492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:fb58e025b700124ef76daf8afef6265e16a27968d6efae18dd3bc79b264ae700

Observation 6e1dc822-8c61-4464-b2ea-73165ac7f0a0 · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities AHELM: A Holistic Evaluation of Audio-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.430007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:7ddb36cd85f5470d6038ffaa705d9c1039d51eb3ba5acae8ba6629762a7a1e26

Observation 0a7d2131-ece3-4fec-9428-112d78267f36 · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages AHELM: A Holistic Evaluation of Audio-Language Models

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.875641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:e6ad1711116853a85426c83ff86ec28819b790a3b83e0aba263be81003edc979

Observation 6d5d8abb-cfc8-46d0-a5f4-cc38b5f97d4c · inbound

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems cites this paper.

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems AHELM: A Holistic Evaluation of Audio-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:37.780637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:37.780637Z digest=sha256:6213ce29354ebb7bd057a71d605b7f9f6fe78aa1640fd2bdc0873cbcd01fb02f