Pith. sign in

Paper Citation Record · LEDGER

AHELM: A Holistic Evaluation of Audio-Language Models

As of 19 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 10 inbound Pith citation observations for arXiv:2508.21376.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21376 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:22:34.997009Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:52:45.801592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:52.874154Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9781f1e-d579-44da-aab7-1f362649c634 · outbound

This paper cites GPT-4 Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.697071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.697071Z digest=sha256:98108e00be6878c134a4acfc3a82d62441d15ef35379b88915627496bb8909e7

Observation 51e0b0aa-3846-40c3-8ef3-af0dd4b33398 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models The Claude 3 model family: Opus, Sonnet, Haiku, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.739040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:33.778200Z digest=sha256:9ab6bc349d3f3feb4fba10aa6102692ec8f5e8de31555baa5526136611c28506

Observation a2426604-3b6c-4168-a590-8cb736b79234 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

AHELM: A Holistic Evaluation of Audio-Language Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.923565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.923565Z digest=sha256:376a819fae59ee3123cb6966194df8764259d99d25e8160c16502ac8be9a5bc2

Observation 2602b30f-96dd-4597-8e93-c1a7b37ea665 · outbound

This paper cites Qwen Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.102234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.102234Z digest=sha256:0dd630e9a3d28cf090bba5998002a3a5a9aa91289380001c4130b4e0a522d53d

Observation 6b69d948-705f-4460-9ef4-0a4967a3ff44 · outbound

This paper cites Towards multimodal sarcasm detection (an _obviously_ perfect paper).

AHELM: A Holistic Evaluation of Audio-Language Models Towards multimodal sarcasm detection (an _obviously_ perfect paper)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.731229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.243183Z digest=sha256:4e7c74c30f5f9354b0a768827ebb3ace6efe90826a379c103774a5339ac4449d

Observation 2b74c9ce-de68-4f7c-b1c5-70410f7cdd73 · outbound

This paper cites Qwen2-Audio Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.352067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.352067Z digest=sha256:a3cb2612c6a5e6f8a4f11a253e1c5fb4268e726faf219e095bd1bc842b3c749f

Observation 41cacf6c-12bb-4b8c-a526-e4abebb8b96f · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

AHELM: A Holistic Evaluation of Audio-Language Models VoxCeleb2: Deep Speaker Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.385152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.385152Z digest=sha256:9bf3a5804a679a6a36223ef6417a636a0b550d1ba307894f1d3c4b165923f2f8

Observation e4269de4-646a-4529-86dc-f4865d5969a2 · outbound

This paper cites Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit.

AHELM: A Holistic Evaluation of Audio-Language Models Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.722887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.540549Z digest=sha256:b218d8e3ce18e51cec5600ccc8682391e3cf037060e0860103769f60f4fee270

Observation 38f7207c-cba3-4f5a-9765-493bc4c10f6b · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech.

AHELM: A Holistic Evaluation of Audio-Language Models FLEURS: Few-shot learning evaluation of universal representations of speech

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.714660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.635896Z digest=sha256:a946d74baab7ff29e36cd0b9a2f9e4a88cd49299c8a8e56f899a3e08aae9c411

Observation c7aecf5c-9e8e-491b-8fa9-a0115a3c3be6 · outbound

This paper cites MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector.

AHELM: A Holistic Evaluation of Audio-Language Models MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:22:35.199235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.753589Z digest=sha256:3adfaace127ab144d128f5b9f7ee264af52430689af1e078467b232a213a469e

Observation d728dde4-d52a-4a2f-b473-9ccde920fbd1 · outbound

This paper cites Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition.

AHELM: A Holistic Evaluation of Audio-Language Models Speech-transformer: a no-recurrence sequence-to- sequence model for speech recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.706024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.775890Z digest=sha256:77a1bcd71fe3152dd85be83700b2c38d6be917dcc2040728654f197f66b48ab5

Observation fe123ce9-b696-4805-9234-726bc74cbbe1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

AHELM: A Holistic Evaluation of Audio-Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.779733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.779733Z digest=sha256:7ea264b8c04db16ae631605e1e6a91f64d0062ef18ec9ac0542d36781ee5f5e3

Observation 9a8d9c12-d1d7-4781-8e23-3bd439c1ee47 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

AHELM: A Holistic Evaluation of Audio-Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.697626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.783925Z digest=sha256:a62a3bdf430d660475b412e420754ce35c60c3e262e4ea8039a7fa6ff9ca99a4

Observation 02a1fa96-c5ad-42d8-b7c0-252adacb5815 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

AHELM: A Holistic Evaluation of Audio-Language Models Clap learning audio concepts from natural language supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.689202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.787398Z digest=sha256:b9fc60fa934ffc0f782c11087efa29ef353a72b49ba36c1990258dc47450307e

Observation 6b2cf18e-ff38-4505-a3d2-fc3396052351 · outbound

This paper cites Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images.

AHELM: A Holistic Evaluation of Audio-Language Models Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:22:35.170240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.790680Z digest=sha256:3fdb1b40d6e8a075cafe08e709957bfdf3936588780cbf733f33cc0916245a78

Observation dd5a1352-ec15-4aed-98d6-51d73912d090 · outbound

This paper cites CSR-I (WSJ0) Complete.

AHELM: A Holistic Evaluation of Audio-Language Models CSR-I (WSJ0) Complete

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.680312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.794039Z digest=sha256:bb66623c3069a0008c12fedc11548abc3c465823361ed8e65ec5a42e66215b95

Observation a2cebc79-52ca-4f7e-bb5c-7b9e40df3c96 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.797552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.797552Z digest=sha256:31110868832ab03abf7aabeb8496bb4f33bd7f835001fe04e4df738559532d8d

Observation 80c80121-1169-4858-a503-2b0fc6c46247 · outbound

This paper cites Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities.

AHELM: A Holistic Evaluation of Audio-Language Models Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.669917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.801817Z digest=sha256:9021417264074eb45cd0a60caef1b9d31ee6eaed8875817209caf56654897867

Observation 0afda4f7-ceff-44f8-9dc6-e6014bebf7d0 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

AHELM: A Holistic Evaluation of Audio-Language Models V ocalsound: A dataset for improving human vocal sounds recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.661754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.805468Z digest=sha256:b534cc6137b5ddacc37ba33287b7e11fd6e80d24ff8a6f6e0e44d33256f886e8

Observation 93ad2fea-62ad-447d-84c6-01898825071b · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

AHELM: A Holistic Evaluation of Audio-Language Models Sequence Transduction with Recurrent Neural Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.808889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.808889Z digest=sha256:614f38804f0111b6779d79ecd80c2350f88cf311d24b407dd45787f6bd30818f

Observation 71697c35-c355-43ba-b779-faeba1af1613 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

AHELM: A Holistic Evaluation of Audio-Language Models Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.653539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.812080Z digest=sha256:ab940868b7ce3c86d705dbe7988e316fcafecd468f1bfdba3cf8cd6aea21e8a7

Observation 6641787a-fc0b-4f71-8245-0d829f09500b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AHELM: A Holistic Evaluation of Audio-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.815275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.815275Z digest=sha256:ea9cb5e07498a6a7d5572ced2d23322cf2dea486ebbddd9a6871ce37ee81979f

Observation 9668280e-98e8-4b15-88ba-4a84c824c22c · outbound

This paper cites Design of a linguistic statistical decoder for the recognition of continuous speech.

AHELM: A Holistic Evaluation of Audio-Language Models Design of a linguistic statistical decoder for the recognition of continuous speech

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.645260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.818670Z digest=sha256:525fd2499eed28f11baae34881e731cc8c8e738cdb9ee6c2ced2cdd1f8a3235c

Observation 366e1f1f-8735-45f1-bd81-92a809687cb2 · outbound

This paper cites Gemini 2.5: Our most intelligent AI model.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini 2.5: Our most intelligent AI model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.636480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.821970Z digest=sha256:32de9d95892ba334752f542038ccc001da2ac7d4777c14326b4bd4b2e4f75e3b

Observation 13b680d9-4fd5-4b52-95f4-db89ba74f8ea · outbound

This paper cites AudioCaps: Generat- ing captions for audios in the wild.

AHELM: A Holistic Evaluation of Audio-Language Models AudioCaps: Generat- ing captions for audios in the wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.628926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.825065Z digest=sha256:8d7accf1ceee1f0d53c64990a04b8c3fe73c31ba0bb85c36962ce39f845269ad

Observation 106c11c1-51a3-405e-9cb2-fed3cd7d70fc · outbound

This paper cites Prometheus-vision: Vision-language model as a judge for fine-grained evaluation.

AHELM: A Holistic Evaluation of Audio-Language Models Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.620586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.828327Z digest=sha256:73248a4ca9c6a2d3c6d9e6078d99c3386233b0b8ef17a5066debea06dc96a34e

Observation 7d73ef13-c850-4071-8564-9fe6da7cf9b3 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.

AHELM: A Holistic Evaluation of Audio-Language Models Vhelm: A holistic evaluation of vision language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.611540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.832092Z digest=sha256:8b7556f4c488efad949e6f8e82db3a88e80c8037ea0e48e5c88bbb49c1ed9a1a

Observation e85f6957-3677-45d6-8bd4-bf9d69372f0f · outbound

This paper cites Holistic evaluation of text-to-image models.

AHELM: A Holistic Evaluation of Audio-Language Models Holistic evaluation of text-to-image models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.601221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.835355Z digest=sha256:6b2b4731d8f9577e008b17d9792c1670b51e3d3b47dff6f0b0fe8ce606211bc1

Observation 779edca4-f6cc-4ca8-a1b6-cf978c979b3e · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.591876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.838505Z digest=sha256:50f19d65922711b64155d5c2b03f7ea3b8cb296acfd0223b57d035725cac5e4c

Observation e0df2478-580a-4889-a915-2f06b3634273 · outbound

This paper cites The next chapter of the Gemini era for developers.

AHELM: A Holistic Evaluation of Audio-Language Models The next chapter of the Gemini era for developers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.582196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.842252Z digest=sha256:bfa981f1424e59b546fef12b2bd2dff2b47b670607e2685634b47a263d1c466f

Observation 280baaad-fbf2-4893-88cf-f56237ef9faa · outbound

This paper cites Hello GPT-4o, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models Hello GPT-4o, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.572830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.845109Z digest=sha256:7f65ff809dddb818b7323faa227a5cc9ca5d77d8e0c03333b7b3d0049b74ce95

Observation 0ed65fc5-1a08-49c4-95dc-657b8e8cfabb · outbound

This paper cites Introducing our next-generation audio models, Mar 2025.

AHELM: A Holistic Evaluation of Audio-Language Models Introducing our next-generation audio models, Mar 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.563178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.848693Z digest=sha256:dc216231b6cb93b75a4a95974b90fde459e0cda3c059cee083915752a1987969

Observation 1d37ba2b-6924-4f40-bb80-fbaad1e10bbf · outbound

This paper cites LibriSpeech: an ASR corpus based on public domain audio books.

AHELM: A Holistic Evaluation of Audio-Language Models LibriSpeech: an ASR corpus based on public domain audio books

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.553483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.851608Z digest=sha256:5dc21ba089c6be1430640137c4e8e0b5bfeed16658553c22fe9f01be1dd64ee4

Observation 0a627daf-dfb1-40e9-82a3-149fe1217b18 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions.

AHELM: A Holistic Evaluation of Audio-Language Models MELD: A multimodal multi-party dataset for emotion recognition in conversa- tions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.545031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.855344Z digest=sha256:b83b1b2cb043b8efdb085572f00211e7b2e90da7bbc100a6b82f9891e4622135

Observation 4a345e82-50e5-47fb-8f4e-f8f82df64bad · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

AHELM: A Holistic Evaluation of Audio-Language Models MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.858777Z digest=sha256:d8379ed0a09f04307850df22c080281a50fe848337a930ec5ebb48c1a47461cf

Observation 0c9017b8-63ac-46df-81f9-2c5038f0525e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

AHELM: A Holistic Evaluation of Audio-Language Models Robust speech recognition via large-scale weak supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.862427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.862427Z digest=sha256:d63d525396cb2407f5df49288deaa2e081618e8d1caa26188525bd3db35f41a8

Observation 20219ee9-5565-425d-80c8-3895efb66f6e · outbound

This paper cites Speech Robust Bench: A Robustness Benchmark For Speech Recognition.

AHELM: A Holistic Evaluation of Audio-Language Models Speech Robust Bench: A Robustness Benchmark For Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.865850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.865850Z digest=sha256:8ee805dad94fb7ed43eac60f6e6824715780d7aa93d94e5a09648592ddc628aa

Observation 90cb00ee-5387-4568-af70-1a4c4781ecf3 · outbound

This paper cites V oice jailbreak attacks against GPT-4o, 2024.

AHELM: A Holistic Evaluation of Audio-Language Models V oice jailbreak attacks against GPT-4o, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.530681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.869568Z digest=sha256:5c1c7b218e52b913132a1fb9c77988660f9c2cae9292baa3aea12137842f3a53

Observation 09d16adf-5107-46f1-bb0e-e5f9d8bf2d42 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

AHELM: A Holistic Evaluation of Audio-Language Models Salmonn: Towards generic hearing abilities for large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.521493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.872720Z digest=sha256:3d60691fb74cd013a9197bb5644157a3acd076f7c5849baf0e9571f420ff2f7e

Observation 4e4f081b-e1f0-4e0c-887e-5c4cd491836a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AHELM: A Holistic Evaluation of Audio-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.875777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.875777Z digest=sha256:7573860f2844a049b0bc833aa512cf2f44f97f7b194e05a967e9a679e057785b

Observation 3abc1829-0a24-4cb6-bc49-48c29cb16690 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

AHELM: A Holistic Evaluation of Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.879225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.879225Z digest=sha256:07319ae5277b6a2bbc5fdb3c43e106a4354a2f68ac0c310be5eee3f4950321f1

Observation b57f1fcb-69fa-49dc-8cfe-fa202f93dd4e · outbound

This paper cites Qwen2.5-Omni Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen2.5-Omni Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.882641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.882641Z digest=sha256:4b7abdcc3711ed494e6f76a9de7f5f634c770395be0a601077b089e51d01701f

Observation 81ee07bf-d9ee-44bb-b7c7-8ea7ded00799 · outbound

This paper cites Qwen3 Technical Report.

AHELM: A Holistic Evaluation of Audio-Language Models Qwen3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.886194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.886194Z digest=sha256:abaf8363e88fbea4eca70a6ce50ce85f8e3861f2b664a29305e4254e61d7af3b

Observation d1b45dbc-f6be-4405-871f-4538421e4457 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

AHELM: A Holistic Evaluation of Audio-Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.889693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.889693Z digest=sha256:1608dc3cfb2a109973394e5d613e187ff3d08485abfc58bf12a2d5e5263ba6b2

Observation 4d5afb4d-7645-48d4-86bc-306b24464242 · outbound

This paper cites Speechlm: Enhanced speech pre-training with unpaired textual data.

AHELM: A Holistic Evaluation of Audio-Language Models Speechlm: Enhanced speech pre-training with unpaired textual data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.512709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.893240Z digest=sha256:7a382713797e2a7d48789dc6b663dbd48702558935a20504c4fead796ef6a34b

Observation e4df7286-3747-4109-a699-fcaec0306b8a · outbound

This paper cites Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese.

AHELM: A Holistic Evaluation of Audio-Language Models Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.896569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.896569Z digest=sha256:22b0bbb20b55dd1c71d5aae35831487e5db2a6d3523a38bdb5dc6f938bcc38fc

Observation f46aaaf8-49d0-4b7f-9471-8852c1804a6d · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.503933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.901477Z digest=sha256:6d4e8990cebf2ff2386131ea257af6e9c7e1693f81bfe7b4c41ab9cdafc2ff85

Observation 6b02af22-b3d2-4244-a430-85bf59d46073 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.495380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.905490Z digest=sha256:93491e862233976c5150f32f4f7c8ef8af21aef6f67a6b1f30494b82b916fe23

Observation 94f55932-df50-4f9f-ae84-315f1317d399 · outbound

This paper cites humorous and imaginative.

AHELM: A Holistic Evaluation of Audio-Language Models humorous and imaginative

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.487336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.908871Z digest=sha256:91327e9ee9d813199ee89414b583d7132bba63399600b42aaba9131476719ae3

Observation 9122540e-17ca-44f8-9f7d-e72828c7eedb · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.478978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.912676Z digest=sha256:9848a2735e8566cfc6c8cd688828ba895ff92e78a62076f78677380aacf63e5a

Observation 36bc4d4e-83e2-4d0e-bf9d-0e3c89ed6915 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.470623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.916173Z digest=sha256:3f8eb04b1354d49ed59d7f9f1b47a9c0286b292704158d87a88018e2b96e37b7

Observation 746c16fa-8b67-4ab6-8ddd-f90736f57628 · outbound

This paper cites E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content.

AHELM: A Holistic Evaluation of Audio-Language Models E.1.1 Obtaining a list of contrasting roles We use the list of roles from PAIRS (replicated in Table A4) to seed the generation of speech content

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.461880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.919399Z digest=sha256:e3cd65cb87dae08c4e6361b595c8f5b5c66346876c4b08eaaa82ec9dacd57fcb

Observation 1953059c-e1a3-4b3a-9680-cad11029ab8e · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.452859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.922630Z digest=sha256:86273e64b47955cc406d0514abe3ead3bcdbef4e64f80c992b4be71dd355d628

Observation 7100bcb0-534f-44b1-853f-e0a1fb4429cb · outbound

This paper cites You should refer to the score rubric.

AHELM: A Holistic Evaluation of Audio-Language Models You should refer to the score rubric

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.445050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.926181Z digest=sha256:702aef3b0028747555c4cb7e5dec8b950bd154428a8ba02ece43cab34a229a80

Observation 91bc551c-76b9-45aa-bc4f-37b5f5d10b12 · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.437420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.929628Z digest=sha256:984d5556d791d2d33b9b5b3f86bd06c6cc7f05a5a5e69df9e837d8c200bbf8b4

Observation fd9eec2a-1908-42c0-81cc-67730c0c68d4 · outbound

This paper cites haha”) or throat clearing (e.g., “ahem.

AHELM: A Holistic Evaluation of Audio-Language Models haha”) or throat clearing (e.g., “ahem

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.429327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.933794Z digest=sha256:d1e1c6ec3d7909e010de2160a8514688641f238b5908414a26cc0d9c29ef7077

Observation cafd4911-0adb-4276-8e8d-1cd50d4a9aeb · outbound

This paper cites an unresolved cited work.

AHELM: A Holistic Evaluation of Audio-Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:22:35.421481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.937749Z digest=sha256:de95671dcbb0b7150ae4ac53eaa79e246fa454bce51eb58ec3aee222d3b6e6ed

Observation d000f3ec-6c48-4317-99a0-4a1f55975eaf · outbound

This paper cites From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash.

AHELM: A Holistic Evaluation of Audio-Language Models From Table A9, we see that Qwen2-Audio Instruct takes the lead in audio knowledge, followed by Gemini 2.5 Pro (05-06 Preview) and then Gemini 2.0 Flash

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.413407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.940870Z digest=sha256:dd6649991c9853e24733cc1458c2a04a54c54b3a62f293ef94ecff14013a4145

Observation 0ad1c622-00ad-4d0f-be3a-b2dc963f88b7 · outbound

This paper cites When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack.

AHELM: A Holistic Evaluation of Audio-Language Models When looking at the safety aspect, we see that OpenAI models are robust to the voice jailbreak attack

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.404930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.944391Z digest=sha256:d50cd5b598cc81d0d44eeecd8550e27226d49f6211d4d272e4dea9c67b0135f9

Observation ce0143f0-edec-44b0-a2c5-f3104c8eec86 · outbound

This paper cites We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5.

AHELM: A Holistic Evaluation of Audio-Language Models We explain our benchmark in Section 3 and describe the experiments in Section 4 and report results in Section 5

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.396940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.947659Z digest=sha256:744635a17d54f69d88025dea362721528dfe1c9a98354971400d4111b9ca5196

Observation fe1d0241-7fb1-4404-878f-8d5d953d958a · outbound

This paper cites Limitations.

AHELM: A Holistic Evaluation of Audio-Language Models Limitations

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.388203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.950813Z digest=sha256:1b487a96848bdbdf5fa58019c95f88528cc40dd817a553163d0ddfda2353715c

Observation 94467b2d-8bef-4025-b527-944ac890cf07 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.379320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.954108Z digest=sha256:cef932746ae77ed26aca32a94f86c3d7ebd4066e6871df850e7bc55417155832

Observation 865d1775-6ed3-449e-a803-01997f63b213 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.371248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.957811Z digest=sha256:f6a57cea8433bf5ec21e1ed32bfabd8b97de75a7d5de776228b6f551afc46fa3

Observation bed55278-f3a2-4c5b-942a-36e402193d8b · outbound

This paper cites com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1.

AHELM: A Holistic Evaluation of Audio-Language Models com/stanford-crfm/helm and the new datasets at https://huggingface.co/ datasets/UCSC-VLAA/PARADE_audio and https://huggingface.co/datasets/ stanford-crfm/CoReBench_v1

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.362856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.961310Z digest=sha256:80487f342b8c4d3a12e74fbd3565590ca78057ae015204563f2b24479d923969

Observation 31b3c5a6-d911-4336-a6d4-ab1b7ac171cd · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.354147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.964523Z digest=sha256:97b9d64afe7acf07e2f8f85f336efafae9b518e0b0793c0b3156a6e408d00b52

Observation c9c61ba0-91ae-45a4-bcfd-6cd5a60f4db5 · outbound

This paper cites But we do not compute error bars for other scenarios.

AHELM: A Holistic Evaluation of Audio-Language Models But we do not compute error bars for other scenarios

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.344723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.967592Z digest=sha256:56dc756148537ea1c3a55b6609fbe7b99cb8eaae018a5f3c16337869cd427879

Observation 2b1cd651-9153-45be-a311-d04b174b3f1d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.335836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.970937Z digest=sha256:330979ebba307b2f64c3919f3f16d9f4a2cfd184e4b6269203cfc6aeb26b09b5

Observation 0ba18208-7e04-4b5c-8bd6-64e4f23efcec · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.327529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.974113Z digest=sha256:57383a647c721894f72da88be4674b2f7178688400c9aca5f51ddcbd71e1a84b

Observation 1aacc2fa-b354-4c1b-82a3-e84f0db5faf4 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.318849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.977488Z digest=sha256:eacb681a96f5b0e5d4398d600beeaca8d45afcb552da74806d0af8afe2dc87d0

Observation 462c6a8a-0ff0-46e1-ae74-e02ca8c24b93 · outbound

This paper cites Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata.

AHELM: A Holistic Evaluation of Audio-Language Models Before transforming tran- scripts to audio, we performed human scrutiny of the audio transcripts to make sure that there is no improper or toxic content in the metadata

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.309287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.980801Z digest=sha256:440ce556cdf803f417268ff940bd1206350d7cbb21fa55cd9a215830fab64819

Observation 979de6e4-0d9a-433a-8dfa-f1b0852bb2da · outbound

This paper cites We cite all the datasets and models used in our work.

AHELM: A Holistic Evaluation of Audio-Language Models We cite all the datasets and models used in our work

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.300092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.983943Z digest=sha256:5e63f200ff29d241b249287bd429f06dd0dfad829bfc7484db7210532f42c506

Observation e119bfdd-abc9-4fae-92af-603f17650782 · outbound

This paper cites PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio.

AHELM: A Holistic Evaluation of Audio-Language Models PARADE is available at https://huggingface.co/datasets/UCSC-VLAA/PARADE_ audio

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.289975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.987392Z digest=sha256:ec361f4bf11723c1ed7a791bc0b577f1e485c301a0b19ff896b8b4df8080b2ae

Observation 85324c3b-c3d3-4ad6-904e-fa33b94ec416 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.279914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.990404Z digest=sha256:8b22c469a788be7e8a519daee3c625f5a8f946021a963c1e74cf1f3c080cf864

Observation 48de1972-aa06-47ee-883c-a22d9a349457 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AHELM: A Holistic Evaluation of Audio-Language Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.269972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.993270Z digest=sha256:afc355547c946d9cb4b9d7e952816e979a45f11a79af4944c5e1b26d30e6b359

Observation 913828ab-683a-4f62-91d4-fe2b0213da28 · outbound

This paper cites Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark.

AHELM: A Holistic Evaluation of Audio-Language Models Answer: [Yes] Justification: As detailed in the Appendix B, we leverage OpenAI’s GPT-4o to create audio transcripts for the curation of PARADE benchmark

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:35.259930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T14:22:34.997009Z digest=sha256:0858684b42ca6f8347d13c4621af9db8e90b844aa8927326695307aa585efb7e

Pith citing papers

Observation 438bb6eb-afa4-4118-9e01-f38ceee8c266 · inbound

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs cites this paper.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs AHELM: A Holistic Evaluation of Audio-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.185900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:81fe07fe89d164da54021022ce58080bd6612413d6ecb5b45df3151f4f07221d

Observation 420de55c-2f6c-49e9-9414-845e6f2c69af · inbound

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics cites this paper.

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics AHELM: A Holistic Evaluation of Audio-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:52:45.801592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:52:45.801592Z digest=sha256:f58bcb3383e769f56f5f91be768b98b0de5609bac79e4d7e09236a29a9a6bb54

Observation b118258d-1bab-4c12-ba6a-d91a2b208abd · inbound

PRiSM: Benchmarking Phone Realization in Speech Models cites this paper.

PRiSM: Benchmarking Phone Realization in Speech Models AHELM: A Holistic Evaluation of Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:31.131370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:31.131370Z digest=sha256:f9db0ca92284989ba3fbdee717fc73f1b0dfbce53b3fc862e470056a06cc2903

Observation 3af163d3-1a39-4b7a-8f15-d33576402f1d · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where AHELM: A Holistic Evaluation of Audio-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:22.088783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:9c8f8ebca77ed6bae42942d4c4cefa4346e43febd0a36e68570f4d5866bd5610

Observation a2e3a273-2a73-4c70-9492-e80937080f18 · inbound

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech cites this paper.

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech AHELM: A Holistic Evaluation of Audio-Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.226416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:46:50.923340Z digest=sha256:cbf75a4421e97170783ad107a32391c6974cb40d7fe5a39cb21141bc4187a59a

Observation c7a02f6c-211f-4b94-ba6b-cd699f7bcfc5 · inbound

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI cites this paper.

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI AHELM: A Holistic Evaluation of Audio-Language Models

Reference 227

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:55:29.197360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T19:27:18.774649Z digest=sha256:cca05994d4c2d00d9b41369176bf84a38046d17296d56e67172141cd6ff16a03

Observation ec1072ce-d1cb-4a1e-97ae-f072a79f6845 · inbound

AudioMosaic: Contrastive Masked Audio Representation Learning cites this paper.

AudioMosaic: Contrastive Masked Audio Representation Learning AHELM: A Holistic Evaluation of Audio-Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:53:28.861492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T01:52:01.164694Z digest=sha256:c149be16561889102bf8e72f21317e5e00dd3a7b5cf6b5e03aae7315bd1d69aa

Observation 6e1dc822-8c61-4464-b2ea-73165ac7f0a0 · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities AHELM: A Holistic Evaluation of Audio-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.430007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:f456e45349b2619a8e3945fd06a61922a602ced136557bb8614d756a3cba9aee

Observation 0a7d2131-ece3-4fec-9428-112d78267f36 · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages AHELM: A Holistic Evaluation of Audio-Language Models

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.875641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:3523568df2a85075491c2c0b14f051a070a28765ddda42548e1cc057b911550f

Observation 6d5d8abb-cfc8-46d0-a5f4-cc38b5f97d4c · inbound

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems cites this paper.

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems AHELM: A Holistic Evaluation of Audio-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:37.780637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:37.780637Z digest=sha256:ac03a25825c9f21311e7e6a88d8c5075f3a0369db677bbedf611e5e586c7616f