Pith. sign in

Paper Citation Record · LEDGER

Voxtral

As of 18 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 51 inbound Pith citation observations for arXiv:2507.13264.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13264 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:33:14.493016Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:55.435027Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8a9fb8ba-1b53-4d07-954f-11600f7cb224 · outbound

This paper cites Numbers should be written in English words rather than Arabic or roman numerals.

Voxtral Numbers should be written in English words rather than Arabic or roman numerals

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:16.306416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.503404Z digest=sha256:42bae646eff34e4cdce484fccb08da53b6ce6b764bc0ea8ced8753ae0798b414

Observation c7897d11-f762-442e-af86-a28e70170e55 · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:16.211766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.569506Z digest=sha256:7fd0cad25532ea9c158bb394bdb3fc751a04700bba231789027df9b36ea8c256

Observation d55d5a29-c116-404b-86f4-ade90004e13d · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:16.115400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.596841Z digest=sha256:fbf6be061361991c8efef31466eb9bc8c3a4508b9ecb56af27d262a9d635c4f0

Observation a57e3fd3-7048-4806-9ed5-3838dac06701 · outbound

This paper cites Only if the bullets start with alphabets, use corresponding alphabets like A, B, C, D or use Option A, Option B, Option C, Option D.

Voxtral Only if the bullets start with alphabets, use corresponding alphabets like A, B, C, D or use Option A, Option B, Option C, Option D

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:16.025785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.640166Z digest=sha256:a3dddf9acdb629ea3b86740c04163d1bb26319d8b387459cadf271289dc050f1

Observation 03c75ec9-b932-4f30-ba1c-6cd6d98bf3b5 · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:15.928827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.694139Z digest=sha256:3a684387ede032b35830d17b5bb7723166526e237739a9d8bb186f26950b0785

Observation a199fcaa-e111-4cfd-b6e3-a41ac343b1b0 · outbound

This paper cites For Eg: ’ffmpeg’ can be broken down to ’F F M P E G’, ’.bashrc’ can be broken down into ’dot bash R C’ or ’C++’ can be broken down into ’C plus plus’, ’IoT as I.

Voxtral For Eg: ’ffmpeg’ can be broken down to ’F F M P E G’, ’.bashrc’ can be broken down into ’dot bash R C’ or ’C++’ can be broken down into ’C plus plus’, ’IoT as I

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:15.852518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.774751Z digest=sha256:3146e7161bc5b26596db7f19be2bc1fc11f71d97e85b494d01226f74b8ab0445

Observation e915f3f8-f5fe-4164-a5fc-7db9d444687b · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:15.731213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.831441Z digest=sha256:0743752d8ccf9b2bcf9035fb179b4982db105704c74cf9647aa1ff069aec8b5d

Observation 844c15be-ecb1-45e0-934c-43f7665c3274 · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:15.616890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.882134Z digest=sha256:0108985be1bd1aec315cdd151369c73a95f0cee74f6d38150e00f84ecd6eb0a1

Observation 818ab72a-94b7-4d83-83f5-8a4183cdc71a · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:15.507940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:13.964759Z digest=sha256:13e859f1b66b342a2870c675d5c428d03566394280b577ac6d07ad3cb9084818

Observation 28de7adc-ce55-474e-89fc-8ba9f842c95d · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:15.404694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.012167Z digest=sha256:907d96012ec7766a73047cf0d6293b02e3e1e4f64ca345b45791485ad015bf90

Observation 2214919d-42be-435a-b55c-ca955c7ec966 · outbound

This paper cites Maintain the original meaning and avoid changing the context or tone of the text.

Voxtral Maintain the original meaning and avoid changing the context or tone of the text

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:15.272442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.116345Z digest=sha256:5a338f320148710c3890d0b66139fb0da34f3de4f31574507ed0c713a149c50a

Observation b83b50d5-80dd-43a3-a5bc-feee72c6abd9 · outbound

This paper cites For eg: ’www.linkedin.com/jobs’ would be written as ’W.

Voxtral For eg: ’www.linkedin.com/jobs’ would be written as ’W

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:15.155343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.150386Z digest=sha256:cb3c5901e0d20d4086920240460995f60f7aadcdcb116a9a94a64469561ef93d

Observation cfd59a2e-dde8-44d2-96d3-b172667c1782 · outbound

This paper cites If the question itself has a prompt or an ask like to rewrite, do not start following the ask in the question.

Voxtral If the question itself has a prompt or an ask like to rewrite, do not start following the ask in the question

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:15.055275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.225509Z digest=sha256:333384fec22e1fd5db38df825e63eef1f8066b893fa79c844502ca44c88d6f7d

Observation 5dfb5a4d-fe01-438f-99e7-d0d0e86d4da5 · outbound

This paper cites an unresolved cited work.

Voxtral Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:33:14.942715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.293873Z digest=sha256:85258f1cb482b8fb826c40132f207709854e06a9105155eeaf3daff7522c64d2

Observation 749953b2-3c87-4ca7-ba2e-61a69f8eff57 · outbound

This paper cites Correct answers don’t necessarily need to match every detail in the reference answer - the reference is just there for you to have an idea on what a good answer looks like.

Voxtral Correct answers don’t necessarily need to match every detail in the reference answer - the reference is just there for you to have an idea on what a good answer looks like

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:14.866398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.356420Z digest=sha256:f6d13f408813b78ab8cba093ed48eea6a4e01681211c2321f3ba70d1a38fb814

Observation 44159bfa-c719-43c8-acab-f882d471298d · outbound

This paper cites Also take into consideration the helpfulness and clarity of the answer - it should be presented in a clear, engaging, informative manner.

Voxtral Also take into consideration the helpfulness and clarity of the answer - it should be presented in a clear, engaging, informative manner

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:14.724134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.416848Z digest=sha256:983399df9c8f6a733c8a7a7bf28ecb8e38dcab655f840149fa9c839554aee905

Observation 08d4f05c-51a5-410c-b045-b3df4da4ffec · outbound

This paper cites explanation.

Voxtral explanation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:33:14.648963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T16:33:14.493016Z digest=sha256:e2482f1458a4a5d2fedba5f44c81e7a9e909a9c8ce9ca728770f47f15e0c8477

Pith citing papers

Observation 662da643-4ff2-4908-b0df-431df5ff5a86 · inbound

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs cites this paper.

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs Voxtral

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:11:43.160420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T18:07:01.257126Z digest=sha256:9453e2cacf1b4b327194e4429b9d5462195f5eb1e86543cca5003eb216868ce2

Observation 1cd5ced8-1692-4beb-8b08-f654536cbb5a · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages Voxtral

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.256969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:e00f9c1bad0e8c2eff185438c95e5af3746dc15d7a30e03aeb0462cef3669eab

Observation a0e3317c-5ce5-44cf-9647-e5ce27401e88 · inbound

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models cites this paper.

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models Voxtral

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:17.928697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T11:08:08.916893Z digest=sha256:951e2d24b3d33fc76b52bb8b6c774796ac9028baddf32c95309284da2f1839a7

Observation 0f6ce27a-0879-40da-ba58-40b297daf860 · inbound

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages cites this paper.

MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages Voxtral

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:18:57.194337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T03:15:04.685150Z digest=sha256:c805d2f253223f1f1435ba6d4a17900866ec61cbfa8fb2301c374846c799017e

Observation aedec695-5e30-49da-8b5f-1e495ca7d192 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Voxtral

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.744592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:e506ef5047fbca0eb35b070431ecdb55fb0f89affb085b0dcfbc6277cf9e3ba2

Observation 6f2c17ec-f531-4595-844d-f988b6c02138 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Voxtral

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:01.060858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:01.060858Z digest=sha256:464715e52b202cdfa0324547f2bddf769e9f699819f01a8c61766e2f470c496c

Observation 7fee6458-3f1a-4934-a8fb-d11077dcf17f · inbound

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus cites this paper.

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus Voxtral

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:31:01.948157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T14:30:46.664861Z digest=sha256:f4f9c9649e1341737f3ab212646e1ddf21708ef504bd01f64fa170ad19b223cf

Observation 6f18f1dd-31ad-4a52-824c-bc3794f08c4b · inbound

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization cites this paper.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Voxtral

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:17:47.177241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.177241Z digest=sha256:3827efc5778c66e1d5fdcff651511c4738ad3a96b4ddfd7826b9570c701ebc50

Observation ead7a03c-a489-4e31-909e-3f4ffdf8286e · inbound

Voxtral Realtime cites this paper.

Voxtral Realtime Voxtral

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:17:07.534451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T02:12:16.230433Z digest=sha256:6a289396583594b1383317bbb7d52ed2328902b9a3fa81dcf0edc17e0a11b32b

Observation 30e63df0-e21b-49e6-987d-3d7423224a24 · inbound

Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding cites this paper.

Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding Voxtral

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T13:58:09.323985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:58:09.323985Z digest=sha256:2892570308d92d74ac867714ae2ab698939dd259144a18f67bbe9ffb7a8d49df

Observation 1ed46d8e-0b15-47c3-a503-5d2f63050ffb · inbound

Voxtral TTS cites this paper.

Voxtral TTS Voxtral

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.904186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:5f13eb0c4130bd4e9fd2aded6249937ded59af5b8122c51139d26237eb11adc7

Observation 5ce1c60b-9d74-46cb-a146-1d958d5f5268 · inbound

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning cites this paper.

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning Voxtral

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:38:02.912734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T17:34:10.089555Z digest=sha256:22ce0085cce3d9b7e475f7edfd44d385ce583ad8440525cb4be6a1e788bab409

Observation ee50c181-d1bb-48e4-8ade-7d783804a987 · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Voxtral

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:30:57.273062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:d2ccb7d61c439a8a0bca00c804ca36765369e7e88f587983c730aff06d73361a

Observation 80077c05-c593-4017-a3fc-d8b955e24824 · inbound

Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference cites this paper.

Pushing the Limits of On-Device Streaming ASR: A Compact, High-Accuracy English Model for Low-Latency Inference Voxtral

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:21.947631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:04:25.418910Z digest=sha256:5305c3445f2097e15fc0f3cf08486dd3e73fb546e1ab17845766d4d790b092e8

Observation 90b0e157-88b6-4957-bf2c-41ebed10af62 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Voxtral

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:52869ec20d0d77e2aff04fd7adee859a0357f29e9cc292b8d551288b422d4807

Observation fd193cf0-bba9-4e8f-8dc3-90a2dbfc38af · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Voxtral

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.928382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:9db5d889d0d7326c8267bbafa91cc5a2260facc57db316c1555d81374969d50f

Observation b62a8db8-b7fa-4358-8b82-bd05cd3a6147 · inbound

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR cites this paper.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Voxtral

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:21:05.956205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T03:48:14.211240Z digest=sha256:c05d62687885e0f6678a49b02679e2f4168784ccdaf210223fa7b832bd378efa

Observation eb5755be-f0df-4074-9c12-38e2410bcec6 · inbound

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR cites this paper.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Voxtral

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T13:21:06.288305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-05T13:15:50.794969Z digest=sha256:c26ec392ffbf555cf906d5a53989a7f5204e549bb73b8b7f35fd6507f483628f

Observation 256912e3-6864-4baf-9ea3-8de8e581005e · inbound

Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps cites this paper.

Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps Voxtral

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:04.863844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T01:56:44.270045Z digest=sha256:df8290078e90d3b8750a860de9e502f902ee8dc1201f3ac4c2fba870c7516106

Observation 4b0de199-c1d6-4d80-96d5-4d692b1707f5 · inbound

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation cites this paper.

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation Voxtral

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:36.207134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T17:39:38.234052Z digest=sha256:b34482649bbc282c0988c9431ea5ced0f9cca914497a9057784b332a59d04c0c

Observation cddb048c-95c5-4ca8-89b1-7f2772e7c382 · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Voxtral

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:12.313110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:c87dba97449aa2c15f94a0b74ec659e78c5ae3c0896c5e3dcd6f0f827f60aacf

Observation 60e21894-abda-4684-9f83-0ce4229d7933 · inbound

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings cites this paper.

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings Voxtral

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:28:19.127591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T13:25:09.024054Z digest=sha256:52c8f285a442dde8b8e070fd3d0e8e892e7a05a9d1f51a13e0f098fa0fe28f03

Observation d3e0dafa-6f9e-4827-918a-b1b55040a066 · inbound

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding cites this paper.

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding Voxtral

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.390433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:56:38.345335Z digest=sha256:289661424d8a9c4ed119a1cb275bf457ffe3da5b4b2b8a30815f0293856a7f65

Observation d0f92b36-1701-424a-996a-d33d1a2d8f79 · inbound

RealityTest: How People Probe AI Identity and Whether Models Disclose It cites this paper.

RealityTest: How People Probe AI Identity and Whether Models Disclose It Voxtral

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:36:08.724825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T22:22:43.432328Z digest=sha256:e31e6cfce46bfee0ea32a45aed4fc841178c8ab5ed707c848e67fe767a8234d7

Observation bbdcde03-ba1e-4cec-ae6b-f9221912c4b9 · inbound

MURMUR: An Efficient Inference System for Long-Form ASR cites this paper.

MURMUR: An Efficient Inference System for Long-Form ASR Voxtral

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:12:25.096494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:06:02.978205Z digest=sha256:4553ffd9fa4e094ca6568a20ce52278a355a71395a3527506f476e0055366e6b

Observation 7584a4a5-0da4-48d6-8531-a78848bbc4cc · inbound

Benchmarking Speech-to-Speech Translation Models cites this paper.

Benchmarking Speech-to-Speech Translation Models Voxtral

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:46:29.444473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:34:10.510906Z digest=sha256:b4652c8bb8e9f1692ede0ccdaa915869a29826f02e973c233abcb3894a60bdd6

Observation a0d2211e-f890-4faf-950f-365b3c96cbf4 · inbound

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech cites this paper.

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech Voxtral

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:37:06.724970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T23:44:42.188662Z digest=sha256:1eccf58d68fb332001a24219a06cd9c4636ae16b3808f7427af3d964ea9a4a1a

Observation aafcd460-72a7-47fd-8edf-415cf5d7ba2e · inbound

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios cites this paper.

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Voxtral

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.045645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T01:11:56.037674Z digest=sha256:fd7060490c6683d1cc55ebfbee361dedbad2ca8dae469272c6b7e715a72a44e1

Observation d44d6b97-d702-4272-b6db-bac2e0c8b265 · inbound

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition cites this paper.

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition Voxtral

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:57.087582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T01:58:13.871771Z digest=sha256:770935a66e38eb9a202f281596de25bd7592a868d99ff74912e7e938617fc653

Observation b6a6fc55-6346-4ddd-a7b6-aaa034707fd8 · inbound

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models cites this paper.

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models Voxtral

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.526077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T19:51:59.265579Z digest=sha256:cd54a1c4acc0c5f9f6ac99fde81e7a69bee14fed3082772a26207d0ee9ed8e7c

Observation 1dd1ca22-5f05-4303-a152-1a539f8dad00 · inbound

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages cites this paper.

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages Voxtral

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.489322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:10:03.282212Z digest=sha256:d7fef697ff6009a5544870e974959d8c3d07a614413252d438e22fd4acb74880

Observation e1e30ead-2568-4317-951a-3ec879112b81 · inbound

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations cites this paper.

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations Voxtral

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-27T07:10:41.472377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T07:06:13.658919Z digest=sha256:54276940c29352c2db9e91d65ff615b9b062c56bb51a24b2955a9ff3ad5939a6

Observation 6725bc1e-3f7f-4e9a-9927-0a6805a368f8 · inbound

LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning cites this paper.

LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning Voxtral

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:13.111737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:13:38.489672Z digest=sha256:73486efb9e7775c8bc29e44d7a2aa6ab0c103ae883aa65fde896fb857d0b08c2

Observation 40abd6fa-fb90-4d9c-aa60-a7a9a732fbf9 · inbound

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning cites this paper.

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning Voxtral

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:03.351616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T22:39:09.967750Z digest=sha256:ad738881409cbebb98d09f75117854a3ce1957e1de713d2c6c0633cbc3afbd20

Observation d77aea3a-6b1f-445a-a6fd-03536a02d0bc · inbound

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages cites this paper.

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages Voxtral

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.345552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T19:14:26.452851Z digest=sha256:027495a82d7a095b9e93d3c595b96f2430a36dad78edeab970fbddba93fb1427

Observation bf1f3ad5-679f-4774-ab11-7bdcb3f241b0 · inbound

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi cites this paper.

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi Voxtral

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:39.279437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T13:07:37.187581Z digest=sha256:ac3156bbb052a92d3b86df1bec2e28336f5e6c968cdd8fc6bc6d8de4507420ea

Observation d8fea7a4-5eef-4f8c-8b54-d05aaf81e144 · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models Voxtral

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.978640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:627da4368ef9af3c674661b49a50006c9296b2b57cda660a089d4a6d135bc909

Observation d2974562-cdae-4ac4-9b9b-b750ba3a5404 · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages Voxtral

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.910091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:67b2ff168c914624161a9c35afd9b553683d5c4471d701c967fd7631a99d6b8a

Observation fe7dcb65-5b71-4ae1-8988-2dca377b95f5 · inbound

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving cites this paper.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Voxtral

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.663262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:5e6f4ad25a9abae6b84e19dbe2b5b2d21b8cff677fd1fba9a3a3817ed65d3da8

Observation 58468181-dd03-488c-b1e6-6a5ecce3a2d3 · inbound

S-DiverSe: Spanish Diverse Speech cites this paper.

S-DiverSe: Spanish Diverse Speech Voxtral

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T04:08:13.798171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:08:13.798171Z digest=sha256:6b486920ae60f866c903f8d3f0109f1f8038a767abc319c2f33e98ff477d37a0

Observation 16a04cd9-a983-4ec9-a239-ed53e31d7c02 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Voxtral

Reference 234

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.545770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:c4a529f0fcdcf930352dd24cd3381dde4f8c0faf8f42d6ccd7886f4a2430e5c1

Observation f1cf07ff-4252-4e63-9d6d-39b3b03ac91f · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Voxtral

Reference 234

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:910a05c7cb8950b5d9bafe8a59b750c83b4f7898c0b507d902d0363b8102c554

Observation 3db9f10c-f72e-4a63-802f-e476e6864eec · inbound

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs cites this paper.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Voxtral

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.607654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:78ec8f04a36f479142917489b6ea3627801e227ca17680358c25e60ab23320a0

Observation 382579c1-70e4-4e99-9485-5309bbe4cacb · inbound

GigaChat Audio: Time-aware Large Audio Language Model cites this paper.

GigaChat Audio: Time-aware Large Audio Language Model Voxtral

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T12:07:39.817307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:07:39.817307Z digest=sha256:9b91fe7ad4915e8352f479c47cff204729aa467ddecb13f5e877a38b992728c7

Observation 608bbfa2-db29-48f0-b7a6-465c09ffa510 · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Voxtral

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.370603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.370603Z digest=sha256:044e27e2a54a32bad2802d284512fd6cd2eb68d8b7eb40420a3bc25422d20259

Observation eee421ae-6f81-4633-954c-6d145337df2e · inbound

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation cites this paper.

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation Voxtral

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:08:03.719567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:08:03.719567Z digest=sha256:0d5a6f85bdae478b9cac46304257e798374d2258393454b3fce5b5a98c78e80a

Observation 8e8acb12-a731-4957-b1d3-2163a6c723c3 · inbound

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge cites this paper.

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge Voxtral

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:42:01.682549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:42:01.682549Z digest=sha256:4dd78043e7f155c1c0ffc9c104cd1503b437ddf6ce382280fbdeb7334d898789

Observation 9ccec7c0-c125-411f-809c-446281cc4a93 · inbound

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions cites this paper.

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions Voxtral

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T16:57:11.564423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:57:11.564423Z digest=sha256:0b918e3bc68711a63182abda400994cf63a46a9814881f2676b7e07f0c19ab19

Observation f66feb1f-3887-4922-a1bb-388b081b4484 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Voxtral

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.795578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.795578Z digest=sha256:6e3ca240ac1606c2bfdf57ed04584794c0fd1a2331837f641f77ca533f0a6ebb

Observation 1a6436a5-9d86-4ce1-a02a-b3b484edd004 · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks Voxtral

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:21.284618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:21.284618Z digest=sha256:0999b646011ba424a0a082e26908611e1dc9df2db1d23ec139f74ebc098e9e09

Observation 0c3b5525-9e95-4a16-acdf-ccbc3b7bc1d4 · inbound

Alignment Drift in Single-Model Speculative Decoding for ASR: Mechanism, Correction, and Cost cites this paper.

Alignment Drift in Single-Model Speculative Decoding for ASR: Mechanism, Correction, and Cost Voxtral

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:55.435027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:18:55.435027Z digest=sha256:cfae2aa0c6cfceb687ab27fce162d3e0830c397b18834396c51327cac9aea72a