Pith. sign in

Paper Citation Record · LEDGER

ALAS: An Automatic Latent Alignment Score for Audio Language Models

As of 13 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2505.19937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19937 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:23.830824Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 117a7724-c508-4d03-8e14-3f0b0bd86268 · outbound

This paper cites GPT-4 Technical Report.

ALAS: An Automatic Latent Alignment Score for Audio Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:21.896626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:21.896626Z digest=sha256:4fc8480155b873b00deb928d530a7aa1af6dce747be7e05fd6fe15b025202bef

Observation 5f76d5ca-c8fc-4a70-b4bf-5d5efddd2488 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

ALAS: An Automatic Latent Alignment Score for Audio Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:24.781925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:07:22.050159Z digest=sha256:49037ac9df618a1cdac13d7aa5146247fb17d19c1a4f84ba665b9a12f985577d

Observation da7506df-fcba-433e-8503-7e8253153d93 · outbound

This paper cites Qwen2-Audio Technical Report.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Qwen2-Audio Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.240948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.240948Z digest=sha256:8e94f9a8fdfad8e34dcd4c58ad0750479e99474705d1060f1f249416fa6a39c0

Observation f34a3591-95b3-4b10-9b4e-d2b3dbe235ee · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

ALAS: An Automatic Latent Alignment Score for Audio Language Models GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.407437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.407437Z digest=sha256:f23b7157c03f12ecb077130f2f95f012ee8a580aea2ebb790580047eb712ecd8

Observation 4ee1ee6e-52f7-4975-adb3-778e897e1cef · outbound

This paper cites Listen, Think, and Understand.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Listen, Think, and Understand

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.512637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.512637Z digest=sha256:2cfc2067ec5146af4e1fe724049a02850526bea32246082df5d4eb7afc02ff21

Observation 838fc321-f817-4d08-ae20-4c03750c6c18 · outbound

This paper cites Less peaky and more accurate ctc forced alignment by label priors.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Less peaky and more accurate ctc forced alignment by label priors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:24.668628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:07:22.654319Z digest=sha256:38e7105903b67dfe85025b68abdee8585e9c2ac6ed2755c6309569062db9bafd

Observation 0b8de5be-e226-4aab-a95d-1bee71a71793 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

ALAS: An Automatic Latent Alignment Score for Audio Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.834050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.834050Z digest=sha256:dc64c7feafeccd8cb3906fe84b566919f924de967b383dcaf96308d7b851c739

Observation 40d0e662-d768-4050-b547-09b89b29de55 · outbound

This paper cites Montreal forced aligner: Trainable text- speech alignment using kaldi.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Montreal forced aligner: Trainable text- speech alignment using kaldi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:24.565292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:07:22.886727Z digest=sha256:0b59dfdf672cd09d9f9b11194987a61cfc2234819a64f69ffa06c96792b45f49

Observation ea1ca65f-c90e-4d56-ba24-7e9ba8b2ebd9 · outbound

This paper cites The 5 ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs kaldi speech recognition toolkit.

ALAS: An Automatic Latent Alignment Score for Audio Language Models The 5 ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs kaldi speech recognition toolkit

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:24.448199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:07:23.016716Z digest=sha256:22be677697ce667df653d03ee7f8f186e76a249cb62340918eb052cb5c2d46d6

Observation 022b5e44-29a8-4f66-be7d-462e045deee4 · outbound

This paper cites SSR: Alignment-Aware Modality Connector for Speech Language Models.

ALAS: An Automatic Latent Alignment Score for Audio Language Models SSR: Alignment-Aware Modality Connector for Speech Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.282165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.282165Z digest=sha256:a4b50d83e94fa2f895bc70f5c00a65f8d5e2919714cbca7c1b05bc069137fa63

Observation e4c78260-337a-4ee4-8ff8-73621a02d77b · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

ALAS: An Automatic Latent Alignment Score for Audio Language Models SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.341592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.341592Z digest=sha256:175d959bbed6534edb062fcd82137a497af26369430c3f29bd28f08ecc8e19d8

Observation 2ad01c76-6895-4f58-90ea-6e42eb139049 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ALAS: An Automatic Latent Alignment Score for Audio Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.446712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.446712Z digest=sha256:ddc93963a136f16cb51867cbdc9f5e146735d815c2cb93af5b28948beb8b40ff

Observation 20831a34-18e7-4c68-9a0e-3efe59b7320f · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

ALAS: An Automatic Latent Alignment Score for Audio Language Models AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.503947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.503947Z digest=sha256:96279fa9ee56bfeaf8748ff7342eaff1ec17d3940962329d735fc77fd9dcaed2

Observation 507516f7-2aaf-42b5-9c23-527be28a2e5d · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.623745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.623745Z digest=sha256:468c7aa3329ae39cee3e32eb46d2a3946b7cb5e518348adc0229b8199ac60163

Observation 272009db-ab9b-447d-a45b-7dacbfc765b0 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

ALAS: An Automatic Latent Alignment Score for Audio Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.738043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.738043Z digest=sha256:c7106c5069c3d6e14f597552e302e7dec4836286bbc1bd9c4509b1a36eaf0caa

Observation 905bd063-3264-4878-a381-fae90558398e · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.830824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.830824Z digest=sha256:1b7f754eb899bcd5b8251bedf771c246c0d4b85d69c480a653d8598036091b88

Observation c481c263-ae91-4edd-b67e-aece03537254 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

ALAS: An Automatic Latent Alignment Score for Audio Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.593442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.593442Z digest=sha256:9ac242d75bf4a44b35db2f9f425feb0ebfc02cb247b9e01115906fe606f288de

Observation 01e6943c-7841-4421-ba24-c3679471688e · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

ALAS: An Automatic Latent Alignment Score for Audio Language Models X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.142203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.142203Z digest=sha256:628d73aae1b5bb3550f61b7c36e80eb6f44a6d97a2cc0d3864c387d0bc0a15b8

Observation 4275cce0-b4bc-449c-ab99-cf883ff0d518 · outbound

This paper cites LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs.

ALAS: An Automatic Latent Alignment Score for Audio Language Models LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:07:24.190776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:07:22.953311Z digest=sha256:373b5afbdcc55a224eb3b4604d0d930d2aa0712322db5f04a095eb599453ee63

Observation 171b07cc-0e63-4c17-adaa-75eadafd4fd2 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.188014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.188014Z digest=sha256:20375d0336fc54a9c6de87adb0a344687a8ab506cb614e2c60f2f6925fd6df75

Observation 797a0878-61ed-43b0-9bde-1b475a03866f · outbound

This paper cites Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.739092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.739092Z digest=sha256:3e0af518398957c315803fcf5dc83a59adac023e425d44497bd56151afa1de36

Observation 653a531d-516d-43c4-a81d-4a0d21144b4e · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.097224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.097224Z digest=sha256:cf704ea917156da7a278c3278357c384199ee58e310384115d480b46f2f1ce33

Observation 4c46a8f5-927f-4f61-8981-28f0fa111913 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:21.971099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:21.971099Z digest=sha256:d52eb6cf66e48b1705ff91828e59efba7e5043321bbbc94111286e26318ebbbf

Observation 09b101c0-a0f2-4d13-ba31-d07cf36a437f · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

ALAS: An Automatic Latent Alignment Score for Audio Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:22.328026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:22.328026Z digest=sha256:b7f0a4828e63b04332d224b73300cbf7516edb1450fd329e61cbf593b73a0643

Pith citing papers

No inbound Pith citation observations are available.