Pith. sign in

Paper Citation Record · LEDGER

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.17019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17019 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:09.446488Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f723c553-dd42-4d7d-b2bf-1f1b8bea7139 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.447360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.447360Z digest=sha256:5fb3f90850399a86112c817fb5fa7d6c6f070411cf4119a09721fa9574bff51a

Observation 88e5864e-a2e5-4be5-a0b6-af6a54601ce3 · outbound

This paper cites arXiv preprint arXiv:2503.10620.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning arXiv preprint arXiv:2503.10620

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.624600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.624600Z digest=sha256:4a22cd3d1f6cc846c4a8994aa60055c5fde7450208151355c7ebd2d86de1ffc3

Observation 7e6a1902-06e1-4559-98fb-e208df2133fc · outbound

This paper cites Qwen2-Audio Technical Report.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.960891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.960891Z digest=sha256:4588e1995fb796cbab52074a4d43357a6d8031394e54475deb95b8a7158ab64e

Observation 483e7c21-5dfc-4b52-bfac-70f87284c879 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.072917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.072917Z digest=sha256:a59e05fcec753b62c3e648c31f8276f29108b6e7839af5f387fb426de2547b30

Observation a3ff050c-a152-471c-940b-a48d59925cba · outbound

This paper cites In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.169768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:07.129512Z digest=sha256:95b8f85bc2c29cef5e23b19f72488c5da2b1dc34d80f6d061c421c31576456cf

Observation 3e69669d-9fc9-4cac-bb93-3d48cdd6be69 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.278388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.278388Z digest=sha256:b4127cf7bb329daf6cae269a03af8f589d50b07a48072ca8c8fac8eca58c5683

Observation 6dcc4476-94e2-48e5-b951-dbf7ab1d4ac7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.562021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.562021Z digest=sha256:4d44b72026a52a9b179362c1b5c20700a587e35346cb6c8cdf1556020a9b42b2

Observation 36ee1ed8-a703-4db4-8cbf-4966191d8a52 · outbound

This paper cites The Llama 3 Herd of Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.676909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.676909Z digest=sha256:268ad4ff256510ad82d40ca4e8c47b621fac37ee5f69086a0dbd5294abbda50a

Observation c7cc3634-92b5-4eea-86da-8264703a0e80 · outbound

This paper cites In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2024, pages 4552–4572, Miami, Florida, USA.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2024, pages 4552–4572, Miami, Florida, USA

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.046606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:07.937565Z digest=sha256:b0e33d4b8ff597f0a9c47354d38324544537ccbc8b10b463256e143a8006b81a

Observation d1bc3b29-c5cb-4d4e-9183-02054e959fba · outbound

This paper cites Speech Translation with Large Language Models: An Industrial Practice.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Speech Translation with Large Language Models: An Industrial Practice

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.081674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.081674Z digest=sha256:6714276a34b1e79b84ae148efbca01757cac38b7bcbc838a10b1d64567b163b8

Observation c2cc3502-ee96-4697-b803-f288712f68e8 · outbound

This paper cites In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.732535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:08.247414Z digest=sha256:adb5c9bccf8e44d4b7bba2ab5205b0f27154db23fa1f7e3fb18ef35cf1a88d3c

Observation 912abaa8-18dd-420b-b3a2-6c4210d234ec · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning EuroLLM: Multilingual Language Models for Europe

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.333794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.333794Z digest=sha256:7eebbe9c4e06f3ecce285a2f2df2052380ffe4dabde86483d484072bfe0e7af3

Observation 52b9ffae-d251-4592-8138-9f2843071aff · outbound

This paper cites an unresolved cited work.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:10.304258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:08.694602Z digest=sha256:a68a3df7df61dd0a094ffa9bdd9de5313ba7522b0c3553bbbf5bae58f8164955

Observation 8be90602-4665-47d1-9e30-66dedbaac60b · outbound

This paper cites In Proceedings of the Seventh Conference on Machine Translation (WMT) , pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid).

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the Seventh Conference on Machine Translation (WMT) , pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.105964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:08.823192Z digest=sha256:2dde0475e3789145b8c0cad3b8053d7d7c2000e846585d79e06a93f8c6b51167

Observation 3769e255-2bcd-4641-a8d3-7107add7efd4 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.892983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.892983Z digest=sha256:d962a9e7ae976f091192474729960a7772504181cc07115d18ec1fb7f1818467

Observation ee251e65-6943-49ba-b351-b7e7d8353639 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.989599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.989599Z digest=sha256:428d8859b7eb2ee91e993e98daaaf35a7fd1d58943156439d31c98ff8abb6156

Observation 9d90560f-7ae2-47ac-a5cf-b99fa37eaee7 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.173525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.173525Z digest=sha256:aabcd09a82d79ca224fcf3d60cdfc6759e56a40a8e0ed5797b8e9238b63f43fc

Observation 2cd2468b-8d69-4899-8a8f-edfd7ca0502b · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.324302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.324302Z digest=sha256:a1a443f4d03319fa7d5b0e2ff826d054812c0663cf16e423a6007244750ac880

Observation cc82fe11-6225-4091-9bae-df733af749d4 · outbound

This paper cites Qwen2.5 Technical Report.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.446488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.446488Z digest=sha256:da82bf7ee09c6b44ca31bf6f5a1e22d5aab1ba76dc5472c5126bb55c9fae2760

Observation 75a282bb-6d7c-4218-9015-e2b28c55ef58 · outbound

This paper cites In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.497929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:08.567919Z digest=sha256:dcd591d8072963e60775e3b2b32ea147f1d617f4fb9bad03da7483d427ee7093

Observation 3cbfb931-1243-401a-9c4d-1451cfd44ed1 · outbound

This paper cites an unresolved cited work.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:10.891849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:08.177645Z digest=sha256:93fc8517ec0c2e10ff0a246f46ca78dea6d4259cf309572b89d5b00bc1ea40d1

Observation da49e256-227c-4898-9ed6-4c9792f3051c · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.813102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.813102Z digest=sha256:8c6642de7ed19b29f85bbe19dad7f2ef3e329de6f6e5a4b2a38e4b7fa23dd4c2

Observation 82dc9500-d820-4672-8a3b-a1b4f3ff12c7 · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.432916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.432916Z digest=sha256:9e2cff696a85378c772004e2535cffd4409c227711e4faa18fe13b302553f886

Observation 97ce2ada-7205-422f-9073-2058f440d569 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.449809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.449809Z digest=sha256:c840b47770a00711d194b9810148e36d3986b8236d6e1fc7e8abdc2ed195750d

Observation 8117339d-05be-41c1-85d4-dc3e9f67baea · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.852611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.852611Z digest=sha256:eeb4f80808f08b543278e51f2a1b7c0c91b49483630e76e00fbb7b96286b32da

Observation cf7879b9-d990-47df-88d3-753e3a0f641b · outbound

This paper cites In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21318–21340, Miami, Florida, USA.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21318–21340, Miami, Florida, USA

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.300091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:06.715685Z digest=sha256:5af00ecb7beb5734a17f8fb2044d0cb25f4efad81966d1ae512943b48dd1cdae

Observation d21fd134-09ea-408a-994e-1be5c6334827 · outbound

This paper cites In Proceedings of the 22nd Interna- tional Conference on Spoken Language Translation (IWSLT 2025), Vienna, Austria (in-person and on- line).

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the 22nd Interna- tional Conference on Spoken Language Translation (IWSLT 2025), Vienna, Austria (in-person and on- line)

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.419483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:34:06.365008Z digest=sha256:3e10ac08d66207d5bc43103db4188f3f511a7398d73d24bfe0c029095cf4fa00

Pith citing papers

No inbound Pith citation observations are available.