Pith. sign in

Paper Citation Record · LEDGER

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.17019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17019 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:09.446488Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f723c553-dd42-4d7d-b2bf-1f1b8bea7139 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.447360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.447360Z digest=sha256:6e08aab90e9477e861ea8f2790a56596f087c7d7fef81293877179edc05de168

Observation 88e5864e-a2e5-4be5-a0b6-af6a54601ce3 · outbound

This paper cites arXiv preprint arXiv:2503.10620.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning arXiv preprint arXiv:2503.10620

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.624600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.624600Z digest=sha256:c97724a721e07138a9a1a7faee810d5dcfa25c3a27854887b076d7e3a62811e6

Observation 7e6a1902-06e1-4559-98fb-e208df2133fc · outbound

This paper cites Qwen2-Audio Technical Report.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.960891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.960891Z digest=sha256:34de8d478fdda016b05583b5e5a461da36bd0b6245733b0072a8137b7fe34420

Observation 483e7c21-5dfc-4b52-bfac-70f87284c879 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.072917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.072917Z digest=sha256:589849e1f31e0d66030d85fccb52c848b7d776cdd4a00606444d620d5e5c1f6d

Observation a3ff050c-a152-471c-940b-a48d59925cba · outbound

This paper cites In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.169768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:07.129512Z digest=sha256:6d3985fb2d27d8581c1e4163862d96b8628a6400de1b5ac888b2f85ea469b64e

Observation 3e69669d-9fc9-4cac-bb93-3d48cdd6be69 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.278388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.278388Z digest=sha256:e240969a852371cbd6947c464878fb9a5ae5569d0fae5a2ccf4f43087d885c43

Observation 6dcc4476-94e2-48e5-b951-dbf7ab1d4ac7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.562021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.562021Z digest=sha256:6759e5407f8e1f51f1c1ce73f35f63cb6542d0792aead82fdc56cd5fc1514f8b

Observation 36ee1ed8-a703-4db4-8cbf-4966191d8a52 · outbound

This paper cites The Llama 3 Herd of Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.676909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.676909Z digest=sha256:73ee0778efe677abd634a369e7e51a3aa34a6c2c9f7b422dea12b3c11bb3b084

Observation c7cc3634-92b5-4eea-86da-8264703a0e80 · outbound

This paper cites In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2024, pages 4552–4572, Miami, Florida, USA.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2024, pages 4552–4572, Miami, Florida, USA

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.046606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:07.937565Z digest=sha256:2b165d2ecccbc4b2190dad5ea97884dcb841a02df701b9951ac0c6ebb97a4ac8

Observation d1bc3b29-c5cb-4d4e-9183-02054e959fba · outbound

This paper cites Speech Translation with Large Language Models: An Industrial Practice.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Speech Translation with Large Language Models: An Industrial Practice

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.081674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.081674Z digest=sha256:c52ca0630c37dc7c7bea4ef1060bd2fb3fd7d6cafe65f93487bd26af961ebaa4

Observation c2cc3502-ee96-4697-b803-f288712f68e8 · outbound

This paper cites In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.732535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:08.247414Z digest=sha256:243bab0b3de87990138955875cf53abf4059a1a4c35858814fbb7b727b70300d

Observation 912abaa8-18dd-420b-b3a2-6c4210d234ec · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning EuroLLM: Multilingual Language Models for Europe

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.333794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.333794Z digest=sha256:fd6831118f9b6c4ae06528724c995650df101d14a29008165770b813a2a3dbec

Observation 52b9ffae-d251-4592-8138-9f2843071aff · outbound

This paper cites an unresolved cited work.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:10.304258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:08.694602Z digest=sha256:7e30686d0f8be11235e941f67634630b69bdeae9757a6f28f49e60947971e988

Observation 8be90602-4665-47d1-9e30-66dedbaac60b · outbound

This paper cites In Proceedings of the Seventh Conference on Machine Translation (WMT) , pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid).

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the Seventh Conference on Machine Translation (WMT) , pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.105964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:08.823192Z digest=sha256:58fceae70621cd5c3c4683b2fe6fa42ff0f1190bf95c5afda46b32a18dd40b86

Observation 3769e255-2bcd-4641-a8d3-7107add7efd4 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.892983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.892983Z digest=sha256:3bf8667effeb6b76e1ab922a821bee682422ad79bdecda8962f40cfce99817d6

Observation ee251e65-6943-49ba-b351-b7e7d8353639 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.989599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.989599Z digest=sha256:860c385b8762eef1623f618b3d3f5ea437dc413519949ec0c82218cf08dce097

Observation 9d90560f-7ae2-47ac-a5cf-b99fa37eaee7 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.173525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.173525Z digest=sha256:1cddd90abe87673b27793280a8b3adf1ce60922026e052e94df6d21223413176

Observation 2cd2468b-8d69-4899-8a8f-edfd7ca0502b · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.324302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.324302Z digest=sha256:ea5b51abca59aec330716a93f8311ecd81cb929d3f4c0f18573b6957536a1edb

Observation cc82fe11-6225-4091-9bae-df733af749d4 · outbound

This paper cites Qwen2.5 Technical Report.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.446488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.446488Z digest=sha256:b0e55349a05d7b6a8eace3fc47c314523660bf38095796d66b55510b8a78935f

Observation 75a282bb-6d7c-4218-9015-e2b28c55ef58 · outbound

This paper cites In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.497929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:08.567919Z digest=sha256:fb8766e49c54e5681b3749f166a7aebc08ebaddb66f13b1398c35b20f24e9d38

Observation 3cbfb931-1243-401a-9c4d-1451cfd44ed1 · outbound

This paper cites an unresolved cited work.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:10.891849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:08.177645Z digest=sha256:ebd0536f0d23e63e7d13c326c93b7a964521713980e0c973dce6149af7bea643

Observation da49e256-227c-4898-9ed6-4c9792f3051c · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.813102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.813102Z digest=sha256:28a36f0ec5f6162919012f75dfb25bbc4bdf0080f46444a3b5ece6f33c147d08

Observation 82dc9500-d820-4672-8a3b-a1b4f3ff12c7 · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.432916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.432916Z digest=sha256:fac2f344d84c172a3d3b02955311bc742b86d5e0998e9f18bb8db9c6c54e46c8

Observation 97ce2ada-7205-422f-9073-2058f440d569 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.449809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.449809Z digest=sha256:e2bf8f6e6a95695ff6bcfaeaa142e8737f134349c1cfba17ce4f251fd31c642e

Observation 8117339d-05be-41c1-85d4-dc3e9f67baea · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.852611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.852611Z digest=sha256:1c9e47e00ba0d1a81a5f8b9c113523eb50ecf16408b874c6ba3066b76cb2eba5

Observation cf7879b9-d990-47df-88d3-753e3a0f641b · outbound

This paper cites In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21318–21340, Miami, Florida, USA.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21318–21340, Miami, Florida, USA

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.300091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:06.715685Z digest=sha256:7351376e598ab84bf93aa4dc5da870bb589f2a6c5b6d09d98e9d93e3606a8b22

Observation d21fd134-09ea-408a-994e-1be5c6334827 · outbound

This paper cites In Proceedings of the 22nd Interna- tional Conference on Spoken Language Translation (IWSLT 2025), Vienna, Austria (in-person and on- line).

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the 22nd Interna- tional Conference on Spoken Language Translation (IWSLT 2025), Vienna, Austria (in-person and on- line)

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.419483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:34:06.365008Z digest=sha256:48125e58654409638ca44308b11fe55b42744f04e4c7688ef59a7734ae45e5dd

Pith citing papers

No inbound Pith citation observations are available.