Pith. sign in

Paper Citation Record · LEDGER

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

As of 15 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2506.06343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06343 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:44.194819Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:50:33.363981Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c806618a-d343-424a-993a-e6ff54626dfd · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Moshi: a speech-text foundation model for real-time dialogue

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.491594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.491594Z digest=sha256:57b814668df28857dcac9c32a0c947e816b6b511b624b0b4a5943b78c9b3dbd0

Observation cb7aac0f-dd7d-4a7b-b1c4-2e3a9c3bc669 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.561463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.561463Z digest=sha256:8cd525b961b188c756d61df9b8b99993c8c57bf8f6d54ac7cd4082ae4bc39858

Observation 4a70dcad-3b11-496f-98e0-1fda8883fab2 · outbound

This paper cites Qwen2-Audio Technical Report.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Qwen2-Audio Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.684392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.684392Z digest=sha256:76e43ecabeef3ca72554abda1c58e2cbf92d6f0ea84d7e3ceb97b97c3fdc8b8f

Observation 631d0733-5633-44b9-8519-626565e9a63e · outbound

This paper cites Baichuan-omni-1.5 technical report,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Baichuan-omni-1.5 technical report,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.810890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.810890Z digest=sha256:83df84f89f93335b1dce87c9b35d98df9e2c282b91faf0f231cfb9650481be79

Observation 076e0b54-13c8-48fa-aa92-1a3b5b45342a · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.924179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.924179Z digest=sha256:61d7d8c7a87bf584700adb0108c5496e637d3530cedbe332bbe4655dea675c8e

Observation 0e806ba7-755f-4f05-b107-ee54be2e91b8 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.015033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.015033Z digest=sha256:dc73c48727715af2e0b97859e388d61c59a33f0693a9c54a638f044a58a14f3b

Observation 123184ec-ddfa-45d5-854a-097c2d554f96 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.157874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.157874Z digest=sha256:83f397d91318c2be3678f2d184db322859694e7843f798acda000c93e064b8cc

Observation bed4a7ab-c49e-4df2-8453-db75df57dc7e · outbound

This paper cites InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.275667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.275667Z digest=sha256:c513d9cd53f782da3405fdc4d6774bf7950ec05ed544496e95267b2eb5bfca69

Observation 3fec8c50-3c76-4ff4-837a-7c90a0fa5985 · outbound

This paper cites Distilling an End-to-End Voice Assistant Without Instruction Training Data.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.386187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.386187Z digest=sha256:a3d8654900eed14aa3a74044ef67c777e83c06bbe47fd818ab8c829b35eecbf5

Observation 6df42d5f-d314-4609-a221-5560619ae578 · outbound

This paper cites Speechless: Speech Instruction Training Without Speech for Low Resource Languages.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:44.627131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:41.531921Z digest=sha256:b0d4c0685aa0d696d2147d9d8ef51780f787d254398778c95659e7c2bafad613

Observation 4e0f11e2-22a8-4db6-94d5-26769db66329 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.673699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.673699Z digest=sha256:c00d71750394df265fd55791de3038220aebe90ea7796c5cbad082929f058645

Observation c6c26aea-4ace-4858-abee-3cebbce44cf4 · outbound

This paper cites SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.805991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.805991Z digest=sha256:02b5f0b7e71a56b00c4a02e11e6b13780b044cf97c3c59d5ec9a9116766b874f

Observation 073b2990-7cce-476c-be3a-76ec126d729b · outbound

This paper cites mSLAM: Massively multilingual joint pre-training for speech and text.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment mSLAM: Massively multilingual joint pre-training for speech and text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.929162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.929162Z digest=sha256:2eea3484f15ef306f7eabe367950fb288b60426d665a1bf85afbce1d6a7a915d

Observation 7cc83ce2-6990-4d94-9e94-e5b4ad163fc3 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.065571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.065571Z digest=sha256:55854beea1f85e1e12fc612b7cf6f30014bd531c3e084e9131a93757cf677484

Observation d1f7fc13-e024-4ffa-882b-2ba265f37872 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.240760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.240760Z digest=sha256:8a2beece76bbec08b53fd5e8d70c8e8ad8478f04d9802ef243d077618066ad51

Observation 4e7d927f-2eff-4969-982b-308d2f6ec336 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language mod- eling for self-supervised speech pre-training,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment W2v-bert: Combining contrastive learning and masked language mod- eling for self-supervised speech pre-training,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.346892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.346892Z digest=sha256:b173a2c42bf0c2bf86d754f122c0a8ff8cf2ef9bad0ce8554a17b014c9b9fcc2

Observation a3c87180-6074-434f-bc2e-2c9d23dd2acd · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.482144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.482144Z digest=sha256:efa76145a798fa0d522d82310e09c933654c52a10977c4bdde87c81a2b43ab99

Observation 6ad31277-d93e-4068-a1cd-f9b30221ff29 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.585526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.585526Z digest=sha256:b334ed79f5114d81688d75c41c7d91f51e728eb790338fc7952730f3e5cbf94a

Observation 87aa7824-7b99-471b-86b6-de5acded7abb · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.752950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.752950Z digest=sha256:4e99f7d9a36e4124c6e1eae71072214ea98ca68fb064360af46d6c4068734c1b

Observation f8818343-d73f-4f41-966a-38e8a1e32526 · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Spirit-lm: Interleaved spoken and written language model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.543567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:42.886123Z digest=sha256:104b9d158843a963138ea410f8766472775adbb4068d5bf64c5aebe8ddf1b4bf

Observation cb7bff6c-c7a8-4840-8086-89d7b676e829 · outbound

This paper cites Dissecting learning and forgetting in language model finetuning,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Dissecting learning and forgetting in language model finetuning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.248933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:43.012070Z digest=sha256:a15f3ef9667a32757c042736b1f2f149573d0e20959c04d9c80940b5706c4610

Observation 12c6ad8f-6c7b-484d-8d59-4e2841806b24 · outbound

This paper cites The Llama 3 Herd of Models.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:43.136188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:43.136188Z digest=sha256:446aafda140672874c77921db32f09a2d8f67f217633ae25eb49042536fac87e

Observation e3aa62cf-f587-4f1a-a40b-15fce72d20b0 · outbound

This paper cites Openwebtext corpus,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Openwebtext corpus,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.929557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:43.312559Z digest=sha256:970847bb579b4bf04700b556813e45123228c034e30e9429d5e8f78771b78504

Observation 09a79cd4-a638-4dfc-a433-35c79d77fbd4 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Enhancing chat language models by scaling high-quality instructional conversations,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.639363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:43.450637Z digest=sha256:6e922435a06c2b13484f87d95a4de645be6c7e00c636eb2e0940ca6613cdee13

Observation 46b02533-0f52-45e5-9613-d44a4c88cc04 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.383713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:43.579967Z digest=sha256:7b7d53ae1d900dfd0cbe2fbb703f8d67b1b5ad9d4984d98b1fe476ec63fddd9d

Observation b10ec70f-cf80-4788-8a44-f29e9e14ff3e · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Can a suit of armor conduct electricity? a new dataset for open book question answering,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:43.752110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:43.752110Z digest=sha256:cb02eebec05714d5e41d8a3c9dc41eca6779f3eee884e685cf13dbf0fb89c27b

Observation dfadb9da-0797-43a0-b1b7-f9913edbc643 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment CommonsenseQA: A question answering challenge targeting commonsense knowledge,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.100191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:58:43.896352Z digest=sha256:fcf8a9617b4597c03ba1d05fdfb66624014ca8dd178031f0cb9dbf115342a24b

Observation bb23e8f6-66ff-44a6-8976-98ff6d1c6755 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Librispeech: an asr corpus based on public domain audio books,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:44.064280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:44.064280Z digest=sha256:4684afec07b584cc7dcc75981a4fdc7366cc721d7d9ca1177edc2f1444339b73

Observation 2a8231c3-cc6a-4139-afe9-79d01fc0f81a · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:44.194819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:44.194819Z digest=sha256:74f99229a154348b6dc36891a908cc39d0ea0987e55eb56c098eae4742635518

Pith citing papers

Observation 1302492f-f91d-49b9-82bd-e80a90afee7b · inbound

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond cites this paper.

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T05:50:33.363981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:50:33.363981Z digest=sha256:15ab1f04e22f7528c8a7c6dc4f5810f55bfd8480216479b4847db76f769dce3d