Pith. sign in

Paper Citation Record · LEDGER

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 7 inbound Pith citation observations for arXiv:2505.02707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02707 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:47:52.778178Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:55.251451Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved53
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e124b35f-b9c8-45e1-a784-0dfd85b264ff · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.506447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.506447Z digest=sha256:6dbc218dcd3081fefad031c3cc121b6cc4e3fd5a5afcf14fa296628f7d3e93e7

Observation d8be03d7-7dcc-4c1d-b70d-58b1911250be · outbound

This paper cites B \'e rub \'e , M.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play B \'e rub \'e , M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.494593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.511006Z digest=sha256:5ce0a4c04ca758b054414f1e4e463c7d5d07eb99312419c4409f9c5daf17d888

Observation a510f1b4-927b-434a-962d-9f3fad62be5c · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.484279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.514268Z digest=sha256:0d3e16c58910b039dd0379cfadd00d2b60ccaa173c14fb5955bd0c43f87fac14

Observation a2033109-87c9-4f83-9d1c-4f044543f15f · outbound

This paper cites Borsos, R.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Borsos, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.517186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.517186Z digest=sha256:efb88699c30bd901ea6fba5c0bc8a73b9c75721a7b8b0bf69335c83cc5749185

Observation 4b35f937-b2db-4fb5-9955-91b3c754920b · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.466867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.521195Z digest=sha256:04fdfb1b894de2458e848e850eab498e2d0acec148fb83b27946998e85184ba9

Observation 17bb2e35-9205-470f-8be0-6d7182b9cc2c · outbound

This paper cites Buyukgoz, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Buyukgoz, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.456685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.525052Z digest=sha256:954c7ba920e5e598e84239434921277314e9e29f735171751374988234124f95

Observation 56c71573-1471-4375-ad9c-1864319638e5 · outbound

This paper cites Casanova, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Casanova, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.444870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.528913Z digest=sha256:a5a209a30ddae8dd0eddc0ea19abeffab8f937198dd6f4d2d5e50641fea80ecc

Observation 8903aae6-1e15-4f3b-abd8-a70f88f14332 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.434951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.531889Z digest=sha256:a14d78b4f756fcdbb053ed14a64835e348ae8b832d95429dc3684e3129c6ef7b

Observation 59d1fd59-0ee8-4a84-8bb0-b17c4195c86e · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.535285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.535285Z digest=sha256:c3650ba6d9d0ca52775b2d0832fc92d09d331df927782f3d94f807c647133209

Observation 519002cb-9b15-4bfb-bb92-b18a5bb4d9b3 · outbound

This paper cites Qwen2-Audio Technical Report.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.538991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.538991Z digest=sha256:8b1cfc4d0d428434bc89658d5efdb16c594b77a692717fb5f294b9fb247c4973

Observation 67784ff6-797e-40cb-aff3-b5e552203ed2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.542207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.542207Z digest=sha256:3bc70794094d878abb9d84cde69e1f4f91e94cb01e26eba25b1fd5bd79f5b722

Observation 46bf2d6a-3dad-426d-bc88-c789309f10ac · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Moshi: a speech-text foundation model for real-time dialogue

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.545213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.545213Z digest=sha256:bbe453ce7a8120f1be71f6df2a2a4f0b2941d21ef5fdf022e46afddf89fcf0ba

Observation c5bcc4a5-56cd-4a66-8135-89fa41d12f68 · outbound

This paper cites High Fidelity Neural Audio Compression.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play High Fidelity Neural Audio Compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.548376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.548376Z digest=sha256:4eb668b8b7264199a8df485001db89ce716c85e37d7f6eca141fd8a538cfbd05

Observation 39adc5f5-de4b-46f5-9e91-8dea42713698 · outbound

This paper cites Faruqui and D.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Faruqui and D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.425302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.552419Z digest=sha256:02827bdba810ea75daa326ab6c3764d82104f981c14c227d0d746f733facd5e8

Observation c3771bc7-2dde-4bdb-b3b9-b4c92714756b · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.415727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.555483Z digest=sha256:6bc437d71c874478cbd565632b8ee6014e2e156419b3be7f628fc0276a1313ac

Observation b95d26c8-5e12-4a1b-94df-07cace989b7b · outbound

This paper cites Grosinger.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Grosinger

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.406167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.559084Z digest=sha256:135e2bc073849a8a6f6b98b781243945666191a1eecaa39dc2e5f297288b7089

Observation 706bec63-9814-477f-ad16-35a6009adc6d · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.395594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.564368Z digest=sha256:626479a0f28e9b6878f78c2e985994493b25fc778a45d7392b17efc1851ccad5

Observation e6ae605f-3df7-4622-a997-8a7a2297bdc2 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.567799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.567799Z digest=sha256:74ff7ad039cf9a01d0417510ef0ad019a2344c00bf374b1040e80e8f47337fb2

Observation b51c8092-1136-4c74-b1a2-3a93f23398aa · outbound

This paper cites Hassid, T.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Hassid, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.386549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.571005Z digest=sha256:4f739eb2aff5192fab83e188d5faf41bdc0f82f6cf75d20fd43642bdfa606a2c

Observation 9e948fc4-7731-49bd-8472-069c295a0b0a · outbound

This paper cites Hendrycks, C.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Hendrycks, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.377707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.574760Z digest=sha256:f8c9f0dfbf855051af27648ab41dc5bd01c30b961f8846e13f29a53695c26da1

Observation aa217a71-8c4e-4921-8de3-4dc61eb4c6c7 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.578112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.578112Z digest=sha256:968c41a378e988a92e43e7850c4c94dffbf7eb6d4f69288f9d8fc2afa823431e

Observation 2ad3e9ab-9dcc-41e0-ae84-a367fa1adab2 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.368341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.581601Z digest=sha256:97c519756d2b9ee0d57714b291a1b595b8d2de9d0588ec28bc2dae526a9127ab

Observation dd19769d-df76-4e6a-b750-8aebb223ab4f · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.584754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.584754Z digest=sha256:9f519136ffc076a8fc636cf8d88f7e11b0dce0313e49b128f7dda5c8279828df

Observation ca06c36e-4cf8-40d8-8e86-b79fe08c4bef · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.588104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.588104Z digest=sha256:921a1ba44c118c450cb22aa6d34a536082ca2cf0c02518c0b686d312f77942c9

Observation 1fecb7db-846f-4aed-b124-fb178a8a67bd · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.591562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.591562Z digest=sha256:ac4befbe929b791abc03abf041accfd1489d4f735ce8ea4bc335b5797f9c21d1

Observation 9ad78b39-1085-43a6-af14-bdb6d5897c9b · outbound

This paper cites GPT-4o System Card.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.595206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.595206Z digest=sha256:f821b9c1baee19200c4cdbaf5d42ac9db8ab0df7cae72798775957de0f43ddda

Observation 65fb27dc-5e98-43e8-9467-071900e7a294 · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.598526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.598526Z digest=sha256:412dddf55016b22287fffe4fe251073eea5e16cb9b1bda215435ab943bf2302e

Observation 4042dfa0-57e4-4d32-81e1-636e573b1533 · outbound

This paper cites Kumar, P.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Kumar, P

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.358299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.602591Z digest=sha256:eb94272a7c3a68cd460b023a2fe44d3580a1d478afa8c59b88cce7ea934c801f

Observation ce7b5bce-bbf4-4000-998e-05a92f83cb3c · outbound

This paper cites Kwiatkowski, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Kwiatkowski, J

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-16T00:47:52.605505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.605505Z digest=sha256:0732923eadda2ec2c91f34d35875c1cdc48ba1fbdc62f6197d4fbb63f9525ddf

Observation 0e8e9f1c-2813-427c-a6a7-182d5e2dcd28 · outbound

This paper cites Generative Spoken Language Modeling from Raw Audio.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Generative Spoken Language Modeling from Raw Audio

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.608982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.608982Z digest=sha256:44e47d8dff0931cb9520d161c5a14017df20ca183fa158f709f67985b5e0bfae

Observation b0a35792-c6de-41db-8eb3-49c802a4c071 · outbound

This paper cites Lebourdais, M.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Lebourdais, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.347076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.611927Z digest=sha256:0552b9139b32c87e7384a15b7460e1caecf99db31e6764724b44f348ba5636d7

Observation 60ed26f4-d025-4287-b213-755b7d048d7a · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.614452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.614452Z digest=sha256:7dffc0fc56aaeb0011034cfa1727315c7588ff4842aac91cc016a31d6a5768fa

Observation 12aa6e90-c54c-4c1d-81fa-8b68fb61d921 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.335224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.618402Z digest=sha256:901ac3a5d6ba5b76e0fb9224d7eb2f64832366ca1098721936279c9a43dc1e3e

Observation 731ece00-b6c7-4959-bacc-3173929f9821 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.325123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.621910Z digest=sha256:5f255e6ae52eb44b6cbf1c0eaa32844d52dff2904e52bfcd74d57b18060a7d9a

Observation 321b423b-9119-4aeb-a979-6bb06a3172e4 · outbound

This paper cites Maiti, Y.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Maiti, Y

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.314597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.624714Z digest=sha256:e66d6d37ccfe53c62cc664ab88d5a4aef51cf51522c2bbeae8cd71b11df06878

Observation 66e86393-b44d-4274-bbff-2a3de08adac7 · outbound

This paper cites Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.627183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.627183Z digest=sha256:ed7eb11821f3522206dfd5e8373148861075411a52e91d95f4adfc26502d9331

Observation df2f3311-6684-4854-b1b5-22349b1cdcbf · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.304127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.629900Z digest=sha256:d9bfd1d745452af59e9dbd6ccc9e763c97cfa6c7fd9a017f0cba6a5c3817ff57

Observation 86f2ce39-1688-49cc-ae51-f7f4d60645bc · outbound

This paper cites PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.632277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.632277Z digest=sha256:d72dc243aaffd58e8148425543efc06e69db80554f7d9967989e6d0d21940858

Observation c602c896-a122-405c-a812-cc008e55911c · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.634671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.634671Z digest=sha256:29226e2f2fb87592264c2ed11ce03d85193df5ea4a0b249defb0968c602bdac0

Observation 040bb85f-0076-4631-938e-9eb39ba930aa · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Spirit LM: Interleaved Spoken and Written Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.638150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.638150Z digest=sha256:724ba03f4a1a12de7b580a1cb5a1caa0dc4ccc99a0e5d5b9273df4656975cd33

Observation 0be1560d-72f5-4df5-9a23-8668bbd4e047 · outbound

This paper cites Ouyang, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Ouyang, J

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.642350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.642350Z digest=sha256:8a46057242370dccf0662fe65b636ef54355dc7794a3d37ff793dc4a47e39e8a

Observation d13bd9c0-4b51-4b27-b020-590eadd590b5 · outbound

This paper cites Panayotov, G.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Panayotov, G

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.287763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.645654Z digest=sha256:adfb4a3df09e518fd136e9e359c627e885fbdf3762801487f91f53cc751620c9

Observation 020236d6-44e6-45b9-a87d-e9494255eace · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.649088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.649088Z digest=sha256:5522b5ac63ba4012eac8108537fb141c6aa40a8f2c14e065b82d2f2f65abe011

Observation 29aa0f4e-35be-494a-89a7-13f78d16c65d · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Robust Speech Recognition via Large-Scale Weak Supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.652682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.652682Z digest=sha256:209ccaee8f65d7bf9ff6e4e8f8c72f72561df368e23367af1ccda6e2a69385a2

Observation a827304e-1275-4db8-8515-77e71f3623c0 · outbound

This paper cites Rekesh, N.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Rekesh, N

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.276218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.655586Z digest=sha256:47c57a20cd1d5451b78e0398d13df6483a9223843d4f5143e2d82126d0bc2f60

Observation f8ba1da9-bc90-422a-85b2-6627000f6fba · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.658941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.658941Z digest=sha256:908e11a7d14c619a5ec3606987ff95aee93151d149d9b1a018cfdb8c83543dd2

Observation cf50e6d5-0777-4971-8ee7-e200be675097 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.266294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.662957Z digest=sha256:a5484ff4f4d8e8e77b990c26cc4edac5fdd832838ae8023b46976cdbf7d611ec

Observation 780d4055-048f-4c61-9f8a-fdb0d592921b · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.666078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.666078Z digest=sha256:ebc2b1ea38083adba95573d8c11f584aaac41542c37affdb8aaab047f141b202

Observation 391b8b51-aa3e-4abb-930b-35c049ae8bf1 · outbound

This paper cites Schroeder and N.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Schroeder and N

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.255410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.669859Z digest=sha256:d3ea34a28cb7f8c2f46d995eb359253fa082661ea7455c5a5b5b3ae9892e93bb

Observation c404d6b6-b12f-41db-a962-fdae744eea6c · outbound

This paper cites Character-LLM: A Trainable Agent for Role-Playing.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Character-LLM: A Trainable Agent for Role-Playing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.673073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.673073Z digest=sha256:78d2c04471d6ec25bea6241b748d17ee17bbb641a9cfcda2fd6adab03806e3d1

Observation 97022753-9f82-4bda-9b23-9a7034fe7dc1 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.676290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.676290Z digest=sha256:4c59de5dc38aa7db1bf843f4c4bfee3a194f6f5e512076236b928a010df29027

Observation 04b6a8bb-7126-4b30-929d-db844060fb51 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.243969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.679284Z digest=sha256:916bf434ee146396aaf55e463a381730f4ea7489d189ce4537232c63b592c7a9

Observation 33e96a23-b2b9-4a47-a0d5-fea449e9efe4 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.233519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.681853Z digest=sha256:c84239456beabab842699c6d5cb1aacd8dfd5d36717e5784f4970bb21eb818e8

Observation 1e8faf03-4548-46eb-8a99-15ad140b7bdb · outbound

This paper cites Introducing hertz-dev, the first open-source base model for conversational audio generation, 2024.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Introducing hertz-dev, the first open-source base model for conversational audio generation, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.224794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.684728Z digest=sha256:74fe8cca76d72319ad2abbdac5ffc319bf7246449df62f9a0ac21e69610daa10

Observation d9205236-29ae-423c-862d-48785d9ae384 · outbound

This paper cites Stivers, N.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Stivers, N

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.214806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.687522Z digest=sha256:b5e02c9e7f5eb50911c54b7e5b6aa60438125591b26860a2c4c1508402260fa0

Observation d5d39c18-155c-4ffb-8298-20d9fa53c0ef · outbound

This paper cites SSR: Alignment-Aware Modality Connector for Speech Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SSR: Alignment-Aware Modality Connector for Speech Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.690476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.690476Z digest=sha256:724788fe6774b9d72ed8124e44d76a470930f681f66fc641d1d7abe3092114cf

Observation 05124a52-219b-4a34-b46a-bf77dd7c50a0 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.693798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.693798Z digest=sha256:55afae19b4f0d0cb58c9e46046f3258a5a6af1d525c62a709a78debb2c7cb04a

Observation 9468cbe7-5613-4786-b2ca-87458060fb61 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.696916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.696916Z digest=sha256:2f45ce09491a86ed1a5e37703ff0b535d117144100f5df8c929c2a19af7061da

Observation d7d929f1-4a1e-4044-b00b-41f099df31a0 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.203398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.700497Z digest=sha256:4709ad701414e1610d4665126be149958944db0407f65df67f569304a5d28236

Observation 8a6fca1c-7c2b-486d-b314-96c269d06ed8 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.704511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.704511Z digest=sha256:c185438664509a2d64bcafad114714a50ecb29f5bb3883b3e94f4ae57751f7bf

Observation 5ede4a8f-2d46-40e6-b28d-b2bb80cfd332 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.712479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.712479Z digest=sha256:9a58d9a52944aa6d3a4589c4fe48cd8cbc81c58ebe6b94deaa11e86614e25f6a

Observation 668f5971-9f1b-4a44-8233-570d9d7e1fc4 · outbound

This paper cites From Role-Play to Drama-Interaction: An LLM Solution.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play From Role-Play to Drama-Interaction: An LLM Solution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.726959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.726959Z digest=sha256:a3cc789f54c50e1a45b9e50773aec496415955c47e4369b35509ed6ebafab075

Observation 4dde856f-4d42-4f7e-ae86-fc7c80882e52 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.756762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.756762Z digest=sha256:bc75189dc264a9f584d767d3ea3f19d92815ed6d53d78f8f6688c719acb00c3d

Observation addf096c-29e5-4612-8417-8eb86aeb6bca · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.186620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.760363Z digest=sha256:ae6bc302f92ee7e3e85129eb2e95e143d9c3f05da18cd416e815fd732202942e

Observation cba341f0-e7f1-4915-b679-65993d3d94ce · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play ReAct: Synergizing Reasoning and Acting in Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.764544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.764544Z digest=sha256:09d8426f708b5fce43cc7577c67442513e3dd74559ee9ae6f4fe2eca8ec9a0ac

Observation 50d1e163-6763-4fb1-b4a0-e8add6acbc77 · outbound

This paper cites Zeghidour, A.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Zeghidour, A

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.767540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.767540Z digest=sha256:84ae20b89468c891f80a7846bc93428d70ec6da40c4cd210d51343b49591efd4

Observation 09397b9d-0bc5-4036-9df7-19890a7270cf · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.770658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.770658Z digest=sha256:d482c95d68246f457c66cf21878d710e56c252a909cedb65a0cff8b8fb7a6b00

Observation f08281d8-d065-4ac6-92cf-81953a788e60 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.774628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.774628Z digest=sha256:37976776fb0d1184e7c62b574fe53c904a22777c7c9004a21b2711d7805ed8e1

Observation 3c77d63f-3d38-4005-bdbd-50969192f301 · outbound

This paper cites Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.778178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.778178Z digest=sha256:3c30e72822ed0c4f6c3f6931bb677bb57ba90db9f7d940d41e7a76532380ff76

Pith citing papers

Observation 8375cedf-f628-4809-be40-bb2805f9e2a9 · inbound

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding cites this paper.

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:33.556834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T02:32:35.333227Z digest=sha256:5e8246ad53c943dbc52c2cb8aea004c4721bb3851b7cf3831378be55085272c5

Observation 8aeddb8c-0870-42e3-8557-840d46f94454 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.088844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:eb6c98836a0d6d27b450542ffef0cd1bb7f459a7bc8cbbc7e069d633eaf6b43b

Observation 57dc6c13-b400-45cb-822b-464e6734da8d · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.349819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:720113d35359c9944274c8cdc69052b4e4a084d560fd600eb70904e457480610

Observation 4d1cc85d-18b3-4337-b6cd-4e2310f524e0 · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:20:57.347831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:bb44f78ee6ceccce92f6af74595d105ff1651b4df5c678547bcb0c55514c8a27

Observation 523fcff5-c0a9-4e53-bd9f-7f9eac891c1c · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:14:14.956765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T03:09:57.069086Z digest=sha256:62c9bfb272944a2bd88cfafb3c8880cec404643efaa9f109b02e97deda366273

Observation 861f0d5d-5929-4867-ac51-c5f99282675b · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.879111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:28:07.670564Z digest=sha256:a528029279a3f951c85f79e85bae157dc1005831b8635bd1453eb54b24ce2aa9

Observation 97497c93-ea4a-4ecd-b488-ea0dfa8c8a84 · inbound

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects cites this paper.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.251451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.251451Z digest=sha256:c6eeb50989fe86f12ef2bb7139f7e443a2e488e8e0141a6bb723a1a64ef48a3b