Pith. sign in

Paper Citation Record · LEDGER

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 7 inbound Pith citation observations for arXiv:2505.02707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02707 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:47:52.778178Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:55.251451Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved53
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e124b35f-b9c8-45e1-a784-0dfd85b264ff · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.506447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.506447Z digest=sha256:af81cbb3e7d53a1738e7e3e294ac7a05d36eb78e82a21e65dd063a7963620072

Observation d8be03d7-7dcc-4c1d-b70d-58b1911250be · outbound

This paper cites B \'e rub \'e , M.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play B \'e rub \'e , M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.494593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.511006Z digest=sha256:56f8caa718adb2adfbaf1f9a13eda92ee96f647761f6a461af3653cdc6766929

Observation a510f1b4-927b-434a-962d-9f3fad62be5c · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.484279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.514268Z digest=sha256:ffb12d748b341701411ac0ce4c1f405ddbbd444fa38f8899996c1551cc159a7c

Observation a2033109-87c9-4f83-9d1c-4f044543f15f · outbound

This paper cites Borsos, R.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Borsos, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.517186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.517186Z digest=sha256:941b5da0feb63e7af705505790a9b420ae2728f4ecc788d9b6b92a0d51dc1af9

Observation 4b35f937-b2db-4fb5-9955-91b3c754920b · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.466867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.521195Z digest=sha256:4295b4c73209c0239b6673888803214ab3a4ee96d7ff4490b040c4c1848ccfb7

Observation 17bb2e35-9205-470f-8be0-6d7182b9cc2c · outbound

This paper cites Buyukgoz, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Buyukgoz, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.456685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.525052Z digest=sha256:8e911a7eb9042aeea03d4eb9113ebd639d1cc63ace21fab7d1610c0a2414e50c

Observation 56c71573-1471-4375-ad9c-1864319638e5 · outbound

This paper cites Casanova, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Casanova, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.444870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.528913Z digest=sha256:f2e8d0c8105366d80c36cd8eebda96902df853d64de1a4d51b18b855397fa9dd

Observation 8903aae6-1e15-4f3b-abd8-a70f88f14332 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.434951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.531889Z digest=sha256:cb34db835c42bb61640af3c625775bfddd68277a0716ec18f148ca8dfd736e73

Observation 59d1fd59-0ee8-4a84-8bb0-b17c4195c86e · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.535285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.535285Z digest=sha256:3410a84ddb253d4bdb5417baea12c12590792959df42451d6f0b2af896134fae

Observation 519002cb-9b15-4bfb-bb92-b18a5bb4d9b3 · outbound

This paper cites Qwen2-Audio Technical Report.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.538991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.538991Z digest=sha256:edd0c182a8c6220623ceec9d2c042a04bdc6cfe4fcc133238fd53ce9e8d95ad6

Observation 67784ff6-797e-40cb-aff3-b5e552203ed2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.542207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.542207Z digest=sha256:781b6e1b80e7c9fc933420e4a6d0c7588c6c0dc26141847c4febddc22f7e2a57

Observation 46bf2d6a-3dad-426d-bc88-c789309f10ac · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Moshi: a speech-text foundation model for real-time dialogue

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.545213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.545213Z digest=sha256:76c086cd3d1c43e6d7e47f0cd70ee01b70fe581869f19c64e10d348574c1f0b6

Observation c5bcc4a5-56cd-4a66-8135-89fa41d12f68 · outbound

This paper cites High Fidelity Neural Audio Compression.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play High Fidelity Neural Audio Compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.548376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.548376Z digest=sha256:2caadd438dd8bb9e48541d3f44db3b7a319bd4db8d2c689c0e560f315268ba3b

Observation 39adc5f5-de4b-46f5-9e91-8dea42713698 · outbound

This paper cites Faruqui and D.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Faruqui and D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.425302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.552419Z digest=sha256:4075a32219c7e2d5bb2c1bef3696955a26880005564b3ce489e4518716ffc7c7

Observation c3771bc7-2dde-4bdb-b3b9-b4c92714756b · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.415727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.555483Z digest=sha256:64bf44e2123168907d962bedce166c7f61e1ba948839f3aec3cd798be3fca12e

Observation b95d26c8-5e12-4a1b-94df-07cace989b7b · outbound

This paper cites Grosinger.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Grosinger

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.406167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.559084Z digest=sha256:0191c1d060d27a3595bad8c73f87df2f18d6e5f99c07c51984b56e555921124b

Observation 706bec63-9814-477f-ad16-35a6009adc6d · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.395594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.564368Z digest=sha256:161318e3dd799e2f2a74073e1edf9db14850c5dbdecde3f333fe44a576af4975

Observation e6ae605f-3df7-4622-a997-8a7a2297bdc2 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.567799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.567799Z digest=sha256:f413fe22dba0274b39528407fd00436c1886c4c6a6ed41d4de71f65189e98a0e

Observation b51c8092-1136-4c74-b1a2-3a93f23398aa · outbound

This paper cites Hassid, T.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Hassid, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.386549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.571005Z digest=sha256:5b85215661b268820f1bc65f6953bd2a1d840f41a10afb94ea4f37647ccf9c83

Observation 9e948fc4-7731-49bd-8472-069c295a0b0a · outbound

This paper cites Hendrycks, C.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Hendrycks, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.377707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.574760Z digest=sha256:03923987de4b39cf66d2f6b1b302fd4c367e7f829bf2bfd84794715ad553b07e

Observation aa217a71-8c4e-4921-8de3-4dc61eb4c6c7 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.578112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.578112Z digest=sha256:e4fac960a9e199ef8b0cb9cf5d5a9b52f1b1f16167350736f3a1fd8ecb0d4c98

Observation 2ad3e9ab-9dcc-41e0-ae84-a367fa1adab2 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.368341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.581601Z digest=sha256:19c0f98033d023d2fa753abca34f1d438c34008a3b14c7e1a79b5a613660113a

Observation dd19769d-df76-4e6a-b750-8aebb223ab4f · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.584754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.584754Z digest=sha256:97754359ca677e131490d3a80f9480237a2787f9ce7b1cd0f1e796d6a0d3ced3

Observation ca06c36e-4cf8-40d8-8e86-b79fe08c4bef · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.588104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.588104Z digest=sha256:413e52dfba20690097ad01f8a360c2d16831ab37422a7becf840f36cc40c07ed

Observation 1fecb7db-846f-4aed-b124-fb178a8a67bd · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.591562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.591562Z digest=sha256:f6cf71e989af3641dc1c8f392090be28355a57d308f1d014850af732762ab13d

Observation 9ad78b39-1085-43a6-af14-bdb6d5897c9b · outbound

This paper cites GPT-4o System Card.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.595206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.595206Z digest=sha256:3e0c32364c0619a4ba65e04497ccb08dd54c9400f4b20ef029a448e2a25f3941

Observation 65fb27dc-5e98-43e8-9467-071900e7a294 · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.598526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.598526Z digest=sha256:8323d10c24baca4230fbb251b2b83f10606ba94e18a054b13e91c4f6822f0a3a

Observation 4042dfa0-57e4-4d32-81e1-636e573b1533 · outbound

This paper cites Kumar, P.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Kumar, P

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.358299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.602591Z digest=sha256:93fe60bc90fa12cddaf529d9eb43e1ba482e0622e1de1a2bf51abd793d931e77

Observation ce7b5bce-bbf4-4000-998e-05a92f83cb3c · outbound

This paper cites Kwiatkowski, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Kwiatkowski, J

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-16T00:47:52.605505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.605505Z digest=sha256:a880dfa550313e5ee4e64f89196321b43b2731c24c88c0de9061181b23aed1b9

Observation 0e8e9f1c-2813-427c-a6a7-182d5e2dcd28 · outbound

This paper cites Generative Spoken Language Modeling from Raw Audio.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Generative Spoken Language Modeling from Raw Audio

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.608982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.608982Z digest=sha256:95eed43169ac97f2cb16feb56a7a97e03801fa6ae5aefcf9a1a3d89f7ebeadc5

Observation b0a35792-c6de-41db-8eb3-49c802a4c071 · outbound

This paper cites Lebourdais, M.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Lebourdais, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.347076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.611927Z digest=sha256:2ead70f9d268c47447fa927f7e48fdd673e83020c7e7eae5fd9f7b50ffae1975

Observation 60ed26f4-d025-4287-b213-755b7d048d7a · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.614452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.614452Z digest=sha256:9b9a024b3d9608a138237d56f28fa7328dd3f4d3de76028338884ebcdc829144

Observation 12aa6e90-c54c-4c1d-81fa-8b68fb61d921 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.335224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.618402Z digest=sha256:481a79e3a5c0af6ea5e53c25351e8850e414e28e2f467fac8126b15a89bd9496

Observation 731ece00-b6c7-4959-bacc-3173929f9821 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.325123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.621910Z digest=sha256:c964a25cb982b3af9fdc975023a5a0623963365c5d8c4bc8bc32d697b8c56b51

Observation 321b423b-9119-4aeb-a979-6bb06a3172e4 · outbound

This paper cites Maiti, Y.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Maiti, Y

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.314597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.624714Z digest=sha256:873e5623f558bf8d2b030ded1cdd2be481e64686409c1579738c63600011be11

Observation 66e86393-b44d-4274-bbff-2a3de08adac7 · outbound

This paper cites Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.627183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.627183Z digest=sha256:bb642e697971d3bb9a680b131412f92dfe5e1c71594f713c43a29b2d5fc6c9f7

Observation df2f3311-6684-4854-b1b5-22349b1cdcbf · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.304127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.629900Z digest=sha256:1f645d58d03b48e20db794c27600151a4b8d664533a803cdd53d7eadcaa233ab

Observation 86f2ce39-1688-49cc-ae51-f7f4d60645bc · outbound

This paper cites PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.632277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.632277Z digest=sha256:60f3959813f9fc50482647805c5e233a185e333b67bd221ee9f279f58dff716c

Observation c602c896-a122-405c-a812-cc008e55911c · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.634671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.634671Z digest=sha256:41fd2594ca7704f70ab7d3da6737b935edd4bf61106907ef5dd0ebe8aa1828e0

Observation 040bb85f-0076-4631-938e-9eb39ba930aa · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Spirit LM: Interleaved Spoken and Written Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.638150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.638150Z digest=sha256:9631faad16fd531f0b5d2555186a9f02acec8d1ca3953e2e4bef9ce4125cc59b

Observation 0be1560d-72f5-4df5-9a23-8668bbd4e047 · outbound

This paper cites Ouyang, J.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Ouyang, J

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.642350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.642350Z digest=sha256:5698632a740fefd0999a725339889215d743c8fd2e5952ee3ca1c6283ffa792b

Observation d13bd9c0-4b51-4b27-b020-590eadd590b5 · outbound

This paper cites Panayotov, G.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Panayotov, G

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.287763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.645654Z digest=sha256:bdd3d021025761aa766c6952ca15620ab516b9bdcab30a726b2f9861b4f613e7

Observation 020236d6-44e6-45b9-a87d-e9494255eace · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.649088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.649088Z digest=sha256:6d0099257a8c5f8efc2bf766251c6af0d9929b7c8d07a3ba5076e8f01b4ac770

Observation 29aa0f4e-35be-494a-89a7-13f78d16c65d · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Robust Speech Recognition via Large-Scale Weak Supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.652682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.652682Z digest=sha256:69427cef6ea5741ea19d7a0ae01bcc3d4cd388367017c979eff892b6df11da51

Observation a827304e-1275-4db8-8515-77e71f3623c0 · outbound

This paper cites Rekesh, N.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Rekesh, N

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.276218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.655586Z digest=sha256:239c0f12cd8a56980b77bdb1d6c05fa48132dd2dcaeb35a0369f769bd974e8f4

Observation f8ba1da9-bc90-422a-85b2-6627000f6fba · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.658941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.658941Z digest=sha256:4e3990f5a71478d649feb799d2b1ec18a8ae381a15d52eb61392876581c7c1e6

Observation cf50e6d5-0777-4971-8ee7-e200be675097 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.266294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.662957Z digest=sha256:7daa7c2b9c042c06c280aca9e8079cdfbd21e6308d702a827a45f2bb169bbd13

Observation 780d4055-048f-4c61-9f8a-fdb0d592921b · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.666078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.666078Z digest=sha256:101f673a4fdf9368685482466953e8c8259dc04846350a14ac995635f6bf6659

Observation 391b8b51-aa3e-4abb-930b-35c049ae8bf1 · outbound

This paper cites Schroeder and N.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Schroeder and N

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.255410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.669859Z digest=sha256:34c536f8bb2b415c68fc5a466a554f92fb8d5d29c6221b2ef64f2991c2fa7e1a

Observation c404d6b6-b12f-41db-a962-fdae744eea6c · outbound

This paper cites Character-LLM: A Trainable Agent for Role-Playing.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Character-LLM: A Trainable Agent for Role-Playing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.673073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.673073Z digest=sha256:7604a86cddb34ca0abd4257ee504cb69963e39ead2fc34876592b8c1e781d074

Observation 97022753-9f82-4bda-9b23-9a7034fe7dc1 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.676290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.676290Z digest=sha256:fdb8ef10a544333ef4e8161a5c4ca65ab1f628a957b836b08954384a28229cfc

Observation 04b6a8bb-7126-4b30-929d-db844060fb51 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.243969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.679284Z digest=sha256:0420e407e0807c3bfc064e3054e1ee7adb2d18bb91dc7993c86a8ebe7afc4f83

Observation 33e96a23-b2b9-4a47-a0d5-fea449e9efe4 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.233519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.681853Z digest=sha256:c041ed4fa7d92423c69a4edc34ad53011e3da781c05eb1b4d7b9a0b9f3ad273d

Observation 1e8faf03-4548-46eb-8a99-15ad140b7bdb · outbound

This paper cites Introducing hertz-dev, the first open-source base model for conversational audio generation, 2024.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Introducing hertz-dev, the first open-source base model for conversational audio generation, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.224794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.684728Z digest=sha256:9ebbe0a462f8e49a38949c9c9cae6abd162761f8d5de43efe82cb93ce03a58c1

Observation d9205236-29ae-423c-862d-48785d9ae384 · outbound

This paper cites Stivers, N.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Stivers, N

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:47:53.214806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.687522Z digest=sha256:0c9f48907c17856e91e07bc98c3af8924709ad4f37013bc18e92eee6380b1c4d

Observation d5d39c18-155c-4ffb-8298-20d9fa53c0ef · outbound

This paper cites SSR: Alignment-Aware Modality Connector for Speech Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SSR: Alignment-Aware Modality Connector for Speech Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.690476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.690476Z digest=sha256:2abe835dc52b89e8275135dda283b113d82d8fa750d54fb8d9d896b358655b0e

Observation 05124a52-219b-4a34-b46a-bf77dd7c50a0 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.693798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.693798Z digest=sha256:4645d1de6efec7f594c139ab6019e320642753ffb347e6936fdac2f39b9b7156

Observation 9468cbe7-5613-4786-b2ca-87458060fb61 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.696916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.696916Z digest=sha256:12a8bafffd8e4a866356aed8165210015686888fede3ccd3a9e806b2d6c5272e

Observation d7d929f1-4a1e-4044-b00b-41f099df31a0 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.203398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.700497Z digest=sha256:c6bb240a88a0668c491ee4773ddc6ca5efb11573f83215978f0f2c2a4e3fe367

Observation 8a6fca1c-7c2b-486d-b314-96c269d06ed8 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.704511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.704511Z digest=sha256:7c1f9f9b7a7b89b6a9c815033e4a24a6686a76f995bd500ab47c50d389e2ad9c

Observation 5ede4a8f-2d46-40e6-b28d-b2bb80cfd332 · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.712479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.712479Z digest=sha256:86ae99fb17df6f07a7d55f85f57cb2b76c553a32414c8d4caf7b86e9ccb03fd0

Observation 668f5971-9f1b-4a44-8233-570d9d7e1fc4 · outbound

This paper cites From Role-Play to Drama-Interaction: An LLM Solution.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play From Role-Play to Drama-Interaction: An LLM Solution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.726959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.726959Z digest=sha256:6660cf523e4a3fd3682e87a6f46fabecada8c96c35a046e69e32513f7d6a0bf2

Observation 4dde856f-4d42-4f7e-ae86-fc7c80882e52 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.756762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.756762Z digest=sha256:09e9303e2f75031c722ceb8c2a2dc2115449252d771041e0b87bc086998e60f3

Observation addf096c-29e5-4612-8417-8eb86aeb6bca · outbound

This paper cites an unresolved cited work.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:47:53.186620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T00:47:52.760363Z digest=sha256:84ca0c2f9a2a5c4b396319c3369cae7cf52f709900273b0a59fc984e0dfc33b5

Observation cba341f0-e7f1-4915-b679-65993d3d94ce · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play ReAct: Synergizing Reasoning and Acting in Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.764544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.764544Z digest=sha256:6791b374ac7a44abf5673ddee4c5a0faea147705fe81b4cda123e9140744b6eb

Observation 50d1e163-6763-4fb1-b4a0-e8add6acbc77 · outbound

This paper cites Zeghidour, A.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Zeghidour, A

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.767540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.767540Z digest=sha256:780a41ff2254e937ffd36a79b6a5c97614e6f5c6169f87a078f1f3aba5d33063

Observation 09397b9d-0bc5-4036-9df7-19890a7270cf · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.770658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.770658Z digest=sha256:b5e3717d993a956d586ce601d7f1d9e91169658470972b544259e75f53f42f03

Observation f08281d8-d065-4ac6-92cf-81953a788e60 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.774628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.774628Z digest=sha256:fd3268af6d84c138d655c63cacd6dd7e02c7c9e58dbeb47d954445f449735f0d

Observation 3c77d63f-3d38-4005-bdbd-50969192f301 · outbound

This paper cites Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.778178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.778178Z digest=sha256:d5e895485644965d4a10e2e2096f141e8600d3ca3cb14c2019daa75dbaa222f1

Pith citing papers

Observation 8375cedf-f628-4809-be40-bb2805f9e2a9 · inbound

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding cites this paper.

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:33.556834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T02:32:35.333227Z digest=sha256:2c68d2471afe550fc1b62daf69e5af61c50fd96f62edd65afb6c00eab06634dd

Observation 8aeddb8c-0870-42e3-8557-840d46f94454 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.088844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:4ceb6e79f2a43657a638635b82582a7e6ea5e2bed0532aacafd8ed2b08b37bf8

Observation 57dc6c13-b400-45cb-822b-464e6734da8d · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.349819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:cf3f57d10bd2a56b726bd5916f446a84f1bde621b6084d86a71648922cba3839

Observation 4d1cc85d-18b3-4337-b6cd-4e2310f524e0 · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:20:57.347831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:dffb5dce3bf1ed96f0b58dfc3a7e43492635e5659eac2a4ddc9485c913728ff9

Observation 523fcff5-c0a9-4e53-bd9f-7f9eac891c1c · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:14:14.956765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T03:09:57.069086Z digest=sha256:e1e4754ed5a201a0a8cf393c7c5a66bb903c50fdf026376dd967e1250399caaf

Observation 861f0d5d-5929-4867-ac51-c5f99282675b · inbound

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices cites this paper.

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.879111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:28:07.670564Z digest=sha256:cc6d9c94dcaf2e61b8a877e9193eabe9fa2678134c2a654c460a19300f1fbd4c

Observation 97497c93-ea4a-4ecd-b488-ea0dfa8c8a84 · inbound

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects cites this paper.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.251451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.251451Z digest=sha256:fd6799a5578dbe71d534814cc1e2134828661ce02ea35d157a46997b025d161e