Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:34.901221Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.03980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:34.901221Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:30.784001Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T15:16:18.173817Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · outbound
Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b51346-a3ec-4558-a0b3-360959c5f71a · outbound
Voice Activity Projection Model with Multimodal Encoders Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa716f32-312b-415a-a2b2-ae805ea34828 · outbound
Voice Activity Projection Model with Multimodal Encoders Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c6cd6bf3-e16e-41ee-9126-6c918ecce3f2 · outbound
Voice Activity Projection Model with Multimodal Encoders Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 224039b8-c633-4b71-a07f-f0d11163178b · outbound
Voice Activity Projection Model with Multimodal Encoders Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4cdeaaba-ed8d-459b-bdc4-8419da09eb7c · outbound
Voice Activity Projection Model with Multimodal Encoders Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3ed24270-7e68-43f5-96e1-a5766682c055 · outbound
Voice Activity Projection Model with Multimodal Encoders Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37]
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 780aeee2-700b-4651-a456-3398cd8e9303 · outbound
Voice Activity Projection Model with Multimodal Encoders For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60c487df-7565-4fb3-ade1-fefc984fdf05 · outbound
Voice Activity Projection Model with Multimodal Encoders By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb5de9e2-133c-41a7-8d17-4924acea615d · outbound
Voice Activity Projection Model with Multimodal Encoders Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62bce116-9f4b-44ee-9713-047f6d3fcf15 · outbound
Voice Activity Projection Model with Multimodal Encoders A simplest systemat- ics for the organization of turn-taking for conversation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5b02f3d6-2f7a-4823-a7dd-4c68fde28c3d · outbound
Voice Activity Projection Model with Multimodal Encoders Turn-taking: A critical analysis of the research tradition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 591fbe38-ee34-4ebf-9c78-a272121bf8a0 · outbound
Voice Activity Projection Model with Multimodal Encoders Some signals and rules for taking speaking turns in conversations,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aec0d1ca-d0be-4ae3-ad90-f445f9f4a695 · outbound
Voice Activity Projection Model with Multimodal Encoders Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e1f26d21-f463-40f4-91fc-2084d25146de · outbound
Voice Activity Projection Model with Multimodal Encoders Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd7c379f-dd7d-47c8-a25b-bde9899b98af · outbound
Voice Activity Projection Model with Multimodal Encoders Using uh and um in spontaneous speaking,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 09b69a46-189d-4ea2-83fb-766c2db89b85 · outbound
Voice Activity Projection Model with Multimodal Encoders Using prosodic clues to decide when to produce back- channel utterances,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ccd90798-8833-4eb4-be36-b1610fea83a1 · outbound
Voice Activity Projection Model with Multimodal Encoders Nonverbal behaviours improving a simulation of small group discussion,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13f5be9a-d41e-4f27-8a59-05e5a9ac65c3 · outbound
Voice Activity Projection Model with Multimodal Encoders Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63bf7315-b13f-4021-ac0c-781f5dc7fbcb · outbound
Voice Activity Projection Model with Multimodal Encoders The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d6ffe9f2-4f0a-47e6-92d8-2091d2c9dec8 · outbound
Voice Activity Projection Model with Multimodal Encoders Now or when? interrup- tion timing prediction in dyadic interaction,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f42f06c1-b3a3-4cfd-bf38-f36ac0dae25d · outbound
Voice Activity Projection Model with Multimodal Encoders How turn-taking strategies influence users’ impressions of an agent,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62270099-7a5e-4217-afb9-d21fc7f21aa8 · outbound
Voice Activity Projection Model with Multimodal Encoders Timing in turn-taking and its im- plications for processing models of language,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb06fffc-88c0-48bf-8da3-211339adf5c2 · outbound
Voice Activity Projection Model with Multimodal Encoders Timing in conversation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df526a09-03c7-476f-95ea-b7f86eab929f · outbound
Voice Activity Projection Model with Multimodal Encoders Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4da0c1f6-3c35-4047-bdf6-26c2e15dce5f · outbound
Voice Activity Projection Model with Multimodal Encoders Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c39117f-2ccc-46af-b5d0-f9a8e5f281eb · outbound
Voice Activity Projection Model with Multimodal Encoders Attentive listening system with backchanneling, response generation and flexible turn-taking,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4397cd86-ad7e-4d7a-98d4-0b499497769d · outbound
Voice Activity Projection Model with Multimodal Encoders V oice activity projection: Self- supervised learning of turn-taking events,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88e17781-a09a-4660-8764-15a4c9979cef · outbound
Voice Activity Projection Model with Multimodal Encoders Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 72879ccf-5500-4aeb-9098-a3004ce9138c · outbound
Voice Activity Projection Model with Multimodal Encoders Predicting turn-taking by compact gazing transition patterns in multiparty conversation,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9c1df9b-1a7d-4722-a582-d1657e5baa05 · outbound
Voice Activity Projection Model with Multimodal Encoders Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cffd3c23-bf87-4e5b-a3c6-4e63e3116661 · outbound
Voice Activity Projection Model with Multimodal Encoders Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27f3612d-ddcc-497f-9b80-59a3d55a0de0 · outbound
Voice Activity Projection Model with Multimodal Encoders Turn-taking and backchannel prediction with acoustic and large language model fusion,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ee2159a5-e41e-429c-b0ab-71b5bfd41980 · outbound
Voice Activity Projection Model with Multimodal Encoders Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 319c46e6-58d3-46f3-b227-99ef7960d4b4 · outbound
Voice Activity Projection Model with Multimodal Encoders Mini-omni: Language models can hear, talk while thinking in streaming,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 38d5a649-549b-4ec6-bb89-2e1661e190ac · outbound
Voice Activity Projection Model with Multimodal Encoders Moshi: a speech-text foundation model for real-time dialogue,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca140c6e-e660-48da-9588-d30903dab38a · outbound
Voice Activity Projection Model with Multimodal Encoders [Online]
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec01cbc1-9b8f-4645-8879-73d67df253f6 · outbound
Voice Activity Projection Model with Multimodal Encoders [Online]
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 881446a4-fe91-4cae-9141-96ada48bfbe2 · outbound
Voice Activity Projection Model with Multimodal Encoders Real-time and continuous turn-taking prediction using voice ac- tivity projection,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fc479d7-856d-45ad-b697-8ddd0a9c8624 · outbound
Voice Activity Projection Model with Multimodal Encoders Multilingual turn-taking prediction using voice activity projection,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e17c12c0-735c-4ac2-b140-e657a73e465a · outbound
Voice Activity Projection Model with Multimodal Encoders How much does prosody help turn- taking? investigations using voice activity projection models,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e2e21d0-e993-4a8e-a422-81c281be1e99 · outbound
Voice Activity Projection Model with Multimodal Encoders OpenFace: An open source facial behavior analysis toolkit,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0a9cae1f-8d18-410a-888c-8affc0869fb9 · outbound
Voice Activity Projection Model with Multimodal Encoders Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6fd51c1-dc02-4a73-b9ac-86fccdf24ffa · outbound
Voice Activity Projection Model with Multimodal Encoders Former-dfer: Dynamic facial expression recognition transformer,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4d281b4e-0199-4e08-b775-6c119c49253e · outbound
Voice Activity Projection Model with Multimodal Encoders Unsu- pervised pretraining transfers well across languages,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 55d37418-757f-4d48-ba23-cd3102fd6f51 · outbound
Voice Activity Projection Model with Multimodal Encoders Dlib-ml: A machine learning toolkit,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f56658d-9e42-4ba9-923b-fac07bf14a51 · outbound
Voice Activity Projection Model with Multimodal Encoders The NoXi database: multimodal recordings of mediated novice-expert interactions,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4cdced7e-395a-4d45-a249-982e7f1df638 · outbound
Voice Activity Projection Model with Multimodal Encoders PyTorch Lightning,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73512c0f-2b7e-4b8b-9385-91b48365d77a · outbound
Voice Activity Projection Model with Multimodal Encoders Discourse as an interactional achievement iii: The omnirelevance of action,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eda6959a-4dde-4ea0-ab68-e3e9ac2710a4 · outbound
Voice Activity Projection Model with Multimodal Encoders Between and within: Alternative sequential treat- ments of continuers and assessments,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ab26f48-cf14-4575-a333-2892af7ef0f5 · outbound
Voice Activity Projection Model with Multimodal Encoders Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fba6514b-15e3-4dd3-90ba-c469c37b185a · outbound
Voice Activity Projection Model with Multimodal Encoders Available: https://github.com/Lightning-AI/ pytorch-lightning
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · inbound
Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 470d4985-fcdd-40ae-960d-58c68db749cc · inbound
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection Model with Multimodal Encoders
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.