Pith. sign in

Paper Citation Record · LEDGER

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.07294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07294 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T15:07:52.715880Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact13
  • verified fuzzy30
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0282f24d-3551-4817-bef1-7b320ee434d1 · outbound

This paper cites A Simplest Systematics for the Organization of Turn-Taking for Conversation,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders A Simplest Systematics for the Organization of Turn-Taking for Conversation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.352789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b7445f872e260ad3d1afbe301b226efaa479b68dcd307f5c5185934c8f593b95

Observation f9129915-712f-4024-8910-f406a887bc7c · outbound

This paper cites Palinko, L.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Palinko, L

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.358292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:8a763bf357197d577712e030bc4847462b2f9f2d2dcf9c96e2c26e0c98e6aabc

Observation 52b7dab9-f39a-45be-b973-b72343cca828 · outbound

This paper cites Towards improving turn-taking in social robots using Visual-Only V oice Activity Detection in multimodal dialogue systems,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Towards improving turn-taking in social robots using Visual-Only V oice Activity Detection in multimodal dialogue systems,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.394045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:3cde5d34ca1ac45629a277ef51c1bef40f335dc2bca067ca489fe5b78e4c72f8

Observation b62a8dfb-9685-4f45-81b0-e8c88c2bdb24 · outbound

This paper cites Design of Social Features for Robot-mediated Cross-cultural Interaction,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Design of Social Features for Robot-mediated Cross-cultural Interaction,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.367807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:2a8e34ab464fc61d543de022314028ee634ca09f3cc34bb794b9eba7468290be

Observation 87a5e57a-9755-4805-aaad-905ec97166d4 · outbound

This paper cites Haru in the Care Network: Stakeholder Perspec- tives on Privacy with Social Robots in Pediatrics,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Haru in the Care Network: Stakeholder Perspec- tives on Privacy with Social Robots in Pediatrics,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.376031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:d94d9275e4115832fb4fab04a8fc789e2f70c6971f43ab108c8e35b8945181b2

Observation ea58cff9-99f7-4cdd-bba2-b13bef42da12 · outbound

This paper cites Building Friendships Across Borders: The Role of Social Robot Haru in Children Group Communication and Con- nection Development,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Building Friendships Across Borders: The Role of Social Robot Haru in Children Group Communication and Con- nection Development,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.345600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:18f0302ed5cb12482dba5cb40ef461c33b93dbc5063e1025609f124150ec1992

Observation db4b203d-461d-42bc-b452-13de60bc5d7d · outbound

This paper cites Multimodal Transformer Models for Turn-Taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction During Cooperative Gameplay,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Transformer Models for Turn-Taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction During Cooperative Gameplay,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.362684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:bcfe22158d70cfd13d0ae4277e4bd32b5d048296dec385d2e4b03df48d0f2fb6

Observation 68bb76cc-eb5f-459f-88f5-3ad4ac8d6b12 · outbound

This paper cites When and How to Express Empathy in Human-Robot Interaction Scenarios,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders When and How to Express Empathy in Human-Robot Interaction Scenarios,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.360215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:59e713384e06cdfbf0a6f9832cf95a7233fbebd40b9c07af7dc5943ca6bda229

Observation 73e92f35-0e57-4eba-8dd0-72374184aaf6 · outbound

This paper cites Visual Cues Enhance Predic- tive Turn-Taking for Two-Party Human Interaction,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Visual Cues Enhance Predic- tive Turn-Taking for Two-Party Human Interaction,

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:16:18.180515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:60acaccd5ca498b0383f164516e599ee9dbf82319ab980dc56c63f05fa2c3fb6

Observation fa1a4aef-7966-4b6b-a468-8273d5052685 · outbound

This paper cites Turn-taking in Conversational Systems and Human- Robot Interaction: A Review,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Turn-taking in Conversational Systems and Human- Robot Interaction: A Review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.382925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:6124a14cdf678624954dfd8dc39e443f91b7764bdfdecec676d1e1ee42422c6a

Observation 9466d6d4-c7cf-44d1-8baa-558880af1735 · outbound

This paper cites Data-driven models for timing feedback responses in a Map Task dialogue system,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Data-driven models for timing feedback responses in a Map Task dialogue system,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.356589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:51925688ef4338c53ffe8725f427ee9726c4d10e967e14759f86155ca8e9ba3a

Observation 147ab622-c0ef-4294-83e1-e929682a9059 · outbound

This paper cites On temporal aspects of turn taking in conversational dialogues,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders On temporal aspects of turn taking in conversational dialogues,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.403311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:70c2e920d91631241fa34f60910d83275680fa4d4508d1e9ec2c1879fb2172b0

Observation 5a3cfe38-9e60-465c-8841-10c1474ca5f6 · outbound

This paper cites Turn-Taking Modelling in Conversational Systems: A Review of Recent Advances,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Turn-Taking Modelling in Conversational Systems: A Review of Recent Advances,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.386459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:84d169a162dac8f6b43d138b5c11c8e73b335e39527ae149158540e5b3ea471f

Observation ad296677-e421-4ec5-a49f-bad88b7a66d7 · outbound

This paper cites Pauses, gaps and overlaps in conversa- tions,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Pauses, gaps and overlaps in conversa- tions,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.388335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:08f25c9698d0624f73179d8e94123b0efe7979e246448c83f97302a9ffa6b018

Observation 1782f32b-0171-4208-a4ca-efc617d8e4a3 · outbound

This paper cites Timing in turn-taking and its impli- cations for processing models of language,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Timing in turn-taking and its impli- cations for processing models of language,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.400964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b90e21d473fdd92f9fa8b54f5296adb7f45e2f1345d6cee871109794fe3a85f8

Observation 52335817-0e8d-4365-930c-919e49bd33b4 · outbound

This paper cites Timed picture naming in seven languages,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Timed picture naming in seven languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.381221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:a31927403051fb47ea6f0f40233845545b0c7347f16ce29a39c18ebb7b10f99c

Observation a5cecb58-38eb-49ad-a568-5d9f1e1f671d · outbound

This paper cites Multimodal Turn Analysis and Prediction for Multi-party Conversations,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Turn Analysis and Prediction for Multi-party Conversations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.377072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b9a97786a77eef2da0ab7e4fb5122ec24202f0e12dd37669b57ec9f6136a35b4

Observation 4bc0f3d5-efba-410c-8251-d83d4b2439f7 · outbound

This paper cites Voice Activity Projection: Self-supervised Learning of Turn-taking Events.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection: Self-supervised Learning of Turn-taking Events

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.178117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:11a24efc9752e06a4a0a156591de1e9f8740432425faeecffea4f5e5167839fe

Observation 8ccd277d-0bec-4730-98fa-10db1f7dcc45 · outbound

This paper cites Multimodal V oice Activity Projection for Turn-Taking and Effects on Speaker Adaptation,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal V oice Activity Projection for Turn-Taking and Effects on Speaker Adaptation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.375030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:c60cd31e460c1f9c92b673dfb965cd3960ad1a4d2db3e6da043ab166b8f259d3

Observation 470d4985-fcdd-40ae-960d-58c68db749cc · outbound

This paper cites Voice Activity Projection Model with Multimodal Encoders.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.175063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:f35eda28a5d463380969468d7b3bcd453583e702a538f00d66fcc06e914f8e7f

Observation 4ad90fe7-4bca-4639-8d5a-dc3befd6f0f8 · outbound

This paper cites Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.187022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:804f0458ac6a9f403564eead0dbd7ccb0e5d31c0d54c99c9e759747b6519a7d6

Observation 8e2c3d74-7b8b-46de-84b1-d06f7598c9ca · outbound

This paper cites Behind the scenes: Mechanistic interpretability of lora-adapted whis- per for speech emotion recognition.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Behind the scenes: Mechanistic interpretability of lora-adapted whis- per for speech emotion recognition

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:16:18.189958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5121472648ae52279088dda3ae65c3565ea0160d4f675e0ae5b0121aa044b029

Observation 81a4e9cf-4fcc-43ee-8751-69b055435dd9 · outbound

This paper cites GRPO- Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multi- modal Emotion Recognition,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders GRPO- Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multi- modal Emotion Recognition,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.381415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:4acc7fa39b0d2793c47ab818c35bd1fd1563ed7b37d9b73974a758f0bb5459ee

Observation 5ffdd3e1-feae-44d4-8b2e-af516f9e1ffa · outbound

This paper cites LoRA- Whisper: Parameter-Efficient and Extensible Multilingual ASR,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders LoRA- Whisper: Parameter-Efficient and Extensible Multilingual ASR,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.396659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:2614ab7fe48de85ea0f7ae3622d610baac1dcfa03eba30a9a689c633a762615c

Observation 627f9364-7480-480c-a8e3-e381cc071047 · outbound

This paper cites M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.189364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:eebcaad16ad050a7dd219c600b913917bbeae872d6207c0e9352d082e47310ee

Observation 60165d4b-3bf1-426a-abff-8715198e1dab · outbound

This paper cites Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.371418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:da4b68647f4e95f1879a7b011c9f3f298c25d99982e912c5697eeb03da6b533c

Observation e56bc495-1146-412f-a0d5-7d6b5ef30de5 · outbound

This paper cites Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.192224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:a7107ee2e40875a4576cee82baa38234f683f4c971c7d705e95ec1d9ac986088

Observation 4692176c-5662-45d6-b5ac-c5fa00ca9c1c · outbound

This paper cites Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.373311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:b26a08bcd1adfffc9be86020477a565af7c9469ce9b39fcce23c3d480818e11d

Observation a6bd654a-fe3d-425b-9f65-b98a84a59c74 · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.339998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:d398809dcf5f0e9a98b657f3b9fe13fd136c79936b7516ed0368571ca207e8fd

Observation 8823cca0-f77c-4553-b411-afb1410635fd · outbound

This paper cites Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.184153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:bc5892b33d16dcd1163fa81e1629a83735a64c736325fea171b966326ff4ac42

Observation f39153ad-0790-4acc-a295-d007ff5aaa2a · outbound

This paper cites Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.166295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:ed583e7099daa2df839f01cf6f9a13348e2df51ae7e92a8ea7f283747ddd6bff

Observation 9eef50a1-4531-47c6-904d-92b43f8c6444 · outbound

This paper cites Triadic Multi- party V oice Activity Projection for Turn-taking in Spoken Dia- logue Systems,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Triadic Multi- party V oice Activity Projection for Turn-taking in Spoken Dia- logue Systems,

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:16:18.195382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5101164a33fa740030e1cf78ba80425bf87dba5dd15121868209c8d478ffe013

Observation 6364f2d9-d724-4985-9d88-3fbf7c07ef6e · outbound

This paper cites Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of V oice Activity Projection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of V oice Activity Projection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.363896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:db829345a38116dfb7dec84c719c9e19d1d2acafa238fcb739e39e789dde7dd2

Observation 798ee106-eef7-402b-ad6e-c44d0ada81c3 · outbound

This paper cites Predicting End-of- turn and Backchannel Based on Multimodal V oice Activity Prediction Model,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Predicting End-of- turn and Backchannel Based on Multimodal V oice Activity Prediction Model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.365940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:92088b95de0ce9d85e92f962ac805204689ea8c0e8bdba4351537ea049136701

Observation efab9e88-132a-48c6-bc93-38cca0e4d9a0 · outbound

This paper cites Multi- lingual Turn-taking Prediction Using V oice Activity Projection,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multi- lingual Turn-taking Prediction Using V oice Activity Projection,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.366437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:777cdf820d925d58d84bb6f7f4963ec638d017336b7c211e144b46d428d98528

Observation deaf6d46-f308-4e15-b733-d528c8ebf0c3 · outbound

This paper cites Multilingual Turn-taking Prediction Using Voice Activity Projection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multilingual Turn-taking Prediction Using Voice Activity Projection

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.174784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:ebeba5e816944a16b2c635ebb79c77e25a77cd1941b13bfa839c82512b672455

Observation 43e6bba6-d75f-40ce-b5e1-dae18d5ae313 · outbound

This paper cites Investigating the Language Independence of V oice Activity Projection Models through Stan- dardization of Speech Segmentation Labels,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Investigating the Language Independence of V oice Activity Projection Models through Stan- dardization of Speech Segmentation Labels,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.353223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:f26ba590e974226acf5a91d05a53de8e6c68d7d0ab8582c3d1ac3eae2599b577

Observation e2e4601e-d5ba-4c72-b71c-0438836b5697 · outbound

This paper cites Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.172565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5cd2a690a825e4025904617ec3b27d228a3db81e669bbdf91102dc105157965d

Observation 15f01bfe-0936-4d00-88a9-dc5b1ca7f709 · outbound

This paper cites Applying General Turn-Taking Models to Conversational Human-Robot Interaction,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Applying General Turn-Taking Models to Conversational Human-Robot Interaction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.349501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:83561406f107b1510fd8caf38c11f2a26719cccb47d93fbcf483e338c4f5c396

Observation 1a211102-c65a-442b-9e0e-168a5460d3fc · outbound

This paper cites A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.383096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:796adee54eb4bf368963a74b86b36069dae9c33ba398364828edddd2f843b735

Observation 7534f11c-3d9f-4270-8692-5ce08e54edea · outbound

This paper cites Multimodal V oice Activity Prediction: Turn-taking Events Detection in Expert-Novice Conver- sation,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multimodal V oice Activity Prediction: Turn-taking Events Detection in Expert-Novice Conver- sation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.392045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:4fc24fe759c7e316aa16836242341083466475cbff2ff66732a5ca7a99f20d34

Observation fbb30dbd-9767-496d-9c1c-d43bcc6e0859 · outbound

This paper cites The NoXi database: multimodal recordings of mediated novice-expert interactions,.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders The NoXi database: multimodal recordings of mediated novice-expert interactions,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:16:18.361970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:4405f7c0904644ded7b931b99c115139b15992d220b02d0fd5d336244ebc5108

Observation 20d8a772-8675-41ad-9df8-c90507514cfb · outbound

This paper cites Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on Engagement.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on Engagement

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.186867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:a706618ef12119a1d58faab52fdd33b5666c176f3db1f10b9e75f545a047d84d

Pith citing papers

No inbound Pith citation observations are available.