Pith. sign in

Paper Citation Record · LEDGER

Voice Activity Projection Model with Multimodal Encoders

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2506.03980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03980 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:34.901221Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:30.784001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.173817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact2
  • verified fuzzy44
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · outbound

This paper cites Voice Activity Projection Model with Multimodal Encoders.

Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:30.784001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:30.784001Z digest=sha256:a689de9f1ac1e0c4877ee7ff059ccad7b80343656e7306ba4232359130beb028

Observation 73b51346-a3ec-4558-a0b3-360959c5f71a · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:43.054931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:30.879347Z digest=sha256:9425f862d01044c07278917ff6df50605314466db7409d73b58b171fc251d51b

Observation aa716f32-312b-415a-a2b2-ae805ea34828 · outbound

This paper cites Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units.

Voice Activity Projection Model with Multimodal Encoders Unlike Onishi’s multimodal model, we in- serted a pretrained encoder for facial images instead of action units

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:30.968030Z digest=sha256:3eaa4793bc2bb9cc48ac8a80664cc01aa45fdaed6c7a861d5f35738924676676

Observation c6cd6bf3-e16e-41ee-9126-6c918ecce3f2 · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:42.792038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.049890Z digest=sha256:72287635f2cfaf9ec5520e60d0ce64194b137e2f0e74d72c1947d0596209819c

Observation 224039b8-c633-4b71-a07f-f0d11163178b · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:42.629601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.162124Z digest=sha256:25599156afe0253bedb29490a8a01f75bc2fda9009f70e68580b49a54954288b

Observation 4cdeaaba-ed8d-459b-bdc4-8419da09eb7c · outbound

This paper cites Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities.

Voice Activity Projection Model with Multimodal Encoders Onishi’s approach involved merging user signals for each modality sepa- rately and then fusing across different modalities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.446772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.240357Z digest=sha256:5bc2e1f191250e977b99661123bc1aa18f959e4cd62345b41fb98a2286cb0fca

Observation 3ed24270-7e68-43f5-96e1-a5766682c055 · outbound

This paper cites Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37].

Voice Activity Projection Model with Multimodal Encoders Dataset: NoXi We used the NoXi multimodal dataset for our experiments [37]

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.334883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.302301Z digest=sha256:7cd4ddde78dd63cb0a99fce072a9ace3e620188d4bf4f180a48ba227ed0a90ab

Observation 780aeee2-700b-4651-a456-3398cd8e9303 · outbound

This paper cites For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder.

Voice Activity Projection Model with Multimodal Encoders For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.153637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.386731Z digest=sha256:c6119d0f6340ab7628711467e041f845549ffe0944685b854157e9e9b3e43e10

Observation 60c487df-7565-4fb3-ade1-fefc984fdf05 · outbound

This paper cites By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models.

Voice Activity Projection Model with Multimodal Encoders By incorporating a face image encoder, the proposed models demonstrated competitive or even superior performance compared to the multimodal state-of-the-art (SoTA) models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:42.009648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.455482Z digest=sha256:df704e260f7183fdcffe29968057f54296dfb864f3f4b9b64628d8dea8d4704e

Observation eb5de9e2-133c-41a7-8d17-4924acea615d · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:41.838076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.536151Z digest=sha256:c634059ba094ebecee3720979f4eac4052267b3143cfb4eaee41411f67bb3d69

Observation 62bce116-9f4b-44ee-9713-047f6d3fcf15 · outbound

This paper cites A simplest systemat- ics for the organization of turn-taking for conversation,.

Voice Activity Projection Model with Multimodal Encoders A simplest systemat- ics for the organization of turn-taking for conversation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.696757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.603372Z digest=sha256:79dff89dfaed7f0787bf3bffbed1c83895e9b749272541b30cd40782d55c7c35

Observation 5b02f3d6-2f7a-4823-a7dd-4c68fde28c3d · outbound

This paper cites Turn-taking: A critical analysis of the research tradition,.

Voice Activity Projection Model with Multimodal Encoders Turn-taking: A critical analysis of the research tradition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.521561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.672711Z digest=sha256:a4c5e060cbb45e32ece18f8d891e95542d13641d9daeddbc61f25da3a4fbf457

Observation 591fbe38-ee34-4ebf-9c78-a272121bf8a0 · outbound

This paper cites Some signals and rules for taking speaking turns in conversations,.

Voice Activity Projection Model with Multimodal Encoders Some signals and rules for taking speaking turns in conversations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.371958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.768806Z digest=sha256:085e0c57cae73a009602dd8fb393480ffe0c613e8e3ecc6ee73db41369abd54e

Observation aec0d1ca-d0be-4ae3-ad90-f445f9f4a695 · outbound

This paper cites Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,.

Voice Activity Projection Model with Multimodal Encoders Interactional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:41.256020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.848267Z digest=sha256:9047ad7db04a924748becc5aca6985dd68777f71ffe12c1bc4da0bfff72fdc3e

Observation e1f26d21-f463-40f4-91fc-2084d25146de · outbound

This paper cites an unresolved cited work.

Voice Activity Projection Model with Multimodal Encoders Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:53:41.114513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.931068Z digest=sha256:b1747062635e0e015f68886906f356138e27f29368a59eed03c1d039ec68147c

Observation cd7c379f-dd7d-47c8-a25b-bde9899b98af · outbound

This paper cites Using uh and um in spontaneous speaking,.

Voice Activity Projection Model with Multimodal Encoders Using uh and um in spontaneous speaking,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.996352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:31.990975Z digest=sha256:52e2c73d92f4a4fa92395c71fc48d631fa2d4046a86c8d04bb94df77d0f65f6c

Observation 09b69a46-189d-4ea2-83fb-766c2db89b85 · outbound

This paper cites Using prosodic clues to decide when to produce back- channel utterances,.

Voice Activity Projection Model with Multimodal Encoders Using prosodic clues to decide when to produce back- channel utterances,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.844132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.071262Z digest=sha256:5a5be314d01251552a87dcee4b90565feffee43a41fdf55487bbbc07040e217f

Observation ccd90798-8833-4eb4-be36-b1610fea83a1 · outbound

This paper cites Nonverbal behaviours improving a simulation of small group discussion,.

Voice Activity Projection Model with Multimodal Encoders Nonverbal behaviours improving a simulation of small group discussion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.712473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.132742Z digest=sha256:5fae81421b420104ccd9e44cd2228b5260e7049050896731bc722247f480026c

Observation 13f5be9a-d41e-4f27-8a59-05e5a9ac65c3 · outbound

This paper cites Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,.

Voice Activity Projection Model with Multimodal Encoders Sustaining interaction dynamics and engage- ment in dyadic child-robot interaction kinesics: lessons learnt from an exploratory study,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.564372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.220226Z digest=sha256:5f90e5e415d8bba8a785993da0552586ed03d7b9ff52f72069b924abeb68c5c1

Observation 63bf7315-b13f-4021-ac0c-781f5dc7fbcb · outbound

This paper cites The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,.

Voice Activity Projection Model with Multimodal Encoders The effects of interrupting behavior on interpersonal attitude and engagement in dyadic in- teractions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.417748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.302198Z digest=sha256:f486d67294bce2648c85d927d911739171730157b993c8b89f313588486e0fd3

Observation d6ffe9f2-4f0a-47e6-92d8-2091d2c9dec8 · outbound

This paper cites Now or when? interrup- tion timing prediction in dyadic interaction,.

Voice Activity Projection Model with Multimodal Encoders Now or when? interrup- tion timing prediction in dyadic interaction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.299939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.384924Z digest=sha256:9a5088e18acc65cd6409c5a8cbd7f62f0d18b261ba064232411b5ab1fe8146bf

Observation f42f06c1-b3a3-4cfd-bf38-f36ac0dae25d · outbound

This paper cites How turn-taking strategies influence users’ impressions of an agent,.

Voice Activity Projection Model with Multimodal Encoders How turn-taking strategies influence users’ impressions of an agent,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.132579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.475213Z digest=sha256:4de4c9b66bf81accf79d7bdfba64fe21bdc6d03d4afd9e5b72254b99ff4b625c

Observation 62270099-7a5e-4217-afb9-d21fc7f21aa8 · outbound

This paper cites Timing in turn-taking and its im- plications for processing models of language,.

Voice Activity Projection Model with Multimodal Encoders Timing in turn-taking and its im- plications for processing models of language,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:40.014896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.568382Z digest=sha256:d4c7e7cd2e38ec4d189bf8a1cec304da2fd77908a83b9f080c7847261708f2b3

Observation bb06fffc-88c0-48bf-8da3-211339adf5c2 · outbound

This paper cites Timing in conversation,.

Voice Activity Projection Model with Multimodal Encoders Timing in conversation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.863966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.636080Z digest=sha256:a3ae6785faf62ca8831b2da5a2c749a0711edb5cd7dc0d93baa6c85374d41e9d

Observation df526a09-03c7-476f-95ea-b7f86eab929f · outbound

This paper cites Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,.

Voice Activity Projection Model with Multimodal Encoders Can I finish?: learning when to respond to incremental interpretation results in interac- tive dialogue,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.721820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.724770Z digest=sha256:55ba6f3e243521e2347616e5cdbc5f6e1685ceaad406ef496b491c31a883ea6a

Observation 4da0c1f6-3c35-4047-bdf6-26c2e15dce5f · outbound

This paper cites Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,.

Voice Activity Projection Model with Multimodal Encoders Investigating fluidity for human- robot interaction with real-time, real-world grounding strategies,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.546358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.843858Z digest=sha256:9282af37dff0240004087a72b6f9d162e3c21f95e5a794c3148ee847e2d2d1c2

Observation 0c39117f-2ccc-46af-b5d0-f9a8e5f281eb · outbound

This paper cites Attentive listening system with backchanneling, response generation and flexible turn-taking,.

Voice Activity Projection Model with Multimodal Encoders Attentive listening system with backchanneling, response generation and flexible turn-taking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.388384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.909303Z digest=sha256:86b7d51e7a0b62d76cefc9d72d3f090c1c7655713e7f19593e4230445d8df1d1

Observation 4397cd86-ad7e-4d7a-98d4-0b499497769d · outbound

This paper cites V oice activity projection: Self- supervised learning of turn-taking events,.

Voice Activity Projection Model with Multimodal Encoders V oice activity projection: Self- supervised learning of turn-taking events,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.187792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:32.990781Z digest=sha256:d592558d0e04175c267a20cad65d58acce3d2dd875e03007a8b8bf972fc4bbe4

Observation 88e17781-a09a-4660-8764-15a4c9979cef · outbound

This paper cites Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,.

Voice Activity Projection Model with Multimodal Encoders Pre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:39.016839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.117340Z digest=sha256:108a05bd8302af97956ed09dd34830995b5877a66399fd05657e717cf038fbfe

Observation 72879ccf-5500-4aeb-9098-a3004ce9138c · outbound

This paper cites Predicting turn-taking by compact gazing transition patterns in multiparty conversation,.

Voice Activity Projection Model with Multimodal Encoders Predicting turn-taking by compact gazing transition patterns in multiparty conversation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.872661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.186043Z digest=sha256:1d75bb53409ace56f3478dc67b2b97342e7a13fe4381d60b86f0f153ec6607d1

Observation f9c1df9b-1a7d-4722-a582-d1657e5baa05 · outbound

This paper cites Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,.

Voice Activity Projection Model with Multimodal Encoders Prediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.700880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.270468Z digest=sha256:ec040c627a6c55c481db55c48729726b24b93314e5f0ee399300f9ce528ba4b2

Observation cffd3c23-bf87-4e5b-a3c6-4e63e3116661 · outbound

This paper cites Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,.

Voice Activity Projection Model with Multimodal Encoders Multimodal voice ac- tivity prediction: Turn-taking events detection in expert-novice conversation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.574119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.345041Z digest=sha256:cb00df862d3ec2efadd0d41398323df6f7eb0547d4061820f9022fb0612dd712

Observation 27f3612d-ddcc-497f-9b80-59a3d55a0de0 · outbound

This paper cites Turn-taking and backchannel prediction with acoustic and large language model fusion,.

Voice Activity Projection Model with Multimodal Encoders Turn-taking and backchannel prediction with acoustic and large language model fusion,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.412562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.427059Z digest=sha256:fc64f1a3a626d88969813985754bef70ecfab2cdd8396f6f1451ac2cd443c31b

Observation ee2159a5-e41e-429c-b0ab-71b5bfd41980 · outbound

This paper cites Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,.

Voice Activity Projection Model with Multimodal Encoders Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot di- alogue,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.272095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.524622Z digest=sha256:7f2c32bd9a6e530d784086af8dae4fa60d3d8b0ff5ffbd58aa3a92a72ea8aa2b

Observation 319c46e6-58d3-46f3-b227-99ef7960d4b4 · outbound

This paper cites Mini-omni: Language models can hear, talk while thinking in streaming,.

Voice Activity Projection Model with Multimodal Encoders Mini-omni: Language models can hear, talk while thinking in streaming,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:38.118055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.597594Z digest=sha256:09c4f8c44a27ea7cbc01a07bf1ac50e223b9afacc3cd12a84bb895e609644cab

Observation 38d5a649-549b-4ec6-bb89-2e1661e190ac · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

Voice Activity Projection Model with Multimodal Encoders Moshi: a speech-text foundation model for real-time dialogue,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.955757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.663036Z digest=sha256:b2f74dd4fb402b1505d1a7cc8fe77303b5db0f3ab4893efb770417be9172d968

Observation ca140c6e-e660-48da-9588-d30903dab38a · outbound

This paper cites [Online].

Voice Activity Projection Model with Multimodal Encoders [Online]

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:53:35.452694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.737389Z digest=sha256:58e44aa584fe11b3fbe20b1f2ce27f1ceb2835f6ccc7b9889095e8e3ba251c75

Observation ec01cbc1-9b8f-4645-8879-73d67df253f6 · outbound

This paper cites [Online].

Voice Activity Projection Model with Multimodal Encoders [Online]

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.793002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.808721Z digest=sha256:6d5c70b7af62bc374cb6d2f21f847937160db22910f5880ef4eebc2b7a2e742f

Observation 881446a4-fe91-4cae-9141-96ada48bfbe2 · outbound

This paper cites Real-time and continuous turn-taking prediction using voice ac- tivity projection,.

Voice Activity Projection Model with Multimodal Encoders Real-time and continuous turn-taking prediction using voice ac- tivity projection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.615970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.886558Z digest=sha256:8741e7e5125220afa9469d8b205743a9b4420a6c3620855e63768fed1bf83e9a

Observation 3fc479d7-856d-45ad-b697-8ddd0a9c8624 · outbound

This paper cites Multilingual turn-taking prediction using voice activity projection,.

Voice Activity Projection Model with Multimodal Encoders Multilingual turn-taking prediction using voice activity projection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.419432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:33.981132Z digest=sha256:5cb57cb39a1e811cea833a6d994216ee362377795b8e69082752a213bf5b9998

Observation e17c12c0-735c-4ac2-b140-e657a73e465a · outbound

This paper cites How much does prosody help turn- taking? investigations using voice activity projection models,.

Voice Activity Projection Model with Multimodal Encoders How much does prosody help turn- taking? investigations using voice activity projection models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.283172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.064405Z digest=sha256:3e73759065367f7eff50569a0b24c6a61bdb27d1b8bff050ada0ac59a01adfb8

Observation 2e2e21d0-e993-4a8e-a422-81c281be1e99 · outbound

This paper cites OpenFace: An open source facial behavior analysis toolkit,.

Voice Activity Projection Model with Multimodal Encoders OpenFace: An open source facial behavior analysis toolkit,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:37.094480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.136834Z digest=sha256:9b175a188aa1c7bee5f30c40af3cfd2cf675ff7d1bc06ace25fb2c8e06830781

Observation 0a9cae1f-8d18-410a-888c-8affc0869fb9 · outbound

This paper cites Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,.

Voice Activity Projection Model with Multimodal Encoders Open- pose: Realtime multi-person 2d pose estimation using part affinity fields,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.970605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.203584Z digest=sha256:3e3b3d6df716eaf4a41c25df592a27850f9734fea8600d1e676aa0d2abf9d4e1

Observation e6fd51c1-dc02-4a73-b9ac-86fccdf24ffa · outbound

This paper cites Former-dfer: Dynamic facial expression recognition transformer,.

Voice Activity Projection Model with Multimodal Encoders Former-dfer: Dynamic facial expression recognition transformer,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.817813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.268089Z digest=sha256:8a60195d747060552fcc1ef9b3abd18ada6ff6d2a679596ba9a61d48085f7daf

Observation 4d281b4e-0199-4e08-b775-6c119c49253e · outbound

This paper cites Unsu- pervised pretraining transfers well across languages,.

Voice Activity Projection Model with Multimodal Encoders Unsu- pervised pretraining transfers well across languages,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.605534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.338299Z digest=sha256:ff12663d1d09a0cbeaad39c2fb6986e7a13f459dd2a4b8ed591e794173d7b6bc

Observation 55d37418-757f-4d48-ba23-cd3102fd6f51 · outbound

This paper cites Dlib-ml: A machine learning toolkit,.

Voice Activity Projection Model with Multimodal Encoders Dlib-ml: A machine learning toolkit,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.456539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.416269Z digest=sha256:adaac1f188a79fc65580449f05297ffcac3d12ae05dbd756592567451e9c365c

Observation 0f56658d-9e42-4ba9-923b-fac07bf14a51 · outbound

This paper cites The NoXi database: multimodal recordings of mediated novice-expert interactions,.

Voice Activity Projection Model with Multimodal Encoders The NoXi database: multimodal recordings of mediated novice-expert interactions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.282903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.513926Z digest=sha256:8af5589ab92998b41bd1d9db156d257ce614ecfafbf9ac78fe4d16d89d5f5453

Observation 4cdced7e-395a-4d45-a249-982e7f1df638 · outbound

This paper cites PyTorch Lightning,.

Voice Activity Projection Model with Multimodal Encoders PyTorch Lightning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:36.137530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.593139Z digest=sha256:224aed48958a3df3cc0ce60bde50f0c7670d5d43a261a129aa2e732af2c3f3d3

Observation 73512c0f-2b7e-4b8b-9385-91b48365d77a · outbound

This paper cites Discourse as an interactional achievement iii: The omnirelevance of action,.

Voice Activity Projection Model with Multimodal Encoders Discourse as an interactional achievement iii: The omnirelevance of action,

Reference 50

Resolution
verified exact
doi, observed 2026-08-07T10:53:35.150121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.751806Z digest=sha256:e4e7733f806dd526c4b9687a191b6987133bcbb7fb1e79170b8a0a78bf105975

Observation eda6959a-4dde-4ea0-ab68-e3e9ac2710a4 · outbound

This paper cites Between and within: Alternative sequential treat- ments of continuers and assessments,.

Voice Activity Projection Model with Multimodal Encoders Between and within: Alternative sequential treat- ments of continuers and assessments,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.828855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.819433Z digest=sha256:035e580e21f7bf6e2594ca5761d8c577935421fd554e28a3ef723fd35192a431

Observation 5ab26f48-cf14-4575-a333-2892af7ef0f5 · outbound

This paper cites Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,.

Voice Activity Projection Model with Multimodal Encoders Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.653190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.901221Z digest=sha256:9da54485c5cf2e8cf91741cc182df2e59b709b677bd2dd3a0499596c1b177cf9

Observation fba6514b-15e3-4dd3-90ba-c469c37b185a · outbound

This paper cites Available: https://github.com/Lightning-AI/ pytorch-lightning.

Voice Activity Projection Model with Multimodal Encoders Available: https://github.com/Lightning-AI/ pytorch-lightning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:35.986850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:34.665269Z digest=sha256:bb4a52e7b9c8e125626c79c8edb602099ec9170623cce11e30e236881fe8c79f

Pith citing papers

Observation 9de12901-6558-4484-b0f2-f1c58c483bb6 · inbound

Voice Activity Projection Model with Multimodal Encoders cites this paper.

Voice Activity Projection Model with Multimodal Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:30.784001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:30.784001Z digest=sha256:a689de9f1ac1e0c4877ee7ff059ccad7b80343656e7306ba4232359130beb028

Observation 470d4985-fcdd-40ae-960d-58c68db749cc · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Voice Activity Projection Model with Multimodal Encoders

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.175063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:5094bb40cfc8f8fa686226b21446264178bf888a1d6885b9afc97e595a30a4c3