Pith. sign in

For SP and BC, our models (Proposed1 and Proposed3) outperformed the baselines, where the differ- ence lies in the use of facial embeddings from the pre-trained encoder

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

method 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

roles

method 1

polarities

use method 1

representative citing papers

Voice Activity Projection Model with Multimodal Encoders

cs.CL · 2025-06-04 · conditional · novelty 4.0

Adding a pretrained facial-expression encoder to a voice activity projection model yields competitive or slightly better turn-taking prediction on the NoXi French subset, though the improvement is not statistically validated.

citing papers explorer

Showing 1 of 1 citing paper.

  • Voice Activity Projection Model with Multimodal Encoders cs.CL · 2025-06-04 · conditional · none · ref 8

    Adding a pretrained facial-expression encoder to a voice activity projection model yields competitive or slightly better turn-taking prediction on the NoXi French subset, though the improvement is not statistically validated.