REVIEW 4 cited by
EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa, a simple yet expressive scheme of solving the ERC (emotion recognition in conversation) task. By simply prepending speaker names to utterances and inserting separation tokens between the utterances in a dialogue, EmoBERTa can learn intra- and inter- speaker states and context to predict the emotion of a current speaker, in an end-to-end manner. Our experiments show that we reach a new state of the art on the two popular ERC datasets using a basic and straight-forward approach. We've open sourced our code and models at https://github.com/tae898/erc.
Forward citations
Cited by 4 Pith papers
-
SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations
SCoPE adds a speaker-conditioned GRU prior gated by predicted emotion shifts to multimodal ERC, reporting state-of-the-art IEMOCAP results and consistent baseline gains.
-
Towards Designing Social Interventions For Online Climate Change Denialism Discussions
In-field Reddit bot interventions using insider climate-denial language and linked evidence produced more positive engagement from climate change deniers and additional evidence from supporters.
-
AtmosERC: Modeling Dialogue-Level Affective Atmosphere for Emotion Recognition in Conversation
A relation-aware conversational graph can extract a reusable affective-atmosphere prior that modestly improves lightweight and LLM-based emotion recognition in conversation.
-
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.
Discussion (0). Continue with ORCID to comment.