Pith. sign in

REVIEW 14 cited by

R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05379 v2 pith:YFLX3INS submitted 2025-03-07 cs.LG cs.CV

classification cs.LGcs.CV
keywords emotionrecognitionmodelrlvraudiocapabilitylanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we present the first application of Reinforcement Learning with Verifiable Reward (RLVR) to an Omni-multimodal large language model in the context of emotion recognition, a task where both visual and audio modalities play crucial roles. We leverage RLVR to optimize the Omni model, significantly enhancing its performance in three key aspects: reasoning capability, emotion recognition accuracy, and generalization ability. The introduction of RLVR not only improves the model's overall performance on in-distribution data but also demonstrates superior robustness when evaluated on out-of-distribution datasets. More importantly, the improved reasoning capability enables clear analysis of the contributions of different modalities, particularly visual and audio information, in the emotion recognition process. This provides valuable insights into the optimization of multimodal large language models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

    cs.AI 2026-07 conditional novelty 7.0 of 10

    A new in-cabin dataset pairs RGB/IR video, audio, and Chinese dialogue text with emotion, fatigue, and distraction labels, and baselines show fusion beats single modalities on the Chinese partition.

  2. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free, query-anchored, modality-decoupled token compression method outperforms unidirectional audio/video compression baselines on Qwen2.5-Omni while cutting prefill cost.

  3. EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

    cs.AI 2026-07 conditional novelty 6.0 of 10

    EmoAgent-R1 combines dynamic agent routing with a token-reweighted GRPO variant (P-GRPO) to reach 77.85% mean on MER-UniBench, exceeding AffectGPT-R1 by 1.90 points.

  4. TimeThink: Reasoning with Time for Video LLMs

    cs.CV 2026-07 accept novelty 6.0 of 10

    TimeThink adds step-wise temporal process rewards (max IoU of referenced intervals) to GRPO for Video-LLMs, improving grounding and reasoning over outcome-only RL baselines.

  5. Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new large-scale facial emotion caption dataset and a global-local contrastive training framework with positive mining improve zero-shot facial expression recognition.

  6. LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free token compression method using semantic connected components in space and time keeps video understanding accuracy high even when retaining only 5-10% of visual tokens.

  7. HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Requiring omni-modal models to summarize context before reasoning, with LLM-judged context and logical rewards, improves human-intent reasoning benchmarks.

  8. GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GRPO-CARE improves answer accuracy and reasoning coherence over standard GRPO on a new video reasoning benchmark, with a 6.7 point gain on the hardest level and a 24.5 point higher consistency rate.

  9. OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

    cs.CV 2026-06 conditional novelty 5.0 of 10

    Rationale-privileged on-policy self-distillation reaches 84.19 mean on MER-UniBench by scoring student rollouts with a local teacher that alone sees frontier-generated multimodal evidence.

  10. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

    cs.AI 2026-02 conditional novelty 5.0 of 10

    EMO-R3, which combines a three-step emotional reasoning prompt with a reward for the model agreeing with its own image–emotion judgments, raises visual emotion-recognition accuracy by about one point over plain GRPO.

  11. Group Relative Policy Optimization for Speech Recognition

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.

  12. Advancing the Foundation Model for Music Understanding

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    MuFun is proposed as a unified music foundation model that jointly handles instrumental and lyrical content, and it is claimed to outperform existing audio language models on the authors' new MuCUE benchmark.

  13. From Emergence to Control: Probing and Modulating Self-Reflection in Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    Self-reflection in LLMs can be steered up or down by a single activation-space vector, improving accuracy when amplified and cutting output length when suppressed.

  14. Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional Data

    cs.DB 2025-08 reject novelty 4.0 of 10

    The submission's title and abstract describe a single-server privacy-preserving ANNS system with distance comparison encryption, but the manuscript body is an unrelated activity-recognition paper, leaving the advertis...

Pith tools