Pith. sign in

REVIEW 4 cited by

Benchmarking Cross-Domain Audio-Visual Deception Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06995 v4 pith:7SWAJTPQ submitted 2024-05-11 cs.SD cs.CVcs.MMeess.AS

classification cs.SDcs.CVcs.MMeess.AS
keywords deceptiondetectionaudio-visualcross-domaindomaingeneralizationperformanceaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine the authenticity of an individual's statements. Nevertheless, recent developments in automated deception detection have demonstrated that multimodal features derived from both audio and video modalities may outperform human observers on publicly available datasets. Despite these positive findings, the generalizability of existing audio-visual deception detection approaches across different scenarios remains largely unexplored. To close this gap, we present the first cross-domain audio-visual deception detection benchmark, that enables us to assess how well these methods generalize for use in real-world scenarios. We used widely adopted audio and visual features and different architectures for benchmarking, comparing single-to-single and multi-to-single domain generalization performance. To further exploit the impacts using data from multiple source domains for training, we investigate three types of domain sampling strategies, including domain-simultaneous, domain-alternating, and domain-by-domain for multi-to-single domain generalization evaluation. We also propose an algorithm to enhance the generalization performance by maximizing the gradient inner products between modality encoders, named ``MM-IDGM". Furthermore, we proposed the Attention-Mixer fusion method to improve performance, and we believe that this new cross-domain benchmark will facilitate future research in audio-visual deception detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    Schema-constrained MLLM reports plus SICS/DMC modules and the T4-Deception dataset yield SOTA in-domain, cross-domain, and cross-cultural multimodal deception detection.

  2. ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    ThinkDeception introduces MLLMs, a multimodal CoT dataset, and VAC-GRPO progressive RL to convert deception detection into interpretable reasoning and claims new SOTA accuracy plus rationale quality.

  3. AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

    cs.MM 2025-08 conditional novelty 5.0 of 10

    AD-AVSR combines dual-stream audio encoding, audio-guided visual refinement, visual-guided noise suppression, and thresholded audio-visual pair selection to improve audio-visual speech recognition word error rates und...

  4. Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings

    cs.CL 2025-06 conditional novelty 5.0 of 10

    An evaluation of 7 LLMs/LMMs on 3 deception datasets shows fine-tuned text LLMs set benchmarks on review spam while multimodal models lag behind video-based baselines.

Pith tools