{"id":"520f3e4f-43dc-4825-b2d5-e9a27b31955d","arxiv_id":"2605.30472","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Multimodal speech models show word error rate differences of up to 4.05 points when the same audio is paired with faces differing in self-declared gender and ethnicity.","lead":"The paper creates videos pairing different faces with identical audio and measures how multimodal speech models change their transcription accuracy. Smart generalists should read it to see that adding visual signals to AI systems can create new quality-of-service gaps across gender and ethnicity.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"WER differences may arise from video generation/lighting artifacts rather than face-modality processing","rationale":"The reader's weakest_assumption directly identifies the same experimental-control gap that is load-bearing for the causal claim. Because the abstract alone does not resolve it and the full text is referenced but not reproduced here, the UNVERDICTED status is unaffected.","tokens_in":1672,"tokens_out":287,"duration_ms":18433,"concrete_test":"Re-generate the face-swapped videos using a single fixed pipeline with explicit controls (identical lighting, background, lip landmarks, codec settings) and re-run the mWhisper-Flamingo and Gemini evaluations; if the WER gaps disappear or fall below statistical significance, the modality-bias interpretation is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that identical audio paired with different faces produces WER changes solely because the model processes the visual face input. This holds only if video synthesis controls every other visual variable (lighting, background, lip-sync precision, resolution, compression). The abstract states videos were created but supplies no description of the generation pipeline, matching procedure, or controls (e.g., same actor body, fixed camera, identical post-processing). Without those, any observed 4.05 WER gap could be an artifact of the synthesis step rather than an internal model bias.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents the first bias evaluation of multimodal speech recognition models. It creates videos that pair different faces (varying by self-declared gender and ethnicity) with identical audio clips, then measures resulting changes in word error rate (WER) for models including mWhisper-Flamingo and Gemini. The central empirical finding is large quality-of-service disparities, with WER drops of up to 4.05 points across gender, ethnicity, and their intersections.","tokens_in":1768,"tokens_out":505,"duration_ms":17080,"significance":"If the WER differences are shown to arise from the models' internal processing of the visual face modality rather than from uncontrolled factors in video synthesis, the result would be significant. It would demonstrate that adding modalities can degrade performance and introduce demographic biases even when audio is held fixed, providing a concrete priority for developers to evaluate multimodal systems for fairness. The work is strengthened by its focus on intersectional effects and by using real self-declared demographic labels.","major_comments":[{"comment":"The experimental setup (Methods section) does not describe the video generation pipeline or the controls used to isolate the face modality. No information is supplied on how faces were synthesized or swapped onto the same audio, nor on whether lighting, background, camera angle, lip-sync precision, resolution, and post-processing compression were held identical across conditions. Without these controls, the observed 4.05 WER gaps cannot be attributed to model-internal face processing rather than synthesis artifacts.","section":"Methods"},{"comment":"Results reporting (Results section) provides no dataset size, number of audio clips per demographic group, statistical tests for the WER differences, or error bars. The claim of 'large quality-of-service differences' therefore rests on point estimates whose reliability and generalizability cannot be assessed from the supplied information.","section":"Results"}],"minor_comments":[{"comment":"The abstract and introduction use 'self-declared gender, ethnicity' without clarifying how these labels were obtained or whether they align with the visual appearance presented to the model.","section":"Abstract"},{"comment":"Notation for the models (mWhisper-Flamingo, Gemini) should include version numbers or exact checkpoints used, as multimodal behavior can change across releases.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their detailed and constructive review. The comments highlight important areas for improving methodological transparency and statistical rigor. We address each point below and will revise the manuscript to incorporate additional details and analyses.","responses":[{"response":"We agree that the Methods section requires expanded description of the video generation pipeline. In the revised manuscript we will add a dedicated subsection detailing the face-swapping procedure (including the specific synthesis tool and parameters), the source of the base videos, and explicit controls ensuring identical lighting, background, camera angle, resolution, and compression across all conditions. Lip-sync was performed with a fixed audio track and verified for consistency; we will report quantitative checks on sync quality. These additions will allow readers to evaluate whether differences arise from model processing of the visual modality.","revision_made":"yes","referee_comment":"[Methods] The experimental setup (Methods section) does not describe the video generation pipeline or the controls used to isolate the face modality. No information is supplied on how faces were synthesized or swapped onto the same audio, nor on whether lighting, background, camera angle, lip-sync precision, resolution, and post-processing compression were held identical across conditions. Without these controls, the observed 4.05 WER gaps cannot be attributed to model-internal face processing rather than synthesis artifacts."},{"response":"We acknowledge the need for fuller statistical reporting. The revised Results section will state the total number of audio clips, the breakdown per demographic group (gender, ethnicity, and intersections), and will include error bars or confidence intervals on all WER values. We will also add appropriate statistical tests (e.g., paired t-tests or Wilcoxon tests with multiple-comparison correction) comparing WER across face conditions for each model. These changes will substantiate the reliability of the reported differences.","revision_made":"yes","referee_comment":"[Results] Results reporting (Results section) provides no dataset size, number of audio clips per demographic group, statistical tests for the WER differences, or error bars. The claim of 'large quality-of-service differences' therefore rests on point estimates whose reliability and generalizability cannot be assessed from the supplied information."}],"tokens_in":1350,"tokens_out":468,"duration_ms":14390,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that pairing different faces with identical audio produces measurable WER shifts in mWhisper-Flamingo and Gemini, up to 4.05 points across gender and ethnicity intersections. That setup is new relative to the single-modality bias papers it cites.\n\nWhat the work does well is extend the fairness question to the multimodal case. The abstract makes a clean point that adding a modality is not automatically an improvement and can create new quality-of-service gaps. The experimental framing—same audio, swapped faces—is direct and easy to understand.\n\nThe main soft spot is exactly the one the stress test flags. The abstract says videos were created but supplies no pipeline description, no controls for lighting, background, lip-sync accuracy, resolution, or compression, and no mention of how faces were aligned to the audio track. Without those, any WER difference could trace to synthesis artifacts rather than the model’s use of the visual input. No dataset sizes, statistical tests, or error bars are referenced either, which leaves the 4.05-point claim hard to evaluate.\n\nThis is for researchers who study bias in multimodal systems. A reader already working on fairness metrics would pick up the experimental idea and the cautionary result. The paper is coherent on its own terms and engages the relevant literature, so it deserves a serious referee even though the current version needs substantial methods clarification. I would send it to review with a request for the missing controls and stats.","headline":"This paper runs the first controlled test of face-induced bias in multimodal ASR and finds WER gaps up to 4 points, but the video generation details are missing so the cause is unclear.","tokens_in":2220,"tokens_out":379,"would_cite":false,"duration_ms":16023,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Multimodal speech models produce higher word error rates when the same audio is paired with faces from certain gender and ethnic groups.","keywords":["multimodal speech recognition","bias evaluation","word error rate","gender bias","ethnicity bias","audio-visual models","quality of service"],"falsifier":"Re-running the exact audio tracks with faces generated under fully controlled and identical conditions and checking whether the word error rate gaps of up to 4.05 points remain.","tokens_in":2572,"feed_emoji":"🎤","tokens_out":432,"duration_ms":23825,"temperature":0.7,"pith_summary":"The paper tests whether adding a visual face modality to speech recognition improves performance equally across groups. Researchers generate videos that keep the audio track fixed while swapping in different faces, then run the same audio through two multimodal models. Transcription accuracy drops by as much as 4.05 word error rate points for some gender-ethnicity combinations. The results indicate that extra modalities can introduce new quality-of-service gaps rather than eliminate them. Developers are urged to measure and address these effects before wider deployment.","feed_headline":"Same audio, different faces yield up to 4-point WER gaps","feed_subtitle":"Multimodal models transcribe fixed audio less accurately when paired with certain gender and ethnic faces, showing extra signals can create","key_machinery":"The controlled face-audio pairing experiment that holds the audio constant while varying only the visual face input to isolate effects on transcription output.","core_discovery":"We create videos pairing different faces with the same audio and measure changes in speech transcription accuracy. We find large quality-of-service differences across mWhisper-Flamingo and Gemini models, with drops of up to 4.05 word error rate points, across self-declared gender, ethnicity, and their intersection.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Different faces same audio cause up to 4 WER gaps in models","Multimodal models show 4 point WER differences from face pairs","Same audio different faces lead to 4 WER point transcription gaps","Face gender and ethnicity create 4 point WER drops in speech AI"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The observed transcription differences are produced by the models' processing of the face visual input rather than by differences in video generation, lighting, or other experimental factors.","fun_headline_variants_meta":{"raw":{"variants":["Different faces same audio cause up to 4 WER gaps in models","Multimodal models show 4 point WER differences from face pairs","Same audio different faces lead to 4 WER point transcription gaps","Face gender and ethnicity create 4 point WER drops in speech AI"]},"model":"grok-4.3","cost_usd":0.004611,"raw_usage":{"total_tokens":2258,"prompt_tokens":612,"num_sources_used":0,"completion_tokens":76,"cost_in_usd_ticks":46112000,"prompt_tokens_details":{"text_tokens":612,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1570,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":612,"tokens_out":76,"duration_ms":13099,"temperature":1.0,"reasoning_tokens":1570,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T07:42:56.426855+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-running the exact audio tracks with faces generated under fully controlled and identical conditions and checking whether the word error rate gaps of up to 4.05 points remain.","supporting_citations":[],"review_version":1}