{"id":"80ef0a58-db06-4e3a-a556-b10b72e42fe0","arxiv_id":"2601.14569","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A three-part benchmark for multimodal LLMs separates question-answering about social videos from holistic and directed scene description, and shows AI judges can approximate human quality ratings.","lead":"This paper introduces SOCIALCAPTION, a scoring framework that tests multimodal AI models on three social-understanding tasks: answering questions about interactions, describing whole scenes, and pulling out details relevant to a specific question. It reports that AI judges can score AI answers about social scenes roughly as reliably as humans, and that smaller open models sometimes match much larger closed models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Judge-alignment claim rests on lenient binary F1 at threshold >=3 compared against averaged human scores; stricter thresholds or single-annotator comparisons could erode 'strong alignment'.","rationale":"The reader's weakest assumption is exactly the load-bearing point: the proof-of-concept for scalable MLLM evaluation depends on the metric used to measure judge-human alignment. Binary F1 at a lenient threshold, with human scores averaged across annotators, can overstate agreement. The table shows only InternVL3 judges exceed human-human F1, while Gemini-2.5-Pro—a key closed-source judge—underperforms human-human on several sub-dimensions. If the alignment metric is changed to a stricter or ordinal one, the 'strong alignment' claim may need to be sharply qualified. This does not invalidate the framework's contribution (the dimensions are well-motivated and the qualitative analysis is informative), but it does mean the paper's headline scalability claim is conditionally supported. The reader already marked this CONDITIONAL; our analysis reinforces that judgment rather than overturning it. We therefore recommend no change to the verdict, while urging the authors to release annotations and perform the sensitivity analysis.","tokens_in":26846,"tokens_out":5374,"duration_ms":63854,"concrete_test":"Using the authors' released annotations (or a fresh human evaluation on 50+ videos), recompute human-judge alignment with ordinal metrics on raw Likert scores (quadratic weighted kappa or Spearman rho) and compare each judge against each individual annotator, not the averaged human score; also report binary F1 at thresholds >=4 and >=3. If InternVL3-8B/78B no longer exceed human-human agreement, or Gemini-2.5-Pro drops below a pre-specified acceptable level (e.g., kappa<0.6), the 'strong alignment' conclusion should be weakened. This single check would resolve whether the scalability claim holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that MLLM judges have 'strong alignment with human scores' (Abstract) is supported primarily by Table 4, which reports binary F1 computed after collapsing 5-point Likert scores at a lenient threshold (>=3 = high, Appendix E). Two methodological choices inflate this alignment. First, human-judge agreement is computed against the mean of two annotators, whereas human-human agreement compares single annotators; averaging reduces noise, so InternVL3 judges exceeding human-human F1 (92.17/92.65 vs 88.85; 92.36/92.58 vs 86.69) may reflect this artifact rather than superior fidelity. Second, the lenient threshold makes most labels 'high'; a judge that simply assigns high scores will trivially achieve high F1. Indeed, Appendix G notes all judges inflate absolute scores, and Gemini-2.5-Pro—the strongest SI model—has agreement below human-human on multiple HSA/DSA sub-dimensions (Table 2: e.g., Individuals 75.09 vs 84.24 HSA; Relevant Interactions 75.37 vs 87.04 DSA). The abstract's unqualified 'strong alignment' is therefore not supported across judges or dimensions. Additionally, the 20-video sample (Appendix D) gives wide confidence intervals; the power analysis (dz>=0.66) addresses model differences, not agreement precision. Since the paper's scalability conclusion depends on judges reliably reproducing human ratings, this metric choice is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SOCIALCAPTION, a three-dimensional evaluation framework for multimodal social understanding: Social Inference (multiple-choice QA accuracy), Holistic Social Analysis, and Directed Social Analysis (open-ended generation scored with six Likert sub-dimensions). Using SOCIAL-IQ 2.0 validation videos and a range of open- and closed-source MLLMs, it reports SI accuracy with and without transcripts, human and MLLM-judge HSA/DSA scores, analyses of model scale/architecture/spoken-context effects, and judge-alignment results. The central positive claim is that MLLM judges—especially InternVL3-8B and InternVL3-78B—align strongly with human HSA/DSA ratings, thereby enabling scalable automated evaluation of social understanding.","tokens_in":27228,"tokens_out":7801,"duration_ms":79236,"significance":"If the judge-alignment result holds, the paper would provide a useful, theory-grounded complement to QA-only social benchmarks. The framework's HSA/DSA sub-dimensions are anchored in an external taxonomy (APRACE), the prompts are given in full, annotator agreement is reported transparently, and the mismatch test is a reasonable control against indiscriminate high scoring. The central claim is falsifiable. However, the evidence currently does not support the abstract's unqualified 'strong alignment' statement: the agreement metric is a binary F1 at a single lenient threshold, the human-judge comparison may not be matched to the human-human baseline, and the HSA/DSA evidence base is 20 videos with judge models selected from the top-SI performers. These issues are load-bearing for the scalability conclusion.","major_comments":[{"comment":"The headline 'strong alignment with human scores' rests on binary F1 at a single lenient threshold (Likert ≥3 = high). Appendix G shows all MLLM judges inflate absolute scores, so most real responses are 'high'; a high-labeling judge can achieve high F1 without human-like discrimination. Table 4 itself shows H-G F1 (86.10/81.99) below H-H (88.85/86.69), and Table 2 shows Gemini-2.5-Pro below H-H in 8 of 12 sub-dimensions (e.g., HSA Individuals 75.09 vs 84.24; DSA Relevant Interactions 75.37 vs 87.04; DSA Answer Detail 71.76 vs 85.06). Please qualify the claim, vary the threshold (e.g., ≥4, =5, ordinal/rank agreement), and report a matched single-annotator baseline rather than an averaged human score (Appendix G states human scores are averaged across annotators).","section":"Section 4.3 / Appendix E / Table 4"},{"comment":"All HSA/DSA and judge-agreement conclusions rest on 20 videos and 10 models (top-7 SI performers plus 3 closed-source models). The power analysis (dz≥0.66) addresses paired model-mean differences, not precision of per-subdimension F1; with 20 videos the subdimension differences in Table 2 have wide confidence intervals. Models with low generation quality are omitted from HSA/DSA (Table 1 footnote), which can bias agreement upward. Please provide bootstrap CIs, analyze or at least characterize excluded models, and release outputs/annotations so the agreement numbers are independently checkable.","section":"Section 3.2.4 / Appendix D / Table 1"},{"comment":"The three judges were selected as the highest-SI models (Gemini-2.5-Pro, InternVL3-8B, InternVL3-78B). The scalability conclusion ('open-source models such as InternVL3 can serve as evaluators') is therefore based on a favorable subset. Report judge-alignment for additional MLLMs (e.g., Qwen2.5-VL or GPT-4o) or justify why SI ranking is a valid criterion for judge selection; absent that, the claim that 'MLLM judges' align with humans is overgeneralized.","section":"Section 4.3 / Appendix G"},{"comment":"The paper's own ranking results reveal judge-human disagreement. Human annotators rank Gemini-1.5-Pro highest in HSA (24.95), while all three MLLM judges rank Gemini-2.5-Pro highest (29.95/30.00/29.90). Binary F1 at a per-item threshold cannot capture this kind of rank-order disagreement. The mismatch test (Appendix H) shows judges penalize unrelated pairs, but it does not show that their relative ordering of valid responses is human-like. Report rank correlation (Spearman/Kendall) between human and judge totals/sub-dimensions.","section":"Section 4.3 / Table 7"}],"minor_comments":[{"comment":"Typographical and formatting issues: 'Y oussouf' in the author list, 'V olume' in the Hoppler reference, 'LLaV A-NeXT' spacing, 'stucture' in Figure 10, and 'adher .' in Table 8 headers.","section":"General"},{"comment":"The ↑/↓ arrows in the HSA/DSA columns are ambiguous: each model row has four judge columns (H, G, I8B, I78B), and it is unclear whether arrows mark the highest/lowest within each column, within each category, or across the whole row. Please define this in the caption.","section":"Table 1"},{"comment":"The footnote says '-' indicates models not included 'due to low generation quality,' while Section 3.2.4 says the top 7 standard-scale models were selected by SI performance. Clarify whether exclusion was based on SI ranking, generation quality, or both.","section":"Table 1 / Section 3.2.4"},{"comment":"The text says the 20 videos were 'randomly selected to ensure representation of videos across contexts.' Random selection does not by itself ensure representation; if a stratified procedure was used, describe it. Figure 11 validates SI representativeness but not HSA/DSA representativeness.","section":"Appendix D"},{"comment":"No data or code availability statement is included. Given that the judge-alignment and HSA/DSA claims rely on human annotations and model outputs, releasing these artifacts (or providing a clear reason not to) would substantially strengthen the paper.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the framework is well motivated, but the headline judge-alignment conclusion is currently overstated. The main issues—binary threshold choice, unclear/unmatched human baseline, 20-video evidence base, and selected judge set—are fixable within the manuscript's scope, but they need to be addressed before the scalability claim can be accepted. I would be willing to reconsider after the authors add sensitivity analyses, rank-based agreement, and a clearer statement of what evidence supports 'strong alignment.'"},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a real measurement contribution, not a flashy model paper. It gives the MLLM evaluation community a structured way to separate QA accuracy (SI) from generative description abilities (HSA and DSA), and it grounds those new dimensions in an external psychology taxonomy (APRACE) rather than inventing categories ad hoc. That is genuinely new relative to SOCIAL-IQ, SOCIALGENOME, and SIV-BENCH, which mostly stay inside the QA-plus-reasoning-trace box. The paper also does the routine work properly: prompts are in the appendix, inter-annotator agreement is reported, there is a power analysis, a mismatch test for judge bias, and self-preference effects are acknowledged. Credit where it is due.\n\nThe central empirical claim—that strong SI does not automatically translate to strong HSA or DSA—holds up. Qwen2.5-Omni is a nice counterexample: good at multiple-choice social inference, weak at generative description. The model-design analysis (LongVA vs Qwen2-VL, backbones, transcription effects) is useful and mostly convincing.\n\nNow the soft spots, in proportion. The headline claim in the abstract that \"MLLM judges have strong alignment with human scores\" is too strong as written. The stress-test note lands: human-judge agreement is computed against the mean of two annotators, while human-human agreement compares single annotators. That averaging artifact likely explains why InternVL3 judges appear to beat human-human F1. Add the lenient binary threshold (Likert >=3 = high) and the fact that all judges inflate absolute scores, and the InternVL3 numbers look flattered. Gemini-2.5-Pro, the strongest SI model, falls below human-human agreement on several HSA/DSA sub-dimensions. So the scalable-judge conclusion is really a proof-of-concept for InternVL3, not a general result. That is still worth reporting, but the abstract should say so.\n\nThe 20-video sample for HSA/DSA is small. The power analysis addresses model-to-model differences, not agreement precision, so the judge-alignment estimates come with wide error bars. And no code, outputs, annotations, or the exact 20-video subset are released, which blocks independent audit. These are fixable in revision: release artifacts, add a threshold-sensitivity analysis, and compare judge alignment against a single-annotator baseline.\n\nWho gets value: anyone building benchmarks for social AI or using MLLMs as evaluators. The framework itself is the contribution, and it is solid enough to build on. I would send this to peer review, with a request to temper the abstract and run the sensitivity analyses. The core idea survives the corrections.","headline":"New multidimensional social-understanding evaluation for MLLMs that is worth engaging; the judge-alignment headline overstates what the data support.","tokens_in":27736,"tokens_out":1780,"would_cite":true,"duration_ms":22168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multimodal models' social understanding is three separate abilities—answering questions, describing whole scenes, and extracting relevant details—and accuracy on the first does not predict the other two.","keywords":["social understanding","multimodal large language models","evaluation framework","video question answering","holistic social analysis","directed social analysis","MLLM judges","APRACE taxonomy"],"falsifier":"Take a larger sample (100+ videos) and compute agreement on the original 5-point Likert scale (e.g., quadratic weighted kappa or average absolute difference) between each MLLM judge and human raters, then compare with human-human agreement. If InternVL3's strict ordinal agreement drops below human-human agreement, or if including lower-performing models as judges reverses the ranking between Gemini-2.5-Pro and InternVL3, the claim that open-source judges can scale human evaluation is not supported.","tokens_in":26723,"feed_emoji":"🧠","tokens_out":5575,"duration_ms":57041,"temperature":0.7,"pith_summary":"The paper is trying to establish that social understanding in multimodal large language models (MLLMs) is not a single ability and cannot be read off from multiple-choice question-answering accuracy. It introduces SOCIALCAPTION, an evaluation framework with three dimensions: Social Inference (choosing the correct answer about an interaction), Holistic Social Analysis (writing a comprehensive description of the scene), and Directed Social Analysis (writing a description containing only the information relevant to a given question). On one-minute real interaction videos, the authors find that a model can be strong at inference while weak at holistic or directed description; that adding spoken-context transcriptions improves inference for every model; and that small open-source models can match or exceed much larger closed-source models on the generative dimensions. They further report that open-source MLLM judges rate HSA and DSA generations in close binary-F1 agreement with human raters, which they offer as a proof-of-concept for scaling automated evaluation of social understanding.","feed_headline":"Social understanding in AI is three skills, not one score","feed_subtitle":"Models that ace social Q&A still miss the scene; open-source judges can rate the gap.","key_machinery":"The load-bearing instrument is the SOCIALCAPTION rubric: HSA and DSA outputs are each rated on six 1-5 Likert sub-dimensions—scene, individuals, topic/context, socio-emotional analysis, answer detail, and prompt adherence for HSA; relevant scene details, key individuals, relevant interactions, relevant context, plus the same two meta-criteria for DSA—giving a maximum score of 30. The sub-dimensions are grounded in the APRACE taxonomy of social interactions (actors, partners, relations, activities, context, evaluation). SI is measured separately by multiple-choice accuracy. Alignment between human and MLLM-judge ratings is computed by binarizing Likert scores at a threshold of ≥3 and reportin","core_discovery":"The central claim is that social understanding should be measured along three separate axes, and that the two generative axes reveal competencies that QA accuracy hides. In the paper's experiments, Gemini-2.5-Pro leads all models on Social Inference but receives lower human scores than Gemini-1.5-Pro on holistic socio-emotional description; Qwen2.5-Omni is competent on inference yet produces the worst HSA and DSA generations of any evaluated model; and InternVL3-8B, an 8-billion-parameter open model, matches or beats several closed-source models on Directed Social Analysis. The paper also claims that MLLM judges can stand in for human raters: InternVL3-8B and InternVL3-78B show binary-F1 agr","pith_inferences":["If the judge-alignment result is used to build reward models, the observed self-preference and absolute-score inflation are likely to distort preference learning; using relative rankings or calibrated thresholds may be necessary.","The alignment claim rests on only 20 videos and binary F1 at a threshold of 3; re-running on a larger sample with ordinal agreement metrics (e.g., weighted kappa on the 5-point scale) would test whether the proof-of-concept survives stricter measurement.","The HSA/DSA rubrics could be turned into training signal: a model fine-tuned to produce structured, rubric-scored social narratives might develop more usable social understanding than one trained solely on QA.","Because the human annotations are English-only and US-based, the framework's sub-dimensions may encode Western interaction norms; a cross-cultural stress test would show whether the dimensions are universal or culturally specific."],"forward_implications":["Evaluation of social understanding should treat QA accuracy, holistic description, and directed relevance as distinct report cards; a single QA score can mis-rank models.","Providing spoken context (transcriptions) is a reliable lever for improving social inference across all tested models, with gains up to 13 accuracy points.","Scale is not the main driver of generative social understanding: 7-8B open models can match or exceed much larger closed models on HSA and DSA.","MLLM judges—especially InternVL3 variants—can serve as scalable evaluators and potential reward models for social-understanding training, provided their score inflation and self-preference are accounted for.","Architectural choices such as temporal video modeling versus treating video as extended images, and low-latency streaming designs, materially change generative social understanding."],"fun_headline_variants":["AI social IQ: three skills, not one score","QA hides AI's social blind spots","Open 8B model beats giants at social scene reading","Machine judges rate AI social skills as well as humans","Social inference isn't social understanding"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that binary-F1 agreement at a hand-set threshold (Likert ≥3 = high quality) on a 20-video sample faithfully represents 'alignment with human scores'; if the threshold, the video sample, or the judge selection changed, the proof-of-concept for automated evaluation could weaken.","fun_headline_variants_meta":{"raw":{"variants":["AI social IQ: three skills, not one score","QA hides AI's social blind spots","Open 8B model beats giants at social scene reading","Machine judges rate AI social skills as well as humans","Social inference isn't social understanding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":998,"prompt_tokens":643,"completion_tokens":355,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":387,"completion_tokens_details":{"reasoning_tokens":285}},"tokens_in":387,"tokens_out":355,"duration_ms":4955,"temperature":1.0,"reasoning_tokens":285,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:07:29.156720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a larger sample (100+ videos) and compute agreement on the original 5-point Likert scale (e.g., quadratic weighted kappa or average absolute difference) between each MLLM judge and human raters, then compare with human-human agreement. If InternVL3's strict ordinal agreement drops below human-human agreement, or if including lower-performing models as judges reverses the ranking between Gemini-2.5-Pro and InternVL3, the claim that open-source judges can scale human evaluation is not supported.","supporting_citations":[],"review_version":1}