{"id":"0ef9b779-288b-4365-98c8-49fb8970ead5","arxiv_id":"2412.12606","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The MDI benchmark evaluates large multimodal models on age-stratified, real-world multiple-choice questions and finds GPT-4o leading at about 79 percent average accuracy.","lead":"This paper introduces the MDI benchmark, a new test of large multimodal AI models built from more than 500 real-world images and 1,298 multiple-choice questions. It finds that the best model tested, GPT-4o, scores about 79 percent, showing that everyday visual question answering still has room to improve.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Age-stratified accuracy gaps are confounded with question difficulty, and the evaluation prompt omits the age tag, so the benchmark does not yet demonstrate that it measures age-group personalization.","rationale":"The reader's weakest_assumption is close but under-specified. The deeper problem is not only that volunteer-authored questions are a proxy; the evaluation protocol never conditions on age, so even a perfect proxy would not measure personalization. Section 4.5's difficulty-ranking interpretation is a direct admission that age bins differ in item difficulty, making the age-stratified comparisons uninterpretable as capability gaps. I agree with the reader's CONDITIONAL verdict direction, but for a sharper reason; the paper should either add age/user-context to the evaluation, provide a human difficulty calibration and chance-level baseline, or drop the personalization/age-alignment framing and present MDI as a real-world multimodal VQA benchmark. The dataset and code release are real assets, so a conditional acceptance with required revisions is appropriate rather than outright rejection.","tokens_in":19324,"tokens_out":5969,"duration_ms":57848,"concrete_test":"Recruit independent raters from the three age brackets (not the original volunteer authors) to answer a stratified sample of 50-100 MDI questions per age group under the same four-option format, and record human accuracy and per-question difficulty. If the human difficulty ranking reproduces the model ranking (middle-aged hardest, young easiest) or correlates with model accuracy per question, the age-stratified gaps are explained by item difficulty rather than by model personalization; if humans show no such ordering, the authors would still need to add age/user context to the prompt to substantiate the personalization claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that MDI-Benchmark measures whether LMMs align with the diverse needs of different age groups. The most load-bearing gap is that the age dimension is never actually manipulated at inference time. Section 3.2 assigns each volunteer-authored question a `[Level]-[Age]-[Scenario]` tag, but the evaluation prompt in Table 4 contains only `Question` and `Option`; no age, user profile, or 'answer for an elderly user' instruction is given to the model. The differences in Table 3 are therefore differences between question sets written by different volunteer groups, not measurements of a model's ability to adapt to users of different ages. This reading is confirmed by the paper's own analysis in Section 4.5, which interprets the summed scores to conclude 'actual difficulty order of questions across age levels: middle-aged > old > young' and explains that middle-aged questions 'require greater logical reasoning and background knowledge.' That is an admission that age-stratified accuracy is confounded with question difficulty. Without a human difficulty calibration, a chance-level baseline (25% for four options), or age-conditioned prompting, the 79% figure and the age-group gaps cannot support the abstract's claim of a comprehensive, objective, and accurate evaluation of real-world personalization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes MDI-Benchmark, a multimodal VQA dataset of 514 newly collected real-world images and 1,298 human-authored multiple-choice questions organized along three dimensions: six everyday-life scenarios (architecture, education, housework, social services, sports, transport), two question-complexity levels (perceptual vs. reasoning), and three author age groups (young, middle-aged, old). The authors evaluate 14 closed- and open-source LMMs and report that GPT-4o achieves the highest final score (78.46; 79.74 on the age dimension), with generally better performance on Level-1 perceptual questions and on education/architecture scenarios. The main contribution claimed is that the benchmark provides a \"comprehensive, objective and accurate\" evaluation of whether LMMs align with the diverse needs of different age groups in real-world scenarios.","tokens_in":19584,"tokens_out":9073,"duration_ms":72818,"significance":"The dataset itself is a concrete artifact: the images are new, the question/answer pipeline involves three age-stratified volunteer groups, cross-validation by all three groups, and expert screening, and the data and evaluation code are publicly released. Evaluating 14 models is a reasonable survey of the state of the art. If the age-stratified scores were shown to measure age-group alignment, the benchmark would fill a real gap. However, the central construct validity is questionable: the age tag reflects who wrote the question, not a condition the model is asked to satisfy, and the authors' own difficulty interpretation (Section 4.5) indicates that the age gaps are at least partly question-difficulty gaps. The paper also provides no human baseline or uncertainty estimates. With these fixed or the claims appropriately reframed, the benchmark's descriptive value would stand.","major_comments":[{"comment":"The evaluation never manipulates the age dimension. Questions are tagged as old/mid/young according to the age of the volunteer who authored them (Section 3.2), but the prompt template in Table 4 contains only the Question and Option fields; no age, user profile, or 'answer for an elderly user' instruction is given to the model. Consequently, the age-stratified accuracies in Table 3 compare question sets written by different volunteer groups, not the model's ability to adapt to users of different ages. The paper's own analysis in Section 4.5 confirms this: it interprets the summed scores as the 'actual difficulty order of questions across age levels: middle-aged > old > young' and attributes it to middle-aged questions requiring 'greater logical reasoning and background knowledge.' This is an admission that age-stratified accuracy is confounded with question difficulty. Without age-conditioned prompting, difficulty calibration, or a human baseline, the abstract's claim that MDI-Benchmark evaluates age-group personalization is not supported. I recommend either (a) re-running evaluation with age-contextualized prompts (e.g., presenting the user's age and asking the model to answer accordingly), or (b) reframing the contributions as a benchmark of age-stratified question performance rather than age alignment.","section":"§3.2, Table 4, §4.5"},{"comment":"All reported accuracies are point estimates without confidence intervals, significance tests, or a human baseline. From Table 1, the per-scenario age-group cells contain roughly 65–80 questions, and the aggregated age-stratified cells contain 426–436 questions; for such sample sizes, differences of a few percentage points are within the expected sampling error (a 5% gap at n=80 has a standard error of roughly 3–4%). The paper, however, makes fine-grained comparative claims, such as GPT-4o having 'smaller performance gaps across all three age-related categories' (Section 4.5) and a '13-point advantage' over the best open-source model, without any indication of variance. Moreover, no human accuracy is reported, so the statement that 79% 'indicat[es] that existing LMMs still have considerable room for improvement' has no absolute reference point. I ask the authors to add confidence intervals (or bootstrap estimates), a human-baseline experiment, and/or explicit disclaimers about noise in the fine-grained comparisons.","section":"§4.1–§4.5, Tables 2, 3, 6"}],"minor_comments":[{"comment":"The 'Total' row lists the number of images as 86, but the column sums to 514; the correct total should be 514, matching the text '514 images'.","section":"Table 1"},{"comment":"The sentence 'we input the scenario dimension information into open-source models (e.g., GPT-4o, Gemini 1.5 Pro) and closed-source models (e.g., LLaVA-NeXT, MiniCPM)' has the open/closed labels reversed: GPT-4o and Gemini 1.5 Pro are closed-source, while LLaVA-NeXT and MiniCPM are open-source.","section":"§3.2"},{"comment":"The model name is inconsistent: Section 4.1 and the leaderboard list 'LLaVA-NeXT-72B' (also spelled 'LLaV A-NeXT-72B'), while the model list in Section 4.1 says 'LLaV A-NeXT-70B' and Table 5 says 'LLaVA-NeXT-72B'; please unify the name and confirm the correct parameter count.","section":"§4.1 / Table 5"},{"comment":"The strict output-format requirement means that any response not matching the template is counted as incorrect. The paper does not report the fraction of responses that violated the format; if non-negligible, the accuracy scores conflate instruction-following with content correctness. Please report format-failure rates and consider a more tolerant answer-matching procedure.","section":"Table 4 / §4.1"},{"comment":"The phrase 'age-related tasks' in the abstract overstates what is measured. The benchmark contains questions written by (and tagged for) different age groups, but the models are not asked to perform any age-specific behavior. Please adjust the wording to avoid implying that the models were prompted with age information.","section":"Abstract / §4.5"},{"comment":"The claim of being 'the first to propose a multi-modal benchmark for real-world personalization' is not fully substantiated; the related-work section cites several personalized-LLM efforts but does not compare directly with any existing multimodal benchmarks that also incorporate user context or demographic tags. Please add a clearer comparison to establish novelty.","section":"Related Work / §2.3"}],"recommendation":"major_revision","confidential_remarks":"I had moderate confidence in the reader's assessment; after reading the paper in full, the age-inference concern is confirmed by the authors' own Section 4.5 interpretation. The dataset has real value and the issue is fixable (either by re-running with age-contextualized prompts or by reframing the contributions). I therefore recommend major revision rather than rejection. Also note the open/closed-source label swap in Section 3.2 and the incorrect image total in Table 1; these are straightforward to correct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the MDI-Benchmark is a real, new dataset—514 new images, 1298 multiple-choice questions covering six everyday scenarios, with question authors from three age bands. That resource is worth having. The paper's advertised finding about age-group alignment, though, is not supported by the experiment as run. The evaluation prompt (Table 4) gives models only the question and options; no age tag or user profile. So the reported gaps between 'old', 'middle-aged', and 'young' question sets are differences in question content written by different volunteer groups—not measurements of a model adapting to a user's age. The paper's own Section 4.5 actually says middle-aged questions require more logic and background knowledge, which is exactly a difficulty confound. Without a human difficulty calibration or a chance baseline, the 79% figure and the age gaps don't do the work the abstract claims.\n\nWhat's genuinely good: the data collection is careful by benchmark standards—volunteers paid, cross-validation of image categories across three groups, expert screening of questions, some balance across scenario and level. The image set is new and not from existing benchmarks, which addresses contamination concerns. The two-level complexity split and the scenario taxonomy are standard but fine. This is a legitimate, if incremental, resource for multimodal evaluation.\n\nSoft spots, in rough order: (1) The central construct—age-tagged question authorship as needs alignment—is thin. A few dozen volunteers per age band are standing in for an entire demographic. (2) No error bars or human baseline, so we can't tell how meaningful the model gaps are. (3) Table 1 has an internal inconsistency: the total image row says 86, while the text says 514 and the rows sum to 514. That looks like a copy-paste error but needs fixing. (4) The 'first to propose' framing overstates novelty; prior work has real-world and some personalization-oriented multimodal benchmarks, and the paper doesn't engage with them.\n\nBottom line: the dataset can be useful for evaluating LMMs on everyday visual QA, but the personalization/age-alignment interpretation needs a redesign—either condition on age in the prompt, or calibrate question difficulty and show age-specific effects beyond difficulty. I'd send it to review, but I'd tell the authors the age claim needs major revision before it's supportable.","headline":"A useful new multimodal QA resource, but the age-personalization claim outruns the evaluation design; the model is never given the user's age.","tokens_in":20120,"tokens_out":2514,"would_cite":false,"duration_ms":22086,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a new benchmark of over 500 real-world images with age-tagged, two-level questions can reveal whether large multimodal models truly serve different age groups, and that current best models like GPT-4o still cap at…","keywords":["large multimodal models","personalization benchmark","age-stratified evaluation","real-world visual question answering","question complexity","multimodal reasoning","human-aligned AI"],"falsifier":"Ask fresh, independent panels from each of the three age groups, people who did not write MDI questions, to judge whether each age-tagged question reflects their own everyday concerns, then rerun the benchmark on questions that pass independent panel agreement; if the panels reject many questions or the model ranking flips, the claim that the benchmark measures age-group alignment collapses.","tokens_in":19171,"feed_emoji":"🖼️","tokens_out":4400,"duration_ms":38140,"temperature":0.7,"pith_summary":"The paper introduces MDI-Benchmark, a collection of 514 newly gathered real-world images and 1,298 human-authored questions, to test whether large multimodal models can meet people's actual needs in everyday life. Its central claim is that existing benchmarks measure technical skills but miss two things: whether a model understands a real situation at all, and whether it can serve different age groups with different concerns. To capture those, every image carries a simple question and a harder reasoning question, and the questions are tagged as coming from young, middle-aged, or older people. The paper reports that even the strongest model tested, GPT-4o, reaches only 79 percent accuracy on age-related tasks, which it reads as evidence that current models still fall well short of real-world personalization. A sympathetic reader would take the paper's contribution to be a new evaluation instrument: if its age-stratified questions are valid, they expose capability gaps that technical benchmarks miss.","feed_headline":"New image benchmark grades AI on age-group needs","feed_subtitle":"Six everyday scenes, two difficulty levels, three age groups: even the best model scores 79 percent.","key_machinery":"The central object is the MDI-Benchmark dataset: 514 new images from six life scenarios (architecture, education, housework, social service, sport, transport), each paired with two levels of questions, Level 1 for perceptual extraction (object detection, OCR, color and position recognition) and Level 2 for analysis and reasoning beyond image content. Each question is tagged by the volunteer's age group, defining young as 10-25, middle-aged as 35-50, and old as 60-75. The scoring metric combines the two levels with equal weight, $Score_{final} = 0.5 \\cdot Score_{L1} + 0.5 \\cdot Score_{L2}$, converting the abstract idea of personalization into accuracy numbers per scenario, per complexity, and per age bracket.","core_discovery":"On the paper's own terms, MDI-Benchmark provides a comprehensive, objective, and accurate evaluation of whether large multimodal models align with diverse human needs in real-world scenarios, and the age-stratified design reveals a difficulty ordering of middle-aged > old > young across all fourteen tested models. The headline evidence is that GPT-4o scores 78.46 overall, with 79.74 average accuracy across age groups, beating the best open-source model by about 13 points and the lowest closed-source model by about 35 points; yet no model is close to saturation. The benchmark further documents that every model loses accuracy when moving from Level 1 perceptual questions to Level 2 reasoning questions, with the sharpest drops in sport and transport, and interprets this as a sign that current training data under-serves everyday-life domains.","pith_inferences":["Beyond the paper: the same volunteer-question protocol could be reused to stratify by occupation, culture, disability status, or other group dimensions, turning MDI's age axis into a template for measuring personalization more broadly.","Beyond the paper: a testable prediction follows, that models fine-tuned on age-specific preference data should improve mostly on the age bracket they were tuned for, not uniformly across all groups.","Beyond the paper: because the paper's difficulty ordering is computed from model scores rather than human judgments, an independent human-pilot rating of the same questions could reveal whether middle-aged questions are intrinsically harder or merely harder for current models."],"forward_implications":["If MDI-Benchmark measures what it claims, then age-stratified question sets become a standard check for real-world alignment alongside existing technical benchmarks.","The 79 percent ceiling for GPT-4o implies that frontier models still leave roughly a fifth of everyday age-group questions unanswered correctly, and open-source models leave far more.","The consistent Level-1 to Level-2 accuracy drop, such as GPT-4o's education score falling from 94.12 to 70.59, implies that reasoning beyond image content, not basic perception, is the binding constraint for real-world usefulness.","The observed difficulty ordering of middle-aged > old > young implies that training and evaluation should treat age-group coverage as a balance objective rather than a single average."],"supporting_citations":[{"why":"The VQA benchmark that MDI contrasts with, establishing the standard image-question evaluation format that MDI extends with complexity and age dimensions.","marker":"(Goyal et al., 2017)"},{"why":"MMBench, a fine-grained multimodal benchmark MDI cites as measuring broad capabilities but not group-specific real-world needs.","marker":"(Liu et al., 2023)"},{"why":"MM-Vet, representing the integrated-capability evaluation approach that MDI positions itself against.","marker":"(Yu et al., 2023b)"},{"why":"MME, a widely used multimodal evaluation benchmark that MDI says neglects diverse individual needs.","marker":"(Fu et al., 2024a)"},{"why":"MMMU, a multi-discipline benchmark that MDI cites as assessing expert-level understanding rather than everyday age-specific alignment.","marker":"(Yue et al., 2023)"},{"why":"GPT-4o, the strongest model evaluated; its reported 79 percent age-task accuracy is the paper's headline empirical result.","marker":"(OpenAI, 2024)"}],"fun_headline_variants":["Even GPT-4o scores just 79% on age-aware benchmark","New benchmark: AI accuracy drops on age-specific real-world questions","Age-stratified AI test reveals big gap in real-world understanding","Multimodal AI benchmark: all models fall short on age-group tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that questions written by a small group of volunteers in each age bracket faithfully stand in for the needs and perspectives of that whole age group, and that a model picking the single correct multiple-choice answer is the same as meeting those needs.","fun_headline_variants_meta":{"raw":{"variants":["Even GPT-4o scores just 79% on age-aware benchmark","New benchmark: AI accuracy drops on age-specific real-world questions","Age-stratified AI test reveals big gap in real-world understanding","Multimodal AI benchmark: all models fall short on age-group tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000791,"raw_usage":{"total_tokens":3504,"prompt_tokens":983,"completion_tokens":2521,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":2445}},"tokens_in":599,"tokens_out":2521,"duration_ms":18153,"temperature":1.0,"reasoning_tokens":2445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:54:47.673860+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask fresh, independent panels from each of the three age groups, people who did not write MDI questions, to judge whether each age-tagged question reflects their own everyday concerns, then rerun the benchmark on questions that pass independent panel agreement; if the panels reject many questions or the model ranking flips, the claim that the benchmark measures age-group alignment collapses.","supporting_citations":[],"review_version":1}