{"id":"0ae87c4d-b796-4a6c-8a10-0fbfdac0cbee","arxiv_id":"2607.01245","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":8.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"OCB introduces File Fidelity Q&A and Domain Q&A tracks for LLM evaluation on native office files, reporting ~59.3% on domain reasoning for top systems with releases of data and tooling.","lead":"The paper introduces the Office Comprehension Bench (OCB), the first public benchmark for jointly evaluating LLMs on native Word, Excel, and PowerPoint file comprehension. Smart generalists might read it to understand current limits of frontier models on professional document tasks and to see how new evaluation methods measure multi-step reasoning in office software.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"LLM judge ensemble reliability for atomic claims is the load-bearing assumption","rationale":"The reader's weakest_assumption directly identifies the same judge-reliability risk that supports the central performance claim. Full-text access does not remove the need for external validation of the automated scoring step, so the UNVERDICTED stance remains appropriate.","tokens_in":1679,"tokens_out":262,"duration_ms":10960,"concrete_test":"Select 200 atomic claims across 20 Domain Q&A items; obtain independent human expert binary labels on the same model responses; compute Cohen's kappa and per-claim disagreement rate between humans and the LLM ensemble. If kappa < 0.75 or directional bias exceeds 10%, the reported scores are unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The 59.3% Domain Q&A figure is produced by decomposing reference answers into atomic binary claims and scoring model outputs via an ensemble of LLM judges. This pipeline is load-bearing: any systematic bias, low inter-judge agreement, or correlation between judge and evaluated model families would directly alter the headline performance numbers and the claim that deeper reasoning within a tier yields no gain. The abstract provides no human calibration, agreement statistics, or ablation of judge choice.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces the Office Comprehension Benchmark (OCB), the first public benchmark for LLM evaluation on native Word, Excel, and PowerPoint files (.docx, .xlsx, .pptx). It defines two tracks—File Fidelity Q&A for structural and visual perception of document elements and Domain Q&A for expert-level multi-step reasoning across 12 professional domains—where reference answers are decomposed into atomic binary claims scored independently by an ensemble of LLM judges. The central empirical claim is that even the strongest frontier system reaches only 59.3% on Domain Q&A in default mode, with negligible gains from increased thinking depth within a tier and only modest gains from higher product tiers. The dataset, evaluation tooling, judge prompt, and leaderboard are released.","tokens_in":1746,"tokens_out":355,"duration_ms":15277,"significance":"If the judge pipeline proves reliable, OCB would be a valuable addition to the field by exposing concrete limitations in current LLMs' handling of complex, multi-document office reasoning tasks. The explicit release of the full dataset, tooling, judge prompt, and public leaderboard is a clear strength that supports reproducibility and follow-on work.","major_comments":[{"comment":"The 59.3% Domain Q&A result and the claims about reasoning depth versus tier effects are produced by decomposing references into atomic claims and scoring via an LLM judge ensemble. No human calibration data, inter-judge agreement statistics, or ablation on judge model choice or prompt variants are reported, making the headline numbers and downstream conclusions load-bearing on an unvalidated component of the evaluation pipeline.","section":"Abstract / Evaluation pipeline description"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for highlighting a key aspect of our evaluation pipeline. We address the major comment below and will incorporate the requested validations in the revised manuscript.","responses":[{"response":"We agree that the LLM judge ensemble is a load-bearing component of the reported results and that the submitted manuscript lacks explicit validation of this pipeline. In the revised version we will add a dedicated subsection under Evaluation Methodology that reports: (1) human calibration results on a stratified sample of 250 atomic claims (with inter-annotator agreement and agreement with the judge ensemble), (2) pairwise and ensemble-level agreement statistics (percentage agreement and Fleiss’ kappa), and (3) ablation tables comparing judge performance across model families and prompt variants. These additions will directly support the 59.3 % Domain Q&A figure and the reasoning-depth versus tier conclusions. We have already begun the human annotation effort and expect to complete it within the revision window.","revision_made":"yes","referee_comment":"[Abstract / Evaluation pipeline description] The 59.3% Domain Q&A result and the claims about reasoning depth versus tier effects are produced by decomposing references into atomic claims and scoring via an LLM judge ensemble. No human calibration data, inter-judge agreement statistics, or ablation on judge model choice or prompt variants are reported, making the headline numbers and downstream conclusions load-bearing on an unvalidated component of the evaluation pipeline."}],"tokens_in":1316,"tokens_out":308,"duration_ms":13169,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"Colleague,\n\nThe main thing to know is that this is the first public benchmark covering native .docx, .xlsx, and .pptx files together, with separate tracks for file fidelity and domain reasoning over real industry documents. They also release the dataset, tooling, judge prompts, and a leaderboard, which is the practical part that matters.\n\nThe atomic claim decomposition for scoring is a clean move that reduces some subjectivity compared to open-ended grading. Releasing everything makes it possible for others to run their own checks or extend the work.\n\nThe soft spot is the scoring pipeline. The 59.3% Domain Q&A result and the claim that extra thinking depth within a tier adds little come from an LLM judge ensemble. The abstract gives no human calibration data, inter-judge agreement stats, or ablation on judge choice. If the judges share biases with the evaluated models or favor certain response styles, both the absolute number and the tier comparison could shift. That assumption is load-bearing, and the stress-test note is right to flag it.\n\nThis paper is for groups building or testing LLMs for office productivity tasks. Anyone running evals in that area will get value from the data release itself.\n\nIt deserves peer review so referees can examine the dataset construction and judge reliability in detail. The core contribution is solid enough to warrant that step.","headline":"OCB is a genuinely new benchmark for native office files with useful releases, but the LLM judge scoring lacks the validation needed to trust the headline numbers.","tokens_in":2330,"tokens_out":352,"would_cite":false,"duration_ms":14409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Even the strongest frontier LLMs reach only 59.3 percent on expert-level reasoning over real office documents.","keywords":["Office Comprehension Benchmark","LLM evaluation","document understanding","Word Excel PowerPoint","Domain Q&A","File Fidelity","atomic claims","LLM judges"],"falsifier":"A new model scoring well above 59.3 percent on Domain Q&A under the same atomic-claim evaluation, or human raters showing large systematic disagreement with the LLM judge ensemble on the same set of responses.","tokens_in":2587,"feed_emoji":"📄","tokens_out":631,"duration_ms":20732,"temperature":0.7,"pith_summary":"The paper introduces the Office Comprehension Bench to test LLMs jointly on Word, Excel, and PowerPoint files in their native formats. One track measures structural and visual perception of elements such as tables, charts, formulas, and speaker notes. The second track measures multi-step reasoning and synthesis across industry documents in twelve professional domains. Reference answers are broken into atomic claims that an ensemble of LLM judges scores independently. Results show top models top out near 59 percent on the reasoning track, with little gain from deeper thinking in the same tier.","feed_headline":"Frontier LLMs reach only 59% on office document reasoning","feed_subtitle":"New benchmark evaluates native file perception and expert synthesis across 12 professional domains in Word, Excel and PowerPoint","key_machinery":"Office Comprehension Bench with its two tracks and atomic-claim scoring by an LLM-judge ensemble","core_discovery":"The Office Comprehension Bench is the first public benchmark for LLM comprehension of Word, Excel, and PowerPoint over native file formats. It consists of File Fidelity Q&A for structural and visual perception of office artifacts and Domain Q&A for expert-level reasoning grounded in real-world documents across 12 professional domains. Each reference answer is decomposed into atomic, binary-gradable claims scored independently by an ensemble of LLM judges. Even the strongest frontier system reaches only about 59.3 percent on Domain Q&A.","pith_inferences":["Specialized parsing or embedding layers for binary office formats may be needed to close the remaining gap.","The atomic-claim decomposition method could transfer to benchmarks for other structured document types.","The modest tier gains suggest that raw scale or inference budget alone will not solve the comprehension limits observed."],"forward_implications":["Increasing thinking depth within a model tier does not move Domain Q&A performance materially.","Moving to a higher product tier yields only modest gains on domain reasoning.","Current systems still have large gaps on multi-step synthesis across native office artifacts.","The public dataset, tooling, and leaderboard enable standardized tracking of progress on office file comprehension."],"fun_headline_variants":["LLMs score 59% on Office Comprehension Bench","OCB tests LLM perception of Word Excel PowerPoint files","Domain Q&A shows 59% for strongest systems","First public benchmark for LLM office file comprehension"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The ensemble of LLM judges produces reliable, unbiased scores when evaluating model responses against independently scored atomic claims derived from reference answers.","fun_headline_variants_meta":{"raw":{"variants":["LLMs score 59% on Office Comprehension Bench","OCB tests LLM perception of Word Excel PowerPoint files","Domain Q&A shows 59% for strongest systems","First public benchmark for LLM office file comprehension"]},"model":"grok-4.3","cost_usd":0.005554,"raw_usage":{"total_tokens":2652,"prompt_tokens":646,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":55537000,"prompt_tokens_details":{"text_tokens":646,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1946,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":646,"tokens_out":60,"duration_ms":12001,"temperature":1.0,"reasoning_tokens":1946,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-04T00:19:08.061679+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new model scoring well above 59.3 percent on Domain Q&A under the same atomic-claim evaluation, or human raters showing large systematic disagreement with the LLM judge ensemble on the same set of responses.","supporting_citations":[],"review_version":1}