Pith. sign in

REVIEW 3 major objections 7 minor 53 references

Fluent deep-research reports often fail the evidence chain: models write well but miss grounded claims and answers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 00:03 UTC pith:JDLW6MFL

load-bearing objection Solid diagnostic benchmark: the report-vs-grounding gap is real and useful, but the headline 4–11% answer-gate rates partly measure match-to-one-gold-graph, not pure reasoning failure. the 3 major comments →

arxiv 2607.25151 v1 pith:JDLW6MFL submitted 2026-07-27 cs.IR

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research

classification cs.IR
keywords deep researchhierarchical evidence aggregationevidence graphmultimodal RAGcitation accuracyclaim verificationprogressive gatingtraceability evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Deep research systems are usually scored on the finished report—how fluent, complete, or citation-looking it is—while the path from raw evidence to intermediate claims to a final conclusion stays hidden. This paper argues that that black-box scoring cannot tell whether a system actually selected the right evidence, linked sources, and built supported claims. It introduces HiEviDR-Bench: 2,000 human-validated questions, each paired with an explicit hierarchical evidence graph spanning open-domain and academic settings, text-only and multimodal inputs. Evaluation breaks into five scored stages plus progressive gates that withhold later credit when earlier grounding fails. Across 16 multimodal models, report quality stays high while citation accuracy, claim construction, and answer correctness fall sharply, with only a small fraction of samples fully activating the graph through to the answer. The practical message is that surface polish is not a proxy for faithful multi-stage evidence aggregation, and the main breaks are early: identifying key evidence and forming intermediate claims.

Core claim

On HiEviDR-Bench, strong surface-level report quality does not imply grounded multi-stage reasoning. Across sixteen multimodal large language models under both RAG and deep-research setups, report scores remain relatively high while citation accuracy, intermediate claim construction, and answer correctness drop markedly; answer-gate pass rates range only from about 3.8% to 11.5%, and the dominant bottlenecks are evidence identification and claim construction rather than final fluency.

What carries the argument

The hierarchical evidence graph: a directed acyclic graph of evidence nodes, intermediate claim nodes, and one conclusion node, with support edges that make selection, cross-source linking, and aggregation inspectable. Scoring is decomposed into five dimensions (report, traceability, citation, claim, answer), and progressive gates zero out later stages unless the preceding grounding requirement is met.

Load-bearing premise

The annotated evidence graphs and the judge-driven gates are treated as the correct gold standard for what counts as faithful hierarchical aggregation, rather than one plausible scaffold among many.

What would settle it

If independent human raters, using the same reports without privileging the annotated graphs, systematically credit models that fail the paper’s citation/claim/answer gates—or if swapping the judge model reverses the ranking of systems on those gated dimensions while report scores stay stable—then the claimed gap between fluency and grounded aggregation would not hold as measured.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Benchmarking deep research must score intermediate evidence-to-claim links, not only final report quality or short answers.
  • Gains from iterative deep-research pipelines should be judged by better evidence filtering and claim support, not only by more retrieval rounds.
  • Systems can look strong on multimodal report rendering while still failing citation accuracy and claim verification under progressive gates.
  • Error localization will concentrate on key-evidence identification and intermediate claim construction before final answer writing.
  • Open-domain versus academic and text versus multimodal splits can expose where hierarchical aggregation, not retrieval volume, is the limiting step.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Training or reward signals that target claim-level support and irrelevant-evidence suppression may move overall scores more than further fluency tuning.
  • If graphs remain the gold standard, future work will need stress tests for alternate valid reasoning paths that the single annotated scaffold misses.
  • Process-level gates could become a practical filter in agent evaluation loops: do not spend judge budget on answer quality until citation and claim gates pass.
  • Multimodal report HTML with image placeholders may inflate perceived quality unless visual citations are checked against the same evidence graph as text.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper introduces HiEviDR-Bench, a benchmark for evaluating hierarchical evidence aggregation in deep research systems. Each of 2,000 human-validated instances pairs a question with a multimodal corpus, a standard answer, and an explicit evidence graph (evidence → intermediate claims → conclusion). Evaluation decomposes into five 20-point dimensions (report quality, evidence traceability, citation, claim, answer), the last three governed by a progressive gating mechanism in which failure at one stage zeroes subsequent scores. Graphs are LLM-constructed over retrieved candidate pools, then rule-filtered, judge-filtered, and human-validated (30,000 → 3,407 → 2,000). Experiments on 16 MLLMs under RAG and Deep Research settings show report quality remains high while citation, claim, and answer scores drop sharply; answer-gate pass rates are only 3.8%–11.5%. The authors conclude the key bottleneck is evidence identification and intermediate claim construction rather than report fluency.

Significance. If the methodology holds, this is a useful contribution: it moves deep-research evaluation from outcome-only scoring to process-level traceability, and the evidence graph gives a concrete mechanism for error localization (evidence selection vs. claim construction vs. answer). The benchmark is sizable and human-validated (2,000 instances, high inter-annotator agreement), covers both open and academic domains and text-only/multimodal settings, and the authors release code and data, document the construction filters, run 16 systems under two paradigms, and provide judge-reliability and seed-robustness checks. However, the headline diagnostic numbers (gate pass rates) currently inherit two unvalidated assumptions — uniqueness of the gold reasoning path and reliability of the LLM judge's gate decisions — so the central claim should be treated as promising but not yet established.

major comments (3)
  1. [§3.2, Fig. 4] §3.2 (Progressive Gating Mechanism): the paper's most quotable result — answer-gate pass rates of 3.80%–11.50% (Fig. 4) — rests on the assumption that the annotated graph G is not merely *a* correct evidence→claim→conclusion decomposition but effectively the unique one. Per §3.3, graphs are LLM-constructed over a pre-selected 10–25-item candidate pool and human-validated for internal validity (Table 4 checks edges/claims for correctness, not exhaustiveness). A model that reaches the correct conclusion through a different, equally valid claim decomposition or a different evidence subset from the corpus fails the Claim/Answer gate and scores zero. The headline numbers may therefore partially measure graph-conformity. A load-bearing addition: on a human-checked sample, have annotators adjudicate gate failures as 'genuine grounding failure' vs 'valid alternative path', and report the split.
  2. [§3.2, Table 6] §3.2 and App. A.4: gate activation itself (whether cited evidence 'adequately supports the corresponding claim', i.e., the binary w_n decisions in Eqs. 11–12) is made by the Qwen3-VL-235B judge, yet Table 6 validates the judge only via Spearman correlation on continuous scores and system-rank consistency — not gate-level precision/recall against human pass/fail decisions. Because gates compound multiplicatively (a Claim-Gate false negative zeroes the Answer stage), even a modest gate false-negative rate materially moves the 3.8%–11.5% figure. Please report human-vs-judge agreement (κ or P/R) on the three binary gate decisions over a sample, ideally including the alternative judges (GPT-5, Gemini-3.1-Pro) already used in Table 6.
  3. [Table 3, §4.2–4.3] §4.2/Table 3 and §4.3: the ranking of 16 models and per-model RAG-vs-Deep-Research comparisons are reported as point estimates with no uncertainty quantification per system. Table 7 shows the *evaluation pipeline* is stable across seeds, but many inter-model gaps in Table 3 are under 1 point (e.g., 34.00 vs 34.58; 38.13 vs 39.09), which may not exceed per-instance variance. Bootstrap CIs over the 2,000 instances, or at minimum significance statements for the headline RAG-vs-DR gaps in Fig. 5, are needed before model-ordering claims (and the 'best/secondary' highlighting in Table 3) can stand.
minor comments (7)
  1. [§3.2, Eq. (6)] Eq. (6) appears to have a normalization error: as written, 20·(1/(5|K_r|))·Σ(r_k−1)/4 gives a maximum of 4 (with |K_r|=6), yet Table 3 reports Report scores ≈18–19. Presumably the intended form is 20·(1/|K_r|)·Σ(r_k−1)/4. Also, Eq. (6) normalizes via (r_k−1)/4 while Eqs. (11)–(13) use r_k/5, so a minimum judge score contributes 0 to Report but 1/5 elsewhere — please make the normalization consistent or explain the asymmetry.
  2. [§3.2, Eq. (9)] The TF-IDF irrelevance threshold τ (used to define E_irr^gold in Eq. 9) is never given a value, nor is its sensitivity analyzed. Since s3 enters S_trace (Eq. 10), τ should be stated and justified for reproducibility.
  3. [Table 3] Table 3 has formatting defects: several cells are missing separators (e.g., '19.805.44' and '3.680.75' in the GPT-5-mini RAG row; '19.61...' crowding in the DR block). The bold/underline best/secondary highlighting is also hard to verify in places.
  4. [Tables 3 and 8] Model naming is inconsistent: 'InternVL3.5-8B' (Table 3) vs 'Intern3.5-VL-8B' (Table 8, indices 11 and 14). Please unify.
  5. [§3.3 vs Table 2] §3.3 states Easy graphs average 5–8 nodes and Hard 9–17, but Table 2 shows Medium Wiki-Text averaging 10.79 and Medium arXiv-Text 10.15, overlapping the stated Hard range. Clarify whether difficulty is binned by node count, layer count, or a combination.
  6. [§3.3] The arXiv corpus is restricted to 'recent 2 year RAG papers' (Fig. 3). Please state the cutoff date and discuss potential data overlap with the pretraining corpora of the evaluated 2025–2026 models, which could inflate evidence identification on the arXiv subset.
  7. [§2] Related work (§2) would benefit from explicitly contrasting with MiroEval [42] and TRACE [2], which also pursue process-level evaluation; Table 5's row structure is useful but the text does not say what those two lack relative to the evidence-graph formulation.

Circularity Check

0 steps flagged

No circular derivation: empirical benchmark scores are measurements under an explicit metric, not predictions forced by their own inputs.

full rationale

HiEviDR-Bench is an evaluation paper, not a first-principles or fitted-theory paper. Its load-bearing claims are experimental outcomes on a newly defined five-dimension score with progressive gates (Eqs. 5–13, §3.2) and reported model numbers (Table 3, Fig. 4). Defining gold evidence graphs G and scoring reports against them is ordinary benchmark construction: the metric is stipulated, then systems are measured. Nothing in the chain reduces a claimed “prediction” to a fitted parameter or to a self-cited uniqueness theorem. LLM-assisted graph construction plus an LLM judge (Qwen3-VL-235B) raises validity questions about whether gates measure unique grounded reasoning versus conformity to one annotated scaffold, but that is metric-validity risk, not circularity by construction. The paper does not smuggle an ansatz via self-citation, rename a known law as a derivation, or treat a fit as an out-of-sample prediction. Self-contained against its own stated evaluation protocol; circularity score 0.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

As a benchmark paper, load-bearing commitments are methodological: what counts as gold hierarchical structure, how irrelevance and gates are defined, and that LLM judges track human notions of report/citation/claim quality. There is no physical constant fitting; free choices are thresholds, retrieval budgets, and difficulty binning by graph size.

free parameters (4)
  • TF-IDF irrelevance threshold τ = predefined threshold (numeric value not specified in main text)
    Used in §3.2 to define E_gold_irr for Irrelevant Evidence Avoid Accuracy; changes which recalled items count as truly irrelevant and thus moves S_trace.
  • Retrieval budgets (top-15 RAG; top-20/keep 25 text+5 images deep research; ≤3 refinement rounds) = top-15; top-20 then 25+5; 3 rounds
    Implementation details in §4.1 fix the evidence pool size and iteration depth, affecting all downstream traceability and gating rates.
  • Difficulty bins by evidence-graph node count/layers = approx. Easy ~5–8 nodes / Hard ~9–17 nodes, up to 4 layers
    Easy/Medium/Hard partitions (§3.3, Table 2) are defined from graph size/density; this is a design choice that structures reported difficulty trends.
  • Per-dimension 1–5 LLM judge rubrics aggregated to 0–20 = each component capped at 20; total /100
    Report/citation/claim/answer scores are linear transforms of subjective rubric scores; the mapping and dimension sets are author-chosen scoring parameters.
axioms (5)
  • ad hoc to paper A directed acyclic evidence→claim→conclusion graph is the right process model of deep research for evaluation.
    Core formulation Eq. 4–5 and §3.1; alternative process models (flat citations, trajectories only) would change what “failure” means.
  • ad hoc to paper Progressive gates should zero all later scores if citation/claim activation fails, even if the final answer text is correct.
    §3.2 gating mechanism; this design choice strongly drives low Answer scores and the bottleneck narrative.
  • domain assumption LLM-as-judge scores on the stated rubrics are adequate proxies for human judgments of report, citation, claim, and answer quality.
    Standard in open-ended generation evaluation; partially supported by Table 6 correlations on a human-validated subset.
  • domain assumption Human validation after aggressive automatic filtering yields reliable gold graphs and QA pairs.
    §3.3 and Appendix A.1; acceptance ~11% auto-filter then dual annotator checks with high κ.
  • domain assumption Standard retrieval-augmented and iterative deep-research pipelines are fair test harnesses for comparing MLLMs on this task.
    §4.1 and Appendix A.3 pipelines; results are conditional on these harnesses and prompts.
invented entities (2)
  • HiEviDR evidence graph (Ne ∪ Nc ∪ {n_conclusion} with support edges) independent evidence
    purpose: Provide explicit hierarchical supervision so evaluation can score selection, linking, claims, and conclusions separately.
    Defined in §3.1–3.3; the main conceptual object of the benchmark. Independent use outside this paper is possible once the dataset is public, but the ontology is author-defined.
  • Progressive Citation/Claim/Answer gates no independent evidence
    purpose: Force stage-wise credit assignment and fine-grained error localization.
    §3.2; gates are evaluation machinery, not empirical discoveries about nature.

pith-pipeline@v1.2.0-grok45-kimik3 · 31144 in / 3565 out tokens · 72404 ms · 2026-07-31T00:03:17.865545+00:00 · methodology

0 comments
read the original abstract

Deep research requires models to retrieve, connect, and synthesize evidence from large-scale heterogeneous sources to answer complex queries and produce analytical reports. Existing benchmarks mainly evaluate final outcomes, such as answer correctness, report quality, or citation alignment, while providing limited visibility into whether evidence is correctly selected, linked, and aggregated into supported claims and conclusions. To address this gap, we introduce HiEviDR-Bench, a benchmark for evaluating Hierarchical Evidence Aggregation in Deep Research. HiEviDR-Bench covers open-domain and academic-domain settings under both text-only and multimodal conditions, and represents each instance with an explicit evidence graph that captures evidence selection, cross-source linking, and aggregation from evidence to intermediate claims and final conclusions. Based on this formulation, we develop a traceability-oriented evaluation framework with five dimensions: report quality, evidence traceability, citation accuracy, claim verification, and answer correctness, together with a progressive gating mechanism for fine-grained error localization. HiEviDR-Bench contains 2,000 human-validated questions with evidence graphs across multiple difficulty levels. Experiments on 16 representative multimodal large language models show that, although many systems achieve strong report quality, their performance drops markedly on citation accuracy, claim construction, and answer correctness. Further analysis shows that the main bottlenecks lie in evidence identification and intermediate claim construction, revealing that strong surface-level report quality does not necessarily imply grounded multi-stage reasoning on our benchmark.

Figures

Figures reproduced from arXiv: 2607.25151 by Bangrui Xu, Chi Chen, Chunyi Peng, Maosong Sun, Sen Mei, Xuanhe Zhou, Yubo Sun, Yukun Yan, Zhenghao Liu.

Figure 1
Figure 1. Figure 1: Stage-wise score of HiEviDR-Bench, ranked by overall score(0-100). [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the evaluation framework in HiEviDR-Bench. Each question is paired with heterogeneous evidence items, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Pipeline for Constructing HiEviDR-Bench. We first build a multimodal corpus from two source domains, then select [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Stage-wise gate pass rates under the progressive evaluation pipeline. Each stacked bar shows the proportion of samples [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Comparison of RAG and Deep Research on four [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Human validation interface. As shown in [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Comparison of RAG and Deep Research across text [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Overview of the RAG and Deep Research pipelines [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Prompt template used in the RAG setting. [PITH_FULL_IMAGE:figures/full_fig_p014_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: A case study from HiEviDR-Bench. The left panel shows the benchmark instance structure, including the question, [PITH_FULL_IMAGE:figures/full_fig_p015_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Prompt template used in the Evaluation of Report Dimension. [PITH_FULL_IMAGE:figures/full_fig_p016_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Prompt template used in the Evaluation of Evidence Citation Dimension. [PITH_FULL_IMAGE:figures/full_fig_p017_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Prompt template used in the Evaluation of Evidence Claim Dimension. [PITH_FULL_IMAGE:figures/full_fig_p018_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Prompt template used in the Evaluation of Answer Dimension. [PITH_FULL_IMAGE:figures/full_fig_p019_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Prompt template used in the Deep Research Plan Stage. [PITH_FULL_IMAGE:figures/full_fig_p020_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Prompt template used in the Deep Research Init Stage. [PITH_FULL_IMAGE:figures/full_fig_p021_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Prompt template used in the Deep Research Update Stage. [PITH_FULL_IMAGE:figures/full_fig_p022_17.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 24 linked inside Pith

  1. [1]

    BiXie. 2024. wikiimage. https://huggingface.co/datasets/BiXie/wikiimage. Hug- ging Face dataset, accessed on 2026-03-24

  2. [2]

    Yanyu Chen, Jiyue Jiang, Jiahong Liu, Yifei Zhang, Xiao Guo, and Irwin King

  3. [3]

    Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, et al. 2025. Browsecomp- plus: A more fair and transparent evaluation benchmark of deep-research agent. arXiv preprint arXiv:2508.06600(2025)

  4. [4]

    Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. 2025. Deepresearch bench: A comprehensive benchmark for deep research agents. arXiv preprint arXiv:2506.11763(2025)

  5. [5]

    Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, and Anthony KH Tung. 2026. IDRBench: Interactive Deep Research Benchmark. arXiv preprint arXiv:2601.06676(2026)

  6. [6]

    Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling large lan- guage models to generate text with citations. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 6465–6488

  7. [7]

    Tiansheng Hu, Yilun Zhao, Canyu Zhang, Arman Cohan, and Chen Zhao. 2026. SAGE: Benchmarking and Improving Retrieval for Deep Research Agents.arXiv preprint arXiv:2602.05975(2026)

  8. [8]

    Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou, Samiul Alam, Xin Wang, Zexin Li, Zhihao Dou, Li Zhu, Jing Xiong, et al. 2026. MMDeepResearch- Bench: A Benchmark for Multimodal Deep Research Agents.arXiv preprint arXiv:2601.12346(2026)

  9. [9]

    Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, et al . 2025. Deep research agents: A systematic examination and roadmap.arXiv preprint arXiv:2506.18096(2025)

  10. [10]

    Minghao Li, Ying Zeng, Zhihao Cheng, Cong Ma, and Kai Jia. 2025. Report- bench: Evaluating deep research agents via academic survey tasks.arXiv preprint arXiv:2508.15804(2025)

  11. [11]

    Mingxin Li, Yanzhao Zhang, Dingkun Long, Chen Keqin, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2026. Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Frame- work for State-of-the-Art Multimodal Retrieval and Ranking.arXiv preprint arXiv:2601.04720(2026)

  12. [12]

    Zhenghao Liu, Pengcheng Huang, Zhipeng Xu, Xinze Li, Shuliang Liu, Chunyi Peng, Haidong Xin, Yukun Yan, Shuo Wang, Xu Han, et al . 2026. Knowledge intensive agents.AI Open(2026)

  13. [13]

    Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al

  14. [14]

    Chunyi Peng, Zhipeng Xu, Zhenghao Liu, Yishan Li, Yukun Yan, Shuo Wang, Yu Gu, Minghe Yu, Ge Yu, and Maosong Sun. 2026. Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation. InProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1440–1450

  15. [15]

    Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen. ai/blog?id=qwen3.5

  16. [16]

    Yiming Ren, Junjie Wang, Yuxin Meng, Yihang Shi, Zhiqiang Lin, Ruihang Chu, Yiran Xu, Ziming Li, Yunfei Zhao, Zihan Wang, et al . 2026. SIN-Bench: Trac- ing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature.arXiv preprint arXiv:2601.10108(2026)

  17. [17]

    2026.Seed 2.0 Model Card: Towards Intelligence Frontier for Real- World Complexity

    ByteDance Seed. 2026.Seed 2.0 Model Card: Towards Intelligence Frontier for Real- World Complexity. Technical Report. Technical report (model card), February

  18. [18]

    Yijia Shao, Yucheng Jiang, Theodore Kanell, Peter Xu, Omar Khattab, and Mon- ica Lam. 2024. Assisting in writing wikipedia-like articles from scratch with large language models. InProceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6252–6278

  19. [19]

    Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al

  20. [20]

    URL https://lf3-static

  21. [21]

    Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu, Zhenghao Liu, Chi Chen, Zhiyuan Liu, and Maosong Sun. 2025. VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation.arXiv preprint arXiv:2510.09733(2025)

  22. [22]

    Gemma Team. 2026. Gemma 4 Technical Report. arXiv:2607.02770 [cs.CL] https://arxiv.org/abs/2607.02770

  23. [23]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805(2023)

  24. [24]

    Ionut Teodor Sorodoc, Leonardo FR Ribeiro, Rexhina Blloshmi, Christopher Davis, and Adrià de Gispert. 2025. Garage: A benchmark with grounding annotations for rag evaluation. InFindings of the Association for Computational Linguistics: ACL 2025. 17030–17049

  25. [25]

    V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, J...

  26. [26]

    Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny. 2023. Hierarchical multi- modal transformers for Multipage DocVQA.Pattern Recognit.144 (2023), 109834

  27. [27]

    UltraRAG. 2025. UltraRAG Benchmark. https://modelscope.cn/datasets/ UltraRAG/UltraRAG_Benchmark

  28. [29]

    Jiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen, Austin Xu, Zixuan Ke, Frederic Sala, Aws Albarghouthi, Caiming Xiong, and Shafiq Joty. 2025. Livere- searchbench: A live benchmark for user-centric deep research in the wild.arXiv preprint arXiv:2510.14240(2025)

  29. [30]

    Qiuchen Wang, Ruixue Ding, Zehui Chen, Weiqi Wu, Shihang Wang, Pengjun Xie, and Feng Zhao. 2025. ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.CoRRabs/2502.18017 (2025)

  30. [31]

    Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. 2025. InternVL3.5: Ad- vancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency. arXiv preprint arXiv:2508.18265(2025)

  31. [32]

    Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, and Chien-Sheng Wu. 2025. Deeptrace: Auditing deep research ai systems for tracking reliability across citations and evidence.arXiv preprint arXiv:2509.04499(2025)

  32. [33]

    xAI. 2026. Introducing Grok 4.5. https://x.ai/news/grok-4-5. Accessed: 2026-07- 21

  33. [34]

    Yilin Xiao, Junnan Dong, Chuang Zhou, Su Dong, Qian-wen Zhang, Di Yin, Xing Sun, and Xiao Huang. 2025. Graphrag-bench: Challenging domain-specific reasoning for evaluating graph retrieval-augmented generation.arXiv preprint arXiv:2506.02404(2025)

  34. [35]

    LLM-Core-Team Xiaomi. 2025. MiMo-VL Technical Report. arXiv:2506.03569 [cs.CL] https://arxiv.org/abs/2506.03569

  35. [36]

    Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese

  36. [37]

    arXiv preprint arXiv:2504.12516(2025)

    Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516(2025)

  37. [38]

    Bangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou, Yeye He, Shihan Yu, Qianqian Xu, Bin Wang, Guoliang Li, Conghui He, et al. 2026. MoDora: Tree-Based Semi- Structured Document Analysis System.arXiv preprint arXiv:2602.23061(2026)

  38. [39]

    Renjun Xu and Jingwen Peng. 2025. A comprehensive survey of deep research: Systems, methodologies, and applications.arXiv preprint arXiv:2506.12594(2025)

  39. [40]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)

  40. [41]

    Lei Xiong, Huaying Yuan, Zheng Liu, Zhao Cao, and Zhicheng Dou. 2026. Paper- Scope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers. InFindings of the Association for Computational Linguistics: ACL 2026. 8015–8040

  41. [42]

    Yuqi Xiong, Chunyi Peng, Zhipeng Xu, Zhenghao Liu, Zulong Chen, Yukun Yan, Shuo Wang, Yu Gu, and Ge Yu. 2026. Lang2act: Fine-grained visual reasoning through self-emergent linguistic toolchains. InFindings of the Association for Computational Linguistics: ACL 2026. 8375–8399

  42. [43]

    Qinhan Yu, Zhiyou Xiao, Binghui Li, Zhengren Wang, Chong Chen, and Wentao Zhang. 2025. MRAMG-Bench: a comprehensive benchmark for advancing mul- timodal retrieval-augmented multimodal generation. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3616–3626

  43. [44]

    Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chengxing Xie, Cunxiang Wang, et al. 2026. GLM-5: from Vibe Coding to Agentic Engineering.arXiv preprint arXiv:2602.15763(2026)

  44. [45]

    Yu Zeng, Wenxuan Huang, Zhen Fang, Shuang Chen, Yufan Shen, Yishuo Cai, Xi- aoman Wang, Zhenfei Yin, Lin Chen, Zehui Chen, et al. 2026. Vision-deepresearch benchmark: Rethinking visual and textual search for multimodal large language models.arXiv preprint arXiv:2602.02185(2026)

  45. [46]

    Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2025. MiniCPM-V: A GPT-4V Level MLLM on Your Phone.Nat Commun 16, 5509 (2025)(2025)

  46. [47]

    Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao, Yibo Wang, Lei Wang, Zhen Zhang, Lu Wang, et al. 2026. MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome.arXiv preprint arXiv:2603.28407 (2026). Conference’17, July 2017, Washington, DC, USA Yubo Sun, Chunyi Peng, Yukun Yan, Zhenghao Liu, Sen Mei, Bangrui Xu, Xuan...

  47. [48]

    Jiajie Zhang, Yushi Bai, Xin Lv, Wanjun Gu, Danqing Liu, Minhao Zou, Shulin Cao, Lei Hou, Yuxiao Dong, Ling Feng, et al. 2025. Longcite: Enabling llms to generate fine-grained citations in long-context qa. InFindings of the Association for Computational Linguistics: ACL 2025. 5098–5122

  48. [49]

    cover">`-multiple `<div class=

    Wenlin Zhang, Xiaopeng Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Huifeng Guo, Yong Liu, and Xiangyu Zhao. 2025. Deep research: A survey of autonomous research agents.arXiv preprint arXiv:2508.12752(2025). HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research Conference’17, July 2017, Washington, DC, USA A Appendix A.1 Human V...

  49. [51]

    Dingling Zhang, He Zhu, Jincheng Ren, Kangqi Song, Xinran Zhou, Boyu Feng, Shudong Liu, Jiabin Luo, Weihao Xie, Zhaohui Wang, et al. 2025. How Far Are We from Genuinely Useful Deep Research Agents?arXiv preprint arXiv:2512.01948 (2025)

  50. [52]

    Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou, Yanzhe Shan, Haishan Lu, Zhiyong Cao, Jiaoyang Chen, Yuqian Han, Zinan Sheng, et al. 2026. BrowseComp- V3: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents. arXiv preprint arXiv:2602.12876(2026)

  51. [2021]

    Webgpt: Browser-assisted question-answering with human feedback.arXiv preprint arXiv:2112.09332(2021)

  52. [2025]

    Openai gpt-5 system card.arXiv preprint arXiv:2601.03267(2025)

  53. [2026]

    TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents.arXiv preprint arXiv:2602.21230(2026)