Pith. sign in

Paper Citation Record · LEDGER

Characterizing Deep Research: A Benchmark and Formal Definition

As of 18 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 7 inbound Pith citation observations for arXiv:2508.04183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04183 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:56:06.963501Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:55:37.784908Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 03dbcef6-8ba6-4ea8-a1b2-6cc82d8934e8 · outbound

This paper cites OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs.

Characterizing Deep Research: A Benchmark and Formal Definition OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:02.786323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:02.786323Z digest=sha256:f77f3907538879aba03abe32ed164aeaf45b96f38ef2361a9b3f3e5c034339a3

Observation cb5ba02a-ce37-453c-b41d-06d2be453c67 · outbound

This paper cites Deerflow: Deep exploration and efficient research flow.

Characterizing Deep Research: A Benchmark and Formal Definition Deerflow: Deep exploration and efficient research flow

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.557746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:02.847850Z digest=sha256:0c7a1ba709f0171a9b04d9844f19f232b3815fd6b8c8951523e04c119e7f4281

Observation 7873f638-d91f-47cd-88a0-e42195d7d263 · outbound

This paper cites Deep research comparator: A platform for fine-grained human annotations of deep research agents.

Characterizing Deep Research: A Benchmark and Formal Definition Deep research comparator: A platform for fine-grained human annotations of deep research agents

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:56:08.115139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:02.921609Z digest=sha256:5cc21aa2aa2704615d1911d89f9d0b82bc51a5d4468552dec6f2056bd905b5d5

Observation a3e13a14-e5ce-4011-adbb-3f79e81848f2 · outbound

This paper cites Patent claims revisited.

Characterizing Deep Research: A Benchmark and Formal Definition Patent claims revisited

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.446452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.018397Z digest=sha256:eb76c33ccae84177fa9dcaf520d9b0ee1433e9aee4a2bcc614506bc34a7a3282

Observation 47392b29-37d1-4e98-92bf-f97b26228752 · outbound

This paper cites Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research, 2025.

Characterizing Deep Research: A Benchmark and Formal Definition Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.130786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.130786Z digest=sha256:169f3f0645bb7c1a2838f0637d21262f2053e6573488fa4b45551d1a3e0ae6dd

Observation 99c6a37d-7d75-4220-bc87-3016c71eba3a · outbound

This paper cites Curie: Evaluating llms on multitask scientific long-context understanding and reasoning.

Characterizing Deep Research: A Benchmark and Formal Definition Curie: Evaluating llms on multitask scientific long-context understanding and reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.294644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.227291Z digest=sha256:3c13d0a0c0e4ce3fa5f154f4592a01828c3f25cdc34c41252259ca0c045a8223

Observation 4e0ec3d9-7757-42af-a5de-e943f08a60a0 · outbound

This paper cites Claim Verification in the Age of Large Language Models: A Survey.

Characterizing Deep Research: A Benchmark and Formal Definition Claim Verification in the Age of Large Language Models: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.359094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.359094Z digest=sha256:a823502cd0d4e3d8c26fc981dd6befb40f03bedf8ee6a0626dc868b5f4ba6c57

Observation 94b9e6c5-084f-4046-ab83-695e851dc49c · outbound

This paper cites DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.

Characterizing Deep Research: A Benchmark and Formal Definition DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.435700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.435700Z digest=sha256:39f5fad542476c3ea9c68b34308c9ed9cbe7c16dd06595c316e10f108aabc448

Observation 6575b478-1be1-47c9-a63a-a3d3dfef46e3 · outbound

This paper cites Eli5: Long form question answering.

Characterizing Deep Research: A Benchmark and Formal Definition Eli5: Long form question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.105840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.500582Z digest=sha256:c65bd6c857e37145ca36696215cd5614c57d72ac97d7d01e945c8c60dd9c8c05

Observation 9eb6e252-b576-4a9b-83d2-4dc6fe297981 · outbound

This paper cites Deep Research Bench: Evaluating AI Web Research Agents.

Characterizing Deep Research: A Benchmark and Formal Definition Deep Research Bench: Evaluating AI Web Research Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.602916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.602916Z digest=sha256:784c148daf502668bf1fdad8d86e6a6b4de5b5aa6973ed4f5a145e5a0405cde1

Observation a22081e9-1dd1-4e54-bebd-feb46c640066 · outbound

This paper cites Analysis of Plan-based Retrieval for Grounded Text Generation.

Characterizing Deep Research: A Benchmark and Formal Definition Analysis of Plan-based Retrieval for Grounded Text Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.722051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.722051Z digest=sha256:f152ecc56cbc448845ebee449ac022323990e7e517ebe914efaad51bfefc7e80

Observation f49b4db0-90cb-4cff-a7fd-0ecf5d611baf · outbound

This paper cites We’re expanding our gemini 2.5 family of models.

Characterizing Deep Research: A Benchmark and Formal Definition We’re expanding our gemini 2.5 family of models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.959359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.837033Z digest=sha256:927c727e950f4f171acd813089a45265bd269e25073af27179ecf326fe220607

Observation d288c18b-dd7a-4337-b9cb-fc1ecd9a7fd9 · outbound

This paper cites Deep research is now available on gemini 2.5 pro experimental.

Characterizing Deep Research: A Benchmark and Formal Definition Deep research is now available on gemini 2.5 pro experimental

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.789187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.905777Z digest=sha256:8414862016f1a0f011e84846054a947c8fe98e46f726225fa396aeccf13d5c5c

Observation 3e3e4e78-0a2f-435f-bf30-ad03d22bb8e4 · outbound

This paper cites Precise information control in long-form text generation.

Characterizing Deep Research: A Benchmark and Formal Definition Precise information control in long-form text generation

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:56:07.743079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.965981Z digest=sha256:44a5153604713aa99082008c8cdba665f853841a0e1aa35d79b70fb9424e868a

Observation a4ba905c-3371-480d-bb91-b30b043d31c5 · outbound

This paper cites CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review.

Characterizing Deep Research: A Benchmark and Formal Definition CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.028115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.028115Z digest=sha256:34b492dc9a7ff5cd2644b42df9dbbdab91dea429eed71d774ed994125bf53710

Observation 4c1007f5-23ca-41f9-a99e-111c513a6e8d · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

Characterizing Deep Research: A Benchmark and Formal Definition Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.628438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.126062Z digest=sha256:e3b2023d6fe7acad48074dac733ffbc9dc6c2c94353d978ee3ec7890f14ce391

Observation f1955f78-e90c-4975-b142-e196a5de0a86 · outbound

This paper cites Bright: A realistic and challenging benchmark for reasoning-intensive retrieval.

Characterizing Deep Research: A Benchmark and Formal Definition Bright: A realistic and challenging benchmark for reasoning-intensive retrieval

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.433449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.258755Z digest=sha256:5eb8a3010b36df77d9de59b65cbf4b0c8a4a457d00adcb8f7960c666478a8ae4

Observation 54aefd0c-7c46-42f1-ac90-1ae705b8b67d · outbound

This paper cites Deep Research Agents: A Systematic Examination And Roadmap.

Characterizing Deep Research: A Benchmark and Formal Definition Deep Research Agents: A Systematic Examination And Roadmap

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.340301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.340301Z digest=sha256:d57582a57578f95dede26953cca67336653ae6ce48df48984320519f831b4847

Observation ea2b89f0-64fc-452c-8df6-a23bb7b4239d · outbound

This paper cites Open-source deepresearch – freeing our search agents.

Characterizing Deep Research: A Benchmark and Formal Definition Open-source deepresearch – freeing our search agents

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.204652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.429658Z digest=sha256:6dbe6a7e720e859207977e692e0bdff28c3062327be25b781962cccca6f1c17b

Observation ab53b0c9-bfea-4277-bba3-a69948a65897 · outbound

This paper cites GPT-4o System Card.

Characterizing Deep Research: A Benchmark and Formal Definition GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.494148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.494148Z digest=sha256:b2f4d78e9c0eba204e73709c03d8b87f4a1ccc0d0d88fa6840067a4c2822838f

Observation 1b959b1b-33f0-4bf6-a975-0f8ce39e5abd · outbound

This paper cites Researcharena: Benchmarking llms' ability to collect and organize information as research agents.

Characterizing Deep Research: A Benchmark and Formal Definition Researcharena: Benchmarking llms' ability to collect and organize information as research agents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.063859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.614465Z digest=sha256:2646f89820c64837437d6a5028caa8ca92ca64579111265c9f75fd6c8b3116b6

Observation 1c0c6023-6e3a-42d0-8237-85a416344917 · outbound

This paper cites Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation.

Characterizing Deep Research: A Benchmark and Formal Definition Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.705727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.705727Z digest=sha256:9a9f87a24f4a1d5071309f576a96aa16271cf75df6a71b51a066d46d17cfea22

Observation e86d8368-f668-45c0-9d9e-86d5a7ff9b2e · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Characterizing Deep Research: A Benchmark and Formal Definition WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.807861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.807861Z digest=sha256:c6262895412d41a95aa83d135107457a8cdc7d5a11e2682539f40b81cff0f361

Observation 162653f0-454a-4c43-a31d-04e6351292a8 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Characterizing Deep Research: A Benchmark and Formal Definition Lost in the middle: How language models use long contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.922012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.922012Z digest=sha256:40ec87ac32fa8f5cdab9555e76c521b77069b761317dacac53cfd4f7615ccc75

Observation 9756f86a-b74a-4967-94e7-8ad6b900c0b6 · outbound

This paper cites Veritrail: Closed-domain hallucination detection with traceability.

Characterizing Deep Research: A Benchmark and Formal Definition Veritrail: Closed-domain hallucination detection with traceability

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.012999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.012999Z digest=sha256:924d0ce8c853b4141bb762e5610a68dd79329a8b6dc324fc0dadc7025d2a4a81

Observation 4798bd25-a7f3-4e9f-a766-58e4bf4991c8 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Characterizing Deep Research: A Benchmark and Formal Definition Gaia: a benchmark for general ai assistants

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.077103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.077103Z digest=sha256:c2006bcfa0dc198dd5e3384913a9e2becda266b73a4a283c2c32ce4e55a5e21d

Observation 6ec2cfc2-eeac-4ba4-8f86-9ce3e30e4d2b · outbound

This paper cites Introducing researcher and analyst in microsoft 365 copilot.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing researcher and analyst in microsoft 365 copilot

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.910345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.166245Z digest=sha256:ab28c42674283716e9393626f7d15770ba4246c7599c0858580ea63c3e84e984

Observation 56e45693-eb36-4204-87b0-1c6c3cbe3d5c · outbound

This paper cites Introducing gpt-4.1 in the api.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing gpt-4.1 in the api

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.711474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.311592Z digest=sha256:66295fb3ca048d8a185a1583fda7465eb55acdc91e184a53101a863767622b38

Observation 1a1c0847-da88-4ed7-a1ac-210ce821edaf · outbound

This paper cites Introducing openai o3 and o4-mini.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing openai o3 and o4-mini

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.502455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.398991Z digest=sha256:fa06de92a784290a7b90d0a69ecf5634915f9a824abe6774a4f537c9ce5ff48b

Observation f3641c16-58db-46f4-a2e3-c1d4b2e1732a · outbound

This paper cites Introducing deep research.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing deep research

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.331658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.457585Z digest=sha256:4d43db2fbb221e95bbb110dcb5cbc0d2e2df6710e2bdb30ae61eeba8ea2b7c4b

Observation 939e8131-e079-4109-9648-ec56bbe920c3 · outbound

This paper cites Perplexity deep research.

Characterizing Deep Research: A Benchmark and Formal Definition Perplexity deep research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.161922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.539963Z digest=sha256:6d9e8e54210a005fb1b7c631ee1b9cb6edb9c0825103f2c9c5e0fb835fbb61fc

Observation 3a430fc2-771c-4c86-be0b-655f4191d9a5 · outbound

This paper cites Sonar pro.

Characterizing Deep Research: A Benchmark and Formal Definition Sonar pro

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.924311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.639565Z digest=sha256:525d999874bc7cb54f5ac756dbd0d8511d30d3738df297a59ce471103794e0c7

Observation 8a87cb7d-bf03-41b7-8375-c8ada03ce137 · outbound

This paper cites Sonar reasoning.

Characterizing Deep Research: A Benchmark and Formal Definition Sonar reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.664108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.752392Z digest=sha256:b6b1928a9bcee27405858f99fdbf82d32b93632f5aff1fc820cfef9c924bcd3c

Observation 1cf3a592-6fcc-481a-943f-6654d7210807 · outbound

This paper cites Humanity's Last Exam.

Characterizing Deep Research: A Benchmark and Formal Definition Humanity's Last Exam

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.839637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.839637Z digest=sha256:083a00a6ccef972f8b4900acaa3145ba3b69beda4ee87056cf28f2969646ea25

Observation 7db41a09-ac7a-44bb-85f7-d93f61f26b30 · outbound

This paper cites Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models.

Characterizing Deep Research: A Benchmark and Formal Definition Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.908536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.908536Z digest=sha256:979d78cb9055e126be309446c5f11cc0860e370b20d330c0a93b4a197c799426

Observation 2c4cd75a-5d68-4eb9-a54d-391f20305271 · outbound

This paper cites Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning.

Characterizing Deep Research: A Benchmark and Formal Definition Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.979388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.979388Z digest=sha256:9f6c49817f0b839a8c8cc9273a4b0103b8372f3477bc28cc5e2f8ddee286466f

Observation 03684cc1-1c48-4208-ac52-c0fc5903f379 · outbound

This paper cites Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks.

Characterizing Deep Research: A Benchmark and Formal Definition Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.062666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.062666Z digest=sha256:433bc640bdb2f2e0ff47e7a15875057f8423bad8cb55f885e9a7152190bdf709

Observation bb0a689d-f746-4cfe-ab35-ab456294fade · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Characterizing Deep Research: A Benchmark and Formal Definition BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.248077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.248077Z digest=sha256:27c46a0a7a9a6d7e8b2da501208767ff73a8f5c092cc50ecef5dcba501db6e81

Observation 8d05bfba-f63d-4428-98c5-7f50e26485d5 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

Characterizing Deep Research: A Benchmark and Formal Definition Grok 3 beta — the age of reasoning agents

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.483853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:06.363778Z digest=sha256:03b97526063f6b94fb6c59e7fcd7ab56637e105edff1f141c0d38753ddbda526

Observation 3f2aab72-d23c-48fd-aeab-4d798dc666c5 · outbound

This paper cites A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications.

Characterizing Deep Research: A Benchmark and Formal Definition A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.477316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.477316Z digest=sha256:0131783506e257a20dd8a66c6a06b31fb669a59cbb86ccc67d66793e107a7e1b

Observation 77053ef7-0ad3-4000-9422-9ef41ea7db02 · outbound

This paper cites ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry.

Characterizing Deep Research: A Benchmark and Formal Definition ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.611023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.611023Z digest=sha256:a9828632148c78156f82fa28cf2a7007b617cd3ce097a0022fc8a732ef53b00f

Observation 93682201-ba35-45b8-bab6-622c2e40b46a · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

Characterizing Deep Research: A Benchmark and Formal Definition Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.728074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.728074Z digest=sha256:c7df4b7b362b075bec4166aef56aeb786489ba03a58ab370d48c2e774187b9c7

Observation cd0aac83-49ef-486e-a35b-20aaa2520bca · outbound

This paper cites Open deep research.

Characterizing Deep Research: A Benchmark and Formal Definition Open deep research

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.296610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T00:56:06.874170Z digest=sha256:625ed265a29c0e8d8285671ee00951ba66f0b9a475396bc7c7509223ed6e8c3b

Observation 1eee2b58-4690-436a-b8aa-539bfd30cc12 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Characterizing Deep Research: A Benchmark and Formal Definition DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.963501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.963501Z digest=sha256:e2777ac7d166bad149101c3949f386b6338e9e2584335fadee669cf4d66e884a

Pith citing papers

Observation 4dacb85f-9cdc-4029-913b-85458931071d · inbound

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents cites this paper.

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Characterizing Deep Research: A Benchmark and Formal Definition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:37.784908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:37.784908Z digest=sha256:ec7452ec0ba643e8e2edc45e9e2599b006b6b4436d69e2f96793355decbc6218

Observation 501f1ced-25ef-438c-9d5e-50fc2476867b · inbound

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation cites this paper.

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation Characterizing Deep Research: A Benchmark and Formal Definition

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T15:15:21.070914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:15:21.070914Z digest=sha256:2285a405a170586e00539cb019f1cca9eeff302dece21923653a5aa0ddf3dcad

Observation 11b1b2bd-db4d-48d4-b763-f577d1a0f2b7 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective Characterizing Deep Research: A Benchmark and Formal Definition

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:19.781040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T18:54:06.144968Z digest=sha256:d04560ade8c419e2c081d625f776757eee37030f73636e521a80f2acb87d8f4b

Observation 5925a9e2-9178-46d4-944b-42d66103eaff · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective Characterizing Deep Research: A Benchmark and Formal Definition

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:19:16.463224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T00:18:32.423103Z digest=sha256:0f8d83ac13360778ebd1b9d61563b586ca4e615107c42965c6f3787dcdfcb969

Observation 3a906274-6149-4fd3-8cf5-f1915fb0cf04 · inbound

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake cites this paper.

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake Characterizing Deep Research: A Benchmark and Formal Definition

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:37.840871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T13:41:35.986601Z digest=sha256:422cdf799ecc9b19fe39654ed81fe8e3ccd927ce351682d203ea2394094b4cfd

Observation 25802b62-2a13-4874-9fc4-744c6dedfc42 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Characterizing Deep Research: A Benchmark and Formal Definition

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.500298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:5c7a279952b6f3437c386cd91305915adf954e316e5f9a29e60e96cfbba68dee

Observation 69a34f5d-1739-45d7-8a99-af30964826e1 · inbound

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution cites this paper.

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution Characterizing Deep Research: A Benchmark and Formal Definition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:31:21.166209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:31:21.166209Z digest=sha256:d54d51eea5cb5e3f1c148658fe73bb5713597e6319187e45524e49fe9813b632