Pith. sign in

Paper Citation Record · LEDGER

Characterizing Deep Research: A Benchmark and Formal Definition

As of 18 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 7 inbound Pith citation observations for arXiv:2508.04183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04183 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:56:06.963501Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:55:37.784908Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 03dbcef6-8ba6-4ea8-a1b2-6cc82d8934e8 · outbound

This paper cites OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs.

Characterizing Deep Research: A Benchmark and Formal Definition OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:02.786323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:02.786323Z digest=sha256:f77f3907538879aba03abe32ed164aeaf45b96f38ef2361a9b3f3e5c034339a3

Observation cb5ba02a-ce37-453c-b41d-06d2be453c67 · outbound

This paper cites Deerflow: Deep exploration and efficient research flow.

Characterizing Deep Research: A Benchmark and Formal Definition Deerflow: Deep exploration and efficient research flow

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.557746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:02.847850Z digest=sha256:fc928bdea977b48b559fa48bcf0615d3ca680c1f1f0f46708607e0e94c55d406

Observation 7873f638-d91f-47cd-88a0-e42195d7d263 · outbound

This paper cites Deep research comparator: A platform for fine-grained human annotations of deep research agents.

Characterizing Deep Research: A Benchmark and Formal Definition Deep research comparator: A platform for fine-grained human annotations of deep research agents

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:56:08.115139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:02.921609Z digest=sha256:79caeb2806e21c67b5d2e323826c5bdcbd33c46007d37ca9af530b08516c30a2

Observation a3e13a14-e5ce-4011-adbb-3f79e81848f2 · outbound

This paper cites Patent claims revisited.

Characterizing Deep Research: A Benchmark and Formal Definition Patent claims revisited

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.446452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.018397Z digest=sha256:cfccc2e91775b8769e3d900d2894314a27174cb73e1c786b5600894aef426fe5

Observation 47392b29-37d1-4e98-92bf-f97b26228752 · outbound

This paper cites Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research, 2025.

Characterizing Deep Research: A Benchmark and Formal Definition Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.130786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.130786Z digest=sha256:169f3f0645bb7c1a2838f0637d21262f2053e6573488fa4b45551d1a3e0ae6dd

Observation 99c6a37d-7d75-4220-bc87-3016c71eba3a · outbound

This paper cites Curie: Evaluating llms on multitask scientific long-context understanding and reasoning.

Characterizing Deep Research: A Benchmark and Formal Definition Curie: Evaluating llms on multitask scientific long-context understanding and reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.294644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.227291Z digest=sha256:dea83376089f17d280d596fffccd2433eaa171869df07dbca39f06394e29f22b

Observation 4e0ec3d9-7757-42af-a5de-e943f08a60a0 · outbound

This paper cites Claim Verification in the Age of Large Language Models: A Survey.

Characterizing Deep Research: A Benchmark and Formal Definition Claim Verification in the Age of Large Language Models: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.359094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.359094Z digest=sha256:a823502cd0d4e3d8c26fc981dd6befb40f03bedf8ee6a0626dc868b5f4ba6c57

Observation 94b9e6c5-084f-4046-ab83-695e851dc49c · outbound

This paper cites DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.

Characterizing Deep Research: A Benchmark and Formal Definition DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.435700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.435700Z digest=sha256:39f5fad542476c3ea9c68b34308c9ed9cbe7c16dd06595c316e10f108aabc448

Observation 6575b478-1be1-47c9-a63a-a3d3dfef46e3 · outbound

This paper cites Eli5: Long form question answering.

Characterizing Deep Research: A Benchmark and Formal Definition Eli5: Long form question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:11.105840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.500582Z digest=sha256:95f395a3a19d44f804136571467885967c6ae4d33b1cc5fcc73f03094558ab7d

Observation 9eb6e252-b576-4a9b-83d2-4dc6fe297981 · outbound

This paper cites Deep Research Bench: Evaluating AI Web Research Agents.

Characterizing Deep Research: A Benchmark and Formal Definition Deep Research Bench: Evaluating AI Web Research Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.602916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.602916Z digest=sha256:784c148daf502668bf1fdad8d86e6a6b4de5b5aa6973ed4f5a145e5a0405cde1

Observation a22081e9-1dd1-4e54-bebd-feb46c640066 · outbound

This paper cites Analysis of Plan-based Retrieval for Grounded Text Generation.

Characterizing Deep Research: A Benchmark and Formal Definition Analysis of Plan-based Retrieval for Grounded Text Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:03.722051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:03.722051Z digest=sha256:3038154782e14caa2f6e2fa435fe808881e291fd8bf086cecf07c80f23cf595c

Observation f49b4db0-90cb-4cff-a7fd-0ecf5d611baf · outbound

This paper cites We’re expanding our gemini 2.5 family of models.

Characterizing Deep Research: A Benchmark and Formal Definition We’re expanding our gemini 2.5 family of models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.959359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.837033Z digest=sha256:d48b0a15b59f1602ecf4bd18693caa112713b672bc04644255704fd8351eef61

Observation d288c18b-dd7a-4337-b9cb-fc1ecd9a7fd9 · outbound

This paper cites Deep research is now available on gemini 2.5 pro experimental.

Characterizing Deep Research: A Benchmark and Formal Definition Deep research is now available on gemini 2.5 pro experimental

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.789187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.905777Z digest=sha256:93c2a9f96f346f6511b693e097b5a9ab98ec4815324b844af9a686d90ff77e75

Observation 3e3e4e78-0a2f-435f-bf30-ad03d22bb8e4 · outbound

This paper cites Precise information control in long-form text generation.

Characterizing Deep Research: A Benchmark and Formal Definition Precise information control in long-form text generation

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:56:07.743079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:03.965981Z digest=sha256:64f3515cf764c310422729040af9abb65cde70a5f7a704724650bd208e3b8e9c

Observation a4ba905c-3371-480d-bb91-b30b043d31c5 · outbound

This paper cites CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review.

Characterizing Deep Research: A Benchmark and Formal Definition CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.028115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.028115Z digest=sha256:34b492dc9a7ff5cd2644b42df9dbbdab91dea429eed71d774ed994125bf53710

Observation 4c1007f5-23ca-41f9-a99e-111c513a6e8d · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

Characterizing Deep Research: A Benchmark and Formal Definition Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.628438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.126062Z digest=sha256:3dd5e6005f35301ae3c90a4734f65e635593c774ce281052fe82a786e16d2dc0

Observation f1955f78-e90c-4975-b142-e196a5de0a86 · outbound

This paper cites Bright: A realistic and challenging benchmark for reasoning-intensive retrieval.

Characterizing Deep Research: A Benchmark and Formal Definition Bright: A realistic and challenging benchmark for reasoning-intensive retrieval

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.433449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.258755Z digest=sha256:7bafa044d076c966814dcced317d573cce47111172ff53d96395401fce8082f6

Observation 54aefd0c-7c46-42f1-ac90-1ae705b8b67d · outbound

This paper cites Deep Research Agents: A Systematic Examination And Roadmap.

Characterizing Deep Research: A Benchmark and Formal Definition Deep Research Agents: A Systematic Examination And Roadmap

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.340301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.340301Z digest=sha256:d57582a57578f95dede26953cca67336653ae6ce48df48984320519f831b4847

Observation ea2b89f0-64fc-452c-8df6-a23bb7b4239d · outbound

This paper cites Open-source deepresearch – freeing our search agents.

Characterizing Deep Research: A Benchmark and Formal Definition Open-source deepresearch – freeing our search agents

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.204652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.429658Z digest=sha256:d1831e983695fd4fc78f39b5d628f16501c14e99636209e70264198b2daeeff5

Observation ab53b0c9-bfea-4277-bba3-a69948a65897 · outbound

This paper cites GPT-4o System Card.

Characterizing Deep Research: A Benchmark and Formal Definition GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.494148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.494148Z digest=sha256:b2f4d78e9c0eba204e73709c03d8b87f4a1ccc0d0d88fa6840067a4c2822838f

Observation 1b959b1b-33f0-4bf6-a975-0f8ce39e5abd · outbound

This paper cites Researcharena: Benchmarking llms' ability to collect and organize information as research agents.

Characterizing Deep Research: A Benchmark and Formal Definition Researcharena: Benchmarking llms' ability to collect and organize information as research agents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:10.063859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:04.614465Z digest=sha256:590e27a00c96afebb50deeace0bfeb9e22a94cf6cf11459983e168573d7f5c6d

Observation 1c0c6023-6e3a-42d0-8237-85a416344917 · outbound

This paper cites Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation.

Characterizing Deep Research: A Benchmark and Formal Definition Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.705727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.705727Z digest=sha256:553de10cc154822bba87c45a5ced061b1b3395e562d4a54c3be006426c0b9f66

Observation e86d8368-f668-45c0-9d9e-86d5a7ff9b2e · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Characterizing Deep Research: A Benchmark and Formal Definition WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.807861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.807861Z digest=sha256:c6262895412d41a95aa83d135107457a8cdc7d5a11e2682539f40b81cff0f361

Observation 162653f0-454a-4c43-a31d-04e6351292a8 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Characterizing Deep Research: A Benchmark and Formal Definition Lost in the middle: How language models use long contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:04.922012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:04.922012Z digest=sha256:40ec87ac32fa8f5cdab9555e76c521b77069b761317dacac53cfd4f7615ccc75

Observation 9756f86a-b74a-4967-94e7-8ad6b900c0b6 · outbound

This paper cites Veritrail: Closed-domain hallucination detection with traceability.

Characterizing Deep Research: A Benchmark and Formal Definition Veritrail: Closed-domain hallucination detection with traceability

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.012999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.012999Z digest=sha256:924d0ce8c853b4141bb762e5610a68dd79329a8b6dc324fc0dadc7025d2a4a81

Observation 4798bd25-a7f3-4e9f-a766-58e4bf4991c8 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Characterizing Deep Research: A Benchmark and Formal Definition Gaia: a benchmark for general ai assistants

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.077103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.077103Z digest=sha256:c2006bcfa0dc198dd5e3384913a9e2becda266b73a4a283c2c32ce4e55a5e21d

Observation 6ec2cfc2-eeac-4ba4-8f86-9ce3e30e4d2b · outbound

This paper cites Introducing researcher and analyst in microsoft 365 copilot.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing researcher and analyst in microsoft 365 copilot

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.910345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.166245Z digest=sha256:d41a72cddbc8bea4ffdd1c66df24b23fdfb19e637773c6df846257c69e674d65

Observation 56e45693-eb36-4204-87b0-1c6c3cbe3d5c · outbound

This paper cites Introducing gpt-4.1 in the api.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing gpt-4.1 in the api

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.711474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.311592Z digest=sha256:e90a37c046ba8a3694a76bbca43f753a3cc03e6c20442a7cc1bb7aba378f2536

Observation 1a1c0847-da88-4ed7-a1ac-210ce821edaf · outbound

This paper cites Introducing openai o3 and o4-mini.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing openai o3 and o4-mini

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.502455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.398991Z digest=sha256:ddf3578721d6b7f43341c6d52f286d60ea2acea5712920bcebb5b1591d732f34

Observation f3641c16-58db-46f4-a2e3-c1d4b2e1732a · outbound

This paper cites Introducing deep research.

Characterizing Deep Research: A Benchmark and Formal Definition Introducing deep research

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.331658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.457585Z digest=sha256:fc88667c712badee3f6bc52daeda698d0c322a7ff5f2cdc8f1b5a84a0830d238

Observation 939e8131-e079-4109-9648-ec56bbe920c3 · outbound

This paper cites Perplexity deep research.

Characterizing Deep Research: A Benchmark and Formal Definition Perplexity deep research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:09.161922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.539963Z digest=sha256:b3fe9e790f01f8d9e410a49de6959850ba3782475f8ad7c3a220ba39ec7c3526

Observation 3a430fc2-771c-4c86-be0b-655f4191d9a5 · outbound

This paper cites Sonar pro.

Characterizing Deep Research: A Benchmark and Formal Definition Sonar pro

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.924311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.639565Z digest=sha256:b59d812e9736bfd69e3942ac5ffe497f3827ac5a0e8460257d3f9ba1acb2a48b

Observation 8a87cb7d-bf03-41b7-8375-c8ada03ce137 · outbound

This paper cites Sonar reasoning.

Characterizing Deep Research: A Benchmark and Formal Definition Sonar reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.664108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:05.752392Z digest=sha256:e3d1001d3d9b075e33e320bc6a70b3a76c89d9a914948bf8c19e9452faf5d2d4

Observation 1cf3a592-6fcc-481a-943f-6654d7210807 · outbound

This paper cites Humanity's Last Exam.

Characterizing Deep Research: A Benchmark and Formal Definition Humanity's Last Exam

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.839637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.839637Z digest=sha256:083a00a6ccef972f8b4900acaa3145ba3b69beda4ee87056cf28f2969646ea25

Observation 7db41a09-ac7a-44bb-85f7-d93f61f26b30 · outbound

This paper cites Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models.

Characterizing Deep Research: A Benchmark and Formal Definition Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.908536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.908536Z digest=sha256:979d78cb9055e126be309446c5f11cc0860e370b20d330c0a93b4a197c799426

Observation 2c4cd75a-5d68-4eb9-a54d-391f20305271 · outbound

This paper cites Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning.

Characterizing Deep Research: A Benchmark and Formal Definition Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:05.979388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:05.979388Z digest=sha256:9f6c49817f0b839a8c8cc9273a4b0103b8372f3477bc28cc5e2f8ddee286466f

Observation 03684cc1-1c48-4208-ac52-c0fc5903f379 · outbound

This paper cites Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks.

Characterizing Deep Research: A Benchmark and Formal Definition Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.062666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.062666Z digest=sha256:433bc640bdb2f2e0ff47e7a15875057f8423bad8cb55f885e9a7152190bdf709

Observation bb0a689d-f746-4cfe-ab35-ab456294fade · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Characterizing Deep Research: A Benchmark and Formal Definition BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.248077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.248077Z digest=sha256:27c46a0a7a9a6d7e8b2da501208767ff73a8f5c092cc50ecef5dcba501db6e81

Observation 8d05bfba-f63d-4428-98c5-7f50e26485d5 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

Characterizing Deep Research: A Benchmark and Formal Definition Grok 3 beta — the age of reasoning agents

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.483853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:06.363778Z digest=sha256:3d322d3fe5603a7e6da3719ea86b9c8f4a52d523680e11da0ae80bf34711f120

Observation 3f2aab72-d23c-48fd-aeab-4d798dc666c5 · outbound

This paper cites A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications.

Characterizing Deep Research: A Benchmark and Formal Definition A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.477316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.477316Z digest=sha256:0131783506e257a20dd8a66c6a06b31fb669a59cbb86ccc67d66793e107a7e1b

Observation 77053ef7-0ad3-4000-9422-9ef41ea7db02 · outbound

This paper cites ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry.

Characterizing Deep Research: A Benchmark and Formal Definition ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.611023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.611023Z digest=sha256:a9828632148c78156f82fa28cf2a7007b617cd3ce097a0022fc8a732ef53b00f

Observation 93682201-ba35-45b8-bab6-622c2e40b46a · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

Characterizing Deep Research: A Benchmark and Formal Definition Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.728074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.728074Z digest=sha256:c7df4b7b362b075bec4166aef56aeb786489ba03a58ab370d48c2e774187b9c7

Observation cd0aac83-49ef-486e-a35b-20aaa2520bca · outbound

This paper cites Open deep research.

Characterizing Deep Research: A Benchmark and Formal Definition Open deep research

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:56:08.296610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T00:56:06.874170Z digest=sha256:24a2914ef23f7081a7dea428cb154e6f89bfca6833892de0ea481ef919bb28fc

Observation 1eee2b58-4690-436a-b8aa-539bfd30cc12 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Characterizing Deep Research: A Benchmark and Formal Definition DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.963501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.963501Z digest=sha256:e2777ac7d166bad149101c3949f386b6338e9e2584335fadee669cf4d66e884a

Pith citing papers

Observation 4dacb85f-9cdc-4029-913b-85458931071d · inbound

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents cites this paper.

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents Characterizing Deep Research: A Benchmark and Formal Definition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:37.784908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:37.784908Z digest=sha256:ec7452ec0ba643e8e2edc45e9e2599b006b6b4436d69e2f96793355decbc6218

Observation 501f1ced-25ef-438c-9d5e-50fc2476867b · inbound

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation cites this paper.

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation Characterizing Deep Research: A Benchmark and Formal Definition

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T15:15:21.070914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:15:21.070914Z digest=sha256:2285a405a170586e00539cb019f1cca9eeff302dece21923653a5aa0ddf3dcad

Observation 11b1b2bd-db4d-48d4-b763-f577d1a0f2b7 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective Characterizing Deep Research: A Benchmark and Formal Definition

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:19.781040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T18:54:06.144968Z digest=sha256:3dd4809a16ff4c7c56594721f67636151a8e8e0f855d29ec06a1f2d494fca24f

Observation 5925a9e2-9178-46d4-944b-42d66103eaff · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective Characterizing Deep Research: A Benchmark and Formal Definition

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:19:16.463224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T00:18:32.423103Z digest=sha256:c6a6a9090f76ec6def2fbce9bb2e8816624eb2da963fce3e17de243eaf02ac41

Observation 3a906274-6149-4fd3-8cf5-f1915fb0cf04 · inbound

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake cites this paper.

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake Characterizing Deep Research: A Benchmark and Formal Definition

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:37.840871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T13:41:35.986601Z digest=sha256:1c8c5c75b0592a0c8b53a1a103455e0ec9fab2f508a0586fdbb0c30c0bf7740a

Observation 25802b62-2a13-4874-9fc4-744c6dedfc42 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Characterizing Deep Research: A Benchmark and Formal Definition

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.500298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:b46a56e73f83ba9a3e00e2cc8d8e68d14b59effd84febc2b688809b4dda8b523

Observation 69a34f5d-1739-45d7-8a99-af30964826e1 · inbound

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution cites this paper.

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution Characterizing Deep Research: A Benchmark and Formal Definition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:31:21.166209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:31:21.166209Z digest=sha256:d54d51eea5cb5e3f1c148658fe73bb5713597e6319187e45524e49fe9813b632