Pith. sign in

Paper Citation Record · LEDGER

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

As of 15 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2607.12252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.12252 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:42:09.641449Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:39:34.778096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-03T12:16:14.848390Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f48fab4e-7d8b-441f-b44a-8bad75e0fe1c · outbound

This paper cites an unresolved cited work.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:07.919219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:07.919219Z digest=sha256:352aa18c9d6562658f506b38197658c47eff8d7e049e74b4e65de58404949308

Observation a65a7ba9-db4b-4d98-aec5-eb1734d1a06d · outbound

This paper cites DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.127436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.127436Z digest=sha256:cb7429e82d394e4841ea67448b0d687a2a8a292763d9c9b67549580cc094beb6

Observation bceb10e3-d39f-418e-8cbf-87f0774d914a · outbound

This paper cites an unresolved cited work.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.241173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.241173Z digest=sha256:9851ce68836fd366b8f10559bdd78481da1a59e45c62295e0d53cfdbc9591971

Observation 76b952e0-6513-4e6e-bcfc-9704500c9b59 · outbound

This paper cites BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.398960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.398960Z digest=sha256:a5c247198ec3b2534efec9859a9928156c368ee080014dc05b1859750fc9d4ea

Observation cd6a4237-faaa-4657-8234-995e5298f0b5 · outbound

This paper cites an unresolved cited work.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.461514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.461514Z digest=sha256:3ebe1066246ef5a1442493f570515acea88c79e8f0e4b6c9c73c9b9ca450875e

Observation e50a4124-0d90-4ddd-9f75-66bc7224786b · outbound

This paper cites HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.570674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.570674Z digest=sha256:0704b63980f1cd633efd07282ed58cdcd438fe8e7bb11964371ed5f77ad49683

Observation 88db5f24-768f-4ed3-a3c2-e1bbdd7e6cc6 · outbound

This paper cites InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 5977–6043.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 5977–6043

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.624206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.624206Z digest=sha256:80a5bd781cbe5ab3091eaff0243cd194bb27e424791b3c9c63ab9a1e3a9d05a7

Observation 581ae318-8efb-42f9-9210-cc8d88f2ac3e · outbound

This paper cites an unresolved cited work.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.693996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.693996Z digest=sha256:d344fbffe5ef82bf1f9752c66afbf040a941ed0bca50b0cbe6e12f0a4c3bacbc

Observation 15f66bc1-57d8-4608-886c-8243edea3907 · outbound

This paper cites FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.783508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.783508Z digest=sha256:74ddcd443eea08861f60517c617e6e285dbc51febb4024a432ee4a2392b44a4b

Observation ae20f553-2ee5-40b2-9c31-b31403181a56 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.873217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.873217Z digest=sha256:56d116887cf89a11a43509f641ece74efc21ac7f11fd5d7edcd7823b5f9a34bc

Observation 4dd093ed-d0ba-4f34-b129-8b254c0c7724 · outbound

This paper cites LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.925768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.925768Z digest=sha256:68cee6bb5c56ed4c97e85433e15775abf1e1ee3ba0188e84ecebcee502ecd1d4

Observation 1e5994d6-ff87-4c8f-81b5-506f258a1959 · outbound

This paper cites A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.001256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.001256Z digest=sha256:913191746d1052d69bb09033a025892a844cd3d0afe46bdccf4549000e9d2b6b

Observation 21db6964-ef75-4378-a602-96d607e0b95a · outbound

This paper cites arXiv preprint arXiv:2601.05111.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality arXiv preprint arXiv:2601.05111

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.172939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.172939Z digest=sha256:51a176fc759513794f63e1703affca680ede20764f374aa5d6fb43acdeb7f7c2

Observation 7d1d83a4-f5b1-497d-b91d-14440922655b · outbound

This paper cites FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.297213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.297213Z digest=sha256:9139450d82e9c31532b4caa0e01acfef37426bb6583100e38e1219fe0babad79

Observation 5d6a9d45-4a8e-4f57-8576-c6ef12e3ec18 · outbound

This paper cites InProceedings of the 2025 Conference on Empirical Methods in Natural Lan- guage Processing, pages 414–431.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality InProceedings of the 2025 Conference on Empirical Methods in Natural Lan- guage Processing, pages 414–431

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.370074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.370074Z digest=sha256:88c3f1608d068ebb1cf1f6258ce802e71cacfd47e38d9b88379c6322248c489b

Observation 646a2981-2a8e-4c1d-b8ba-c341df0c83e2 · outbound

This paper cites an unresolved cited work.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.489574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.489574Z digest=sha256:c965089acf702f557eee9afa9514969d4386da5a5af278a196419acaf6d2034b

Observation 86b67f9f-38ad-4783-8a98-836a59ac5cd9 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality Agent-as-a-Judge: Evaluate Agents with Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.641449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.641449Z digest=sha256:bc3b7ec1395ac3d42ea337928449ae7bf8f83e7d7a934c5263b4c870a517b161

Observation ad4b7ce1-d9f0-4115-ac33-cc3cbb4b4932 · outbound

This paper cites InProceedings of the 2018 Conference on Empirical Methods in Natural Language Process- ing, pages 2369–2380.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality InProceedings of the 2018 Conference on Empirical Methods in Natural Language Process- ing, pages 2369–2380

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:09.070946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:09.070946Z digest=sha256:c53f05fc9fe665e21015ffae23510582d9b704219cb4759b5585cb2c21ab3c91

Observation bdc2b554-74e2-4efd-ab94-b560e9542a8e · outbound

This paper cites DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.297818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.297818Z digest=sha256:0255dfa080996c110dd5fce4ac392923b70f86139460fde9f7d3eb14d2f22e87

Observation ebf59298-5f97-40a3-b4a8-ad9f5f064476 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality AgentBench: Evaluating LLMs as Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:08.350415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:08.350415Z digest=sha256:f166becd290a896f512ec5d19af6fc0bb9470f227feff0493605a2449b996e96

Observation 5695399d-7498-4a8f-88e3-b65b6faf7ca4 · outbound

This paper cites LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:07.646102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:07.646102Z digest=sha256:c7cc07231bd8a45dbb7b2e626576f64958942881bca4fec68da03baa61e6969b

Observation c68b6e74-dfea-4c26-8820-e258fedee60d · outbound

This paper cites DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:07.800790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:07.800790Z digest=sha256:05cdb4799a973d70c0547e08651982c1b53103b132be8f012fc1a0875a7568ae

Observation bcdbd8bc-9523-4a0c-a55d-61fb0dff0982 · outbound

This paper cites BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation.

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:07.998224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:07.998224Z digest=sha256:946f50324b84743191627b499a9d59c834d42e32368e98fcfe1f3ee53a6fcd7d

Pith citing papers

Observation 40b7bff8-2aa2-4a06-b5be-46f5a3f9164a · inbound

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation cites this paper.

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:48:03.586532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:48:03.586532Z digest=sha256:2259f45407f56a1b1b47551e3d1a0373b48cf83c0dae9404b4f317791f2363b0

Observation 0fb45fc1-202e-410e-ada6-66e1282fb769 · inbound

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation cites this paper.

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-08-03T10:48:30.603281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-03T10:48:03.665912Z digest=sha256:10a560ee6659a00a60671c501ade9f9473e479c43899fb71ae1e9786c93e9e19

Observation 9b3690fa-dd86-4197-a962-8c27a9431167 · inbound

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables cites this paper.

FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:39:34.778096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:39:34.778096Z digest=sha256:5b781336fc9de86f744d96f4c0af7b242c960515054c4ab95f81ce8421e5fa4c