Pith. sign in

Paper Citation Record · LEDGER

INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2412.18174.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18174 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:51:57.987439Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:40:08.185850Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d4442445-3340-4eb0-b89c-21c0f3028ba4 · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:08.963610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:e3343c14bcf43c423e022ecfe02f0e58b292300b5cf05a2dab32e47c6e879faa

Observation cf249527-93ee-4a66-a201-37dc7df469d4 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.987439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.987439Z digest=sha256:5288c7b6d1a98bedfe195f91e4c84bb04aa8216f277afae521d0ec91a4eafd08

Observation 1ac203ab-100b-414c-968e-71096f40a0db · inbound

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model cites this paper.

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:51.742478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:51.742478Z digest=sha256:0a58ad09fa47f64cb44f13523ecb56463cf9b53ed9fc3d899d3b5b2e5acbd203

Observation 368b79ad-531d-43cc-9132-731873f30f02 · inbound

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design cites this paper.

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:42.544819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:42.544819Z digest=sha256:0d0bb300a43beb1de4aa8ccb4a9f36236bd83e7cda4defe65379cc2db94363dd

Observation 161354e1-cd85-4b15-904e-6e4535b2e86b · inbound

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets cites this paper.

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:26.139668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:26.139668Z digest=sha256:887d178a5a442eb21f8ae046a6197b2affdfd09005d7bc267d829bda09c6d5d5

Observation b25a384f-dec2-4731-8ad1-ffa168260518 · inbound

Emergent Social Intelligence Risks in Generative Multi-Agent Systems cites this paper.

Emergent Social Intelligence Risks in Generative Multi-Agent Systems INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.890900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:45:04.625084Z digest=sha256:856d6603627ff814f4753b968cf45172f77a4beb631a25d3c062ee965337f3c0

Observation 2a4a0303-81ed-413d-b22c-b63fdde27eb1 · inbound

QRAFTI: An Agentic Framework for Empirical Research in Quantitative Finance cites this paper.

QRAFTI: An Agentic Framework for Empirical Research in Quantitative Finance INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:02.783297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:06:14.794251Z digest=sha256:309cc50858c10f5f7a7e69988805ad2cd6d5d8bc98b9b8375ac7ddf1374780ac

Observation 0227aee5-b5aa-4057-b03b-ab8d9c4a34b5 · inbound

FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation cites this paper.

FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:32:54.618299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T00:28:48.871454Z digest=sha256:aa1a57243b4d241844fc1f42be0d1e581aebc711c0687c2a45cf0b9cdabb26f8

Observation ea7f6d7b-6150-4963-aa25-322f851e4b11 · inbound

Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems cites this paper.

Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.668483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:35:11.289439Z digest=sha256:cbb8fa9621616d0f2c2196e110ee423aa2fc86d208a7fb8cde4031cab697f3d3

Observation 2dce1e8f-2a9e-497f-8e63-6483e390c9f0 · inbound

Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking cites this paper.

Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:45.694961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:14:31.166883Z digest=sha256:ad89984fcb23f8d0b01080c4189af02a55133ff8c449ff1da2bb6f58718590d4

Observation 45087736-8178-4cbd-aa13-1fd8440445cf · inbound

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy cites this paper.

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:40:08.187168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T19:51:45.701059Z digest=sha256:27373daf9dd2066a05114a8a69d575a943569970f6a798e44ffc762eb16f4c56

Observation 608489da-f9d6-4c5b-9ed4-aabaeca9e5ba · inbound

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy cites this paper.

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T12:08:16.614140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:08:16.614140Z digest=sha256:2da96d2e477341b04c7ea40f03eeb27cf731de667d8a3c0f26654609b2f1628b

Observation ee267393-70d5-40a8-9b8a-87993d363b8a · inbound

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents cites this paper.

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:18.634074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:43:48.382708Z digest=sha256:148bfa6e4b976101076926cc5f1620a8bafc751477e59de7599d3ca2c21e011f

Observation fdf4290a-e3b4-48d4-9971-9a02e155d931 · inbound

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents cites this paper.

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:39:36.957884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:39:36.957884Z digest=sha256:742c216755c16cdc306072775439b8875448b53b2c81b79715f3e063d434d258

Observation dd5e2d37-11b6-42f7-a32d-c1b6fe0b33d8 · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.599846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:79bb4adfea3dbe0a97fcd888c6a3cf5dc500d5a80a13c6c3cbc7c7845baf7bb8

Observation 3801a31e-0b28-41a9-af50-0914ff282248 · inbound

NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management cites this paper.

NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T06:46:36.394796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:46:36.394796Z digest=sha256:4423331e79b22921c21819713dffa3e12bea51f9528e8f2304914262eee5f95e

Observation 210feb0b-e54d-4725-b93b-ce240cb97796 · inbound

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning cites this paper.

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:31:52.564336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:31:52.564336Z digest=sha256:3dbded8a2e0a56af8c4f938de4b1d756f6040766750a0dc2ab4e8885656d4748

Observation 66d1c2ac-48ef-4000-b0bf-b1faaee4fafe · inbound

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning cites this paper.

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:50.852973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:49:50.852973Z digest=sha256:ef56a5ff40e6320add2739d52922eef8e5441e20e9407ba59f63d0070c236dd9

Observation 903e35be-f651-40be-b0b7-454b982553a8 · inbound

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable cites this paper.

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:48:36.690873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:48:36.690873Z digest=sha256:9ed74d21effa0b485faaa30bfab260907440fd833de2ad2b0d89f5a5f6d9145c

Observation dcad189c-03cc-4051-a82d-e708f0bba3a4 · inbound

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? cites this paper.

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T10:42:35.065327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:42:35.065327Z digest=sha256:16cfca2b10ef90d7a54d5e7e99911c4c0f95a69c6a418b67e287e5dd88c6b8a4