Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:54.197640Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 20 inbound Pith citation observations for arXiv:2506.23719.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:54.197640Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:53:23.455387Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
51 of 51 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 4d315fcc-0848-49ac-96e1-6050c8eafbb1 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c41d26e-68ac-488d-8c30-1c7a2a243d09 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53a705c0-6657-455f-819c-9013f35389f5 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Claude 3.5: Next-generation language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c216f380-c155-4219-a7c5-bb70cd4309b3 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Claude 3.7: Advancements in language understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13d8d062-c7cf-449e-bfc8-21c54f844013 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dbeba95d-0dd7-46bf-9482-91eef5049960 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c594346-2d69-4e4f-a8d5-abcdcee3746b · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c164cf-b273-49ef-92df-ad7218e1e9ff · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be278e08-16aa-4dc6-8875-c0a78999081d · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 25d5083f-57c2-4d75-8493-b73cee719fda · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Gemini 2.5: Our most intelligent ai model, 2025
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f42eff18-f0fa-41d3-80dc-ac02d16b4a38 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6923dfeb-86cc-480d-8a21-377f4f3458b5 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd09a51-46ec-42ea-b264-43336721becc · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Text-to-SQL in the wild: A naturally-occurring dataset based on stack exchange data
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d60b1519-3c14-4765-a5db-57d57519602f · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Measuring mathematical problem solving with the MATH dataset
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e367110-4d52-4d25-adb3-e8549ba96535 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning InfiAgent-DABench: Evaluating agents on data analysis tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1977c770-4d5b-46ae-a20c-82fba8066dfc · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Mlagentbench: evaluating language agents on machine learning experimentation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87af90cf-60bb-4edd-87e6-13c81dee47b3 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning DA-code: Agent data science code generation benchmark for large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6b73b48-8231-49a3-9207-fffff11dea19 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Financebench: A new benchmark for financial question answering, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffd4f8f-cb5e-426f-814c-8932093f66a5 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad541e7-06b3-471f-85c6-5da7590a891b · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning DSBench: How far are data science agents from becoming data science experts? In The Thirteenth International Conference on Learning Representa- tions, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3da2933-c194-4152-9803-8d353273e5cd · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Ds-1000: A natural and reliable benchmark for data science code generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccad99ba-4c64-4542-9986-c7e4cd043976 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning KaggleDBQA: Realistic eval- uation of text-to-SQL parsers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad14efe0-2847-4d41-90b9-c4c0929c0bac · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Spider 2.0: Evaluating language models on real-world enterprise text-to-SQL workflows
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f69aea46-4d3c-4d41-ad35-33788897ee2f · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Can llm already serve as a database interface? a big benchmark for large-scale database-grounded text-to-sqls
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88d65af3-85e1-4685-92c3-21a1cf1abceb · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19c53cf2-4f48-41f8-b39f-d0811b0c152a · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning DeepSeek-V3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3a561c-799a-4c41-9755-3b739a34c96c · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Agentbench: Evaluating LLMs as agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6874c40-368f-4d2e-a4bf-5b036e725406 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Gaia: a benchmark for general ai assistants
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ec8070a-dae3-43fa-add3-29b17f4fa92b · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7bca586-d0ee-4c66-9e7c-3b1d8c93405e · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Introducing gpt-4o: Multimodal capabilities and efficiency
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78f6950f-26f5-47ab-929c-f87abf9397ac · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Gpt-4o mini: advancing cost-efficient intelligence
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 73e40c86-e70f-4a96-82ba-d258d2dd2bcd · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Openai o1 and new tools for developers
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfc5149b-3e5c-420a-8654-26d83a5f34dc · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Gpt-4.1 technical overview
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 517e6610-118a-4502-bf78-2fbe3128c30b · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Introducing openai o3 and o4-mini
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 399e59f8-ea89-421d-b6e7-504fe3faaa0b · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning On the Difficulty of Evaluating Baselines: A Study on Recommender Systems
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd058908-c099-4d70-abbe-6dbc1cf6f90c · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning smolagents: A smol library to build great agentic systems, 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 430eae7b-a90d-4ab3-828a-8f9fa0b44d0e · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Let me speak freely? a study on the impact of format restrictions on large language model performance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 224fe5b7-773b-43ce-b753-7d3f9b33d125 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Attention is all you need
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2ad18714-dea3-48c9-9525-941ad833c1ce · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d799cd1-a372-4fb0-ab24-903b849b45f2 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Text-to-sql generation for question answering on electronic medical records
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44fab7f5-6366-4e6f-a27a-ae3086d26246 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Won- derbread: A benchmark for evaluating multimodal foundation models on business process management tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b3f0674-ec5c-47d2-a125-bb7ff1a64c01 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1349b8da-5657-4278-9702-4815ef78a55e · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b9a81d-d68c-4d9f-8bd1-90d789826100 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Intercode: Standard- izing and benchmarking interactive coding with execution feedback
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c985aa84-40ad-4478-ac04-5495149a8b6e · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning React: Synergizing reasoning and acting in language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d723746-67cf-4626-afb7-cda7350050d1 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Natural language to code generation in interactive data science notebooks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77c023b4-86c9-4cae-9759-955d778af1a0 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a5446d4-801b-45d5-87c1-bca8f374a499 · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Benchmarking data science agents
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba97e827-8b90-41de-ad74-e8d97748cfce · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Judging llm-as-a-judge with mt- bench and chatbot arena
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a061464-ef55-4cea-8158-caa7e362e69a · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3f61fc-2d65-4c4c-9e34-3b2b93ff5e2a · outbound
DABstep: Data Agent Benchmark for Multi-step Reasoning Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb3c064e-2200-47d6-b291-ba24dc973c63 · inbound
DSBC : Data Science task Benchmarking with Context engineering DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3acc66b7-f3e8-4d21-bf40-9a6157ddbd12 · inbound
FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791502ba-261b-4ae2-9413-05d8da1a7934 · inbound
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61775897-6a07-44a1-94e7-af61958cd931 · inbound
Structure-Grounded Knowledge Retrieval via Code Dependencies for Multi-Step Data Reasoning DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 460546be-bfd7-4a7a-b7ef-6f93de34232d · inbound
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f49fc793-54bd-435d-a09e-278b60ad2148 · inbound
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b1f780e-650b-4329-ae03-04d19bb71f99 · inbound
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ccecea1c-66f2-4f47-b2d0-bd498e94c771 · inbound
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c5a8580-4aa7-4c60-a654-f250986eebe3 · inbound
PrepBench: How Far Are We from Natural-Language-Driven Data Preparation? DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e073fb7-07c5-4eb4-8c3c-8895a3ed457b · inbound
Text Analytics Evaluation Framework: A Case Study on LLMs and Social Media DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f99a9bc-e4d2-4d84-9274-6d901d19469c · inbound
Agentic AI Workload Characteristics DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 96694f58-b1a6-4369-bcc7-b9e9f682099e · inbound
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 886e156f-5d39-488e-870b-2fbb2dc4bde2 · inbound
VESTA: Visual Exploration with Statistical Tool Agents DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dcd5dbcc-62f6-4266-b7f2-0a5ca6772369 · inbound
Unsupervised Skill Discovery for Agentic Data Analysis DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c333f4e-6b57-4919-93b9-576aafae6069 · inbound
GRACE-DS: a Guarded Reward-guided Agent Correction Environment in Data Science DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cdcd41e0-5419-459c-81eb-a60d331ce251 · inbound
CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4841dbd9-2b56-4d6f-86b1-d410a2e6897b · inbound
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476bc867-a9d9-4d3b-afc6-a0fe2646229b · inbound
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c22f85-25f8-41b1-8fae-261cc4c9ba4c · inbound
UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b6306c8-17af-4a00-bd30-9d2331ff07a6 · inbound
DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness DABstep: Data Agent Benchmark for Multi-step Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.