Pith. sign in

REVIEW 3 major objections 82 references

Can Agentic Trading Systems Pay for Their Own Intelligence?

T0 review · 3 major / 0 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read LLM trading agents are viable only when dynamic decisions create enough timing profit to cover the intelligence costs they incur.

desk verdict Useful audit toolkit that joins Brinson-style profit attribution with LLM cost accounting; the timing residual is imperfect but the diagnostic framing is still worth engaging. read the letter →

arxiv 2607.10286 v1 pith:VZDIPMZL submitted 2026-07-11 cs.AI cs.MA

classification cs.AIcs.MA
keywords agentictradingsystemsLLMagentsprofit-costviabilityTradeLensperformanceattributionintelligence-to-profitconversioncost-awareevaluationtimingeffect
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that evaluations of large language model trading agents usually stop at gross profit or risk-adjusted returns and therefore miss a stricter economic test: whether continual model reasoning, tool use, and reallocation produce enough incremental profit to pay for the costs those decisions induce. Gross numbers can look acceptable while the system still fails to beat passive market exposure or a static hold of its own starting portfolio, and while inference, commissions, and infrastructure absorb the margin. The authors introduce TradeLens, a toolkit that rebuilds trajectories from trading records, runtime traces, and deployment configs, then attributes profit into market exposure, initial asset selection, and timing, and attributes cost into model, trading, infrastructure, and residual terms. Across backbones, capital scales, decision frequencies, and architectures, they find that viability turns on intelligence-to-profit conversion: models fail in different ways (for example weak selection versus negative timing), while capital, frequency, and architecture mainly amplify or degrade decision-attributed timing value rather than independently fixing economics. A reader who cares about real deployment gains a way to ask not only which agent ranks highest, but whether a given system pays for its own intelligence and why.

What carries the argument

TradeLens, a trace-grounded diagnostic toolkit, operationalizes profit–cost viability by reconstructing trajectories and defining agentic viability as timing profit minus dynamic costs. Timing is the residual portfolio-value gain after a market-exposure baseline and a static hold of the agent’s initial allocation; that residual is treated as the component most directly attributable to agentic intervention.

What would settle it

Freeze the agent’s portfolio weights after the initial allocation and re-run the identical accounting window; if TradeLens still reports large nonzero timing profit or changes viability labels, the residual does not isolate dynamic intervention.

Watch

Extended reading notes

Core claim

Viability of agentic trading systems hinges on intelligence-to-profit conversion rather than on gross performance alone. System viability asks whether end-to-end profit covers all deployment costs; agentic viability asks whether timing profit from dynamic LLM-mediated reallocation covers decision-induced costs. Models exhibit distinct failure patterns, such as poor asset selection or negative timing, while capital scale, trading frequency, and architecture matter mainly by amplifying or degrading that decision-attributed timing value.

Load-bearing premise

The leftover profit after removing market movement and a static hold of the agent’s own starting weights is treated as a clean measure of the value created by later dynamic decisions.

Editorial extensions

If this is right

  • Trading-agent evaluations should report system viability and agentic viability separately, not only gross profit or Sharpe-style metrics.
  • Backbone choice should be diagnosed by failure mode—selection versus timing—rather than ranked solely by cumulative return.
  • Raising capital or decision frequency is not a free fix; both can enlarge losses when timing value is already negative.
  • Architectural complexity is justified only when added reasoning and coordination produce measurable timing gains that cover the extra cost.
  • Deployments can be audited from ordinary trading records, runtime traces, and cost configurations without assuming one agent design.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same profit–cost split could diagnose other paid LLM agents whose claimed value must offset inference and tool costs, not only trading systems.
  • Builders may need separate remedies for selection skill and timing skill instead of treating token reduction as a universal fix.
  • Because cost-saving suggestions can reduce fees without restoring strategy edge, viability toolkits should pair cost audits with regime-shift tests of predictive signal quality.
  • Training objectives that directly optimize agentic (timing-minus-dynamic-cost) profit, rather than gross return, are a natural next step once the attribution is trusted.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper argues that LLM-based trading agents should be evaluated not only by gross profit or risk-adjusted returns, but by whether they pay for the intelligence they consume. It defines system viability (net profit after end-to-end costs) and agentic viability (timing profit after decision-induced costs), using nested counterfactuals: market exposure via an S&P 500 proxy, asset selection via static hold of the agent’s initial weights, and timing as the residual dynamic path (Eqs. 1–10). TradeLens operationalizes this with an accounting layer over trading records, runtime traces, and deployment configs, plus an LLM diagnosis layer that produces financial and system reports. Controlled backtests on AI-Trader over a two-month window (Dec 2025–Jan 2026) compare ten backbones, capital scales, daily vs hourly frequency, and CoT / AI-Trader / DeepFund architectures. The central empirical claim is that viability hinges on intelligence-to-profit conversion—models show distinct failure modes (e.g., poor selection for DeepSeek-V3.2, negative timing for GLM-4.7)—while capital, frequency, and architecture mainly amplify or degrade decision-attributed timing value (Tables 1–2, Figs. 3–5).

Significance. If the framing holds, the paper usefully shifts evaluation of agentic trading systems from capability ranking to economically grounded, trace-based diagnosis. Distinguishing system vs agentic viability, joining classical-style attribution with explicit LLM/trading/infra cost accounting, and shipping an open toolkit with practitioner feedback are concrete contributions. The multi-axis experiments (backbone, capital, frequency, architecture) and regime appendix make the diagnostic lens more than a single-run case study. Strengths include reproducible code, explicit cost taxonomy, and evidence-grounded reports rather than opaque leaderboards. The work is best read as a methodology and diagnostic toolkit paper, not as a definitive ranking of model trading skill.

major comments (3)
  1. §3.1 Eqs. (2)–(3) and agentic viability (Eqs. 8–10): P_timing = V_dyn_T − V_base_T is treated as the profit component most directly attributable to dynamic LLM intervention, and the RQ1–RQ4 takeaways rest on that residual. V_base freezes the agent’s own initial weights, so first-period selection skill/luck is removed from “agentic” value, while subsequent path dependence, fill quality, universe drift, and execution frictions remain inside P_timing. The Limitations already note S&P 500 proxy bias. The manuscript should either (i) strengthen identification (e.g., alternative benchmarks, randomized or fixed initial portfolios, execution-quality controls) or (ii) substantially soften claims that capital/frequency/architecture “matter only by amplifying or degrading decision-attributed timing value,” and present P_timing as a residual diagnostic, not a pure skill measure.
  2. §5.1 and primary results (Tables 1–2, Figs. 3–5): The headline patterns are estimated on a two-month, mostly sideways window with a fixed 10-name liquid universe and $100k default capital. Appendix D.5 shows profit sources flip dramatically across bear/sideways/bull while costs stay relatively flat, yet the abstract and takeaways still generalize that timing conversion dominates. Either expand the primary evaluation horizon/regimes or reframe all main claims as diagnostic observations under the stated window (as the Limitations begin to do), and avoid language that reads as general deployability conclusions.
  3. §3.1–3.2 and Appendix C/D.1: Agentic cost C_dyn excludes static infra and includes user-configured stochastic Uniform noise and a success-rate adjustment ρ. R_agent can change sign under modest redefinition of what is “decision-induced.” The paper should report sensitivity of SystemViable/AgenticViable to κ, ρ, commission schedule, and C_sto bounds (or set C_sto=0 as a best-case baseline), and make clear that viability labels are conditional on these accounting choices rather than model-invariant facts.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: viability criteria are explicit accounting definitions applied to observed trajectories, not predictions forced by construction or self-citation.

full rationale

TradeLens defines system and agentic viability from reconstructed portfolio paths and cost logs (Eqs. 1–10): P_1:T is split via two nested counterfactuals into market, selection, and timing; R_agent = P_timing − C_dyn is then an accounting residual, not a fitted forecast. Empirical claims (backbone failure modes; scale/frequency/architecture as amplifiers of timing) come from backtests across models and settings (Tables 1–2, Figs. 3–5), not from parameters fit to the same targets and re-labeled as predictions. Classical Brinson-style attribution is cited externally (Brinson et al. 1986, 1991); self-citations (e.g., AI-Trader, DeepFund) supply experimental systems under audit, not uniqueness theorems that force the result. User-configured cost terms and the S&P 500 proxy affect validity of attribution, not circularity of the derivation chain. No equation reduces the headline claim to a normalization identity or self-citation load-bearing step.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

The central claim rests on an adapted performance-attribution identity, a four-part cost taxonomy with several user-chosen constants, and the interpretive step that residual timing equals agentic intelligence value. No new physical entities; the invented pieces are the viability predicates and the toolkit interface.

free parameters (4)
  • infra cost κ / fixed data subscription = ~$208.20 infra block in tables; $0.20/run + $100/mo
    Set to ~$0.20 per run plus ~$100/month subscription in experiments; user-configured rather than measured end-to-end.
  • LLM success-rate adjustment ρ = 0.98
    Token costs scaled by ρ=98% success rate from provider metrics; affects C_llm.
  • stochastic cost upper bound = max ~$0.5/run
    Uniform(0,n) residual term with example max ~$0.5 per run; additive uncertainty chosen by authors.
  • commission / trading fee schedule = provider schedule (not single scalar)
    Retail brokerage-style per-execution fees taken from Interactive Brokers documentation and applied as C_trd.
assumptions (4)
  • domain assumption Gross profit decomposes as P_sys + P_asset + P_timing via nested market and static-allocation counterfactuals (adapted Brinson logic at portfolio-value level).
    §3.1 Eqs. 2–3; treats residual after market and initial weights as agentic timing value.
  • ad hoc to paper Dynamic decision-induced costs are C_llm + C_trd + C_sto, excluding static infrastructure that does not vary with agent decisions.
    §3.2 Eq. 8; classification of which costs count as agentic is a modeling choice central to AgenticViable.
  • domain assumption S&P 500 return over the window is an adequate proxy for systematic market exposure of the liquid US equity universe.
    §5.1 and Limitations; authors note possible attribution bias.
  • domain assumption Transaction-cost economics framing: intelligence is paid inference realized through fee- and impact-sensitive execution.
    §3.1 / Appendix C citing Dow/Williamson-style transaction costs.
invented entities (2)
  • SystemViable / AgenticViable indicators
    purpose: Binary deployability predicates separating whole-pipeline break-even from timing-minus-dynamic-cost break-even.
    Defined in §3.2 Eqs. 7 and 10; not standard finance metrics.
  • TradeLens toolkit (accounting + diagnosis layers) independent evidence
    purpose: Trace-grounded middleware that reconstructs trajectories, attributes P/C, and emits financial + diagnosis reports.
    Core artifact of the paper; independent use requires the released code and user inputs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can Agentic Trading Systems Pay for Their Own Intelligence?." pith.science (2026). https://pith.science/paper/VZDIPMZL

@misc{pith2026260710286,
  author       = {Pith},
  title        = {Pith review of: Can Agentic Trading Systems Pay for Their Own Intelligence?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VZDIPMZL}},
  note         = {Machine review of arXiv:2607.10286}
}
read the original abstract

Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but rarely examine agentic viability: whether dynamic LLM-mediated decisions convert their induced costs into measurable incremental profit. To apply this criterion, we introduce TradeLens, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion. Our results show that viability hinges on intelligence-to-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek-V3.2 and negative timing in GLM-4.7, while capital scale, trading frequency, and architecture matter only by amplifying or degrading decision-attributed timing value. These findings reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion. Our code is available at https://anonymous.4open.science/r/TradeLens.

Figures

Figures reproduced from arXiv: 2607.10286 by the authors.

Figure 1
Figure 1. Motivation. A profitable agentic trading sys [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The pipeline overview. Retail traders can utilize the toolkit to evaluate the profit–cost viability of their own agentic systems under realistic assumptions by providing trading results, runtime traces, and basic configurations. Consequently, they can gain iterative improvements and achieve promising results. calls, retries, latency, and execution events, used to recover cost drivers and system activities. Sys￾tem c… view at source ↗
Figure 3
Figure 3. Viability across backbone models. (a) Sys￾tem viability and (b) agentic viability: profit (y-axis) versus cost (x-axis) for 10 backbone LLMs under iden￾tical trading configurations. hourly decision frequencies using marginal profit minus marginal cost in one month horizon. RQ4 explores the architecture reasoning complexity by placing AI-Trader as a moderate design. We uti￾lize the typical Chain-of-Thought prompting … view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Viability across capital scale. Net Profit and Agentic Profit under varying initial cash investment scale for GPT-5.2 and DeepSeek-V3.2. tic profit but negative net profit, indicating that ac￾tive improvement is insufficient to overcome total system cost. Overall, back…
Figure 6
Figure 6. Figure 6: Forward diagnostic analysis of trading fre [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

82 extracted references · 5 canonical work pages

  1. [1]

    Neural Information Processing Systems , pages =

    Liu, Jiawei and Xia, Chunqiu Steven and Wang, Yuyao and Zhang, Lingming , title =. Neural Information Processing Systems , pages =. 2023 , publisher =

  2. [2]

    Regime changes and financial markets , author =. Annu. Rev. Financ. Econ. , volume =. 2012 , publisher =

  3. [3]

    Frugalgpt: How to use large language models while reducing cost and improving performance , author =. Trans. Mach. Learn. Res. , year =

  4. [4]

    Engineering , year =

    Advancing financial engineering with foundation models: progress, applications, and challenges , author =. Engineering , year =

  5. [5]

    arXiv.org , doi =

    Chen, Yanxu and Yao, Zijun and Liu, Yantao and Ye, Jin and Yu, Jianing and Hou, Lei and Li, Juanzi , year =. arXiv.org , doi =

  6. [6]

    arXiv.org , doi =

    Ding, Qianggang and Shi, Haochen and Liu, Bang , year =. arXiv.org , doi =

  7. [7]

    Conference on Empirical Methods in Natural Language Processing , volume =

    Large language model agents in finance: A survey bridging research, practice, and real-world deployment , author =. Conference on Empirical Methods in Natural Language Processing , volume =. 2025 , doi =

  8. [8]

    arXiv.org , year =

    Cost-of-Pass: An Economic Framework for Evaluating Language Models , author =. arXiv.org , year =

Show all 82 references
  1. [9]

    IEEE Symposium Series on Computational Intelligence , doi =

    An improved epsilon constraint handling method embedded in MOEA/D for constrained multi-objective optimization problems , author =. IEEE Symposium Series on Computational Intelligence , doi =. 2016 , organization =

  2. [10]

    arXiv.org , doi =

    Fan, Tianyu and Yang, Yuhao and Jiang, Yangqin and Zhang, Yifei and Chen, Yuxuan and Huang, Chao , year =. arXiv.org , doi =

  3. [11]

    Quantitative finance , volume =

    No-dynamic-arbitrage and market impact , author =. Quantitative finance , volume =. 2010 , publisher =

  4. [12]

    International Conference on Computing, Communication and Automation , author =

    Cost-. International Conference on Computing, Communication and Automation , author =. 2025 , journal =. doi:10.1109/ICCCA66364.2025.11325717 , publisher =

  5. [13]

    Gemini 3 Flash: frontier intelligence built for speed , year =

  6. [14]

    Gundlach, Hans and Lynch, Jayson and Mertens, Matthias and Thompson, Neil , year =. The. arXiv.org , doi =

  7. [15]

    arXiv.org , year =

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning , author =. arXiv.org , year =

  8. [16]

    arXiv.org , doi =

    Guo, Taian and Shen, Haiyang and Huang, Jinsheng and Mao, Zhengyang and Luo, Junyu and Chen, Zhuoru and Liu, Xuhui and Xia, Bingyu and Liu, Luchen and Ma, Yun and others , year =. arXiv.org , doi =

  9. [17]

    Frontiers of Information Technology & Electronic Engineering , volume =

    FinSphere: a real-time stock analysis agent with instruction-tuned large language models and domain-specific tool integration , author =. Frontiers of Information Technology & Electronic Engineering , volume =. 2025 , number =. doi:10.1631/FITEE.2500414 , publisher =

  10. [18]

    Proceedings of the VLDB Endowment , year =

    Thriftllm: On cost-effective selection of large language models for classification queries , author =. Proceedings of the VLDB Endowment , year =. doi:10.14778/3749646.3749702 , publisher =

  11. [19]

    arXiv.org , year =

    Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning , author =. arXiv.org , year =

  12. [20]

    Finance Research Letters , volume =

    Can ChatGPT improve investment decisions? From a portfolio management perspective , author =. Finance Research Letters , volume =. 2024 , publisher =

  13. [21]

    European Journal of Operational Research , volume =

    An efficient, adaptive parameter variation scheme for metaheuristics based on the epsilon-constraint method , author =. European Journal of Operational Research , volume =. 2006 , publisher =

  14. [22]

    International Conference on AI in Finance , pages =

    Large language models in finance: A survey , author =. International Conference on AI in Finance , pages =. 2023 , journal =. doi:10.1145/3604237.3626869 , publisher =

  15. [23]

    Conference on Empirical Methods in Natural Language Processing , doi =

    Cryptotrade:. Conference on Empirical Methods in Natural Language Processing , doi =. 2024 , pages =

  16. [24]

    International Conference on Artificial Intelligence and Statistics , year =

    Towards Cost Sensitive Decision Making , author =. International Conference on Artificial Intelligence and Statistics , year =

  17. [25]

    The Web Conference , year =

    HedgeAgents: A Balanced-aware Multi-agent Financial Trading System , author =. The Web Conference , year =. doi:10.1145/3701716.3715232 , publisher =

  18. [26]

    Annual Meeting of the Association for Computational Linguistics , pages =

    Investorbench: A benchmark for financial decision-making tasks with llm-based agent , author =. Annual Meeting of the Association for Computational Linguistics , pages =. 2025 , journal =. doi:10.18653/v1/2025.acl-long.126 , publisher =

  19. [27]

    arXiv.org , doi =

    Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking , author =. arXiv.org , doi =

  20. [28]

    arXiv.org , year =

    Fingpt: Democratizing internet-scale data for financial large language models , author =. arXiv.org , year =

  21. [29]

    arXiv.org , year =

    Deepseek-v3 technical report , author =. arXiv.org , year =

  22. [30]

    , year =

    Liu, Jiayu and Qian, Cheng and Su, Zhaochen and Zong, Qing and Huang, Shijue and He, Bingxiang and Fung, Yi R. , year =. arXiv.org , doi =

  23. [31]

    arXiv.org , doi =

    Luo, Yichen and Feng, Yebo and Xu, Jiahua and Tasca, Paolo and Liu, Yang , year =. arXiv.org , doi =

  24. [32]

    Introducing Llama 4: The next generation of multimodal intelligence , date =

  25. [33]

    MiniMax M2.1: Significantly Enhanced Multi-Language Programming, Built for Real-World Complex Tasks , date =

  26. [34]

    arXiv.org , year =

    Ministral 3 , author =. arXiv.org , year =

  27. [35]

    International Conference on Machine Learning , pages =

    Reinforcement learning for optimized trade execution , author =. International Conference on Machine Learning , pages =. 2006 , journal =. doi:10.1145/1143844.1143929 , publisher =

  28. [36]

    arXiv.org , year =

    xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning , author =. arXiv.org , year =

  29. [37]

    The Journal of Finance , volume =

    Pareto optimality and competition , author =. The Journal of Finance , volume =. 1981 , publisher =

  30. [38]

    Neural Information Processing Systems , volume =

    Trademaster: A holistic quantitative trading platform empowered by reinforcement learning , author =. Neural Information Processing Systems , volume =. 2023 , doi =

  31. [39]

    International Conferences on Computing Advancements , pages =

    Syed, Toqeer Ali and Alshahrani, Abdulaziz and Ullah, Ali and Akarma, Ali and Khan, Sohail and Nauman, Muhammad and Jan, Salman , year =. International Conferences on Computing Advancements , pages =. doi:10.1109/ICCA66035.2025.11430760 , publisher =

  32. [40]

    arXiv preprint arXiv:2507.20534 , year =

    Kimi k2: Open agentic intelligence , author =. arXiv preprint arXiv:2507.20534 , year =

  33. [41]

    Efficient

    Wang, Ningning and Hu, Xavier and Liu, Pai and Zhu, He and Hou, Yue and Huang, Heyuan and Zhang, Shengyu and Yang, Jian and Liu, Jiaheng and Zhang, Ge and others , year =. Efficient. arXiv.org , doi =

  34. [42]

    Applied and Computational Engineering , year =

    Financial analysis: Intelligent financial data analysis system based on llm-rag , author =. Applied and Computational Engineering , year =. doi:10.54254/2755-2721/2025.22221 , publisher =

  35. [43]

    Harnessing the

    Wang, Rui and Wang, Hongru and Xue, Boyang and Pang, Jianhui and Liu, Shudong and Chen, Yi and Qiu, Jiahao and Wong, Derek Fai and Ji, Heng and Wong, Kam-Fai , year =. Harnessing the. arXiv.org , doi =

  36. [44]

    Handbook of industrial organization , volume =

    Transaction cost economics , author =. Handbook of industrial organization , volume =. 2018 , publisher =

  37. [45]

    arXiv.org , year =

    Bloomberggpt: A large language model for finance , author =. arXiv.org , year =

  38. [46]

    arXiv.org , doi =

    Xiao, Yijia and Sun, Edward and Luo, Di and Wang, Wei , year =. arXiv.org , doi =

  39. [47]

    Neural Information Processing Systems , volume =

    Pixiu: A comprehensive benchmark, instruction dataset and large language model for finance , author =. Neural Information Processing Systems , volume =. 2023 , doi =

  40. [48]

    ACM Transactions on Management Information Systems , volume =

    Designing heterogeneous llm agents for financial sentiment analysis , author =. ACM Transactions on Management Information Systems , volume =. 2024 , publisher =

  41. [49]

    arXiv.org , year =

    FinRobot: an open-source AI agent platform for financial applications using large language models , author =. arXiv.org , year =

  42. [50]

    AAAI Conference on Artificial Intelligence , year =

    BAMAS: Structuring Budget-Aware Multi-Agent Systems , author =. AAAI Conference on Artificial Intelligence , year =

  43. [51]

    arXiv preprint arXiv:2505.09388 , year =

    Qwen3 technical report , author =. arXiv preprint arXiv:2505.09388 , year =

  44. [52]

    International Conference on Learning Representations , year =

    Metamath: Bootstrap your own mathematical questions for large language models , author =. International Conference on Learning Representations , year =

  45. [53]

    Neural Information Processing Systems , volume =

    Fincon: A synthesized llm multi-agent system with conceptual verbal reinforcement for enhanced financial decision making , author =. Neural Information Processing Systems , volume =. 2024 , doi =

  46. [54]

    IEEE Transactions on Big Data , year =

    Finmem: A performance-enhanced llm trading agent with layered memory and character design , author =. IEEE Transactions on Big Data , year =

  47. [55]

    International Conference on Learning Representations , year =

    Cut the crap: An economical communication pipeline for llm-based multi-agent systems , author =. International Conference on Learning Representations , year =

  48. [56]

    arXiv.org , year =

    FinWorld: An All-in-One Open-Source Platform for End-to-End Financial AI Research and Deployment , author =. arXiv.org , year =

  49. [57]

    arXiv.org , year =

    FR-LUX: Friction-Aware, Regime-Conditioned Policy Optimization for Implementable Portfolio Management , author =. arXiv.org , year =

  50. [58]

    arXiv.org , doi =

    Zhao, Tianjiao and Lyu, Jingrao and Jones, Stokes and Garber, Harrison and Pasquali, Stefano and Mehta, Dhagash , year =. arXiv.org , doi =

  51. [59]

    arXiv.org , doi =

    GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models , author =. arXiv.org , doi =. 2025 , eprint =

  52. [60]

    arXiv.org , year =

    A survey of financial ai: Architectures, advances and open challenges , author =. arXiv.org , year =

  53. [61]

    AAAI Conference on Artificial Intelligence , volume =

    Training on the benchmark is not all you need , author =. AAAI Conference on Artificial Intelligence , volume =. 2024 , journal =

  54. [62]

    QFINANCE Calculation Toolkit , volume =

    The sharpe ratio , author =. QFINANCE Calculation Toolkit , volume =. 1994 , publisher =

  55. [63]

    Vicinagearth , volume =

    A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges , author =. Vicinagearth , volume =. 2024 , publisher =

  56. [64]

    arXiv.org , year =

    Large language model agent in financial trading: A survey , author =. arXiv.org , year =

  57. [65]

    arXiv.org , year =

    Economic Evaluation of LLMs , author =. arXiv.org , year =

  58. [66]

    arXiv.org , year =

    Plan and Budget: Effective and Efficient Test-Time Scaling on Large Language Model Reasoning , author =. arXiv.org , year =

  59. [67]

    Financial Analysts Journal , volume =

    Determinants of portfolio performance , author =. Financial Analysts Journal , volume =. 1986 , publisher =

  60. [68]

    Financial Analysts Journal , volume =

    Determinants of portfolio performance II: An update , author =. Financial Analysts Journal , volume =. 1991 , publisher =

  61. [69]

    The Journal of Financial Data Science , doi =

    Active portfolio management , author =. The Journal of Financial Data Science , doi =. 2000 , publisher =

  62. [70]

    Financial Analysts Journal , volume =

    Performance attribution: Measuring dynamic allocation skill , author =. Financial Analysts Journal , volume =. 2010 , publisher =

  63. [71]

    The Journal of Business , volume =

    Does it all add up? Benchmarks and the compensation of active portfolio managers , author =. The Journal of Business , volume =. 1997 , publisher =

  64. [72]

    Applied Mathematical Finance , volume =

    Outperformance and tracking: Dynamic asset allocation for active and passive portfolio management , author =. Applied Mathematical Finance , volume =. 2018 , publisher =

  65. [73]

    Journal of Risk and Financial Management , volume =

    Portfolio optimization constrained by performance attribution , author =. Journal of Risk and Financial Management , volume =. 2021 , publisher =

  66. [74]

    Financial Analysts Journal , volume =

    Portfolio constraints and the fundamental law of active management , author =. Financial Analysts Journal , volume =. 2002 , publisher =

  67. [75]

    2006 , publisher =

    Global asset management and performance attribution , author =. 2006 , publisher =

  68. [76]

    International Conference on Machine Learning , year =

    Deep functional factor models: forecasting high-dimensional functional time series via Bayesian nonparametric factorization , author =. International Conference on Machine Learning , year =

  69. [77]

    arXiv preprint arXiv:2403.06779 , year =

    From factor models to deep learning: Machine learning in reshaping empirical asset pricing , author =. arXiv preprint arXiv:2403.06779 , year =

  70. [78]

    Conference on Empirical Methods in Natural Language Processing , pages =

    Reasoning in token economies: budget-aware evaluation of LLM reasoning strategies , author =. Conference on Empirical Methods in Natural Language Processing , pages =. 2024 , journal =. doi:10.48550/arXiv.2406.06461 , publisher =

  71. [79]

    arXiv preprint arXiv:2605.09104 , year =

    Token Economics for LLM Agents: A Dual-View Study from Computing and Economics , author =. arXiv preprint arXiv:2605.09104 , year =

  72. [80]

    Advances in Neural Information Processing Systems 36 , volume =

    Reflexion: Language agents with verbal reinforcement learning, 2023 , author =. Advances in Neural Information Processing Systems 36 , volume =. 2023 , pages =. doi:10.52202/075280-0377 , publisher =

  73. [81]

    Conference on Empirical Methods in Natural Language Processing , pages =

    AgentDiagnose: An Open Toolkit for Diagnosing LLM Agent Trajectories , author =. Conference on Empirical Methods in Natural Language Processing , pages =. 2025 , journal =. doi:10.18653/v1/2025.emnlp-demos.15 , publisher =

  74. [82]

    arXiv preprint arXiv:2510.04550 , year =

    TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use , author =. arXiv preprint arXiv:2510.04550 , year =

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.