Pith. sign in

REVIEW 13 cited by

PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.05443 v1 pith:FSYRE356 submitted 2023-06-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords financialdatatasksbenchmarkinstructionevaluationllmscritical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although large language models (LLMs) has shown great performance on natural language processing (NLP) in the financial domain, there are no publicly available financial tailtored LLMs, instruction tuning datasets, and evaluation benchmarks, which is critical for continually pushing forward the open-source development of financial artificial intelligence (AI). This paper introduces PIXIU, a comprehensive framework including the first financial LLM based on fine-tuning LLaMA with instruction data, the first instruction data with 136K data samples to support the fine-tuning, and an evaluation benchmark with 5 tasks and 9 datasets. We first construct the large-scale multi-task instruction data considering a variety of financial tasks, financial document types, and financial data modalities. We then propose a financial LLM called FinMA by fine-tuning LLaMA with the constructed dataset to be able to follow instructions for various financial tasks. To support the evaluation of financial LLMs, we propose a standardized benchmark that covers a set of critical financial tasks, including five financial NLP tasks and one financial prediction task. With this benchmark, we conduct a detailed analysis of FinMA and several existing LLMs, uncovering their strengths and weaknesses in handling critical financial tasks. The model, datasets, benchmark, and experimental results are open-sourced to facilitate future research in financial AI.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    A ticker-identity-free Chronos-augmented encoder plus MoE PPO policy and 76-parameter LoRA personalization claims +2.93% 14-day alpha vs equal-weight while serving six investment objectives.

  2. FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

    cs.CR 2026-03 unverdicted novelty 6.5 of 10

    Page-granular Flex-Mem and switchable Flex-NPU cut TrustZone LLM TTFT by ~10× vs a CMA strawman and ~2.4× vs a pipelined secure-NPU strawman on RK3588.

  3. Can Large Language Models Execute Parent Orders?

    cs.CE 2026-07 conditional novelty 6.0 of 10

    PACE, a planner–executor LLM framework, beats TWAP, Almgren–Chriss, and ML execution baselines by up to ~0.65 bps on Shenzhen parent orders without task-specific training.

  4. Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Serialization format and in-context examples change both accuracy and gender fairness of LLM loan approvals, with finance-tuned models often showing larger disparities.

  5. Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval

    cs.IR 2026-07 conditional novelty 5.0 of 10

    For multilingual financial exam questions, no single model or reasoning strategy wins across languages: Qwen3-14B regresses 22.9 points on English, Llama-3.1-8B wins Greek, and chain-of-thought collapses Greek accurac...

  6. FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Fine-tuning small language models on synthetic financial errors yields high detection and editing scores, but the evaluation is limited to synthetic data from the same pipeline.

  7. Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Agentar-Fin-R1, an 8B and 32B financial LLM family, reports top scores on FinEval, FinanceIQ, and a new Finova benchmark while keeping general reasoning near its Qwen3 base.

  8. FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios

    cs.CE 2025-07 conditional novelty 5.0 of 10

    A four-agent LLM pipeline trained with role-specific data improves human preference on comprehensive Chinese financial analysis tasks.

  9. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  10. PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading

    cs.CL 2025-06 reject novelty 5.0 of 10

    MAS traders using Reddit sentiment from PulseReddit beat traditional baselines in the reported bull-market backtests, but the gains are small, most runs lose money, and the evaluation has critical flaws.

  11. Interpretable LLMs for Credit Risk: A Systematic Review and Taxonomy

    q-fin.RM 2025-06 conditional novelty 4.0 of 10

    A systematic review and taxonomy that organizes LLM-based credit risk research by model architecture, data modality, explainability mechanism, and application domain.

  12. Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications

    cs.CL 2025-07 reject novelty 3.0 of 10

    FinGPT matches GPT-4 on financial sentiment and headline classification, lags on QA and NER, and shows a bullish bias in stock movement prediction.

  13. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools