REVIEW 6 cited by
Synthetic Data Applications in Finance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Synthetic data has made tremendous strides in various commercial settings including finance, healthcare, and virtual reality. We present a broad overview of prototypical applications of synthetic data in the financial sector and in particular provide richer details for a few select ones. These cover a wide variety of data modalities including tabular, time-series, event-series, and unstructured arising from both markets and retail financial applications. Since finance is a highly regulated industry, synthetic data is a potential approach for dealing with issues related to privacy, fairness, and explainability. Various metrics are utilized in evaluating the quality and effectiveness of our approaches in these applications. We conclude with open directions in synthetic data in the context of the financial domain.
Forward citations
Cited by 6 Pith papers
-
RaMark: Radioactive Watermarking for Generated Tabular Data
A sinusoidal dependency embedded as part of the tabular distribution remains detectable after generative retraining and data-modification attacks while utility is preserved.
-
Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data
Static-distribution fidelity is a poor proxy for temporal fidelity in synthetic sequential tabular data; measuring timestamp, trajectory, cross-sectional, and relational structure over time changes model rankings.
-
Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
Verifier-filtered synthetic retraining improves linear-regression estimates in the short term but converges to the verifier's knowledge center, so sustained improvement requires an unbiased verifier.
-
Ensembling Membership Inference Attacks Against Tabular Generative Models
No single membership inference attack dominates across tabular generative models, and unsupervised ensembles of attacks achieve better average rankings.
-
Synthetic CVs To Build and Test Fairness-Aware Hiring Tools
A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.
-
InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior
InvestAlign fine-tunes LLMs on datasets generated from closed-form solutions of a simplified optimal investment problem, improving alignment with real investor decisions under herd behavior by 45-61% in MSE.
Discussion (0). Continue with ORCID to comment.