Synthetic network generators preserve cross-flow correlations enabling source-level membership inference, shown via the TraceBleed attack across five datasets and six generators.
Realtabformer: Generating realistic relational and tabular data using transformers
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 3polarities
background 3representative citing papers
ICL in LLMs shows a sharp ceiling on categorical distributions for high-cardinality tabular data, failing to reproduce rare classes despite examples, while numerical fidelity improves.
HH-SAE factorizes manifolds into nested contextual (L0), atomic (f1), and compository (f2) tiers, achieving 0.9156 cross-domain zero-shot AUC in fraud detection and +9.9% AUPRC lift in steered synthesis.
Concordia aligns synthetic table generation with federated validation utility via client-side utility scorers and group-relative policy optimization to improve LLM adaptation on non-IID tabular tasks.
MT-MIA uses heterogeneous graph neural networks under a No-Box model to expose user-level membership leakage in synthetic relational data that single-table attacks underestimate.
NetNomos is a multi-stage framework that extracts, filters, and enforces first-order logic rules in generative ML models for networking tasks including telemetry imputation, traffic forecasting, and synthetic trace generation.
LLM-TabLogic extracts inter-column logical constraints using LLMs and conditions a score-based latent diffusion model on them to generate synthetic tabular data that preserves those relationships.
Permutation-invariant fine-tuning (PI-FT) randomizes field order and applies dropout during embedding model training to eliminate sensitivity to serialization order, reducing order-change penalty from 7.4 to 0.2 nDCG@10 on a generated multilingual DevDataBench while outperforming zero-shot baselines
Semantically invariant row and column permutations in tables can cause LLMs to output incorrect answers, and a gradient-based attack called ATP efficiently finds such permutations that degrade performance across many models.
RDDG is an in-context learning system with dynamic guidance and automatic quality feedback that synthesizes high-fidelity relational data to improve imbalanced classification.
Seq. RC-TGAN adds a spectral envelope loss to RC-TGAN and uses VGM discretization plus simulated benchmarks with known envelopes to generate relational time series that better match frequency-domain features.
PSyGenTAB is a constrained-optimization framework that generates privacy-preserving synthetic clinical tabular data while preserving clinical relationships and downstream model performance.
citing papers explorer
-
Cross-Flow Correlations Survive Synthesis: Measuring Source-Level Privacy Leakage in Synthetic Network Traces
Synthetic network generators preserve cross-flow correlations enabling source-level membership inference, shown via the TraceBleed attack across five datasets and six generators.
-
Categorical Prior Lock-in: Why In-Context Learning Fails for Structured Data
ICL in LLMs shows a sharp ceiling on categorical distributions for high-cardinality tabular data, failing to reproduce rare classes despite examples, while numerical fidelity improves.
-
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
HH-SAE factorizes manifolds into nested contextual (L0), atomic (f1), and compository (f2) tiers, achieving 0.9156 cross-domain zero-shot AUC in fraud detection and +9.9% AUPRC lift in steered synthesis.
-
Concordia: Self-Improving Synthetic Tables for Federated LLMs
Concordia aligns synthetic table generation with federated validation utility via client-side utility scorers and group-relative policy optimization to improve LLM adaptation on non-IID tabular tasks.
-
Finding Connections: Membership Inference Attacks for the Multi-Table Synthetic Data Setting
MT-MIA uses heterogeneous graph neural networks under a No-Box model to expose user-level membership leakage in synthetic relational data that single-table attacks underestimate.
-
Making Logic a First-Class Citizen in Generative ML for Networking
NetNomos is a multi-stage framework that extracts, filters, and enforces first-order logic rules in generative ML models for networking tasks including telemetry imputation, traffic forecasting, and synthetic trace generation.
-
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion
LLM-TabLogic extracts inter-column logical constraints using LLMs and conditions a score-based latent diffusion model on them to generate synthetic tabular data that preserves those relationships.
-
Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval
Permutation-invariant fine-tuning (PI-FT) randomizes field order and applies dropout during embedding model training to eliminate sensitivity to serialization order, reducing order-change penalty from 7.4 to 0.2 nDCG@10 on a generated multilingual DevDataBench while outperforming zero-shot baselines
-
The Power of Order: Fooling LLMs with Adversarial Table Permutations
Semantically invariant row and column permutations in tables can cause LLMs to output incorrect answers, and a gradient-based attack called ATP efficiently finds such permutations that degrade performance across many models.
-
Self-Reinforcing Controllable Synthesis of Rare Relational Data via Bayesian Calibration
RDDG is an in-context learning system with dynamic guidance and automatic quality feedback that synthesizes high-fidelity relational data to improve imbalanced classification.
-
Sequential RC-TGAN: Generating Relational Time Series with Spectral Envelope Loss
Seq. RC-TGAN adds a spectral envelope loss to RC-TGAN and uses VGM discretization plus simulated benchmarks with known envelopes to generate relational time series that better match frequency-domain features.
-
PSyGenTAB: A Privacy-Preserving Framework for Synthetic Clinical Tabular Data Generation via Constrained Optimization
PSyGenTAB is a constrained-optimization framework that generates privacy-preserving synthetic clinical tabular data while preserving clinical relationships and downstream model performance.