REVIEW 4 cited by
GenoTEX: An LLM Agent Benchmark for Automated Gene Expression Data Analysis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in machine learning have significantly improved the identification of disease-associated genes from gene expression datasets. However, these processes often require extensive expertise and manual effort, limiting their scalability. Large Language Model (LLM)-based agents have shown promise in automating these tasks due to their increasing problem-solving abilities. To support the evaluation and development of such methods, we introduce GenoTEX, a benchmark dataset for the automated analysis of gene expression data. GenoTEX provides analysis code and results for solving a wide range of gene-trait association problems, encompassing dataset selection, preprocessing, and statistical analysis, in a pipeline that follows computational genomics standards. The benchmark includes expert-curated annotations from bioinformaticians to ensure accuracy and reliability. To provide baselines for these tasks, we present GenoAgent, a team of LLM-based agents that adopt a multi-step programming workflow with flexible self-correction, to collaboratively analyze gene expression datasets. Our experiments demonstrate the potential of LLM-based methods in analyzing genomic data, while error analysis highlights the challenges and areas for future improvement. We propose GenoTEX as a promising resource for benchmarking and enhancing automated methods for gene expression data analysis. The benchmark is available at https://github.com/Liu-Hy/GenoTEX.
Forward citations
Cited by 4 Pith papers
-
Discovery of Disease Relationships via Transcriptomic Signature Analysis Powered by Agentic AI
Agentic AI processing of 1,384 transcriptomic disease-condition pairs yields gene- and pathway-level disease similarity networks, presented as hypotheses for comorbidity and drug repurposing.
-
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
SciToolAgent uses a knowledge graph of over 500 scientific tools to help LLMs select and chain tools, reaching 94% accuracy on a new 531-question benchmark.
-
MSA at SemEval-2025 Task 3: High Quality Weak Labeling and LLM Ensemble Verification for Multilingual Hallucination Detection
An LLM ensemble that extracts hallucinated spans and votes on them, followed by fuzzy matching, achieved top ranks in Arabic and Basque at the Mu-SHROOM shared task.
-
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.
Discussion (0). Sign in to comment.