REVIEW 2 cited by
NL2Formula: Generating Spreadsheet Formulas from Natural Language Queries
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Writing formulas on spreadsheets, such as Microsoft Excel and Google Sheets, is a widespread practice among users performing data analysis. However, crafting formulas on spreadsheets remains a tedious and error-prone task for many end-users, particularly when dealing with complex operations. To alleviate the burden associated with writing spreadsheet formulas, this paper introduces a novel benchmark task called NL2Formula, with the aim to generate executable formulas that are grounded on a spreadsheet table, given a Natural Language (NL) query as input. To accomplish this, we construct a comprehensive dataset consisting of 70,799 paired NL queries and corresponding spreadsheet formulas, covering 21,670 tables and 37 types of formula functions. We realize the NL2Formula task by providing a sequence-to-sequence baseline implementation called fCoder. Experimental results validate the effectiveness of fCoder, demonstrating its superior performance compared to the baseline models. Furthermore, we also compare fCoder with an initial GPT-3.5 model (i.e., text-davinci-003). Lastly, through in-depth error analysis, we identify potential challenges in the NL2Formula task and advocate for further investigation.
Forward citations
Cited by 2 Pith papers
-
Tabularis Formatus: Predictive Formatting for Tables
Tafo, a neuro-symbolic system, predicts spreadsheet conditional-formatting rules including colors with no user input, and its authors report it matches user-applied formatting better than all tested baselines.
-
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
A generated benchmark for Excel runtime-error repair and a baseline LLM evaluation, but with weak human-LLM judge agreement.
Discussion (0). Continue with ORCID to comment.