{"id":"71aa739c-a2c4-4654-8e19-2dee13bc1d41","arxiv_id":"1906.09317","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Creates datasets and an extraction model that identifies task-dataset-metric-score information in NLP papers to support automatic leaderboard construction.","lead":"The paper builds two new datasets and a framework called TDMS-IE to automatically pull task, dataset, metric, and numeric score mentions out of NLP research papers. A smart generalist might care because this is an early attempt to automate the creation of scientific leaderboards that track which methods actually work best.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Performance gains shown only on author-created datasets; no test of whether extracted tuples support accurate leaderboards without human checks","rationale":"The reader's weakest assumption is exactly the point at which the experimental claim fails to transfer to the stated application. Because the full text was not needed to locate this gap, the UNVERDICTED verdict with LOW confidence remains appropriate.","tokens_in":1578,"tokens_out":331,"duration_ms":26866,"concrete_test":"Sample 100 papers from ACL Anthology 2015-2018 not used in the original datasets; run the released model to extract tuples; have two independent annotators score precision of the (task, dataset, metric, score) 4-tuples against the paper text; if micro-averaged precision falls below 65 % the outperformance result does not yet support the leaderboard-construction claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline experimental claim is that TDMS-IE outperforms baselines by a large margin on the two newly created datasets. This claim is load-bearing for the broader goal of automatic leaderboard construction only if (a) the datasets reflect the distribution of result-reporting styles across real NLP papers and (b) the extracted (T,D,M,S) tuples are sufficiently accurate that downstream leaderboards would not require substantial manual correction. The paper provides no inter-annotator agreement figures, no external validation of the annotation scheme, and no end-to-end evaluation measuring leaderboard fidelity on held-out papers. Therefore the reported margin could be an artifact of the authors' own annotation conventions rather than evidence of deployable extraction quality.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces two new datasets for extracting (task, dataset, metric, score) tuples from NLP papers and presents the TDMS-IE framework for this extraction task. Experiments on the author-created datasets show that TDMS-IE outperforms several baselines by a large margin, positioned as an initial step toward automatic scientific leaderboard construction.","tokens_in":1720,"tokens_out":354,"duration_ms":18134,"significance":"If the extraction quality generalizes and the tuples prove sufficiently accurate for leaderboard use, the work could reduce manual effort in tracking NLP results. The current evaluation, however, provides no evidence that the reported gains support deployable leaderboards without substantial human correction.","major_comments":[{"comment":"Dataset construction: no inter-annotator agreement is reported for either of the two new datasets, leaving the reliability of the gold annotations used to train and evaluate TDMS-IE unquantified.","section":"Dataset construction"},{"comment":"Experiments: the evaluation contains no end-to-end test measuring how accurately the extracted tuples reconstruct leaderboards on held-out papers; the large-margin claim therefore does not yet demonstrate that downstream leaderboards would require only minimal manual verification.","section":"Experiments"},{"comment":"Abstract and Experiments: performance is reported exclusively on author-annotated data with no external validation set or papers using different result-reporting conventions, so it remains unclear whether the margin reflects genuine extraction robustness rather than annotation conventions specific to the authors.","section":"Abstract and Experiments"}],"minor_comments":[{"comment":"The abstract would benefit from stating the sizes and annotation guidelines of the two datasets.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"Thank you for the constructive feedback on our manuscript. We address each major comment below, indicating planned revisions where appropriate.","responses":[{"response":"We agree that inter-annotator agreement should be reported to quantify annotation reliability. We will add a second annotator for a subset of the papers, compute agreement metrics such as Cohen's kappa, and include these results in the revised manuscript.","revision_made":"yes","referee_comment":"Dataset construction: no inter-annotator agreement is reported for either of the two new datasets, leaving the reliability of the gold annotations used to train and evaluate TDMS-IE unquantified."},{"response":"The paper explicitly frames TDMS-IE as an initial step and does not claim the results support fully deployable leaderboards without human oversight. The evaluation targets extraction accuracy. We will add a limitations discussion on the absence of end-to-end leaderboard reconstruction and note that such an evaluation would require further annotation effort beyond the current scope.","revision_made":"partial","referee_comment":"Experiments: the evaluation contains no end-to-end test measuring how accurately the extracted tuples reconstruct leaderboards on held-out papers; the large-margin claim therefore does not yet demonstrate that downstream leaderboards would require only minimal manual verification."},{"response":"The datasets follow a documented annotation protocol, and evaluation uses held-out papers from the same collection. External validation on independently created data is not available for these new resources. We will revise the abstract and add a limitations section clarifying the evaluation scope and the potential influence of annotation conventions.","revision_made":"partial","referee_comment":"Abstract and Experiments: performance is reported exclusively on author-annotated data with no external validation set or papers using different result-reporting conventions, so it remains unclear whether the margin reflects genuine extraction robustness rather than annotation conventions specific to the authors."}],"tokens_in":1207,"tokens_out":416,"duration_ms":21754,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The one thing to take away is that the authors built two fresh datasets and a dedicated extraction system aimed at automatic NLP leaderboards, then showed their model beats baselines on those sets. That combination is the actual new piece rather than a routine IE extension. The framing around tracking empirical progress is also direct and practical, which is better than papers that just add another layer to existing scientific IE without a clear use case. They get credit for releasing the data and for targeting a real pain point in the field. The evaluation, however, stays narrow. All reported gains come from the datasets the authors annotated themselves, with no inter-annotator agreement numbers, no comparison to existing leaderboards, and no end-to-end test of whether the extracted tuples would require heavy human fixes before they could be trusted. The stress-test note is right on this: the margin could simply reflect the authors' labeling conventions rather than extraction quality that generalizes. Without those checks, the claim that this advances automatic leaderboard construction rests on an untested assumption about representativeness and downstream fidelity. The work is aimed at people who build tools for scientific information extraction or who maintain leaderboards in fast-moving areas like NLP. A reader who needs new annotated data for tuple extraction tasks would get concrete value from the resources. It is worth sending to peer review because the datasets are new and the goal is stated clearly; referees can then press on annotation quality and ask for the missing validation steps. I would not cite it in its current form, but the resources alone justify a proper review.","headline":"The paper releases two new datasets and a TDMS-IE model for pulling task-dataset-metric-score tuples from NLP papers, but the experiments stay inside the authors' own annotations and never check whether the output would produce usable leaderboards.","tokens_in":2198,"tokens_out":400,"would_cite":false,"duration_ms":17504,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Standard NLP information-extraction pipeline with transformer entailment models; no contact with J-cost, distinction forcing or RS structural theorems","alignment":"orthogonal","rationale":"The paper's machinery (DocTAET/SC representations, DocTAET-TDM and SC-DM transformer classifiers trained on NLP-TDMS/ARC-PDN, NLI-style hypothesis matching for TDM triples) is conventional supervised IE for leaderboard construction. It contains none of the RS primitives (single distinction, J(x)=½(x+x⁻¹)−1, φ-ladder, 8-tick periodicity, parameter-free constant derivation) and makes no claims about recognition cost or reality_from_one_distinction. Hence orthogonal to every module in the RS corpus.","tokens_in":51455,"confidence":"high","tokens_out":169,"duration_ms":6397,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A framework extracts tasks, datasets, metrics and scores from NLP papers to enable automatic leaderboard construction.","keywords":["information extraction","NLP leaderboards","scientific summarization","task dataset metric score","natural language processing","automatic evaluation","research tracking"],"falsifier":"Manual expert review of leaderboards built from the model's extractions on a fresh collection of NLP papers reveals frequent incorrect or incomplete task-dataset-metric-score tuples.","tokens_in":2494,"feed_emoji":"📊","tokens_out":408,"duration_ms":16155,"temperature":0.7,"pith_summary":"The paper tackles the growing difficulty of tracking results across the expanding set of NLP tasks and datasets by building an automated extraction system. The authors created two dedicated datasets and introduced the TDMS-IE framework to pull out task-dataset-metric-score tuples directly from research papers. Experiments show the model beats several baselines by a large margin. If the approach holds, it would let the community maintain up-to-date leaderboards without relying solely on manual curation as publication volume increases.","feed_headline":"Framework pulls tasks, datasets, metrics, scores from NLP papers","feed_subtitle":"TDMS-IE and two new datasets outperform baselines as a step toward automatic scientific result leaderboards","key_machinery":"TDMS-IE framework for extracting task-dataset-metric-score tuples from scientific papers.","core_discovery":"The authors build two datasets and develop a framework (TDMS-IE) aimed at automatically extracting task, dataset, metric and score from NLP papers, towards the automatic construction of leaderboards. Experiments show that their model outperforms several baselines by a large margin. Their model is a first step towards automatic leaderboard construction, e.g., in the NLP domain.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["TDMS-IE extracts tasks datasets metrics scores from NLP papers","Two datasets support TDMS-IE for NLP leaderboard construction","TDMS-IE outperforms baselines in extracting paper results data","Framework targets automatic leaderboard building from NLP papers","TDMS-IE parses NLP papers for tasks datasets metrics and scores"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The two newly created datasets are representative enough of real NLP papers and the extracted tuples can be directly used to construct accurate leaderboards without substantial human verification.","fun_headline_variants_meta":{"raw":{"variants":["TDMS-IE extracts tasks datasets metrics scores from NLP papers","Two datasets support TDMS-IE for NLP leaderboard construction","TDMS-IE outperforms baselines in extracting paper results data","Framework targets automatic leaderboard building from NLP papers","TDMS-IE parses NLP papers for tasks datasets metrics and scores"]},"model":"grok-4.3","cost_usd":0.00494,"raw_usage":{"total_tokens":2364,"prompt_tokens":561,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":49399500,"prompt_tokens_details":{"text_tokens":561,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1724,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":561,"tokens_out":79,"duration_ms":11226,"temperature":1.0,"reasoning_tokens":1724,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T18:35:14.755583+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Manual expert review of leaderboards built from the model's extractions on a fresh collection of NLP papers reveals frequent incorrect or incomplete task-dataset-metric-score tuples.","supporting_citations":[],"review_version":1}