Pith. sign in

REVIEW 3 major objections 5 minor 42 references

FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read FinRank, a benchmark over 10-K and 10-Q filings, shows that hand-curated hard negatives cut every tested system's pairwise ranking accuracy by 13.0 to 20.5 percentage points compared with random negatives.

desk verdict FinRank is a carefully built, genuinely useful benchmark whose novel asset—per-question curated hard negatives—is real, but its headline hardness premium rests on single-annotator labels, so treat the 13–20.5 pt gap as provisional until a double-annotation check lands. read the letter →

arxiv 2608.07400 v1 pith:7IULBE5K submitted 2026-08-07 cs.AI cs.DBecon.GNq-fin.ECq-fin.GN

classification cs.AIcs.DBecon.GNq-fin.ECq-fin.GN
keywords financialquestionansweringinformationretrievalhardnegativesSECfilingsbenchmarkretrieval-augmentedgenerationevidencegrounding10-Kand10-Q
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

FinRank is a financial question-answering benchmark built from 1,185 manually written question–answer pairs over the 10-K and 10-Q filings of 22 companies. Its central claim is that the hardest part of financial QA over SEC filings is not composing an answer but identifying the correct evidence, because near-identical disclosures recur across companies, reporting periods, and filing types. To make that measurable, FinRank releases, for every question, gold supporting passages and human-curated hard negatives—plausible but wrong passages drawn mostly from competitors' filings. Baselines on the benchmark show the point: every evaluated model loses 13.0 to 20.5 percentage points of pairwise ranking accuracy when random negatives are replaced with the curated ones, and even a 7-billion-parameter embedder finds only 44.8% of gold evidence in its top ten. The paper argues that existing financial QA benchmarks, which score answer correctness over supplied snippets, miss this provenance-sensitive failure mode entirely.

What carries the argument

The mechanism that carries the argument is the hard-negative contrast embedded in the benchmark design: for each record $r$, the in-record candidate set $L_r = G_r \cup HN_r$ (gold supporting passages plus hand-curated hard negatives) isolates reranking from first-stage recall, while the global pooled corpus $C$ (5,230 deduplicated passages) defines first-stage retrieval. The decisive quantity is the hardness premium—the drop in pairwise accuracy on (positive, hard-negative) pairs compared with (positive, random-negative) pairs—which is what shows the curated distractors are genuinely confusable and justifies shipping them as a first-class benchmark asset.

What would settle it

Re-annotate a stratified random sample of about 200 records with independent financial analysts and adjudicate disagreements; if a substantial share of gold supporting passages are judged irrelevant or a substantial share of hard negatives are judged relevant, the reported hardness premium would be partly an artifact of label noise rather than model confusion.

Watch

Extended reading notes

Core claim

FinRank claims to be the first financial QA benchmark to release, for every question, a dedicated set of human-selected, semantically confusable hard-negative passages, and to score discrimination against curated versus random distractors directly. The benchmark comprises 1,185 records over 10-K and 10-Q filings of 22 companies; each record pairs a question and reference answer with gold supporting passages, rich metadata, and on average 5.08 hard negatives. Executed baselines show that ranking within a curated pool of confusable filing passages is hard: the strongest evaluated system, a 7B instruction-tuned embedder, reaches 44.8% Recall@10 on the 5,230-passage pooled corpus, sub-billion-parameter encoders gain at most 3.5 points over BM25, and a finance-adapted embedder trails BM25 by 9.7 points. In the per-record reranking task, pairwise accuracy falls 13.0–20.5 percentage points when random negatives are replaced with curated hard negatives, with 79.9% of the hard negatives drawn from same-industry, different-company filings. The authors conclude that evidence discrimination rather than answer composition is the primary bottleneck in financial QA over regulatory filings.

Load-bearing premise

The load-bearing premise is that the gold labels are correct—that each question's supporting passages really support the answer and each hard negative really is non-supporting—since records were single-authored with only sampled review and no formal inter-annotator agreement, and 442 hard negatives are byte-identical to other records' supporting passages.

Editorial extensions

If this is right

  • Retrieval, reranking, and hard-negative discrimination become separately measurable tasks, so a system that produces a plausible answer from the wrong evidence is exposed at the ranking stage rather than masked by answer-correctness metrics.
  • Existing financial QA benchmarks that score numerical or end-to-end correctness over supplied snippets likely miss the dominant failure mode of provenance-sensitive filing questions.
  • Metadata pre-filtering by ticker, year, and document type removes 92.9% of hard negatives and lifts BM25 Recall@10 from 32.1 to 55.0, yet still misses nearly half of gold evidence at k=10, so within-filing semantic discrimination remains a real bottleneck.
  • Aggregate scores on FinRank primarily reflect the dominant strata (10-K, 2025, qualitative, pharmaceutical and oil-and-gas records), so per-stratum reporting is needed to avoid hiding weakness on 10-Q, multi-passage, and quantitative questions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is using FinRank's curated hard negatives as training data: the literature it cites shows hard negatives improve dense retrievers, so the benchmark's value may be as much in training as in evaluation.
  • Because gold labels are single-annotator with sampled review and no inter-annotator agreement statistic, part of the measured hardness gap could be label noise; a stratified double-annotation study would show whether the gap shrinks once labels are adjudicated.
  • The pooled corpus $C$ contains only annotator-selected positives and curated distractors, so FinRank scores are not directly comparable to full-document retrieval; an exhaustive filing-level chunking pass would bridge the benchmark to production RAG settings.
  • The 10-Q records are all from oil and gas and predominantly 2025, so document-type gaps are confounded with sector, year, and annotator; controlled splits would be needed before attributing the 10-Q deficit to document type.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces FinRank, a benchmark of 1185 question–answer records over SEC 10-K and 10-Q filings from 22 companies. Each record includes a reference answer, gold supporting passages, and human-curated hard negatives drawn from confusable disclosures in other filings, other periods, or other sections. The benchmark defines separate retrieval, reranking, and hard-negative discrimination tasks, and the paper reports baselines spanning TF-IDF, BM25, dense encoders, and a cross-encoder. The headline empirical result is that replacing random negatives with the curated hard negatives lowers pairwise ranking accuracy by 13.0–20.5 percentage points across all evaluated models (Section 7.3, Figure 3B). The release includes normalization scripts, split files, a repair log, and a hard-negative taxonomy, and the paper is explicit that the retrieval corpus is a curated pool of passages rather than full filing text.

Significance. If the gold labels hold up, FinRank fills an actual gap: it is, to my knowledge, the first financial QA benchmark to release per-question curated hard negatives and to score retrieval and reranking directly against them. The stratified metadata (topic, difficulty, reasoning type, evidence scope), the query-rewrite field, the deterministic normalization pipeline, and the reproducible baseline harness are genuine strengths, as is the candid discussion of the corpus definition's scope. The empirical baseline results are useful reference points for the community. However, the central claim that shipped hard negatives 'demonstrably degrade' ranking rests entirely on single-annotator gold labels with no inter-annotator agreement statistic, and the headline hardness premium is reported without confidence intervals. These issues must be addressed before the benchmark's main contribution can be considered established.

major comments (3)
  1. [Sections 3.5, 9, and Appendix A.10] The gold labels (supporting passages and hard negatives) are each authored by a single annotator with only sampled author review and no formal inter-annotator agreement statistic, and Appendix A.10 explicitly states that the methodology does not certify the correctness of any individual question, answer, supporting passage, or hard negative. The central claim of Section 7.3 and Figure 3(B) that curated hard negatives are 'genuinely harder' and cause a 13.0–20.5-point pair-wise accuracy drop assumes these labels are correct; if even a small fraction of hard negatives actually support their own question, the drop is inflated, and if gold passages are misidentified, the drop is mismeasured. The manuscript honestly lists this as a limitation, but the abstract and Section 7.3 present the hardness premium as an established result. Please add a stratified double-annotation study with reported agreement on a representative sample, or, failing that, rephrase the headline claim as conditional on single-annotator labels and show that the hardness premium survives on a verified subset.
  2. [Section 7.3 and Figure 3(B)] The hardness premium is reported as a single point estimate per model with no confidence intervals, and the random-negative baseline is a single draw at seed 42 (Section 6.2). The 13.0–20.5-point gaps are large, but without error bars the reader cannot judge whether the relative ordering of models by premium is stable or whether the gap could shrink under resampling. Please report bootstrap confidence intervals over records (or queries) and, ideally, repeat the random-negative sampling with multiple seeds and report the variance of the premium.
  3. [Section 3.7 and Section 7.3] The paper reports that 442 hard negatives (about 7.3%) are byte-identical to supporting passages of a different record, and notes that such passages 'may legitimately be relevant to more than one question.' The pairwise-accuracy metric in Section 7.3 nonetheless treats every hard negative as non-supporting for its own question. Because these overlaps could include passages that are actually relevant to the query, the hardness premium should be recomputed on the subset of records after removing or relabeling these overlapping hard negatives (using the released hn_taxonomy.json), or the paper should demonstrate that the premium is unchanged when this subset is excluded.
minor comments (5)
  1. [Figure 1 caption and figure body] The caption cites 'the 6021 hard negatives' while the figure itself reports 'N = 6,052'; please reconcile these counts.
  2. [Section 4 vs. Section 9] Section 4 says 88% of records come from 10-K filings, while Section 9 says 87%; Table 14 gives 1048/1185 = 88.4%, so Section 9 should be corrected.
  3. [Section 3.7] The percentage for the 442 overlapping hard negatives is given as '~7.3%'; please compute this percentage against the same total used in Figure 1, since the caption and the figure currently report different totals.
  4. [Section 7.3] The statement that 'records are averaged with equal weight' is slightly ambiguous; please clarify that pairwise accuracy is first computed within each record and then averaged across records, rather than pooling all pairs globally.
  5. [Section 3.6] The global pool C is said to contain 5230 unique passages, but the average counts of 1.96 supporting passages and 5.08 hard negatives per record imply larger raw totals before deduplication; please state the deduplication effect explicitly for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: FinRank is a benchmark release whose headline numbers are empirical measurements on released data, not derived predictions.

full rationale

FinRank's central claims are benchmark-construction claims and empirical measurements, not derived predictions. The hardness premium in Section 7.3 compares measured pairwise accuracy on curated versus random negatives; the curated negatives were hand-selected by annotators, not fitted to the evaluated models, and the paper reports no fitted parameter renamed as a prediction. The only self-citation (Mansouri et al., 2026, discussed in Section 2.1) contrasts an MCP-based approach with FinRank's retrieval setting and carries no load-bearing argument. The concern that single-annotator gold labels (Sections 3.5, 9, and Appendix A.10) could affect the hardness premium is a data-correctness and validity limitation, not circularity: the benchmark numbers are exactly reproducible from the released jsonl and baselines, and no equation or theorem derives its conclusions from its own inputs. The query-rewrite ablation is explicitly labeled an oracle upper bound, and the pooled-corpus retrieval numbers are explicitly stated not to be full-document retrieval, so there is no disguised fit or renamed known result.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No numerical parameters are fitted to data in this paper. The benchmark construction targets (30/40/30 difficulty, 40/30/30 passage scope, balanced reasoning type) are fixed ex ante in the criteria document and reported against realized deviations in Table 2, not tuned to produce the headline results. The load-bearing assumptions are about annotation correctness and evidence granularity.

assumptions (3)
  • domain assumption The reference answers and gold supporting passages in each record are correct.
    The entire benchmark presupposes label correctness; the paper states that records were not independently adjudicated and that no formal inter-annotator agreement statistic was computed (Sections 3.5 and 9).
  • domain assumption The hand-selected hard negatives are not supporting evidence for their own record's question.
    The hardness premium depends on this. The paper removes 11 duplicates within a record and flags 442 cross-record overlaps that may be legitimately relevant to another record's question (Section 3.7).
  • domain assumption Passage-level text is a sufficient representation of evidence.
    Tables and figures are represented as text, so evaluations requiring true multimodal reasoning are out of scope; the paper states this as a limitation (Section 9).

how reviews work

0 comments
Cite this review

Pith. "Pith review of FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings." pith.science (2026). https://pith.science/paper/7IULBE5K

@misc{pith2026260807400,
  author       = {Pith},
  title        = {Pith review of: FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7IULBE5K}},
  note         = {Machine review of arXiv:2608.07400}
}
read the original abstract

Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.

Figures

Figures reproduced from arXiv: 2608.07400 by the authors.

Figure 1
Figure 1. Taxonomy of the 6021 hard negatives, by the [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Composition of FinRank across the 1185 records: (a) sector, (b) difficulty, (c) reasoning type, (d) evidence [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Executed baselines on FinRank. (A) First-stage retrieval over the global pool: Recall@k for the evaluated systems. Even the strongest, a 7B instruction-tuned embedder, reaches only 44.8% Recall@10, and sub-billion￾parameter dense encoders gain little over BM25. (B) The hardness premium: pairwise ranking accuracy of each model against its curated hard negatives versus random negatives. The 13.0–20.5 point gap is the … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

42 extracted references · 28 canonical work pages

  1. [1]

    Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

    Chen, Zhiyu and Chen, Wenhu and Smiley, Charese and Shah, Sameena and Borova, Iana and Langdon, Dylan and Moussa, Reema and Beane, Matt and Huang, Ting-Hao and Routledge, Bryan and Wang, William Yang , title =. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

  2. [2]

    Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL) , year =

    Zhu, Fengbin and Lei, Wenqiang and Huang, Youcheng and Wang, Chao and Zhang, Shuo and Lv, Jiancheng and Feng, Fuli and Chua, Tat-Seng , title =. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL) , year =

  3. [3]

    Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

    Chen, Zhiyu and Li, Shiyang and Smiley, Charese and Ma, Zhiqiang and Shah, Sameena and Wang, William Yang , title =. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

  4. [4]

    arXiv preprint arXiv:2311.11944 , year =

    Islam, Pranab and Kannappan, Anand and Kiela, Douwe and Qian, Rebecca and Scherrer, Nino and Vidgen, Bertie , title =. arXiv preprint arXiv:2311.11944 , year =

  5. [5]

    arXiv preprint arXiv:2303.17564 , year =

    Wu, Shijie and Ouyang, Zheng and Papay, Michael and Nair, Sameer and Zhang, Rui and Shah, Sameena and Lewis, Patrick and Zhao, Peikun and Huang, Ting-Hao and Wang, William Yang , title =. arXiv preprint arXiv:2303.17564 , year =

  6. [6]

    arXiv preprint arXiv:2504.15800 , year =

    Choi, Chanyeol and others , title =. arXiv preprint arXiv:2504.15800 , year =

  7. [7]

    Retrieval-Augmented Generation for Knowledge-Intensive

    Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and K. Retrieval-Augmented Generation for Knowledge-Intensive. Advances in Neural Information Processing Systems (NeurIPS) , year =

  8. [8]

    Dense Passage Retrieval for Open-Domain Question Answering , booktitle =

    Karpukhin, Vladimir and O. Dense Passage Retrieval for Open-Domain Question Answering , booktitle =

Show all 42 references
  1. [9]

    Foundations and Trends in Information Retrieval , volume =

    Robertson, Stephen and Zaragoza, Hugo , title =. Foundations and Trends in Information Retrieval , volume =

  2. [10]

    Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

    Reimers, Nils and Gurevych, Iryna , title =. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

  3. [11]

    arXiv preprint arXiv:1901.04085 , year =

    Nogueira, Rodrigo and Cho, Kyunghyun , title =. arXiv preprint arXiv:1901.04085 , year =

  4. [12]

    Transactions on Machine Learning Research (TMLR) , year =

    Izacard, Gautier and Caron, Mathilde and Hosseini, Lucas and Riedel, Sebastian and Bojanowski, Piotr and Joulin, Armand and Grave, Edouard , title =. Transactions on Machine Learning Research (TMLR) , year =

  5. [13]

    and Artzi, Yoav , title =

    Zhang, Tianyi and Kishore, Varsha and Wu, Felix and Weinberger, Kilian Q. and Artzi, Yoav , title =. International Conference on Learning Representations (ICLR) , year =

  6. [14]

    Text Summarization Branches Out , year =

    Lin, Chin-Yew , title =. Text Summarization Branches Out , year =

  7. [15]

    Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

    Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy , title =. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

  8. [16]

    and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav , title =

    Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-W...

  9. [17]

    and Zettlemoyer, Luke , title =

    Joshi, Mandar and Choi, Eunsol and Weld, Daniel S. and Zettlemoyer, Luke , title =. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL) , year =

  10. [18]

    Computational Linguistics , year =

    Rashkin, Hannah and Nikolaev, Vitaly and Lamm, Matthew and Aroyo, Lora and Collins, Michael and Das, Dipanjan and Petrov, Slav and Tomar, Gaurav Singh and Turc, Iulia and Reitter, David , title =. Computational Linguistics , year =

  11. [19]

    Bohnet, Bernd and Tran, Vinh Q. and Verga, Pat and Aharoni, Roee and Andor, Daniel and Soares, Livio Baldini and Eisenschlos, Julian and Fierro, Constanza and Gor, Maor and Krishna, Kalpesh and Singh, Tania and Sun, Si-Qing and Sutton, Charles and Wieting, John and Williams, A...

  12. [20]

    Companion Proceedings of the The Web Conference 2018 , year =

    Maia, Macedo and Handschuh, Siegfried and Freitas, Andr. Companion Proceedings of the The Web Conference 2018 , year =

  13. [21]

    Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

    Shah, Raj Sanjay and Chawla, Kunal and Eidnani, Dheeraj and Shah, Agam and Du, Wendi and Chava, Sudheer and Raman, Natraj and Smiley, Charese and Chen, Jiaao and Yang, Diyi , title =. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP...

  14. [22]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =

    Reddy, Varshini and Koncel-Kedziorski, Rik and Lai, Viet Dac and Krumdick, Michael and Lovering, Charles and Tanner, Chris , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =. 2024 , eprint =

  15. [23]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    Chen, Jian and Zhou, Peilin and Hua, Yining and Loh, Yingxin and Chen, Kehui and Li, Ziyuan and Zhu, Bing and Liang, Junwei , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2024 , eprint =

  16. [24]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    Zhao, Yilun and Li, Yunxiang and Li, Chenying and Zhang, Rui , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2022 , eprint =

  17. [25]

    Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages =

    Deng, Yang and Lei, Wenqiang and Zhang, Wenxuan and Lam, Wai and Chua, Tat-Seng , title =. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages =. 2022 , eprint =

  18. [26]

    Proceedings of the 6th ACM International Conference on AI in Finance (ICAIF '25) , publisher =

    Choi, Chanyeol and Kwon, Jihoon and Lopez-Lira, Alejandro and Kim, Chaewoon and Kim, Minjae and Hwang, Juneha and Ha, Jaeseon and Choi, Hojun and Yun, Suyeol and Kim, Yongjin and Lee, Yongjae , title =. Proceedings of the 6th ACM International Conference on AI in Finance (ICAI...

  19. [27]

    Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    Strich, Jan and Isgorur, Enes Kutay and Trescher, Maximilian and Biemann, Chris and Semmann, Martin , title =. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2026 , eprint =

  20. [28]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

    Zhao, Suifeng and Jin, Zhuoran and Li, Sujian and Gao, Jun , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =. 2025 , eprint =

  21. [29]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

    Wasserman, Navve and Heinimann, Oliver and Golbari, Yuval and Zimbalist, Tal and Schwartz, Eli and Irani, Michal , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =. 2025 , eprint =

  22. [30]

    2024 , note =

    Choi, Chanyeol and Sohn, Jy-Yong and Kwon, Jihoon and Kim, Jin and Pang, Subeen and Ha, Jaeseon and Ryoo, Hoyeon and Choi, Hojun and Lee, Yongjae , title =. 2024 , note =

  23. [31]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

    Tang, Yixuan and Yang, Yi , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =. 2025 , eprint =

  24. [32]

    Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , pages =

    Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , pages =. 2023 , eprint =

  25. [33]

    Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS) , year =

    Thakur, Nandan and Reimers, Nils and R. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS) , year =. 2104.08663 , archivePrefix =

  26. [34]

    and Ahmed, Junaid and Overwijk, Arnold , title =

    Xiong, Lee and Xiong, Chenyan and Li, Ye and Tang, Kwok-Fung and Liu, Jialin and Bennett, Paul N. and Ahmed, Junaid and Overwijk, Arnold , title =. International Conference on Learning Representations (ICLR) , year =. 2007.00808 , archivePrefix =

  27. [35]

    Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages =

    Qu, Yingqi and Ding, Yuchen and Liu, Jing and Liu, Kai and Ren, Ruiyang and Zhao, Wayne Xin and Dong, Daxiang and Wu, Hua and Wang, Haifeng , title =. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langu...

  28. [36]

    Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , publisher =

    Zhan, Jingtao and Mao, Jiaxin and Liu, Yiqun and Guo, Jiafeng and Zhang, Min and Ma, Shaoping , title =. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , publisher =. 2021 , eprint =

  29. [37]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =

    Gao, Tianyu and Yen, Howard and Yu, Jiatong and Chen, Danqi , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =. 2023 , eprint =

  30. [38]

    Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations , pages =

    Es, Shahul and James, Jithin and Espinosa-Anke, Luis and Schockaert, Steven , title =. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations , pages =. 2024 , eprint =

  31. [39]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =

    Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang Wei and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =. 20...

  32. [40]

    arXiv preprint arXiv:2603.20316 , year =

    Mansouri, Sasan and Pilla, Edoardo and Wahrenburg, Mark and Woebbeking, Fabian , title =. arXiv preprint arXiv:2603.20316 , year =. 2603.20316 , archivePrefix =

  33. [41]

    Findings of the Association for Computational Linguistics: ACL 2026 , publisher =

    Zhou, Yixi and Zhang, Fan and Chen, Yu and Zhang, Haipeng and Nakov, Preslav and Xie, Zhuohan , title =. Findings of the Association for Computational Linguistics: ACL 2026 , publisher =. 2026 , doi =. 2601.06992 , archivePrefix =

  34. [42]

    Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '26) , year =

    Jiang, Yidong and Chen, Junrong and Makri, Eftychia and Chen, Jialin and Li, Peiwen and Maatouk, Ali and Tassiulas, Leandros and Brenner, Eliot and Xiang, Bing and Ying, Rex , title =. Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '2...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.