REVIEW 2 cited by
Koala: An Index for Quantifying Overlaps with Pre-training Corpora
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In very recent years more attention has been placed on probing the role of pre-training data in Large Language Models (LLMs) downstream behaviour. Despite the importance, there is no public tool that supports such analysis of pre-training corpora at large scale. To help research in this space, we launch Koala, a searchable index over large pre-training corpora using compressed suffix arrays with highly efficient compression rate and search support. In its first release we index the public proportion of OPT 175B pre-training data. Koala provides a framework to do forensic analysis on the current and future benchmarks as well as to assess the degree of memorization in the output from the LLMs. Koala is available for public use at https://koala-index.erc.monash.edu/.
Forward citations
Cited by 2 Pith papers
-
ProDS: Preference-oriented Data Selection for Instruction Tuning
ProDS picks instruction-tuning data by matching training-sample gradients to preference gradients from DPO, achieving slight gains over prior selection methods on MMLU, TYDIQA, BBH, and Alpaca-style tests.
-
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
RICo scores instruction examples by their in-context perplexity effect on an assessment set, then trains a lightweight selector to pick top-scoring data, achieving better benchmark results from 5% to 15% of the original data.
Discussion (0). Continue with ORCID to comment.