Pith. sign in

Paper Citation Record · LEDGER

Data Efficacy for Language Model Training

As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.21545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21545 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:31:40.416807Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T04:49:28.598430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.413967Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3472e36a-d530-47ef-ae56-b2e31262b75a · outbound

This paper cites Training language models to follow instructions with human feedback.

Data Efficacy for Language Model Training Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:19.795365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:19.795365Z digest=sha256:dfee265f107f3159d2639928c7db5b08ed85f5e23ee278d843966143fcdf2eea

Observation 34907434-1239-45b4-88c2-d0f1c4fee130 · outbound

This paper cites GPT-4 Technical Report.

Data Efficacy for Language Model Training GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.789743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.789743Z digest=sha256:1e04b56edb6454fd06a7b59228dccb942bae36363ccd44c3a3d64d5dfb6dbdba

Observation c20a7caf-e52b-4419-ab88-69ffb3476f62 · outbound

This paper cites The Llama 3 Herd of Models.

Data Efficacy for Language Model Training The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.895488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.895488Z digest=sha256:40910541e1f5f2baab5ce2832da5b1a84b5eef10c40f627daf1c56dbac5a33ac

Observation 109ce30e-a715-4be8-9bde-903bf3f6e665 · outbound

This paper cites Advances in natural language processing.

Data Efficacy for Language Model Training Advances in natural language processing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:46.230969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.076029Z digest=sha256:934406b6f1efcf5738daf3f6ed1d510e1a0507178217b466262349266a06c8eb

Observation 43b7d785-8f9e-4852-8966-d9661c2e5f0c · outbound

This paper cites Exploring Sentiment Analysis Techniques in Natural Language Processing: A Comprehensive Review.

Data Efficacy for Language Model Training Exploring Sentiment Analysis Techniques in Natural Language Processing: A Comprehensive Review

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:40.922314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.157846Z digest=sha256:f37194657c5e8e27053340d843083bb2ee0664d43b3b48a58a4a52ef6f5f50d4

Observation 13d9fd78-e418-4d90-ab5a-2a45644dba8a · outbound

This paper cites Natural language reasoning, a survey.ACM Computing Surveys, 56(12):1–39, 2024.

Data Efficacy for Language Model Training Natural language reasoning, a survey.ACM Computing Surveys, 56(12):1–39, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.320880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.320880Z digest=sha256:2b04218b84ae545a99e906c41fc04a38496640ca8c4b94783b95d7e8c14a82c3

Observation af796198-10d4-423a-b3d0-5532553f81bc · outbound

This paper cites Ai- based conversational agents: a scoping review from technologies to future directions.

Data Efficacy for Language Model Training Ai- based conversational agents: a scoping review from technologies to future directions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:46.065682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.434001Z digest=sha256:a770122ce47676dd1f1b9cfb826fbb7fecaff5c24794f5a8bda959b00878a2c3

Observation a1927fc7-0ceb-4778-b125-e5bfacb792d9 · outbound

This paper cites A Survey on Data Selection for Language Models.

Data Efficacy for Language Model Training A Survey on Data Selection for Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.530387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.530387Z digest=sha256:0d11c0114f12ded87f1deacb3466718040973c575610f64ea6ade4f4e61f53fc

Observation ef3463bf-d78b-476a-ac3e-401d1550cb6f · outbound

This paper cites Data selection for language models via importance resampling.

Data Efficacy for Language Model Training Data selection for language models via importance resampling

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.895241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.679951Z digest=sha256:c4410df17384d2a3201af4d7311cd7c23d8ea84afd51d101d795c3a83762424a

Observation e3ec3028-3dba-451f-a437-18ddd3f97bb0 · outbound

This paper cites Data Selection via Optimal Control for Language Models.

Data Efficacy for Language Model Training Data Selection via Optimal Control for Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.732689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.732689Z digest=sha256:a0f209d9a06d8199131292f0d88ae42353dc8151c996bfd0d97f937cd4f74b51

Observation 9afb825a-41b3-4b1a-ab8e-df850b14d8db · outbound

This paper cites Curriculum learning for language modeling.

Data Efficacy for Language Model Training Curriculum learning for language modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.794833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.794833Z digest=sha256:199bd7d553546eabf9d45f82ad973ebd950a856dd32d97cca373de9f5b864437

Observation 5455404c-f873-465e-8537-443144bf71ca · outbound

This paper cites A survey on curriculum learning.IEEE transactions on pattern analysis and machine intelligence, 44(9):4555–4576, 2021.

Data Efficacy for Language Model Training A survey on curriculum learning.IEEE transactions on pattern analysis and machine intelligence, 44(9):4555–4576, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.871785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.871785Z digest=sha256:9056678f057624c1e591b782873044af714c9b6160a4a73c72a12d26f7230743

Observation 181826f8-e99f-458a-b806-4e797e939955 · outbound

This paper cites hello-gpt-4o.

Data Efficacy for Language Model Training hello-gpt-4o

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.681972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.032874Z digest=sha256:5126f17f185f6a793a30a560d76b0b3d2b861936ef1f68de583a62d47916d2f2

Observation 732c55b2-97bc-4c44-84b0-dcb59c421d78 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Data Efficacy for Language Model Training Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.095933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.095933Z digest=sha256:08ccc261100ffac43c6d29775917a4f7ab36e3c73302bb4526cc0e839304d659

Observation bfed6da5-5d01-4875-bc9d-f1b78466d95d · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Data Efficacy for Language Model Training Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.144310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.144310Z digest=sha256:4866804e26f489e6205efc416b32f537d17225302c40317afbd5a018c739ee03

Observation 688f2516-476c-4595-b6e0-f95ff308ca9e · outbound

This paper cites Long short-term memory.

Data Efficacy for Language Model Training Long short-term memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.249863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.249863Z digest=sha256:1af9d2770f34cd4920b9056bbc9bdf58d5d8507cd55fce171a86dd01c43cb9b8

Observation 6d60b150-431e-4a24-a423-184016f8a3e5 · outbound

This paper cites Scaling Laws for Neural Language Models.

Data Efficacy for Language Model Training Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.341314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.341314Z digest=sha256:29fdf83e3d27a14a7f8b11db24a76b1c3f39cbcc0081a9680906e5202ab3ed42

Observation 0858127b-aa58-4d30-8628-25ec4970a281 · outbound

This paper cites Scaling laws for data filtering–data curation cannot be compute agnostic.

Data Efficacy for Language Model Training Scaling laws for data filtering–data curation cannot be compute agnostic

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.455618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.403199Z digest=sha256:029289807c16a590ec52e342bd71e5455a5a962116d0809e9e3e40efb2f58a07

Observation 02bebf18-ee54-457b-b079-53b8c8b972d8 · outbound

This paper cites KenLM: Faster and smaller language model queries.

Data Efficacy for Language Model Training KenLM: Faster and smaller language model queries

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.203086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.509845Z digest=sha256:f9cb4bdc0701ae5ad65fff0bd64d5390027cdc36e5ef0575dde6565f53f37325

Observation 8c87e9eb-75b3-4a0a-9f11-8afdcd28682a · outbound

This paper cites Claude 3 haiku: our fastest model yet.

Data Efficacy for Language Model Training Claude 3 haiku: our fastest model yet

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.984977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.629087Z digest=sha256:d29df22f738db25d139f64b0d68cbe9dbd1c7107dfb131791efaac1c91da8117

Observation 1252c827-a720-4254-8352-bdd6bde866cf · outbound

This paper cites Paml 4: phylogenetic analysis by maximum likelihood.

Data Efficacy for Language Model Training Paml 4: phylogenetic analysis by maximum likelihood

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.857188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.737006Z digest=sha256:142ce8500a730cb879b86a0776c6a5da711fcef7352565e0f24e49df3e86a1a9

Observation edcc5ac3-4dc4-4523-a9fe-b72b283f75ac · outbound

This paper cites Common crawl – building an open web-scale crawl using hadoop, 2010.

Data Efficacy for Language Model Training Common crawl – building an open web-scale crawl using hadoop, 2010

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.730683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.842436Z digest=sha256:8c9860a476d5044b6012462b4bf872044310374fae696b60b9ea986b0caf91b9

Observation 2a3c4282-7a05-45db-aa0a-e57051d3bb3b · outbound

This paper cites Project gutenberg, 2004.

Data Efficacy for Language Model Training Project gutenberg, 2004

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.626745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.925600Z digest=sha256:33bf41b6d4b14253aba2404a7af090724b21d247b00f45e297d57a86df22c4ef

Observation 1800c862-3d77-4e66-8e4d-b2ff96e97ab8 · outbound

This paper cites Synthetic data for deep learning , volume 174.

Data Efficacy for Language Model Training Synthetic data for deep learning , volume 174

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.504460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.983223Z digest=sha256:fee9b16d3c9091b9fdeddc2be9ebb4c5589310efeb8ddcf8e36b9b5aebba6a17

Observation e1b7d25e-5cca-4c8c-8ae9-ab772e92dc82 · outbound

This paper cites Virtual sensors: Abstracting data from physical sensors.

Data Efficacy for Language Model Training Virtual sensors: Abstracting data from physical sensors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.406796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.055947Z digest=sha256:475b4bca91ddd54323f376c3182055cc82d8aed380fc1a1cda2bb806c0da7a10

Observation ef5ba448-e607-4b88-95f1-db06e6d7af76 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Data Efficacy for Language Model Training Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.178879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.178879Z digest=sha256:f48299278833e39c1ba5e7c6142572252fd138d1f928accd01a61d553ba53ea6

Observation 53478cd4-9bc4-4938-9b6e-df65e110288a · outbound

This paper cites The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only.

Data Efficacy for Language Model Training The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.320180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.243725Z digest=sha256:fcf0e1be5c7a5eabc4e2e643f980fc36b6df2f53be6287a495ef9dc082aedb86

Observation 61cb4c60-bc85-4f79-8533-cfae6e3eb1ff · outbound

This paper cites Redpajama: an open dataset for training large language models.

Data Efficacy for Language Model Training Redpajama: an open dataset for training large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.336209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.336209Z digest=sha256:ab546946ef924aa74794abc76d69f10b0bcc78d171dd710dfa1189a969add304

Observation a9af43c7-a9b7-451e-b517-6d0514dd2b55 · outbound

This paper cites RedStone: Curating General, Code, Math, and QA Data for Large Language Models.

Data Efficacy for Language Model Training RedStone: Curating General, Code, Math, and QA Data for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.409614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.409614Z digest=sha256:7ebc955597f516cbc6f40a8a32a800488c958756297433309564f7de3db68a4b

Observation a2af7a3f-b91f-433b-9166-60a72cc70ff1 · outbound

This paper cites Mates: Model-aware data selection for efficient pretraining with data influence models.

Data Efficacy for Language Model Training Mates: Model-aware data selection for efficient pretraining with data influence models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.202949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.516009Z digest=sha256:c39302cb8f35aa3fa439a5ba047746127ecf0c3aebbf281f47a80d33f8e006ab

Observation e7f4349a-1d10-4a6f-a369-fef8925df5bf · outbound

This paper cites Training-free dataset pruning for instance segmentation.

Data Efficacy for Language Model Training Training-free dataset pruning for instance segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.095050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.572752Z digest=sha256:52b0882d20ad80d02ef6a56c8454baad4fd3802a532695f42712e600bb798fd2

Observation 5ad266e4-2031-4862-899b-a22d6be16703 · outbound

This paper cites P-diff+: Improving learning classifier with noisy labels by noisy negative learning loss.

Data Efficacy for Language Model Training P-diff+: Improving learning classifier with noisy labels by noisy negative learning loss

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.990600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.620112Z digest=sha256:c702a677adaab1366f4322bc24c223c751d0409ca47fcb2a444d3d73a3e8586c

Observation 3fb4ae5a-4f96-48e0-9d3f-a804118be1dd · outbound

This paper cites P-diff: Learning classifier with noisy labels based on probability difference distributions.

Data Efficacy for Language Model Training P-diff: Learning classifier with noisy labels based on probability difference distributions

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.881483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.665766Z digest=sha256:990fb6eba39b64e68c5b20c12440b3b0aa8065f9b3c82a52faeb0516c2f58303

Observation ea0935cd-49b8-421c-b4b8-883d969337d8 · outbound

This paper cites SemDeDup: Data-efficient learning at web-scale through semantic deduplication.

Data Efficacy for Language Model Training SemDeDup: Data-efficient learning at web-scale through semantic deduplication

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.703358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.703358Z digest=sha256:ba639f97a735aad14e1620ecd72ca0f2be8c8f39e245d31ac4ffe7e835d1ff16

Observation e655be23-c353-4bf2-8cc8-0f06301c35cc · outbound

This paper cites D4: Improving llm pretraining via document de-duplication and diversification.

Data Efficacy for Language Model Training D4: Improving llm pretraining via document de-duplication and diversification

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.768535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.798234Z digest=sha256:735d8cc89e8acb058053c1abc8db6b4c58cdcbf6ab2edb049acd008b1bcd2ce7

Observation 13b16161-46e5-45cc-b23a-cb4b4c4a5f00 · outbound

This paper cites Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning.

Data Efficacy for Language Model Training Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.853648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.853648Z digest=sha256:b3923491bfa7a5532264b3cc5a196029699ebdaacdcdb0542249fcfb9cf885dc

Observation 4b295c33-2616-43ae-a09a-71f2730342b9 · outbound

This paper cites Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum Learning.

Data Efficacy for Language Model Training Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:40.710309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.971943Z digest=sha256:ae24d8d85d3f445377d51a440151ed14a504e627aaa9b50bd3519919407da228

Observation 8c0be7cb-f617-4dd3-8ec5-df298404a13f · outbound

This paper cites DoReMi: Optimizing data mixtures speeds up language model pretraining.

Data Efficacy for Language Model Training DoReMi: Optimizing data mixtures speeds up language model pretraining

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.673371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:38.143000Z digest=sha256:a00e0416cfbe3b9da9a6ecee0c7b913c9d5f287c87ab6d14aa9e53fd09ffcb4a

Observation 0b2a68e7-4a68-4024-9700-cf5ce660c34e · outbound

This paper cites Lima: Less is more for alignment.

Data Efficacy for Language Model Training Lima: Less is more for alignment

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.563723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:38.258542Z digest=sha256:e3231047daecae71841797a318e514d38590b8a19b96c9a69c7bc54bbf691a5f

Observation ca58f8fa-b5bd-49ec-963f-0d6d9132a0b2 · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text, 2023.

Data Efficacy for Language Model Training Openwebmath: An open dataset of high-quality mathematical web text, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.370698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.370698Z digest=sha256:f4ae87a374d763dbbdc1be1519bff533b8c66713edd5bc11ad0e30534f022689

Observation 3eb84262-2755-4f47-9c16-039dbd257e40 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

Data Efficacy for Language Model Training MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.509344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.509344Z digest=sha256:03685b0edb1f17df46d5682389d4c77ec57606bb720fa8d430ee166b89bb6350

Observation b18f4e11-793c-4db5-b71e-5e668be63f4a · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Data Efficacy for Language Model Training StarCoder 2 and The Stack v2: The Next Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.679922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.679922Z digest=sha256:e1af66321ca9e8afb674a398e201343a1d97dff390e8090c9f6b7f6d7b2faa9d

Observation f1c0a048-7337-4760-b423-606723280f2d · outbound

This paper cites Epicoder: Encompassing diversity and complexity in code generation.

Data Efficacy for Language Model Training Epicoder: Encompassing diversity and complexity in code generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.794928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.794928Z digest=sha256:dc759242d68ca22fb191af5ebcc7a3e9f6b1215e1788c4676339fa1e80104bb6

Observation 7eaf93f6-b84b-47a9-af78-a170358f2de2 · outbound

This paper cites Mistral 7B.

Data Efficacy for Language Model Training Mistral 7B

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.889964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.889964Z digest=sha256:91c2d74a06ec917426cbf7baf48bfd5db70e0fd631956f162715db7063499359

Observation e5ee4398-ee2c-41cb-a1c9-f5e6ec0c7bf8 · outbound

This paper cites Qwen Technical Report.

Data Efficacy for Language Model Training Qwen Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.997852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.997852Z digest=sha256:15038ca8071656e17be9be5855411297b9d43e1abaab3070ec4294dd6a82448c

Observation 3bb581e0-945e-4cf0-bbb4-8478a8f754fd · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Data Efficacy for Language Model Training OLMo: Accelerating the Science of Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.060536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.060536Z digest=sha256:15efc7bdb0d71dc4baf9e1ff1ee2bf061057614b37ab52dbb3863c408ebc5811

Observation 0e32dec9-e955-44f6-b0b1-f9a89ceda0eb · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of ACL, 2019.

Data Efficacy for Language Model Training Hellaswag: Can a machine really finish your sentence? In Proceedings of ACL, 2019

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.416290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.121670Z digest=sha256:9bb1508b85b5cbcedb87a6da31aefdee42d622bef05430771f825e3e2df11b3a

Observation d30f3ec4-a3b1-498c-b6e9-edfd1ef90b55 · outbound

This paper cites The winograd schema challenge.

Data Efficacy for Language Model Training The winograd schema challenge

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.234992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.177098Z digest=sha256:bf1ad6c5f6ff6183e856f325374dae24f3b8897c928f0eb76509fd1d4d432458

Observation 1f834b77-8d9d-473d-b3fb-7b758a0a26b8 · outbound

This paper cites The lambada dataset: Word prediction requiring a broad discourse context.

Data Efficacy for Language Model Training The lambada dataset: Word prediction requiring a broad discourse context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.085204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.220870Z digest=sha256:b0e4deba9c568fe53c07f8b70ba072c2e4757a4cbbc90535e9b4971aaf80ec43

Observation 6aa43bff-0821-4ecb-9854-5b3586e92bc1 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Data Efficacy for Language Model Training Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.927964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.290035Z digest=sha256:f9ec97d62fcd7e42a396d0a8b1a84041120c03c11a6069ab71f3ed2148ffa2f0

Observation a029a61f-bf3a-4d63-8253-7a5d63fb8916 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Data Efficacy for Language Model Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.363914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.363914Z digest=sha256:4dba7125305b259a0dc856fb8bb153bdba66b5ab779219d7e541b840a2b06f4d

Observation fc872b53-e93a-4a05-a77d-65bd721bfd29 · outbound

This paper cites Piqa: Reasoning about physical common- sense in natural language.

Data Efficacy for Language Model Training Piqa: Reasoning about physical common- sense in natural language

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.710576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.444920Z digest=sha256:4dc053653a3e3c3d2af1cac071ce11b2693e0428d4771fc1a3ad6b26c8c30254

Observation f407e16b-dd00-4152-aab2-32f538352ad5 · outbound

This paper cites Liu, and Matt Gardner.

Data Efficacy for Language Model Training Liu, and Matt Gardner

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.527336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.518996Z digest=sha256:a6a596335c54f2b1fa9e525b0fa6a95b19e2d5be6ce559b93e6eed8b0f3db9a1

Observation 451ce2a6-f0a4-45f6-9ee0-4059911f2b09 · outbound

This paper cites BoolQ: Exploring the surprising difficulty of natural yes/no questions.

Data Efficacy for Language Model Training BoolQ: Exploring the surprising difficulty of natural yes/no questions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.323259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.599850Z digest=sha256:74403322361c3c3041fbe54d44896b728048629eeb4c1a442af12380c94fd43c

Observation fa2ec6e9-33ea-4869-8e1b-9ea6d509a64e · outbound

This paper cites an unresolved cited work.

Data Efficacy for Language Model Training Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.689642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.689642Z digest=sha256:fc1c8a57ff8806d6bb1f0f0b681df46bef24a194a6a54213a45896fe14668ce2

Observation 17350254-b8c4-48a4-8322-bfa482950b43 · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019.

Data Efficacy for Language Model Training Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.764554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.764554Z digest=sha256:ab761c836d68a5d39a8747f07579fe4366867b3d0afb7db923ff76cb9eeda393

Observation 9caa4a74-9795-4b26-ae42-fdb2e2fb3e54 · outbound

This paper cites an unresolved cited work.

Data Efficacy for Language Model Training Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:31:42.162493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.849319Z digest=sha256:2a13f3f493cfef8fbd15e0e983029d99c1d9c4442594135e53ef317b86f91d44

Observation ed73ddf4-2c92-400d-a2e8-1fc17282ccd3 · outbound

This paper cites Program Synthesis with Large Language Models.

Data Efficacy for Language Model Training Program Synthesis with Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.955344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.955344Z digest=sha256:9e983aa4a9794538315639c36f33dd00fc62103b688fe9a73a52e3e4c51cd6fb

Observation 752b0d3d-0c92-47dd-aad2-1212e2992466 · outbound

This paper cites Efficient large scale language modeling with mixtures of experts.

Data Efficacy for Language Model Training Efficient large scale language modeling with mixtures of experts

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.973753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:40.039407Z digest=sha256:8e2fbbcb34e1c8ee77edd1cf19200578e0a83fc3970d830bc4a88497460b675d

Observation bd642587-8300-4094-999f-0cd460e8c61d · outbound

This paper cites Decoupled weight decay regularization.

Data Efficacy for Language Model Training Decoupled weight decay regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.786740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:40.140435Z digest=sha256:224079a23b0c970da91c3fb36b26df22252b50aa5467ca94d9861bbe79aacc6f

Observation 2bc119d4-5638-4b74-be28-6607582b63c6 · outbound

This paper cites Comparing the pearson and spearman correla- tion coefficients across distributions and sample sizes: A tutorial using simulations and empirical data.

Data Efficacy for Language Model Training Comparing the pearson and spearman correla- tion coefficients across distributions and sample sizes: A tutorial using simulations and empirical data

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.578495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:40.240414Z digest=sha256:e3235fc8bcde4cda8c28b99d9ea6ebc27531de22c740369ccafd793bcc30d59b

Observation 6d65bff9-d365-460a-af49-54a81ef4cd17 · outbound

This paper cites As described in algorithm 1, we apply a linear transformation to the mean-pooled representations of instances along the sequence length.

Data Efficacy for Language Model Training As described in algorithm 1, we apply a linear transformation to the mean-pooled representations of instances along the sequence length

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.369377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:40.328843Z digest=sha256:5b2eb0f105895060343543557345b260a99e5aa13705f6bc8bb5584bec1becb7

Observation dc04d5f7-4d10-422b-bf68-e651558f3866 · outbound

This paper cites I am overpowered by the discovery of my own genius for management.

Data Efficacy for Language Model Training I am overpowered by the discovery of my own genius for management

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.149390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:40.416807Z digest=sha256:a9d6d8c5109fc9321d04f8483bb8804831e966192cc24409c7c305ba6c7db379

Pith citing papers

Observation e2a8aeb6-0d71-470b-aa72-6e671db70eb7 · inbound

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning cites this paper.

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning Data Efficacy for Language Model Training

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T04:49:28.598430Z digest=sha256:7fd23344f0c30d988f90996b6590833e0b814061e826b10d1616ff26197606a0