Pith. sign in

Paper Citation Record · LEDGER

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2502.06733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06733 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:36:36.594376Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:09:50.090492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:09:51.295780Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 64c56565-2fec-4d9c-9e1a-8cea83f807e5 · outbound

This paper cites write newline.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.371171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.371171Z digest=sha256:9e20386cb1d859e672b0cd1d22971e0b89ceb1a3ec87f2492185c959a0380be6

Observation 145c6055-4bcf-4d83-b6ee-c758d2e238b9 · outbound

This paper cites GPT-4 Technical Report.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.377163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.377163Z digest=sha256:4e3eabb02dcf37ed055142df788cc44022ebf96f9fa226f98c0c3b93ed8dbb23

Observation b200ef2d-00e0-475e-af3e-b5dc128e5c58 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Piqa: Reasoning about physical commonsense in natural language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.382732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.382732Z digest=sha256:e323bf001f11635c9c1fe2d298bfdc5aa1716582ed1636fbcd694c4097e7dd7e

Observation 7a369ffc-4bd7-4a41-82e9-40ba5d586606 · outbound

This paper cites Language models are few-shot learners.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.387751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.387751Z digest=sha256:c7f85429766cf46d6373a2fcb7c60611d1eb4ae4501d8693888521929c6a81ef

Observation 037c0e44-3ff5-4e1a-8e51-fc7a9e100b6e · outbound

This paper cites Skill-it! a data-driven skills framework for understanding and training language models.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Skill-it! a data-driven skills framework for understanding and training language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.575749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.392796Z digest=sha256:5e2a3b0d896fe13c4ccbcbe4e13de543a79e0a2da15f3b8ac88e016a427cd4e3

Observation 9311650d-fc56-431a-ad92-bf9903a64d29 · outbound

This paper cites Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.397638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.397638Z digest=sha256:2c6344b138b5c01d514cdc55a87fac13415193b14e45e38ed75ffdd91304a4d2

Observation 4053a7a8-7ef9-44c9-9130-7e24b0bcd021 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Palm: Scaling language modeling with pathways

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.402751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.402751Z digest=sha256:1e0631702b857c8351a68ae719d4555c91e35b29d223f7b113457b010bad6ef4

Observation 32bddf59-41d5-4333-b82c-361002398ebb · outbound

This paper cites A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.407927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.407927Z digest=sha256:fe72a1fdc413bd5b5487a9a2f16a4a4aee2e0bf75dc6d6d3bae0fa1a7086eb04

Observation d8b248d9-d51d-4735-a035-09d1ebc902fc · outbound

This paper cites The Llama 3 Herd of Models.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.413017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.413017Z digest=sha256:1db8410bc65fa1446d07d5c0e1d6321d9d63eaf8bb66fa732b2fb7dc2e978116

Observation f164f619-804d-4e77-b050-cb185a0ffaa6 · outbound

This paper cites Irreducible Curriculum for Language Model Pretraining.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Irreducible Curriculum for Language Model Pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.418212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.418212Z digest=sha256:4c97fbdfca22acfc79ca7602e6a1990baf73863191ce3d8b219dcf1e84378f52

Observation 7f18b473-c5e8-4fcf-9a23-c12d131e717c · outbound

This paper cites DOGE : Domain reweighting with generalization estimation.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining DOGE : Domain reweighting with generalization estimation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.479585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.423775Z digest=sha256:f39f2b8bc74d4188854babe832f2f4a60b1843ee03f29a2d3943ce35dec28a00

Observation dc59a251-160d-4d96-b914-4ccd88e57a78 · outbound

This paper cites Rethinking importance weighting for deep learning under distribution shift.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Rethinking importance weighting for deep learning under distribution shift

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.428397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.428397Z digest=sha256:41af0eb877bb47fa04cc42c495b2bd2e2973c2bba6935c6204ed88e660536bbe

Observation cd800386-1655-4ed8-b3d9-ed3703bc65d7 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.433253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.433253Z digest=sha256:99e58d48ef1a5e37edc22f6905e0af7b2dc908107f7d040ceb374b46a6e54d73

Observation 7351fd65-d779-4697-ba4c-ce4641107993 · outbound

This paper cites Adaptive Training Distributions with Scalable Online Bilevel Optimization.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Adaptive Training Distributions with Scalable Online Bilevel Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:36:36.972736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.438299Z digest=sha256:253a59a8a9a95a0359987b2275a920267fcb1619caf14b4a97a4528f84df9969

Observation a60e61d2-3a92-46d6-910b-d1bdddf903c2 · outbound

This paper cites Fedexp: Speeding up federated averaging via extrapolation.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Fedexp: Speeding up federated averaging via extrapolation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.394485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.443123Z digest=sha256:50c6d8a6f3d4f79df5f742a5f39e431bb9d78318104df84403ba1bbadece24c7

Observation a908a46f-4f5f-44e6-9726-cf66d8072390 · outbound

This paper cites Accelerating Deep Learning by Focusing on the Biggest Losers.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Accelerating Deep Learning by Focusing on the Biggest Losers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.447693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.447693Z digest=sha256:5ac540882da9ba9a35eef3dd72ec7da97c5cddf2b9349c766f268c93790ef54e

Observation 6c359732-9b22-4939-a8e4-d99ef841ffa8 · outbound

This paper cites Importance Weighting Can Help Large Language Models Self-Improve.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Importance Weighting Can Help Large Language Models Self-Improve

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.452596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.452596Z digest=sha256:55480fddaef10ede8d079efd8a33c233d6a7e9e887d040a5fcc9b97879c4b4cb

Observation 07d8b422-8c1e-4fc3-88e9-0fcd14d12602 · outbound

This paper cites Instance weighting for domain adaptation in NLP.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Instance weighting for domain adaptation in NLP

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.315830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.457381Z digest=sha256:58bca9ede600929bd863f4703a0d2134891f79bb1c22af210621d1ef2a8ad691

Observation 7734a31c-4f42-4cd9-b793-740ca928c358 · outbound

This paper cites Not all samples are created equal: Deep learning with importance sampling.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Not all samples are created equal: Deep learning with importance sampling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.462282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.462282Z digest=sha256:eafd14f72b68094f5785fc65e4c512da7974db1f6b156563b7fcbd48fbca6d68

Observation 2387da00-b5a9-424e-9a12-cb218ebc7a90 · outbound

This paper cites Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:36:36.920322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.466643Z digest=sha256:d6d2590821a911e52c7d2475e4b58fdc9b103878d2b01ee4fbadda98df10523b

Observation a3df622a-a821-48a8-9992-8b8a25502890 · outbound

This paper cites Probabilistic margins for instance reweighting in adversarial training.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Probabilistic margins for instance reweighting in adversarial training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.269915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.471610Z digest=sha256:f30347f28f5981fa6145cc6d775dd0c7c1d11260ca5f63ae1279d011985efe55

Observation 5f15e577-aaa0-48e7-8940-52f416cb07cd · outbound

This paper cites Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.476014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.476014Z digest=sha256:1f5574365867220e69c54e39945a988bc95e9b37d77cdb9710a66500ace351a7

Observation fa7f26f8-24d7-4ea3-bf6b-19e8538d0901 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.480431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.480431Z digest=sha256:17f9c27336f9cd612339d424526bb795d2e44297dd95bb2b1ecbe0ddf06cb295

Observation 3516fb05-27ec-4f9e-b457-255a998579dd · outbound

This paper cites A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.485151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.485151Z digest=sha256:4928803f129b794d7fce71a9a571fb05d764f5cc7b830b6c3507ec5584e92105

Observation 49dd77b7-6a50-44d1-8ab5-0364242ec0d8 · outbound

This paper cites Online Batch Selection for Faster Training of Neural Networks.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Online Batch Selection for Faster Training of Neural Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.489866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.489866Z digest=sha256:37fec19013e86f7f48bf869f21313b510e472063e6571c77935a5f40550d846e

Observation bc97de34-1176-4300-aac6-3827a5853cb0 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.494040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.494040Z digest=sha256:4e99c1744c6f4df088c39d94917668f7973c61cc22081436473304743ad5950a

Observation 9ffc1f42-882c-4357-b311-0b4d9beea84b · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.499623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.499623Z digest=sha256:c6a2f485f391cbbe2bd5d991124e6a88a82b93d848d3afa6823147a76c82b898

Observation 975a3e06-4a62-4cfd-b667-09b58438ee5f · outbound

This paper cites An online method for a class of distributionally robust optimization with non-convex objectives.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining An online method for a class of distributionally robust optimization with non-convex objectives

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.244956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.504056Z digest=sha256:1e898e411834788a92c657df8aa232c32d817ff961e78f8bdc6ff3bb76af4f59

Observation 878d2deb-6e38-4f9b-852b-67819459e719 · outbound

This paper cites Robust optimization over multiple domains.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Robust optimization over multiple domains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.229574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.508337Z digest=sha256:346b6be8177b9848c2ea616618c39b3d36ff26d1fbebff489d22ca8db36c8bde

Observation c1fefab7-baa5-48d1-b45e-86375429c6e8 · outbound

This paper cites Language models are unsupervised multitask learners.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Language models are unsupervised multitask learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.513078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.513078Z digest=sha256:151fff5afd38ca7212facbab4199d310a1ca21906da9aef5d95dd014730ffa48

Observation 65025dcb-73ea-4396-90f3-70901d96ec03 · outbound

This paper cites Overparameterized neural networks implement associative memory.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Overparameterized neural networks implement associative memory

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.517595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.517595Z digest=sha256:1f4602a184c1854a56bcd4393d4cfabc4c01f971efbb61c9221eb23e0d1a1411

Observation d42889cc-31d8-4cbb-a88a-abef68ed65e8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.521933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.521933Z digest=sha256:5a0c8f620e86a08e9bccdbfca2bf3bb5da5a7fa70ceb05cae50d753170ed7435

Observation c5ae3537-a782-48c7-bfaf-3c5036a19fd4 · outbound

This paper cites Learning to reweight examples for robust deep learning.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Learning to reweight examples for robust deep learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.186740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.526334Z digest=sha256:2e4ef9182f705cfc9a5bbd1623f191b508ed4e270a36cf2c90e00116efc6008f

Observation 13e9e73f-5f13-4786-ad4c-071a30b06062 · outbound

This paper cites Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.171130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.530538Z digest=sha256:043548b0f2f31a6fb1ab0342351e0f4310c95642e8e77901ffe1c66e20c4315a

Observation e0419639-5862-4bba-b4f4-6718335ada6e · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.155346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.535047Z digest=sha256:ee967a25cd65b124971f288bb05d4e40239d4e8f0ba2d8f63bd19a21247dbb42

Observation 1b499208-2bbf-46ad-92e4-da3ff31f1400 · outbound

This paper cites Doubly Robust Instance-Reweighted Adversarial Training.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Doubly Robust Instance-Reweighted Adversarial Training

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:36:36.823339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.539574Z digest=sha256:b8e44d7fc96ad07dd1cc38482daffe04e3a352132f433118cef7b2046a36d9c2

Observation 067775a3-e8a9-4d8c-97a7-8b0dc06906b9 · outbound

This paper cites Self-Influence Guided Data Reweighting for Language Model Pre-training.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Self-Influence Guided Data Reweighting for Language Model Pre-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.544083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.544083Z digest=sha256:319efce7c84fd68cc0f31d5e8eaaa6c8b44da633e8027d32ce96e65d3991a6a9

Observation 91fc5ff7-dcfe-4697-8dfd-916f31101db7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.548756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.548756Z digest=sha256:b0d05439ca0160e3220f52304fa36e35183dd77c8525a22fc14eacbb2c383cb7

Observation e40a2f10-2db2-4ca9-ae72-dd7e5256a23d · outbound

This paper cites Attention is all you need.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.553280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.553280Z digest=sha256:13892e85609225f7ab6a72d075b9b40b6c4a54d144435a326ffbd99c76e1c200

Observation 14980b0b-66f3-4c25-b84e-ab3672c9b73e · outbound

This paper cites Liu, and Matt Gardner.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Liu, and Matt Gardner

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.130683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.557827Z digest=sha256:5afffc1dd5b6b67da37c5ca3852a78ed916a469a432f78cb3ef2daa2a2db8df3

Observation 6e524103-9ed9-4afc-8b46-379e36ffe1f5 · outbound

This paper cites Q u R ating: Selecting high-quality data for training language models.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Q u R ating: Selecting high-quality data for training language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.115261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.562124Z digest=sha256:df8a6b051bc06ff2ca9c172a403d51669c69209e003b98f071765462d18caf93

Observation f0addaf2-1bbc-44bb-a1ac-08568136d1bd · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.566504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.566504Z digest=sha256:966ad59f8965f4c0ec0f1b1598910219b67501653afcde358624ff5aaf1b879a

Observation 292adc5b-079c-420b-9b21-1ad503d3cbc8 · outbound

This paper cites Reweighting Augmented Samples by Minimizing the Maximal Expected Loss.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Reweighting Augmented Samples by Minimizing the Maximal Expected Loss

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:36:36.755381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.571036Z digest=sha256:c5a0a07263e8cafbd01777c3d93f794aa55c3df05fd621a3624c3cab74bbd509

Observation 0d86ab4b-182f-4ed0-b62d-bbcbe860286c · outbound

This paper cites Are adversarial examples created equal? a learnable weighted minimax risk for robustness under non-uniform attacks.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Are adversarial examples created equal? a learnable weighted minimax risk for robustness under non-uniform attacks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:36:37.100186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T14:36:36.576280Z digest=sha256:e9041c2bf6ea85513510101feb13802252e5322473352b667873e332bd799a3f

Observation a5493920-4563-4eb5-95f9-ec3f4c696e37 · outbound

This paper cites Geometry-aware Instance-reweighted Adversarial Training.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Geometry-aware Instance-reweighted Adversarial Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.580744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.580744Z digest=sha256:7fdaa399d070843bda0a2a97fc2f709e9384fa5973a51eb7f3492db5202658e6

Observation 2b7bc779-e2e8-42af-8d5c-9d3a65f6e1c7 · outbound

This paper cites @esa (Ref.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.585174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.585174Z digest=sha256:0c821d379d6a710f433f6e314f95b7b773d3535e5f28bb112594151a443ca8d5

Observation 0b37c6c4-73d3-4613-b482-42a83b623b7c · outbound

This paper cites an unresolved cited work.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.590108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.590108Z digest=sha256:ab276d5672b44c123d63ddce47d1e930def6e0d4781dfd6eeb4186774cade174

Observation fecb48e0-38df-4660-bc77-e3154feb128c · outbound

This paper cites 誁ڞYBq0 w @ \˂bF 3`^) f ` WޏTb]tQ Z Ы;]2Zj8ݐlW c kX֏'y S! QfĊ GYs) W 4 P0,Or (W ; )C NƢ). ; W `m GN 徐ݏիhF ސWW xu3ur]!54#VO=? ±* ^p rx ]K f hFIf QY b g+ <mcOt3Gܘ tX.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining 誁ڞYBq0 w @ \˂bF 3`^) f ` WޏTb]tQ Z Ы;]2Zj8ݐlW c kX֏'y S! QfĊ GYs) W 4 P0,Or (W ; )C NƢ). ; W `m GN 徐ݏիhF ސWW xu3ur]!54#VO=? ±* ^p rx ]K f hFIf QY b g+ <mcOt3Gܘ tX

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.594376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.594376Z digest=sha256:f06c7d440ca91aeab467b520ec02a65d28f2697a98bd582352df07062feab3fd

Pith citing papers

Observation 457b51ea-2390-4ed0-8f05-fefe992ccef0 · inbound

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining cites this paper.

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:09:51.406696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:09:50.090492Z digest=sha256:81c695f52b7127fda2d32747b3b93360961f8cdee107a8e953326ed13805be52