Pith. sign in

Paper Citation Record · LEDGER

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

As of 15 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.20380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20380 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:24.977197Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.176059Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 891c50f9-c251-4812-a846-2013608f0c80 · outbound

This paper cites The pile: An 800gb dataset of diverse text for language modeling, 2020.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining The pile: An 800gb dataset of diverse text for language modeling, 2020

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.381295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.381295Z digest=sha256:35282fa3028e4d614fd98f2c8d5e943b82a57b0ccbbf7ea0e662fceb55908e7c

Observation a3914e3c-bc38-4ed8-a2e0-a02e4569b9f7 · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, 2023.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Redpajama: An open source recipe to reproduce llama training dataset, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:26.325372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:22.438328Z digest=sha256:bdfef744055a37dc72e0f2cdb8290a84529b064765881bbe5fa8b947aa14ade5

Observation b4303d09-6850-428d-9ed4-32fa963c15ef · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.550042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.550042Z digest=sha256:ebec8a721c5f17af556d016ae00ea1bb61a3e8e838b30dc9ecc5996712ce620b

Observation 5845024c-8104-4c50-90dd-1a03c46a31a9 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.629012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.629012Z digest=sha256:bb96c0959cdd59501b25ecc1c92a6e6920861fa499f62a498e03993e3390b637

Observation 59c3bb8f-c3bd-49c5-977d-081837ee815c · outbound

This paper cites DoGE: Domain Reweighting with Generalization Estimation.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoGE: Domain Reweighting with Generalization Estimation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.784397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.784397Z digest=sha256:09baf5040f24a27be6a008e8f8614cd3fe86a23be405968fc7ef112b66127588

Observation 4514bdcb-4843-496a-b304-2ba5fc63bd0c · outbound

This paper cites Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.901349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.901349Z digest=sha256:8625626f1a8674948ba06547113b9b9afd26833a4e9d5f7d28d2f8ccad29d693

Observation 9e77b75e-a9a1-4abf-8af6-eed623f1219b · outbound

This paper cites Dynamic Gradient Alignment for Online Data Mixing.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Dynamic Gradient Alignment for Online Data Mixing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.974581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.974581Z digest=sha256:c1c653e209c6249e5d81b6627b8277743ac766942e4796c5ff418f615d79bc8a

Observation ddf230ea-6079-4599-bf9e-7fea37ee7b7d · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Gradient Surgery for Multi-Task Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.075091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.075091Z digest=sha256:d6c39b74a5fdbd2bdf12db378d68e6c645d906397c4ff7b85d6ed45f12ded595

Observation fc54a80d-348a-4890-8693-3559ff35ebce · outbound

This paper cites FAMO: Fast Adaptive Multitask Optimization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining FAMO: Fast Adaptive Multitask Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.171525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.171525Z digest=sha256:90b87aeb659011c08faeb3f825b2ccbd5a8935b9c8ac24c2a587be574817822b

Observation 92c0370f-7511-416a-a66f-6593e8b22641 · outbound

This paper cites Task Weighting through Gradient Projection for Multitask Learning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Task Weighting through Gradient Projection for Multitask Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:04:25.628062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:23.239267Z digest=sha256:dd09d61350eb0554289783d35d4d87dc2d25225a3e3083a68e1e1d2a78e28b87

Observation 890473d1-52e4-421e-9104-c33da6a6e0ee · outbound

This paper cites Learning Models with Uniform Performance via Distributionally Robust Optimization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Learning Models with Uniform Performance via Distributionally Robust Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.320223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.320223Z digest=sha256:fd7ac2cfe5112f43c53b81dd7497d776201f98545a1e7cf02b292d36857fee59

Observation 01494077-0de8-451a-9b61-b332c00f151f · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.384622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.384622Z digest=sha256:b42e3a3682703b2033981be1cbaaef7e813d445a5a9957ad6d90b405f4ef9cb0

Observation 2161bc10-f925-448e-8759-e77f382bd826 · outbound

This paper cites An Online Method for A Class of Distributionally Robust Optimization with Non-Convex Objectives.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining An Online Method for A Class of Distributionally Robust Optimization with Non-Convex Objectives

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:04:25.474251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:23.474114Z digest=sha256:663558de9c81ac33d7f5f013266dfe0e2ee56cef0a667e6c3ec36d5d201c34f4

Observation 48ac4f27-bcf8-43ea-98d2-556ad10af6f0 · outbound

This paper cites Stochastic gradient methods for distributionally robust optimization with f-divergences.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Stochastic gradient methods for distributionally robust optimization with f-divergences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:26.145954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:23.584016Z digest=sha256:b2f5a3c8067ab83b1353489dcbd563ab1f513684f599802588f4b755a1cac065

Observation 6428e2da-6ee0-4c9d-a343-0a49a53c2441 · outbound

This paper cites Attention Is All You Need.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Attention Is All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.690591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.690591Z digest=sha256:55e1b05be9386ef93b4fcf28e46817aa1af95cfae3d0f5c39b6d3aafd3be1e5e

Observation 8d2d5629-0069-4604-8384-21c8d0164403 · outbound

This paper cites Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.796108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.796108Z digest=sha256:870497c579e002f49a8e8b5f589aedc198855204302fd6a76b1636e63a2879ca

Observation b47ec271-8c07-4c8d-b297-fe2d808ee1c9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.913426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.913426Z digest=sha256:c344d47d2e1df470dcbc856059abf56c6e3f3510233de3b98ae0b786d89075b3

Observation f8c2bc0f-7a6c-4fa5-893d-9d91a140078c · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Crowdsourcing Multiple Choice Science Questions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.035893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.035893Z digest=sha256:97c0f67171c102bf35e1d816cffe2cc7774709c611fe42fe0c04a5535b7e43d9

Observation 948a9799-9802-45c3-8935-abbffd40b00a · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.118899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.118899Z digest=sha256:0e146cf9dc82f2c461db356e379929444dd2aa0af23998ffa7410780d4f7d938

Observation 5d2d7f5b-e799-49fe-bfe4-cc6e5db02d67 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.217182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.217182Z digest=sha256:fa9d64d73556025b17f788c7c82f4ea73a3c979759b07a3b07154a356cae787b

Observation 4f8b13c0-bbd2-4f9c-bf77-18f58daf0c7d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.262249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.262249Z digest=sha256:1daf2f432d9dce7678c4e53a4f342780a318df82352598622c021d5d12bb0c7d

Observation 3f4148fc-f73e-403c-b182-5b53eb5ad82f · outbound

This paper cites W iki-40 B : Multilingual language model dataset.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining W iki-40 B : Multilingual language model dataset

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:25.971942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:24.331842Z digest=sha256:ed59a4a1b96c964e5a133af144845fb3db6ce0d364f1a759bdc4007ecaf12c70

Observation 32058b09-fae7-4ce8-bfc0-e861a5801d4b · outbound

This paper cites Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.403499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.403499Z digest=sha256:3ed4217e6785d56deb083a6be05a7facc93e019773435335cc64d1f84a1181ad

Observation 1e101994-f756-4a34-9730-29927a129382 · outbound

This paper cites Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:04:25.253706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:24.473199Z digest=sha256:0549571afa3790658a2ef73e86f34383680e068ed85bb9b7d0ed5ef50900485c

Observation 4f784259-f93c-48de-97e0-b32d974601ba · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Efficient Online Data Mixing For Language Model Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.571162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.571162Z digest=sha256:fa777dd5c86524582373e351b242f72715dd61c7eee6cfad4edd568e66345c3f

Observation f8a9d79b-2215-444b-a291-5b99e7ab11d2 · outbound

This paper cites Conflict-Averse Gradient Descent for Multi-task Learning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Conflict-Averse Gradient Descent for Multi-task Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.642321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.642321Z digest=sha256:9cce881f4436b71efe0bf68ea28ca1ef27f9f8bbaa8d20811a579a379aacc795

Observation a31677a6-6ff4-49b5-a0c2-2baae3a7fd06 · outbound

This paper cites GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.716563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.716563Z digest=sha256:6a343a34dc14f5ef5e0d1432d35dbd8a90a3568539a0ce31b9b568b37751bb8f

Observation 1fa80b2f-9c0c-4312-ad17-8504c1af612b · outbound

This paper cites Multi-Task Learning as a Bargaining Game.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Multi-Task Learning as a Bargaining Game

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.809400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.809400Z digest=sha256:c8a70867452721001590a922c68812004cb174b5ab1c575ee701718f23da5d9b

Observation 3fbc81cf-d05b-4848-ae99-bc8f97745d6f · outbound

This paper cites Multiple-gradient descent algorithm ( MGDA ) for multiobjective optimization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Multiple-gradient descent algorithm ( MGDA ) for multiobjective optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:25.848930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T14:04:24.912025Z digest=sha256:862ab50fb4dac74a74b4dfcfe42e7db0f5dbc29903ae1b719d90f4cabf427cc3

Observation 131c85fc-1716-4469-97cf-1d194ba1e879 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.977197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.977197Z digest=sha256:b7f68d9fb7f102c3b7416bff1b5262f615c54eaf4b17e60b16731f6baf41048c

Pith citing papers

Observation c1292400-923a-4c76-bc8e-077a045b9afa · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.176059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.176059Z digest=sha256:a197ed9c68fc0370c7d7d89fac7901e8aa83dd7f8a23c289437ce7414d420254