Pith. sign in

Paper Citation Record · LEDGER

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2308.12032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.12032 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:26:06.907720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.119173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 08e40cce-7785-46dc-a878-6922821ff89f · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:06.907720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:06.907720Z digest=sha256:5111a40a4240d29355ee132e8951da7b1a17c28c8773c53b845baebb9bac73c0

Observation 57921b79-6d7b-40e2-a015-2d587118a636 · inbound

A Comprehensive Survey on Imbalanced Data Learning cites this paper.

A Comprehensive Survey on Imbalanced Data Learning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 251

Resolution
unresolved
no resolver link, observed 2026-08-07T23:07:19.450634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:07:19.450634Z digest=sha256:e951e0464507d2cca55765f2f2cb38bf2f3ed78dd63a1c71e8e843769fa1d66a

Observation 755a65f1-c9ee-420e-b1de-c12ed84ad7d4 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.882285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:cedc10ea18afbb8fcf5cbc2dddeee8f62e8e79bcf6ec895326fb55be9ac39435

Observation 8185dbc5-4548-4b35-b7aa-1c461762f290 · inbound

Muon is Scalable for LLM Training cites this paper.

Muon is Scalable for LLM Training From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.179429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:11661faff1f5842b5bc665c171ed50a71d6c34778aa7d048858042a5f34bba7f

Observation afec0ce0-e291-4622-9b92-5ed6c9e9203b · inbound

Merge to Mix: Mixing Datasets via Model Merging cites this paper.

Merge to Mix: Mixing Datasets via Model Merging From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:35.884777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:35.884777Z digest=sha256:1476cad96824065a12d342a1abfe135756b0f7664c9f17015e7d88118a422b2d

Observation fad3992c-dfb2-4810-af74-e6080a8d9bd5 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 239

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.357494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.357494Z digest=sha256:d0654271579939c66b19cc2e9f5a68386602cf724e86960a206c8f275136ca31

Observation 12b92e94-51eb-4fd0-80d3-53314c049101 · inbound

GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis cites this paper.

GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:54.951220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:30:54.951220Z digest=sha256:c51fd1e422031c91b33e821245406870116e7caab34a2fccf3acae79a80a2840

Observation 21e7f305-db1c-408e-8639-e83c44e0de5f · inbound

Efficient Data Selection at Scale via Influence Distillation cites this paper.

Efficient Data Selection at Scale via Influence Distillation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:18.225316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:18.225316Z digest=sha256:a4e2a3a28bda382a64e16663ec5abac7409d3f34efe213158e15aedb2cd49d19

Observation f461a272-eee8-4f66-b90f-edd17cfd895f · inbound

Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models cites this paper.

Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:23.921553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:23.921553Z digest=sha256:d3334bb3722c231bc7275d345ea5bb528469cd47b1346d5d879db7511ecf7697

Observation 2ca8c4a2-3841-4b2c-9632-0b03d8b467b3 · inbound

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics cites this paper.

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:04.732253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:59:04.732253Z digest=sha256:133b50385106b363874b02b47ece5f39f28b867cb841b0ab8d3bfe3d893b2c73

Observation 88b2e17a-7fa5-4ce3-9db6-c89937a0496f · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.422873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.422873Z digest=sha256:4b13a5715a9b7a96a9150c63696631eaf59a9df214ddac74c9e6dd6a09e7a610

Observation 4339969f-0b5d-44a1-a95b-4cfccfd83260 · inbound

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models cites this paper.

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.933306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T02:38:27.510872Z digest=sha256:56350593d7b482749099ce6a8505c3a39155e156423fa58765c8695fc59b74b2

Observation 20b8d4f3-aa50-4441-930d-1de62a6c5da9 · inbound

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation cites this paper.

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:22.473706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:22.473706Z digest=sha256:883f5cc2c87cbe2d200abbc51a25f001cb8e397b7e05410d51c7f01a1e94b361

Observation 0b869c00-2632-44ba-8032-938821e53c52 · inbound

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms cites this paper.

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:47.522508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:17:47.522508Z digest=sha256:8e76ab52fdc1b5a777c511d45e3b732682c58da95b3216fa4986c64849b8d18f

Observation 43e43177-ac86-4792-b8bc-86467cb66d3a · inbound

Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains cites this paper.

Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:02:40.415820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:02:40.415820Z digest=sha256:7023892c1b2a1849d1fd663840ca97f4503e98408a230d596ce619e3926eeabf

Observation 31e4a7f6-9428-49c7-96e0-9c0c97af855a · inbound

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning cites this paper.

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:20.079039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:46:20.079039Z digest=sha256:cb71cb6272e8454955432725f607cb6c555f140c5dc44188b1fcb24fe961ec71

Observation 9b88b3bd-52f3-4602-aa42-203c095036dd · inbound

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning cites this paper.

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:50.120744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:32:53.558370Z digest=sha256:875820f8a9f7ced0a584c9b8364cd416cbb44bbd9f05261ed8a6abf71b68ea8a

Observation 77b9def2-5c09-4433-9cca-407f5d1d7264 · inbound

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement cites this paper.

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T13:54:03.879024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:54:03.879024Z digest=sha256:fc0eed8c77e6e7783dfa16bae1200acf1b126765e6022fa3115171483bde9914