Pith. sign in

Paper Citation Record · LEDGER

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2308.12032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.12032 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:03.196830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.119173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 06e5d5b4-6403-4720-b35d-f0979fd130fd · inbound

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning cites this paper.

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:03.196830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:03.196830Z digest=sha256:8389901356286f3e8372ac5798443e4cbd32d9508712bad9e251434d32a03635

Observation 23f02d9e-5609-4298-812c-1babb8b6ee4c · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.779354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.779354Z digest=sha256:3d4851b88096a691dd54f05abd2dffd0948ad6f441326da15c9f666bec71c5af

Observation bac7a134-0a52-4a39-b3fc-5a1e5f2ff06d · inbound

Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud cites this paper.

Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:13:58.117083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:13:58.117083Z digest=sha256:eb6296f66da7fc277f7bcb8d3e0254ebf111807382257b2cf7ef7b08636992af

Observation b731a121-028d-4a7b-ad2f-7c82c79c5cda · inbound

Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud cites this paper.

Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:13:58.122390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:13:58.122390Z digest=sha256:bd0670c1e0af7f3f6716457b47ef50b23f5a91666182b13faa79a4bc269a33fb

Observation d678e190-52ed-4812-b55f-fa3ce86f956b · inbound

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness cites this paper.

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:54:51.985808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:54:51.985808Z digest=sha256:e23e423b01ed34fbcdbe9ebf9e022379894378f7d5bbcfe807c8b25d03eaeeb8

Observation b61a3ac2-5ae1-43d3-8e0f-7cbe76ded128 · inbound

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models cites this paper.

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:15.446512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:44:15.446512Z digest=sha256:171864700218ea8f3af84e2a3e74d6316943aae50122b9eb1b55fe04d8f0dab9

Observation a4b61e5c-e1fd-4943-b860-3eee95e3bf3a · inbound

Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs cites this paper.

Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:18.870117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:18.870117Z digest=sha256:465f34d5f0a54fbc09f468c168deba1a75cc157c774b581a6eeea5fe4c0b7113

Observation d2bf98cd-5eae-40bb-ad2b-4f92178147ca · inbound

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis cites this paper.

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:59:28.303988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:59:28.303988Z digest=sha256:56abab38e6909d19b09da30e5963cb035672c6fa53d4b612315b84764a392174

Observation 1c5596d7-b0e7-4e38-b72f-5ab3e205b463 · inbound

Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline cites this paper.

Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:17:27.266132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:17:27.266132Z digest=sha256:af51e04e6a4cfac406290d677ffcfa95dc8198f23e51bb8fcd5deb61f1767604

Observation 67d16506-8c00-4edb-9dcd-180883516aa8 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.159936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.159936Z digest=sha256:edabd82dee9a5b7df9397aaa52aef0755c992ba0e25142acaadbd4f284fa163f

Observation 75b67c2c-9d08-495d-b4ca-e5924485b7ea · inbound

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions cites this paper.

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 206

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:05.085746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:05.085746Z digest=sha256:657e4b97919d7553007a8755e4280a4986953dbb549be08513f20181897d3828

Observation 8e352668-96c1-4b6d-92a0-f96ef55bdbfe · inbound

R.I.P.: Better Models by Survival of the Fittest Prompts cites this paper.

R.I.P.: Better Models by Survival of the Fittest Prompts From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T23:02:23.333055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T23:02:23.333055Z digest=sha256:c5d663ec4249a7c5f628df2553ea14c841b9ffb97545aab2e201ac39586065df

Observation 08e40cce-7785-46dc-a878-6922821ff89f · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:06.907720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:06.907720Z digest=sha256:51c12e563de05e9b3dca52304edc0322523f218d8ebd128432808ace6d83c222

Observation 57921b79-6d7b-40e2-a015-2d587118a636 · inbound

A Comprehensive Survey on Imbalanced Data Learning cites this paper.

A Comprehensive Survey on Imbalanced Data Learning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 251

Resolution
unresolved
no resolver link, observed 2026-08-07T23:07:19.450634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:07:19.450634Z digest=sha256:79e9d81f2f09af2222d4aa69df6d882634a74982efc5fb035e8fd21e6ffaeb3e

Observation 755a65f1-c9ee-420e-b1de-c12ed84ad7d4 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.882285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:f6d2cee630b83208a96220a5cb3ae6304b45b8ae2b896d1acd330dd755d380e8

Observation 8185dbc5-4548-4b35-b7aa-1c461762f290 · inbound

Muon is Scalable for LLM Training cites this paper.

Muon is Scalable for LLM Training From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.179429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:f76d867ce9a27910e1e006bfcd223171ef74ceb5a5bd2cde2b0f33f14f3cbea6

Observation afec0ce0-e291-4622-9b92-5ed6c9e9203b · inbound

Merge to Mix: Mixing Datasets via Model Merging cites this paper.

Merge to Mix: Mixing Datasets via Model Merging From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:35.884777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:35.884777Z digest=sha256:d8814e378cb6f24bd93504c157581745b5a81394cf6c988e11e1dd0a66654931

Observation fad3992c-dfb2-4810-af74-e6080a8d9bd5 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 239

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.357494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.357494Z digest=sha256:5dbafb2e70d7fee7fdf198a6adfc1acfec49740ba9d915dcf53a146e9592a6ee

Observation 12b92e94-51eb-4fd0-80d3-53314c049101 · inbound

GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis cites this paper.

GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:54.951220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:30:54.951220Z digest=sha256:04ea5d9a84f8bd156e6c002219eaad3e17d94110c03fa7f36b87746151be3991

Observation 21e7f305-db1c-408e-8639-e83c44e0de5f · inbound

Efficient Data Selection at Scale via Influence Distillation cites this paper.

Efficient Data Selection at Scale via Influence Distillation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:18.225316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:18.225316Z digest=sha256:928306abdf7f4ba31e9e2479c87bc5bfacc6b36c828e8fbab9c4de07b4f7a774

Observation f461a272-eee8-4f66-b90f-edd17cfd895f · inbound

Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models cites this paper.

Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:23.921553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:23.921553Z digest=sha256:0f3d62a284540e00e5bbe59615cff23b5de842f3741a51798b00f0da707c0f06

Observation 2ca8c4a2-3841-4b2c-9632-0b03d8b467b3 · inbound

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics cites this paper.

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:04.732253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:59:04.732253Z digest=sha256:5a88ce947c7f1d991b23c082187d6cf7d55926fa5bc83ba24b5ce8958408ec50

Observation 88b2e17a-7fa5-4ce3-9db6-c89937a0496f · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.422873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.422873Z digest=sha256:dcbe5c72b2a561e79f13c438b65d3a82c0496e92df0930f57fabd87da9654540

Observation 4339969f-0b5d-44a1-a95b-4cfccfd83260 · inbound

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models cites this paper.

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.933306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T02:38:27.510872Z digest=sha256:4f1abec8e0c6869264d28e5226dd40767cea0eaf1bd10694d1e717f86a72bd61

Observation 20b8d4f3-aa50-4441-930d-1de62a6c5da9 · inbound

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation cites this paper.

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:22.473706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:22.473706Z digest=sha256:81b3ed126777c298857412e05d4f7a8dc8bd56bf228b41238183b435169c8e8f

Observation 0b869c00-2632-44ba-8032-938821e53c52 · inbound

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms cites this paper.

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:47.522508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:17:47.522508Z digest=sha256:cb7d4cbbb9acac6100b5599fca6a29871f2d43cdbbeb46d735209fc59fca460c

Observation 43e43177-ac86-4792-b8bc-86467cb66d3a · inbound

Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains cites this paper.

Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:02:40.415820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:02:40.415820Z digest=sha256:103f39d67a7dae852d5e5c8d13a698793045044f43f6ad60db3b45311641e5a4

Observation 31e4a7f6-9428-49c7-96e0-9c0c97af855a · inbound

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning cites this paper.

VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:20.079039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:46:20.079039Z digest=sha256:47035399f719a810fe5d31ed2d98804bcc11bf1f673be99245e17322102980a3

Observation 9b88b3bd-52f3-4602-aa42-203c095036dd · inbound

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning cites this paper.

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:50.120744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T18:32:53.558370Z digest=sha256:2a92705b5105332d8a79b66efc5b73e2aabd78f2a294ba043d1a8e0d0feef677

Observation 77b9def2-5c09-4433-9cca-407f5d1d7264 · inbound

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement cites this paper.

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T13:54:03.879024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:54:03.879024Z digest=sha256:cd5d0126a571ef66f1f64f655f00e727770a7ee0df2c50d767b3a7a5fb289f9a