Pith. sign in

Paper Citation Record · LEDGER

Small Language Model as Data Prospector for Large Language Model

As of 12 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2412.09990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09990 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:33:10.502177Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-23T01:03:26.037233Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T01:05:16.527345Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1022d41-beaa-4678-9a5a-b45b3f5d2ae6 · outbound

This paper cites Qwen Technical Report.

Small Language Model as Data Prospector for Large Language Model Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.419660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.419660Z digest=sha256:087b96eb7b63dedc0cb68f96f7e8b0ca001a220e892fcd1e9f4223ffed8d8428

Observation 30cf924f-98f6-467b-a5ff-a8233da790d3 · outbound

This paper cites COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning.

Small Language Model as Data Prospector for Large Language Model COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.425160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.425160Z digest=sha256:05c8877281a1a828e0b57c2c78a1f218ce1063f85cd5e13df115f6028a0cdf56

Observation 4a31ad4a-7136-45e8-8b5c-14b1d658da55 · outbound

This paper cites Instruction Mining: Instruction Data Selection for Tuning Large Language Models.

Small Language Model as Data Prospector for Large Language Model Instruction Mining: Instruction Data Selection for Tuning Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.430133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.430133Z digest=sha256:f5e09d08c71be21f4fb232bd6d288a21a80eb58d4a1ebba77b450f40e851c798

Observation 1c82ed51-a496-4f82-a8e8-828c50636227 · outbound

This paper cites an unresolved cited work.

Small Language Model as Data Prospector for Large Language Model Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:33:10.802930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:33:10.435486Z digest=sha256:70e2d9d926db030a6a65747f6d466d7c392fb58c9c28fc2df1c7c3f4022a288d

Observation ae95652c-81bf-4f9f-a2cb-2022723abd26 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Small Language Model as Data Prospector for Large Language Model Scaling Instruction-Finetuned Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.440064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.440064Z digest=sha256:bd722281814d88ec8f67781b929bbd1984875000ec9d26343ef866be74da24c9

Observation 372eb0b4-0541-4972-a825-6f81ed2cc410 · outbound

This paper cites Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers.

Small Language Model as Data Prospector for Large Language Model Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.445070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.445070Z digest=sha256:f31c102c796678694c132fa8d140c00082890b8ffc2c6766faf6ebb85c523827

Observation 4d87c0c2-8c69-4dd3-9cd8-25ecea6af80d · outbound

This paper cites PaLM 2 Technical Report.

Small Language Model as Data Prospector for Large Language Model PaLM 2 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.450379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.450379Z digest=sha256:86a9c1557f30dab73da7147e3e29e049b0c448610c938a9c4e278f16c8ff7366

Observation 0e65850f-50da-44d9-8bcf-c04404a94663 · outbound

This paper cites Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases.

Small Language Model as Data Prospector for Large Language Model Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.455171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.455171Z digest=sha256:89b7ce0272d2a534a6d53d34ecb217a1b1ff60716a9497f88966bd7d7a0ee052

Observation bf7444fb-3f1a-43e5-aceb-aac927384bea · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Small Language Model as Data Prospector for Large Language Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.459872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.459872Z digest=sha256:9f96c5d6911c754f0d96e75f286f33943299dac50e5b7942018a6762677c1639

Observation 88ddba79-5a3c-4c24-861a-3aa4551ea2d6 · outbound

This paper cites Hashimoto.

Small Language Model as Data Prospector for Large Language Model Hashimoto

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.464326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.464326Z digest=sha256:34194bcba85e61d94936096dbc79d28194f1803aeb3a1e65d98855b0bb6760c0

Observation a74213f2-4c1e-49db-85d1-f24027926658 · outbound

This paper cites One-Shot Learning as Instruction Data Prospector for Large Language Models.

Small Language Model as Data Prospector for Large Language Model One-Shot Learning as Instruction Data Prospector for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.468628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.468628Z digest=sha256:b28d1e22ec0d9c354469bc916daf42734d10e5ac2792007526415bfe21a354c4

Observation d7d1c42a-40ac-4c76-8cea-a155904ec5c9 · outbound

This paper cites GPT-4 Technical Report.

Small Language Model as Data Prospector for Large Language Model GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.473237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.473237Z digest=sha256:4ae8cad7f59a4d44823cfca3a2ff23de3261b4afc421d9d5b69ef3857e473993

Observation 08ff23b2-a104-42fd-873b-e155250f089a · outbound

This paper cites How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources.

Small Language Model as Data Prospector for Large Language Model How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.477351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.477351Z digest=sha256:e50c7ef0f4c8394b0ac243929112118821789dea491da48f0fc43feb2b15be05

Observation 60808d7a-8a8b-4ab3-a1e8-ff1df1d03c4f · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Small Language Model as Data Prospector for Large Language Model Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.481306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.481306Z digest=sha256:e5cb91f697fe3666accec2f73963fb5b371542ff9179d74aa41d34c813ce1252

Observation d7fd0e2b-d464-4121-bb39-b64c89e2f3ff · outbound

This paper cites an unresolved cited work.

Small Language Model as Data Prospector for Large Language Model Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:33:10.778711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:33:10.485389Z digest=sha256:93877644f09ec9b6e4feafbe4bd5e5283aa4ee1e78f91d5972b2ae6921db4ab6

Observation ed9cf9c0-a0b9-4a09-b52d-3908cd101443 · outbound

This paper cites Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data.

Small Language Model as Data Prospector for Large Language Model Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.489111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.489111Z digest=sha256:b82cf05154ec5063e61f8752366cd8f237d8d141b334e1b985adb71bd335d3a0

Observation 3326f2e4-3031-4a5d-a011-cd7b9ee69340 · outbound

This paper cites LIMA: Less Is More for Alignment.

Small Language Model as Data Prospector for Large Language Model LIMA: Less Is More for Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.493125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.493125Z digest=sha256:5bc8d8c7931a6f0b4272d87dee568989be34456ac86fd0d52c0fa7d0899046a9

Observation 817bd649-a8e5-4347-9d74-12a800037960 · outbound

This paper cites online" 'onlinestring :=.

Small Language Model as Data Prospector for Large Language Model online" 'onlinestring :=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.497512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.497512Z digest=sha256:cbc858db03fd1c295d1decdc3a0c133efef8d242be2a63a1998b50a906cbbffe

Observation 933b90b6-ed8b-41d4-9cc3-08fbee6d50c7 · outbound

This paper cites write newline.

Small Language Model as Data Prospector for Large Language Model write newline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.502177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.502177Z digest=sha256:21ba0dffb2c7940c88b19dc303f9f7b820ff253d7a1a736f9af06ee1c66b28fe

Pith citing papers

Observation ccdda23a-90e3-4cc8-ba52-d2b53fc1f73d · inbound

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices cites this paper.

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices Small Language Model as Data Prospector for Large Language Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:05:16.529736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T01:03:26.037233Z digest=sha256:0dfdebf44f8bbb61b696065d61a8d375c5733ad47b64b813eca1d39d0df503a0