Pith. sign in

Paper Citation Record · LEDGER

The MiniPile Challenge for Data-Efficient Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2304.08442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.08442 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:52:07.511745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T14:03:20.499861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ed6fd1fb-cf23-465d-a8f6-d8ba531c2c82 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models The MiniPile Challenge for Data-Efficient Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:17.035055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:02b1f6ba54ffc20ae5870a6de46a4a01ed7743172787e6c7fb5daa4daab03680

Observation fb9a4ed6-9ec1-47c4-a061-9f419192f01c · inbound

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing cites this paper.

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing The MiniPile Challenge for Data-Efficient Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:52:07.511745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:52:07.511745Z digest=sha256:fdef0e213d7ea969e6975b861dc2a48305b8d9b3ab1f3c881b5723ba4a3f86e7

Observation 7ee45e5a-4790-48c3-9e16-75edeca7d294 · inbound

Sparsified State-Space Models are Efficient Highway Networks cites this paper.

Sparsified State-Space Models are Efficient Highway Networks The MiniPile Challenge for Data-Efficient Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:45.649233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:45.649233Z digest=sha256:309fe23b23b77beb66c9b7d3e25739353519847eaec5cb4632ea349ca554106b

Observation 05fcb154-0386-485f-ba5a-f8a97e565261 · inbound

Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation cites this paper.

Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation The MiniPile Challenge for Data-Efficient Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:59.498544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:59.498544Z digest=sha256:3e6f8af8d2214aa9b61f3ee2dc0a52e3b0ce1127a47e21ddf31365656113eb44

Observation dc16a7db-3828-4ed7-92ae-c9be9a81ac67 · inbound

Causal Estimation of Tokenisation Bias cites this paper.

Causal Estimation of Tokenisation Bias The MiniPile Challenge for Data-Efficient Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:42.992139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:15:42.992139Z digest=sha256:fbf0a0336ff3b547a0cf1726ab7ed26a6a739a41a09620c8fd195c1079297eab

Observation a69bde22-87cd-4b7b-9563-486cfafb349e · inbound

ByteSpan: Information-Driven Subword Tokenisation cites this paper.

ByteSpan: Information-Driven Subword Tokenisation The MiniPile Challenge for Data-Efficient Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.816970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.816970Z digest=sha256:1ea315229966428d7fc4fb900031b0946d5601e762775e5e27ba85ca615a9f55

Observation f96c99fe-5eda-4be2-9f39-2bc3de2d429d · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding The MiniPile Challenge for Data-Efficient Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:26.663156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:26.663156Z digest=sha256:722059b5187a90df945eec8721659672676056adb6e3dd71b73281de820cad49

Observation e32e8cf1-5355-4bbf-8076-3d0e8174fe62 · inbound

Dataset Ownership Verification for Pre-trained Masked Models cites this paper.

Dataset Ownership Verification for Pre-trained Masked Models The MiniPile Challenge for Data-Efficient Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:04:35.614370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:04:35.614370Z digest=sha256:7590ee567feb392454c314c8ef78b54c856e253a78c9d1c6276a9234172bbade

Observation ff07b9ad-9655-48f1-862f-f70633ea58ac · inbound

Faster Superword Tokenization cites this paper.

Faster Superword Tokenization The MiniPile Challenge for Data-Efficient Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:52.364277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:51:20.043728Z digest=sha256:c210711b70971cd72ba4eb749c35efa0ae5d51a996db49f440b502ede39d248f

Observation 12e9c6f7-2127-4056-bb17-b3408965877c · inbound

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer cites this paper.

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer The MiniPile Challenge for Data-Efficient Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:16.118183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:31:33.619232Z digest=sha256:ce575c819f82d5b1053dafff88f01f7b528d2823be720e961de9a71b3ac51a7d

Observation 98073453-e083-40c0-afba-e9e15e0897d5 · inbound

LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models cites this paper.

LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models The MiniPile Challenge for Data-Efficient Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:03:20.502894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:58:55.899958Z digest=sha256:f54d18d533ef4f37518498837a7f32c200d244b428ebeeda215658ca0bfae767

Observation 5414b1d7-748e-44f5-81c1-a6a9601f70c5 · inbound

Joint Optimization for Greedy Longest-match Tokenization cites this paper.

Joint Optimization for Greedy Longest-match Tokenization The MiniPile Challenge for Data-Efficient Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T23:47:05.969487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:47:05.969487Z digest=sha256:d67a9b44e57e8057e4ffec93b186a5e7451aa0d94bb301ca1eac303b2902b67e