Pith. sign in

Paper Citation Record · LEDGER

Data Engineering for Scaling Language Models to 128K Context

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2402.10171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10171 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:26:55.365808Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:07:33.537106Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation baf6d7cf-8aaf-4064-8e1b-4d447055843d · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Data Engineering for Scaling Language Models to 128K Context

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:47:27.840811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:75a2e73661a3dfdcde8e953fd10c5301993c99e9a95e9e9f684a0c45e5b09492

Observation b72ba8dd-ffa1-46c3-8ede-13d4fa5b5679 · inbound

Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining cites this paper.

Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining Data Engineering for Scaling Language Models to 128K Context

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:45:56.152590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T02:44:01.340415Z digest=sha256:6a87862cea70a53bcfd22fba17e4acfd5236675db138569bf84b13a51b7c4b78

Observation e1181f8f-1386-4869-85a9-5b1aae4af19d · inbound

RULER: What's the Real Context Size of Your Long-Context Language Models? cites this paper.

RULER: What's the Real Context Size of Your Long-Context Language Models? Data Engineering for Scaling Language Models to 128K Context

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:20.448654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T03:55:20.355345Z digest=sha256:7de82f90a818e4a905366048a0c154f77497dfb7f0edd7820e83f5cc99bf95aa

Observation d8481eb6-6ebf-4745-8035-4da3a1df440a · inbound

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention cites this paper.

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention Data Engineering for Scaling Language Models to 128K Context

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:17:00.229562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T18:17:00.157129Z digest=sha256:cfaae8b02c33f4491ebaf927f9c80939aa9e5309a2a4c3b5d0c5643b596dbcef

Observation 228085f8-6fa4-4317-81d2-77805279be3e · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Data Engineering for Scaling Language Models to 128K Context

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.552862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:0d9bff1f1a80c1bff2ee20dc49a29fd21713f3a0a2b321ffb0089b587b9cfeb2

Observation 7664ed1c-a9c6-4c5d-892f-bf13a0359881 · inbound

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference cites this paper.

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Data Engineering for Scaling Language Models to 128K Context

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:46:47.037769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T17:46:46.845424Z digest=sha256:1aee5ebcce4d5916f5705a9dd1ab0e19141d894cb079619475b6c1ddcad607c6

Observation 51c02580-c25f-4903-81e3-50abda7004a2 · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation Data Engineering for Scaling Language Models to 128K Context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.365808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.365808Z digest=sha256:adc311a8a0ffd5049e63b56ed7f4e0643d38165964a2a57d55ad353810663840

Observation 7c1730a4-5ea6-49fe-9e82-f1906eeea3b6 · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate Data Engineering for Scaling Language Models to 128K Context

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.309704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:22b23bd2cb19e8b41620f2d39da660e3b93a4385e1fa168f83eaeb19af5aa1a3

Observation d0d416bc-4c4c-4280-851d-1d139362df12 · inbound

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions cites this paper.

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions Data Engineering for Scaling Language Models to 128K Context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:59.981503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:59.981503Z digest=sha256:d0b972d99626ce919d391d96f2d945282e972dec67d3bc89a647859fbf94f6fe

Observation 74e21fdf-8cc7-4f77-9d34-24ad798e1e8a · inbound

Curse of High Dimensionality Issue in Transformer for Long-context Modeling cites this paper.

Curse of High Dimensionality Issue in Transformer for Long-context Modeling Data Engineering for Scaling Language Models to 128K Context

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.956205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:23:11.956205Z digest=sha256:e97dff18e4049d922fd5c985606ac177fe442ce4c9265d1e1e56f6c1afa873e2

Observation 4da642ac-0601-46e6-993a-1e492297e921 · inbound

Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems cites this paper.

Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems Data Engineering for Scaling Language Models to 128K Context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:14.441793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:14.441793Z digest=sha256:386008fdd7aa6e1dc056204a0c31dd7f690b22555351248a481aea1e31c76a6e

Observation ebf2ab01-97cd-4065-b5c1-a70bf9fa2800 · inbound

SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models cites this paper.

SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models Data Engineering for Scaling Language Models to 128K Context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:06.057643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:06.057643Z digest=sha256:6692b5d601ff5c64e2614c4ee20d938e0f9c0479ca7f609d4e8c1ee28371be97

Observation 638a5d41-4de1-4dad-b484-b1bd8adac545 · inbound

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks cites this paper.

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks Data Engineering for Scaling Language Models to 128K Context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:16.874877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:16.874877Z digest=sha256:13c428f942eecfa58c350548715dd0f4dcd02c1a3fb1282c83423424c017963a

Observation 62d3bfaa-5308-43fd-9496-0a6d31eda402 · inbound

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models cites this paper.

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models Data Engineering for Scaling Language Models to 128K Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:33:19.999164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:33:19.999164Z digest=sha256:87c56e2304f2a8c819b74feae792a747eb2638848dd214734a31f9161a88b5be

Observation 6089dcf4-55b7-4407-a7b9-99669a33a215 · inbound

Towards Compute-Optimal Many-Shot In-Context Learning cites this paper.

Towards Compute-Optimal Many-Shot In-Context Learning Data Engineering for Scaling Language Models to 128K Context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:38.626324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:38.626324Z digest=sha256:cc94663f8f6d2a27e22f7d3f5a3755749c20135b380596efd0b8dea3f7193a9f

Observation c10f6fcb-9b05-4cee-bbf8-4f4bc853c38b · inbound

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference cites this paper.

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference Data Engineering for Scaling Language Models to 128K Context

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:17:54.895716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T13:16:31.568604Z digest=sha256:d97645fdcaf195e8a7938d59ac4d9edb26726fe75cacd0b5d355ca1b48366863

Observation eae3db87-d56d-40e4-a451-d00d9b48d877 · inbound

Stacked from One: Multi-Scale Self-Injection for Context Window Extension cites this paper.

Stacked from One: Multi-Scale Self-Injection for Context Window Extension Data Engineering for Scaling Language Models to 128K Context

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:00:10.310750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T16:57:09.401220Z digest=sha256:d5dfaf194e184b50bf8b76efafaf2cff0e014807a3760fe8a09d9360e9bc6ab3

Observation be511320-5e39-417e-a4b0-7f4fb0f36a44 · inbound

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants cites this paper.

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants Data Engineering for Scaling Language Models to 128K Context

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.095872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:34:16.075235Z digest=sha256:b9f4cc0846bdcfea1586c7557f558297e77c92ad4a50cbc33f909e1e3c5081fd

Observation 586d3202-9047-477c-b34e-4d37ef6f3161 · inbound

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation cites this paper.

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation Data Engineering for Scaling Language Models to 128K Context

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:45:28.128889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:42:00.440049Z digest=sha256:cd45ba3ec0419f9fc2192d610203222c640999ca39402bc36e6f901ab6ac682b

Observation bb100be8-316f-402e-bc65-6b7f57f491e7 · inbound

Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation cites this paper.

Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation Data Engineering for Scaling Language Models to 128K Context

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:43:12.141839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:40:27.457107Z digest=sha256:46508cca38835a0bba2acf8b03174443212908f5e50f96837b7fc12fdd99f398

Observation 136f9d8d-4ea0-4272-a21e-4865812045e3 · inbound

Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models cites this paper.

Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models Data Engineering for Scaling Language Models to 128K Context

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:01:15.271490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T02:01:14.533941Z digest=sha256:fd5db65e5f0ce7ad32a2ad816cdf4ddec2309b38c1c4c4cb702dcde2a97022ec

Observation d86c350b-3bcc-4dc3-a58f-d1261ada812e · inbound

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing cites this paper.

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing Data Engineering for Scaling Language Models to 128K Context

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:29.149836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:52:45.320454Z digest=sha256:ffec9eb8e9d79e3110dfcff2bcf264ee425350a393dab0624499f80a6bb0824a

Observation 56bbecfb-1a13-470d-a521-412f18dfe504 · inbound

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context cites this paper.

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Data Engineering for Scaling Language Models to 128K Context

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.134001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:16:07.851098Z digest=sha256:30c6e5c891f1ad96c153c34c626e83a2afcc5129d3fc5c6f75540a01c89c1421

Observation b9745d4b-f2bd-479d-a873-55cf963f0097 · inbound

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably cites this paper.

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably Data Engineering for Scaling Language Models to 128K Context

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:37:37.316022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T15:34:09.006815Z digest=sha256:a73f9fa7163fce6af7d091f83b7a1e033a124a7d212b175c61241316c50fabea

Observation 43f4273a-6f40-4463-89ff-61c983e0ccb0 · inbound

WorkBench Revisited: Workplace Agents Two Years On cites this paper.

WorkBench Revisited: Workplace Agents Two Years On Data Engineering for Scaling Language Models to 128K Context

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.394049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T22:32:38.005254Z digest=sha256:e1b61e5755bacc8c41015c0b3cd3d0713b196d3d3928f391544bf976a3028675

Observation 05367672-8f9c-451e-adb3-b39e7e5022cb · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Data Engineering for Scaling Language Models to 128K Context

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.622643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:56c066c194ed941cb6b94d2eb7bf026f3df0d0b3ec4ee01e329ef1a5e4d4a911

Observation 6acecacd-d726-497b-b7b1-5f52e49b19e9 · inbound

Test-Time Training with Next-Token Prediction cites this paper.

Test-Time Training with Next-Token Prediction Data Engineering for Scaling Language Models to 128K Context

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:09:37.401933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T13:52:15.078658Z digest=sha256:9d078630cbc635f6bef302e0697f87e1df90bd3a8174499fe2999eb7403ba817

Observation b9ba3de5-15ad-49cf-b591-ea94623b2897 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Data Engineering for Scaling Language Models to 128K Context

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.443744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:442f39fe712256d27d3897fdc9ae8e8c2280dfa66058d522c3ad06d2fc59e3cc

Observation 9e310bf0-13cb-4690-8005-c75c3db0102e · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Data Engineering for Scaling Language Models to 128K Context

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:16.336372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:16.336372Z digest=sha256:1b670a8910705a1b4d267bf1be29d38d919ffbe8699139d0ecc33dd6f92f5d01

Observation 482814d7-bdb6-475c-962a-6d183153fa9e · inbound

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE cites this paper.

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Data Engineering for Scaling Language Models to 128K Context

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:07:33.538561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T19:59:46.713277Z digest=sha256:5589938c6430b68c8e0ea665a36c2b48572d3f257e1df5fb4f45feffb596d861

Observation 8dd90981-1a13-4c1e-9f5e-b09ab1b7cc2d · inbound

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE cites this paper.

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Data Engineering for Scaling Language Models to 128K Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T06:47:09.927626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:47:09.927626Z digest=sha256:e46b62b8eafe97130ee70533efdd24c86f76bd5261e79bca190723071ffb60b7