Pith. sign in

Paper Citation Record · LEDGER

Data Engineering for Scaling Language Models to 128K Context

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.10171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10171 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:59.981503Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:07:33.537106Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation baf6d7cf-8aaf-4064-8e1b-4d447055843d · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Data Engineering for Scaling Language Models to 128K Context

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:47:27.840811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:494d7c6ce01bb46e83171169235ec527ac123e2dd417f253c58a1e4d8b648259

Observation b72ba8dd-ffa1-46c3-8ede-13d4fa5b5679 · inbound

Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining cites this paper.

Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science with Table-Specific Pretraining Data Engineering for Scaling Language Models to 128K Context

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:45:56.152590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T02:44:01.340415Z digest=sha256:75c63e4a012d804a963f93c377b8840ea906b403a757c6b0006f894423da32ae

Observation e1181f8f-1386-4869-85a9-5b1aae4af19d · inbound

RULER: What's the Real Context Size of Your Long-Context Language Models? cites this paper.

RULER: What's the Real Context Size of Your Long-Context Language Models? Data Engineering for Scaling Language Models to 128K Context

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:20.448654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T03:55:20.355345Z digest=sha256:cba0a6ac214121ebc3fb3c307606095a6c022dc08466e0ff78925be45c7e5cc4

Observation d8481eb6-6ebf-4745-8035-4da3a1df440a · inbound

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention cites this paper.

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention Data Engineering for Scaling Language Models to 128K Context

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:17:00.229562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T18:17:00.157129Z digest=sha256:f93fb4c957527215c76b8f9f7923ac287e285f41c6e5fcfe3d9571cda6dfbf7b

Observation 228085f8-6fa4-4317-81d2-77805279be3e · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Data Engineering for Scaling Language Models to 128K Context

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.552862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:ba9baad66a2f5e9b06acdac50bfe49494a8c02318ec7eb5ba8587869daa79663

Observation 7664ed1c-a9c6-4c5d-892f-bf13a0359881 · inbound

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference cites this paper.

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Data Engineering for Scaling Language Models to 128K Context

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:46:47.037769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T17:46:46.845424Z digest=sha256:7743ded9d45fc0304a1c11cd557892982e227c6ed4f88ff76214d72ccc1c670e

Observation 7c1730a4-5ea6-49fe-9e82-f1906eeea3b6 · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate Data Engineering for Scaling Language Models to 128K Context

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.309704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:31fa8ab65b6c97c181822bc64840e56fbffb2b5314746fbe40216e491fc0be08

Observation d0d416bc-4c4c-4280-851d-1d139362df12 · inbound

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions cites this paper.

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions Data Engineering for Scaling Language Models to 128K Context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:59.981503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:59.981503Z digest=sha256:d0b972d99626ce919d391d96f2d945282e972dec67d3bc89a647859fbf94f6fe

Observation 74e21fdf-8cc7-4f77-9d34-24ad798e1e8a · inbound

Curse of High Dimensionality Issue in Transformer for Long-context Modeling cites this paper.

Curse of High Dimensionality Issue in Transformer for Long-context Modeling Data Engineering for Scaling Language Models to 128K Context

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.956205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:23:11.956205Z digest=sha256:3503b75e2baa82e846a9bba43c350978a581a8e10dd2c62fd42a75b657324809

Observation 4da642ac-0601-46e6-993a-1e492297e921 · inbound

Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems cites this paper.

Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems Data Engineering for Scaling Language Models to 128K Context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:14.441793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:14.441793Z digest=sha256:0c689de9627c8daf5b7854f759283d51663aefa51ff568f96036b80661a4b151

Observation ebf2ab01-97cd-4065-b5c1-a70bf9fa2800 · inbound

SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models cites this paper.

SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models Data Engineering for Scaling Language Models to 128K Context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:06.057643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:06.057643Z digest=sha256:6692b5d601ff5c64e2614c4ee20d938e0f9c0479ca7f609d4e8c1ee28371be97

Observation 638a5d41-4de1-4dad-b484-b1bd8adac545 · inbound

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks cites this paper.

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks Data Engineering for Scaling Language Models to 128K Context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:16.874877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:16.874877Z digest=sha256:13c428f942eecfa58c350548715dd0f4dcd02c1a3fb1282c83423424c017963a

Observation 62d3bfaa-5308-43fd-9496-0a6d31eda402 · inbound

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models cites this paper.

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models Data Engineering for Scaling Language Models to 128K Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:33:19.999164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:33:19.999164Z digest=sha256:87c56e2304f2a8c819b74feae792a747eb2638848dd214734a31f9161a88b5be

Observation 6089dcf4-55b7-4407-a7b9-99669a33a215 · inbound

Towards Compute-Optimal Many-Shot In-Context Learning cites this paper.

Towards Compute-Optimal Many-Shot In-Context Learning Data Engineering for Scaling Language Models to 128K Context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:38.626324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:38.626324Z digest=sha256:cc94663f8f6d2a27e22f7d3f5a3755749c20135b380596efd0b8dea3f7193a9f

Observation c10f6fcb-9b05-4cee-bbf8-4f4bc853c38b · inbound

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference cites this paper.

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference Data Engineering for Scaling Language Models to 128K Context

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:17:54.895716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:16:31.568604Z digest=sha256:f0f5efb1984a2a948db764521b25d373ee11861f1ccbaed61289495d4a01e671

Observation eae3db87-d56d-40e4-a451-d00d9b48d877 · inbound

Stacked from One: Multi-Scale Self-Injection for Context Window Extension cites this paper.

Stacked from One: Multi-Scale Self-Injection for Context Window Extension Data Engineering for Scaling Language Models to 128K Context

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:00:10.310750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:57:09.401220Z digest=sha256:91ecf91824de98c1eedee77d3bf319d31e4c735319a83eaa6458bb971fac62bc

Observation be511320-5e39-417e-a4b0-7f4fb0f36a44 · inbound

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants cites this paper.

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants Data Engineering for Scaling Language Models to 128K Context

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.095872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:34:16.075235Z digest=sha256:55697174949d59cddc8458d7b786a9b80dfc9b4b987fb5327543ade99121dee0

Observation 586d3202-9047-477c-b34e-4d37ef6f3161 · inbound

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation cites this paper.

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation Data Engineering for Scaling Language Models to 128K Context

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:45:28.128889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:42:00.440049Z digest=sha256:b0b8a7c8c65da36e269553d7fd3f4e980ea26564b15f7fa99fbec550bb78eaa2

Observation bb100be8-316f-402e-bc65-6b7f57f491e7 · inbound

Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation cites this paper.

Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation Data Engineering for Scaling Language Models to 128K Context

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:43:12.141839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:40:27.457107Z digest=sha256:885e6e425dd31e09eacf69d726f5e8de7ebc9a74882770ca7e94ccb59375b282

Observation 136f9d8d-4ea0-4272-a21e-4865812045e3 · inbound

Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models cites this paper.

Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models Data Engineering for Scaling Language Models to 128K Context

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:01:15.271490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T02:01:14.533941Z digest=sha256:5204af2e11c5f8dd83186f017d3579263aadb71643dc5159df92e505239c092f

Observation d86c350b-3bcc-4dc3-a58f-d1261ada812e · inbound

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing cites this paper.

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing Data Engineering for Scaling Language Models to 128K Context

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:29.149836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:52:45.320454Z digest=sha256:f555da74d31ca0f4dde10d74282df85575086e7ce26ae5eb6040a46dbaba2515

Observation 56bbecfb-1a13-470d-a521-412f18dfe504 · inbound

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context cites this paper.

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Data Engineering for Scaling Language Models to 128K Context

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.134001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:16:07.851098Z digest=sha256:52dd0acc3ef46cf76159cc44b4c69a8e7697989b4b86e4a0148ddb757e8f71ff

Observation b9745d4b-f2bd-479d-a873-55cf963f0097 · inbound

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably cites this paper.

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably Data Engineering for Scaling Language Models to 128K Context

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:37:37.316022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T15:34:09.006815Z digest=sha256:bf5424e54881badf67891b91fa771acfa4eacdd69c6240102eaace977a48472d

Observation 43f4273a-6f40-4463-89ff-61c983e0ccb0 · inbound

WorkBench Revisited: Workplace Agents Two Years On cites this paper.

WorkBench Revisited: Workplace Agents Two Years On Data Engineering for Scaling Language Models to 128K Context

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.394049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T22:32:38.005254Z digest=sha256:6f222050c4ef57029cdfc48ef68fa94c93ebef54ee4a449711d1286d6487dcf8

Observation 05367672-8f9c-451e-adb3-b39e7e5022cb · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Data Engineering for Scaling Language Models to 128K Context

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.622643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:33389d2bc69189be4df0d5b2fb8cdb03d90724fcde83f66bfe1289a43d853fa9

Observation 6acecacd-d726-497b-b7b1-5f52e49b19e9 · inbound

Test-Time Training with Next-Token Prediction cites this paper.

Test-Time Training with Next-Token Prediction Data Engineering for Scaling Language Models to 128K Context

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:09:37.401933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T13:52:15.078658Z digest=sha256:d58a9ffae6a57241205aa8472b1a06d85b53d26a5451b5ba9a267b0fdd4df5d2

Observation b9ba3de5-15ad-49cf-b591-ea94623b2897 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Data Engineering for Scaling Language Models to 128K Context

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.443744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:01b866b085788b9096688566c6ef438808db9155353fdd87e92a3a96f2974669

Observation 9e310bf0-13cb-4690-8005-c75c3db0102e · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Data Engineering for Scaling Language Models to 128K Context

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:16.336372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:16.336372Z digest=sha256:1b670a8910705a1b4d267bf1be29d38d919ffbe8699139d0ecc33dd6f92f5d01

Observation 482814d7-bdb6-475c-962a-6d183153fa9e · inbound

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE cites this paper.

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Data Engineering for Scaling Language Models to 128K Context

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:07:33.538561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T19:59:46.713277Z digest=sha256:8487a885d818b20fcb374edb211e4a3644f3cbbd9de77742d5f3e28302d34231

Observation 8dd90981-1a13-4c1e-9f5e-b09ab1b7cc2d · inbound

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE cites this paper.

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Data Engineering for Scaling Language Models to 128K Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T06:47:09.927626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:47:09.927626Z digest=sha256:e46b62b8eafe97130ee70533efdd24c86f76bd5261e79bca190723071ffb60b7