Pith. sign in

Paper Citation Record · LEDGER

From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.06214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.06214 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:16:00.996148Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.180712Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dfb3ad48-2a43-47d1-9fef-2f118a452604 · inbound

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention cites this paper.

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:00.996148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:00.996148Z digest=sha256:a000757d3eeab08f6dd55ca658e9ebdcc30ea5e6ff99a2e997698772a8e2b9f6

Observation dd34b158-bdcf-4517-8278-3005e3f6e66a · inbound

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation cites this paper.

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:45:28.120415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:42:00.440049Z digest=sha256:4fea74eeb800a5a257c4a9d4b8eada660eeea17bc7e6e1202c5fe534894c5640

Observation 20ccfe7c-cfdc-409b-ad11-236c93f7e9b3 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:05.154164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:5d4069562fd1d01b46b082218d6de39e139ff503da99f799e4dfc0a7740cb985

Observation 2c63f0d9-3753-42a0-b754-98863efae634 · inbound

How Many Different Outputs Can a Transformer Generate? cites this paper.

How Many Different Outputs Can a Transformer Generate? From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:11:12.744720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T07:09:23.107309Z digest=sha256:a85a5445dd80758972a79bca24c5be21dc908bd5621050b0dd318d8f7ead808f

Observation ea8533d5-77ec-4638-b380-70359f522517 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:03.182027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:eaa77784ec0dcfca413aaea6167b346202163dd5bf7ad63b6b52cbda99ec61af

Observation 62c3ea62-a6fd-4bec-9716-acb0b19113e2 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:15:58.911311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:0ae87e8d78bd83e672baf8d050e3387b92997e8a50bd90578088fcdd6edcfe66

Observation 8532207a-9dd6-4a16-8ac1-feafc98ff1d3 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:09.146522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:09.146522Z digest=sha256:c6fb12e727a59220a3271b01ccbb4432d2185bd7fd60d530c2142c97112b7aa7