Pith. sign in

Paper Citation Record · LEDGER

Core Context Aware Transformers for Long Context Language Modeling

As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2412.12465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12465 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:11:39.748585Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7a503cc-f96d-407b-9d26-69efd1364aca · outbound

This paper cites Same as demonstrated in existing methods (Beltagy et al., 2020; Xiao et al., 2024b).

Core Context Aware Transformers for Long Context Language Modeling Same as demonstrated in existing methods (Beltagy et al., 2020; Xiao et al., 2024b)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.105767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.731601Z digest=sha256:b6c5ac366d4dfd946982415f74781ab038a399e24a58db98e721e7c8a0b70dd1

Observation b396aef2-ebad-4247-940c-1fdd7b936e8a · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Core Context Aware Transformers for Long Context Language Modeling Extending Context Window of Large Language Models via Positional Interpolation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.602020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.602020Z digest=sha256:c526db7506d64eef5fba769153ab41b0a06aaf4c6ea8d72cc8ffc2503b33a77a

Observation f6846c3a-590f-4ab2-be44-6dd2f6cead4c · outbound

This paper cites Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers.

Core Context Aware Transformers for Long Context Language Modeling Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.611527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.611527Z digest=sha256:011c2bc316ecc17c33aef6657a1337ece8781d909416786d8a41873a7ea02227

Observation 2d528daf-9b2b-4e23-8339-e90b4599b224 · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

Core Context Aware Transformers for Long Context Language Modeling LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.616699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.616699Z digest=sha256:376150b6b215a48e9108f230c0fb91e49e6261d451155c2b3dd8f8e9edadd05d

Observation f06184fe-5287-4e72-8eca-abde9a2d60e2 · outbound

This paper cites Data Engineering for Scaling Language Models to 128K Context.

Core Context Aware Transformers for Long Context Language Modeling Data Engineering for Scaling Language Models to 128K Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.622030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.622030Z digest=sha256:dd3e972258458900e098a928807986f4b42a37530ff5424bea4409f324251208

Observation 5aa06aaf-5711-4cca-85b1-fdfbe526fe3f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Core Context Aware Transformers for Long Context Language Modeling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.626849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.626849Z digest=sha256:4b083a9a125db7ce6b365d415bf1353e6ffb4cf843148cadef2c8b130b07fe53

Observation 64ff5deb-44bb-45a3-beb8-45b7e5b24b31 · outbound

This paper cites s 256 512 1024 2048 4096 PPL ↓ 2.98 2.92 2.86 2.79 2.73 Latency ↓ (ms) 457.4 460.1 461.4 462.8 473.1 Effect of Different Updating Strategies.

Core Context Aware Transformers for Long Context Language Modeling s 256 512 1024 2048 4096 PPL ↓ 2.98 2.92 2.86 2.79 2.73 Latency ↓ (ms) 457.4 460.1 461.4 462.8 473.1 Effect of Different Updating Strategies

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.070137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.743922Z digest=sha256:2ad8829c3baacd724763ccc7976e2c3b43defe61cea445471d94b5f76f086816

Observation c2fed794-ad54-46dc-85a7-6ccf3308cfd2 · outbound

This paper cites an unresolved cited work.

Core Context Aware Transformers for Long Context Language Modeling Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:40.248786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.650230Z digest=sha256:a1f145d74059c26f6d1b0ae000831b9d2c3e4878f67360015ff5505cf291f5ab

Observation 1601e86a-3514-4c59-8df0-9c6a8056f98f · outbound

This paper cites GPT-4 Technical Report.

Core Context Aware Transformers for Long Context Language Modeling GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.654390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.654390Z digest=sha256:5f24319d42bc27720ec00257c57b4b2e8f6af16cacd067cce16240c8ef6791a7

Observation 5ff27d7e-ce5e-4b83-9e6b-e3a11d3f3680 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Core Context Aware Transformers for Long Context Language Modeling LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.668374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.668374Z digest=sha256:2b709d216ea75805b13095815812100cedc230c5bfb066ba5df9807f3a13d5cf

Observation 3935e5ac-b689-4170-93a0-c63a56593019 · outbound

This paper cites Qwen2.5 Technical Report.

Core Context Aware Transformers for Long Context Language Modeling Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.673740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.673740Z digest=sha256:78d1ab0c3908eb1b4516ad5cb62336213b5dbbdf2e82e456a1807427ac1bb8fa

Observation a55eb01a-948a-456f-8c96-6ffa2b6848ba · outbound

This paper cites The MMLU benchmark spans 57 diverse subjects, ranging from elementary mathematics to professional law.

Core Context Aware Transformers for Long Context Language Modeling The MMLU benchmark spans 57 diverse subjects, ranging from elementary mathematics to professional law

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.232662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.687129Z digest=sha256:925da7cbb6b3c8b08d8ca796396afce68130252c58df2f1387049475247b5d85

Observation d37893a3-37fc-46b9-acb0-76cc17b11a98 · outbound

This paper cites base frequency.

Core Context Aware Transformers for Long Context Language Modeling base frequency

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.216143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.691363Z digest=sha256:8338dfeeddb5563b47cde9b3d79427e508f5ed7f654cb2a4be1c2748cda900aa

Observation 064f819b-70ea-4a43-9ac2-460d48dc0ec2 · outbound

This paper cites This enables us to integrate our CCA-Attention as a standalone, cache-friendly operator, effectively eliminating redundant computations.

Core Context Aware Transformers for Long Context Language Modeling This enables us to integrate our CCA-Attention as a standalone, cache-friendly operator, effectively eliminating redundant computations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.198829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.697066Z digest=sha256:d73091e0532dda3269de1334d55ec6961d540e1c76accbad53347c7eea8d9eec

Observation 3518ab91-f9e6-4a8b-9662-049c403156cb · outbound

This paper cites Our experiments are based on the LLaMA-2 7B model fine-tuned on sequences of length 32K and 80K (Fu et al., 2024).

Core Context Aware Transformers for Long Context Language Modeling Our experiments are based on the LLaMA-2 7B model fine-tuned on sequences of length 32K and 80K (Fu et al., 2024)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.173335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.706475Z digest=sha256:e1fc6fa93f652bdcb21119b53f434cdadc552fa01004b9c11b69781f766d9c4e

Observation da97c928-c70e-4f82-b2f3-acba2bf192c7 · outbound

This paper cites 16 Core Context Aware Transformers for Long Context Language Modeling C.

Core Context Aware Transformers for Long Context Language Modeling 16 Core Context Aware Transformers for Long Context Language Modeling C

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.138276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.719498Z digest=sha256:872eace8c0bc1d7afc85fb383cb18cc501b14583d94e360f5dfc8a1bfb6d8318

Observation f413f8ae-88d1-4e42-88b5-ddf611e7213e · outbound

This paper cites an unresolved cited work.

Core Context Aware Transformers for Long Context Language Modeling Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:40.122478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.725236Z digest=sha256:26eb52c1ab44ca7d54f4fa4427b43481194ba0093981c8e710c8facae6af506a

Observation f7b4178f-21bc-4802-9862-1a136bc3393e · outbound

This paper cites Strategy Mean Pooling Max Pooling CCA-Attention (Ours) PPL ↓ 2.99 2.99 2.85 Effect of Group Size g.

Core Context Aware Transformers for Long Context Language Modeling Strategy Mean Pooling Max Pooling CCA-Attention (Ours) PPL ↓ 2.99 2.99 2.85 Effect of Group Size g

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.087172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.738189Z digest=sha256:10065b590515bc5c99d13e14364f2c5e809058332b6bd489024e6fc361628fa6

Observation ed662da6-82bc-4bd2-8177-f90118a375c6 · outbound

This paper cites The perplexity rapidly converges within approximately the first 100 iterations and remains stable over 1,000 iterations.

Core Context Aware Transformers for Long Context Language Modeling The perplexity rapidly converges within approximately the first 100 iterations and remains stable over 1,000 iterations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.051901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.748585Z digest=sha256:4b48ce786472b78dfc7727f92b64ad7f3eb018b4dddd1413065eff795e36c377

Observation 0e2e3f84-3068-4f03-9a30-9a7649d6668a · outbound

This paper cites an unresolved cited work.

Core Context Aware Transformers for Long Context Language Modeling Unresolved cited work

Reference 2000

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:40.152196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.715353Z digest=sha256:0c473b2cd1eaf404019ad51d00deecdb0d1485e11ddb2b77ed5fac70b42cd570

Observation 23c79c42-8f28-4294-b24e-edb9a870f05a · outbound

This paper cites DeepSeek-V3 Technical Report.

Core Context Aware Transformers for Long Context Language Modeling DeepSeek-V3 Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.644086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.644086Z digest=sha256:64452a0bc5330527a5d35f1bc1b56df6cb1290baaa08d8fed1d1388960ad9d22

Observation c43a855a-abd8-40e3-a063-3ab388aad145 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Core Context Aware Transformers for Long Context Language Modeling Retentive Network: A Successor to Transformer for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.663274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.663274Z digest=sha256:6562bfd6c6bd1c47282d3900fcd66af71f2e339e47a6bbfd35c472e168dcc12c

Observation 970488b0-c9e3-4845-b982-5811d154f108 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Core Context Aware Transformers for Long Context Language Modeling RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.638818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.638818Z digest=sha256:87b125b2bf435c01b328804f29a247d30d7a923fc8047a6e6a08391905396fc8

Observation 6c96ee85-bc28-43e9-b57d-72b3dbe75c84 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Core Context Aware Transformers for Long Context Language Modeling LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.585247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.585247Z digest=sha256:7749fd688b66bae9fa998c56a4de1266ebd921ac7177a4fb1b57b484ee24ea52

Observation 8e5a5456-13cc-4b3b-a090-19c543b30e14 · outbound

This paper cites Longformer: The Long-Document Transformer.

Core Context Aware Transformers for Long Context Language Modeling Longformer: The Long-Document Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.590791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.590791Z digest=sha256:3ee4a3c8cf2955e56f3646dbd21e018af8913d7f8da060778c1303571ba0ac20

Observation 780698b2-f707-4bcb-a74d-c9aeed4745fe · outbound

This paper cites Chang, Y ., Wang, X., Wang, J., Wu, Y ., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y ., Ye, W., Zhang, Y ., Chang, Y ., Yu, P.

Core Context Aware Transformers for Long Context Language Modeling Chang, Y ., Wang, X., Wang, J., Wu, Y ., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y ., Ye, W., Zhang, Y ., Chang, Y ., Yu, P

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.270750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:11:39.596635Z digest=sha256:8fc8c3138eabc97274632cae38972f2a3727f6c9ea1b0357a561e77b504b3182

Observation 81609661-bbeb-4667-b9fa-27153bd0abb3 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

Core Context Aware Transformers for Long Context Language Modeling LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.633204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.633204Z digest=sha256:100da87fec9610321eb32952ab8af067d347f33fd417fb68fe27ea2af9c22bd4

Pith citing papers

No inbound Pith citation observations are available.