Pith. sign in

Paper Citation Record · LEDGER

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.12419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12419 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:13.242455Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10ecf9ea-2e27-4767-99a6-8c3b2fde81fe · outbound

This paper cites ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.103205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.103205Z digest=sha256:d8a8636ff1c37475d37078e8f0da60baf648e0ad8457ae2cd43c4b006e668f88

Observation 0ffdd637-eb79-4e81-9bfe-41a66ce05bf8 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Generating Long Sequences with Sparse Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.124172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.124172Z digest=sha256:d7d46ad6398eadb154e0154e70f2551ed1421ebddb905a29062df7488f5f6e40

Observation d8e9b861-ffff-48dd-aa7c-c523e445bb0d · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.139255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.139255Z digest=sha256:2c7c846bfd299fc0d00001a6dce9f2b2c4852462862332fda351c89512cfb7c3

Observation b7c9018b-0724-4c60-876b-770efc6c2e16 · outbound

This paper cites Mixtral of Experts.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Mixtral of Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.149309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.149309Z digest=sha256:aa494c8580b4c71a3369d6563a09c1c8e461b7f19a104db7bcc03127d0749e48

Observation f1277a8b-720f-455f-93c4-adf05d28d50b · outbound

This paper cites LM2: Large Memory Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining LM2: Large Memory Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.154209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.154209Z digest=sha256:ec11a3794768eab356aff7ce59f9d270a1fd530632fd16bfe6c047290d83f875

Observation a88d323c-1b27-44f5-9edc-13d59e00b0c0 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Cmmlu: Measuring massive multitask language understanding in chinese

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.819645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:13.159388Z digest=sha256:973146d98e169ce8c6c67af1ab9b707160bd6db3a810b8e1d11956bc09e7d7fe

Observation d4f6923e-5c28-47b0-88b4-46c00825abb2 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.164298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.164298Z digest=sha256:3078185e04894df74ce86cff45e9be23a0396f92b1981faf340315821da4d2ef

Observation d674f108-fe07-4546-b148-6356289e6d4f · outbound

This paper cites YAYI 2: Multilingual Open-Source Large Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining YAYI 2: Multilingual Open-Source Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.168892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.168892Z digest=sha256:d6139d6d6e24813d2390d9a5df2c66ef14e7409c58bd94d332bfbf41237c1610

Observation 5bbb0316-3e14-4ae9-b64f-1085d797167f · outbound

This paper cites and Lin, S.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining and Lin, S

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.173742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.173742Z digest=sha256:d56f9d46efc7880ce99ec0a6925876812cd5a5e30654b7ca828d7553d77a4700

Observation fbefe8db-0d73-4a7d-bc6c-a1a9659ed2ef · outbound

This paper cites D., Man, H., Ngo, N.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining D., Man, H., Ngo, N

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.801749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:13.178158Z digest=sha256:11b619cc4099c265671e995a27faf81ee55299a697eeb428a4007f1b96697b42

Observation 3972f525-b459-4225-aedb-07acd82cc0a9 · outbound

This paper cites GPT-4 Technical Report.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.183149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.183149Z digest=sha256:ca107e39526df6fa0eb2121a294cdb731c699238c15e95158aa39ac9d178a292

Observation 083e52af-51bf-4f5d-9672-fc0c05802e3e · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining MemGPT: Towards LLMs as Operating Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.188096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.188096Z digest=sha256:1098afc1d2808f3282ad742b18ec9f6f9a16c87b29db2a685f8abfcf44286af4

Observation f5c7008d-2b31-435d-bf08-e1f8ea2e907f · outbound

This paper cites Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.192754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.192754Z digest=sha256:479f234e74444f5ea242cca4f72f42d783eb31e6e5838b3a34544bdb525890be

Observation 1e57663f-f5c3-4323-8c67-1a5822d3e045 · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.197652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.197652Z digest=sha256:be2d75c8d771c0bec04d144de2ed0b35d3902b1c4fbc1f363eae1e8f42d60afd

Observation a1af0f49-2fba-435d-94f6-2d4a7a6c8f49 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.202478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.202478Z digest=sha256:daf0a4058a6f05f902a748fdf028d63ac978f4238de5a3235cb6a03f0906f9ae

Observation bc1feda2-20a5-42ac-9a25-f641f915325e · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Skywork: A More Open Bilingual Foundation Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.207202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.207202Z digest=sha256:37b2a93d3ff764fb144263c685806e79c402069a07f06f23a846adf25c8c2d20

Observation 7e3d6708-b9fc-4208-b9f9-f5e2fda05e47 · outbound

This paper cites Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.211915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.211915Z digest=sha256:539c094e9f9536de82712e1948944252985270a6ed2e77a90d147b5180cb0cc8

Observation 9c22f5d8-e225-45e1-ba7e-79e2b7b284ce · outbound

This paper cites Qwen2.5 Technical Report.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.216593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.216593Z digest=sha256:683ac201c8a56569cc01b82f55e7f9db07db957db0b06a033046595efe73652b

Observation 825b28e8-4f0f-47bc-b8b7-d613da9a8ad1 · outbound

This paper cites MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.221795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.221795Z digest=sha256:540c20aad26351efaea0310dfef0923a0baf59e409e2d5dc3a6e67317091864c

Observation 7d9df4fd-26f9-4f69-aff4-d974594451fb · outbound

This paper cites an unresolved cited work.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:37:13.771403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:13.232419Z digest=sha256:5cac618b8595927867c9e3ffb616f8581f9695aa19bdfff338bd66ddb5226891

Observation a366ad2c-0efd-4a46-9fbb-8d00dbd114e0 · outbound

This paper cites It contains 7,473 training and 1,319 hand-written test questions, each requiring two to eight sequential reasoning steps.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining It contains 7,473 training and 1,319 hand-written test questions, each requiring two to eight sequential reasoning steps

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.755433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:13.237384Z digest=sha256:f2eb75b266031ae62699a8de3d2a43fc70ff951b68286e5b93a5fe4d1ae23e66

Observation caef4196-23c6-4ca0-a02d-71256b7aca81 · outbound

This paper cites Newton,” “Calculus,.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Newton,” “Calculus,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.738655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:13.242455Z digest=sha256:09e575ef0f1386c476b6c2790cd52c1a7ecc059d3b1e0aac6516e8e82b6b1ad0

Observation 6953884a-d2cc-4c52-8824-39dea6a68f0a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.134132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.134132Z digest=sha256:166c556b3ee1b73c3f9757af27e7abc67049ad09faf5bcc0af85ed50452d2901

Observation 1a5b64e7-082e-4341-af5e-da75f2a443b5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.129282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.129282Z digest=sha256:82321e41b60e297a2540d591f4d53c25aa15d95d84b8cfea0435234198d75352

Observation 220ca2b1-b40f-44bf-a4fb-4d2a54465cd1 · outbound

This paper cites InternLM2 Technical Report.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining InternLM2 Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.097937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.097937Z digest=sha256:f3b02e6a675352328fcc558ac8a1458dddda277a4698ec0d82bac7f0ff11aa55

Observation 94cda244-ea9a-499e-a5a3-27ddf39a7204 · outbound

This paper cites ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.144294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.144294Z digest=sha256:b4ddc45e640bc09a447eb356e88ae787d89f20d941892cd68ef892bfe88cecf6

Observation af8ef260-a15d-4e14-91de-e2b3839592e9 · outbound

This paper cites We organize our supplementary as follows.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining We organize our supplementary as follows

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.786856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:13.226816Z digest=sha256:d048cebb8ce695641af0c9a0f0c7d88bfbce8878fc2323f28b84326e7d0f8d5a

Observation 84fed58e-b464-4c71-8c58-4dac96563597 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Evaluating Large Language Models Trained on Code

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.108617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.108617Z digest=sha256:5b2dd79df41ff0a40746b7606dec634e65bdd720311c872061932da38719e58c

Observation 8a340e0f-f967-49f9-b52a-5820bd73ee6e · outbound

This paper cites Longformer: The Long-Document Transformer.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Longformer: The Long-Document Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.091654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.091654Z digest=sha256:87c8eefdafe4c2042ba121d84d4de6af4d643812f40281b4c1c55e944dcfc56f

Observation 40db4f9e-3cbf-42d5-a60e-9c119a3c28d1 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.113611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.113611Z digest=sha256:5473420593c2edb2de63c7033be7ad67549f964dc24cf05abd038cce9476188f

Observation 6ff30059-21b8-4350-8253-f25a27c34997 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.118951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.118951Z digest=sha256:8cd56d8cc7303ac7e3f385ee8369ae04180b6b13acb8782e1dbf7ec70dbd94cb

Pith citing papers

No inbound Pith citation observations are available.