Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Model Learning for Language Model

As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2505.11820.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11820 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:52:41.284896Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:37:02.506800Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:37:02.896712Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved31
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abbbde7f-0195-4720-956f-1adc82cd8210 · outbound

This paper cites GPT-4 Technical Report.

Chain-of-Model Learning for Language Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.857376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.857376Z digest=sha256:8d8642d3a66cf236e13845c2e22f958ecac01c31d28b3d7934bc10b6361fc5c0

Observation d5a25b4b-1cf6-4733-804b-e57b12b96e02 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Chain-of-Model Learning for Language Model Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.872328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.872328Z digest=sha256:bb7cdb65b621cc10dfa14da23804d9a68c0fd8c54061a32df0a1949f3ea8ff30

Observation d97bbf1d-5b57-4b0a-901c-b555979ce623 · outbound

This paper cites The Llama 3 Herd of Models.

Chain-of-Model Learning for Language Model The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.881113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.881113Z digest=sha256:539d7b6f80f08c7c0ee2bd172ace6109f6ed079ae780d898b20a991b9619e679

Observation 0fd15a3f-89b4-4fcb-b3b8-df6d8d05f0a9 · outbound

This paper cites Qwen2.5 Technical Report.

Chain-of-Model Learning for Language Model Qwen2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.890396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.890396Z digest=sha256:a29c51234714704bfcb20c9f5260ba76cf387b020efaa25ac2c1cf822833c60b

Observation 75b4661d-f670-41aa-961b-41c10a2a96d4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Chain-of-Model Learning for Language Model DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.898446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.898446Z digest=sha256:0db0d4c1a3d4ccb0f985a4e9ac0593cb716208d513ab3b3bb643442df810b7e6

Observation 10b05d0d-0a7a-4285-9cba-57a7fe2d5fe0 · outbound

This paper cites Mixtral of Experts.

Chain-of-Model Learning for Language Model Mixtral of Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.908323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.908323Z digest=sha256:4800860a7ae04ff1b8e190276a68ca902d4d6601def658a7c0cc4ecd95a317c5

Observation 5c433023-c089-4e30-a138-4955302f6d6f · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Chain-of-Model Learning for Language Model Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.930877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.930877Z digest=sha256:f8f291af11a594135fd97fa19ed6c5f861ff9bdf749e0456c70d6614747644ad

Observation 722e1e4f-e051-44f6-b4d2-b81d7110771c · outbound

This paper cites Scaling Laws for Neural Language Models.

Chain-of-Model Learning for Language Model Scaling Laws for Neural Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.940057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.940057Z digest=sha256:8ce65e9f1a78e82227d6751a3bb4766fa9b5933de50a9cd8d52029684d3e0f79

Observation 80fc95e4-6d0f-4a11-a28b-7bba91b5a4b1 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Chain-of-Model Learning for Language Model Training Compute-Optimal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.946608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.946608Z digest=sha256:b57039c295d69aac52f887195e08b1fd0a916d18fa680f689164f7b625f306aa

Observation a462056e-8faf-4cf3-b667-d72b1767e66d · outbound

This paper cites Once-for-all: Train one network and specialize it for efficient deployment.

Chain-of-Model Learning for Language Model Once-for-all: Train one network and specialize it for efficient deployment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.858644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:40.955446Z digest=sha256:179afb1e269fc2432cfddb7a1e63b868b2bc87c1cbee63e9a8bdeb20742bda6b

Observation 56ea14f7-3399-4f68-90f6-65c739fea9cf · outbound

This paper cites MatFormer: Nested Transformer for Elastic Inference.

Chain-of-Model Learning for Language Model MatFormer: Nested Transformer for Elastic Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.964405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.964405Z digest=sha256:363297712865d71677489de2d7874b4c75a29994216cfbc492675da53b90fb42

Observation c664ac5c-2a18-4a52-91d3-f15de5f217b0 · outbound

This paper cites Progressive Neural Networks.

Chain-of-Model Learning for Language Model Progressive Neural Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:40.974548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:40.974548Z digest=sha256:81d3909ccd25aac2c268268f0a79980c1d53758bd1e699855e9ed7548f8a918b

Observation 1a6842ff-5772-4a5d-a2d8-20b87a464afa · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.

Chain-of-Model Learning for Language Model A comprehensive survey of continual learning: Theory, method and application

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.832987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:40.985919Z digest=sha256:9b945773bf94e6e5f35f3dfd62a6cd606c2e9330f7bbbcec2454862be858a812

Observation 4df3d8fe-a01f-4f1b-a5a3-8ab1482e5bc9 · outbound

This paper cites Jacobs, Michael I.

Chain-of-Model Learning for Language Model Jacobs, Michael I

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.802286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:40.993911Z digest=sha256:88353002353d8e8c7eca60bc4908048cad3c066a9a900cb0b9250809b171fcbe

Observation 23d9df06-707f-40a3-8885-7a6357ad3fc4 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Chain-of-Model Learning for Language Model Overcoming catastrophic forgetting in neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.002135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.002135Z digest=sha256:a46a02c83148dea3aa26809eef0bef313cca10937ed0652bb54c55ae976c7111

Observation b99c5e76-b320-4892-b12c-79ef3e8816b9 · outbound

This paper cites an unresolved cited work.

Chain-of-Model Learning for Language Model Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:42.776990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.013740Z digest=sha256:f741fb86599c3879cddc227791c49d23e9c9ff894848dbeb16e9a07a8b7e5752

Observation 3ad68125-142b-4f70-8211-bb970947fced · outbound

This paper cites GLU Variants Improve Transformer.

Chain-of-Model Learning for Language Model GLU Variants Improve Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.020695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.020695Z digest=sha256:05344fb788a41326a739152ce366955f773992a4ed14489d45191e0f6da491ac

Observation f59da94c-ec05-458e-9f08-f5330899dd55 · outbound

This paper cites Layer Normalization.

Chain-of-Model Learning for Language Model Layer Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.031025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.031025Z digest=sha256:43554d965a54b96f5f67f746d053fad52d2e4b6150cdf4600b5a70bd2686517c

Observation 165e8c0c-3819-41cf-af88-d0e802ca05e9 · outbound

This paper cites Root mean square layer normalization.

Chain-of-Model Learning for Language Model Root mean square layer normalization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.753794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.038632Z digest=sha256:e2c5ed6decac9a8d9378ca404fa9a5476a28c278c386e03e791f4bc094b325cc

Observation 11306399-5191-48d4-94a9-93e790fe9433 · outbound

This paper cites GQA: training generalized multi-query transformer models from multi-head checkpoints.

Chain-of-Model Learning for Language Model GQA: training generalized multi-query transformer models from multi-head checkpoints

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.726061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.048140Z digest=sha256:c4b7efca6786bf317b8cf2a4392e2ee0fef7535f61dc2f23cc61ac4c5a9fa23b

Observation b84c8d85-6079-4baf-ba37-827b7a874cae · outbound

This paper cites Kakade, Prateek Jain, and Ali Farhadi.

Chain-of-Model Learning for Language Model Kakade, Prateek Jain, and Ali Farhadi

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.694462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.056049Z digest=sha256:c697d92cd66aa6870db5e33dd05f19545dbe012bf824c8d8611e8aeded3fa269

Observation a8916ea4-a03a-4525-9dd7-94411dce0994 · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

Chain-of-Model Learning for Language Model SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.668097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.063644Z digest=sha256:248792ee427ef0e57c56795c644885406c793c4d6d0015076fa74aa858d3ab51

Observation faea6002-b360-4004-a777-0409526b251f · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Chain-of-Model Learning for Language Model TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.071086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.071086Z digest=sha256:f31036659abe157fed30cb4795ce99b09e25c0f28fda9ff40986768b213dc250

Observation 6976c0f0-0768-4cef-8f5b-b1639c950ae4 · outbound

This paper cites Kingma and Jimmy Ba.

Chain-of-Model Learning for Language Model Kingma and Jimmy Ba

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.082611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.082611Z digest=sha256:e282e157be6d922d425b1a009e6e01032ad095bdd15556e7ed1be14dbda08302

Observation ea2a1343-7a2f-402d-b2af-3d5c27c27176 · outbound

This paper cites The language model evaluation harness, 07 2024.

Chain-of-Model Learning for Language Model The language model evaluation harness, 07 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.090652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.090652Z digest=sha256:112e82726b5f3994a4bd224ec036b4d1ec06734f3689903e76a3dffdc8bec627

Observation d5b5fde5-b2ba-46f3-a580-b752d45c2c50 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Chain-of-Model Learning for Language Model TinyLlama: An Open-Source Small Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.102903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.102903Z digest=sha256:d1f8dde3d31489752d86b50d0d8092aa753039675437341ce873d8163f516275

Observation cedbfe17-6b84-4a8a-af8a-5f80bd7f3e3a · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Chain-of-Model Learning for Language Model SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.111749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.111749Z digest=sha256:0f81e7963ada1a651c4e94394ce8b80c20b5f1c2d9be20ba2908c098e4dd9543

Observation befb0232-1991-4ea8-8f33-1d4024107bf9 · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.

Chain-of-Model Learning for Language Model Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.601767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.118452Z digest=sha256:e72d7455ba2a5c5079ce567eaac2786cd84136c636b5b47fa2619ce30ee3cdb2

Observation af2df12f-44dc-4c7b-9368-83fc4f263ddf · outbound

This paper cites an unresolved cited work.

Chain-of-Model Learning for Language Model Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:42.573099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.127077Z digest=sha256:b4b705c2fb434d7bebd1cc7b0d5ef9bf29f79c4860b1bb12a190fe3d66e77ad3

Observation 18736cb7-b63f-40f5-ae19-faeeaf988e73 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Chain-of-Model Learning for Language Model Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.138122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.138122Z digest=sha256:1af4430b5cf6cfa13928c18f55de2742a600fb8a54aaf22cf4d6672c1c0408db

Observation 7b07eda3-dffb-4262-9a38-5466b3e85294 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Chain-of-Model Learning for Language Model Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.150842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.150842Z digest=sha256:2ac08e8a633839f6f17831cc32f32c5758d62664f146c4368e11b5b9b54393c8

Observation 847c4791-6ec5-44ea-9c95-47730ee00f13 · outbound

This paper cites Wind, Stanislaw Wozniak, Zhenyuan Zhang, Qinghua Zhou, Jian Zhu, and Rui-Jie Zhu.

Chain-of-Model Learning for Language Model Wind, Stanislaw Wozniak, Zhenyuan Zhang, Qinghua Zhou, Jian Zhu, and Rui-Jie Zhu

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.522934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.162809Z digest=sha256:22f7018cae1dca64da0e4d8cca260b1c891359a48360d6f26fcad43c1a4a0d60

Observation ef9a3191-9eac-4563-a0d9-4905a2f6e81f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Chain-of-Model Learning for Language Model Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.174532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.174532Z digest=sha256:fa060e0994a56e30db71955ea14c93b985352742a33b12f128c2a10ad9e0a9da

Observation d4f4e4a3-dbaf-47e2-a4ca-28144f2aae66 · outbound

This paper cites MatMamba: A Matryoshka State Space Model.

Chain-of-Model Learning for Language Model MatMamba: A Matryoshka State Space Model

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:52:41.527132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.182571Z digest=sha256:339e39bb2abc0f510d4980fc36da0f107bf16aafd18f7fc03372de39511de63d

Observation a30bbcc3-cdce-4ad3-baca-28c84e8f40ab · outbound

This paper cites SOLAR 10.7b: Scaling large language models with simple yet effective depth up-scaling.

Chain-of-Model Learning for Language Model SOLAR 10.7b: Scaling large language models with simple yet effective depth up-scaling

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:52:42.495716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.190178Z digest=sha256:56dbcb1878cd3ea534bed94decd21b7d3a7baea1c23506f469d76ea5174b2768

Observation ebf54bf6-4610-48a1-9bc6-4278e32c8a38 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Chain-of-Model Learning for Language Model Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.200323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.200323Z digest=sha256:43bf6d0647323545d634b94d409729d89b46aedf2a28a56be386202c74f9d683

Observation 6f563703-b9d1-4090-8860-27e75cfc7e83 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

Chain-of-Model Learning for Language Model Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.436336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.205348Z digest=sha256:9e63ed563879c26c19ca09a19c45a567fa42d9ddf98a0b6bca3236c264bdd016

Observation 2e28bc46-2212-415c-a9ce-940a07d4df11 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Chain-of-Model Learning for Language Model Splitwise: Efficient generative LLM inference using phase splitting

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.213034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.213034Z digest=sha256:05591290c1ae6d3d7fb0951a45d596cbdd1ea391d607315f851e9b1e1150d740

Observation 11b1bb0b-6f09-491e-b884-8ec82646ce65 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Chain-of-Model Learning for Language Model Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.222257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.222257Z digest=sha256:32ea266fe54d0b05f80783066c218995f0845a0580ee87dcd218efb7f8901cff

Observation 3b9ddda9-d54d-4d83-a1ad-ac306117dec9 · outbound

This paper cites Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich.

Chain-of-Model Learning for Language Model Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.369859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.227463Z digest=sha256:37c176b11eccbce810247e222d428f34119dcaac119e7ecf5034caea6f973e5b

Observation a20dabb0-5582-4279-9859-739d64e7499b · outbound

This paper cites Reducing transformer key-value cache size with cross-layer attention.

Chain-of-Model Learning for Language Model Reducing transformer key-value cache size with cross-layer attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.332296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.233175Z digest=sha256:8842c4d44e3f15cc492e5de7b2f68caae7cb0e9b0abe8a3854035b33a083016e

Observation a9710453-8852-4674-b3bc-4e910f7e6aba · outbound

This paper cites MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding.

Chain-of-Model Learning for Language Model MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.239135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.239135Z digest=sha256:c1010e23301908789acae208801ec40551d592ff9a08073ec3955c52d1efd1ac

Observation 3e0cde71-0ec7-494f-a06f-fb168f2098bb · outbound

This paper cites You only cache once: Decoder-decoder architectures for lan- guage models.

Chain-of-Model Learning for Language Model You only cache once: Decoder-decoder architectures for lan- guage models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.297450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.247577Z digest=sha256:0473e1d0babccc224af472500ad15a6f303e8de4156d7a69438379af1e8755a0

Observation 9e697198-509e-48fe-8225-0f2ef22e9333 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Chain-of-Model Learning for Language Model Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.254335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.254335Z digest=sha256:fe052d5562a477a5dc260dfb20383abbbf965b964cfd421984198e17bc71e18a

Observation f92af436-3ca2-40ec-8f50-e9f62706aab5 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Chain-of-Model Learning for Language Model GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:41.260286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.260286Z digest=sha256:09ddbbf594fd25bc1f653dfe504587193b2baa879d38ba30aada2cd35a0967fb

Observation 5370875a-42b7-4f5a-aa9e-12db040f7b8e · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Chain-of-Model Learning for Language Model DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-15T20:52:41.270994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:41.270994Z digest=sha256:c01fdbe57825c59a581804a8000b5babeff851d1db03e3f820788c83e462937f

Observation cee2a940-4c94-40f0-b3c0-c9b8447aa4ef · outbound

This paper cites an unresolved cited work.

Chain-of-Model Learning for Language Model Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:42.269462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.278882Z digest=sha256:8a35c4e81832b7cc63cc179412999dac442fdc40b3f1e2c1e7134bb588a8e7fb

Observation 7b62d4f2-27c0-482c-9101-bfc985536330 · outbound

This paper cites "" dim: Dimension of input features hidden_dim: Dimension of output features chains: Base of each chain bias: whether using bias or not.

Chain-of-Model Learning for Language Model "" dim: Dimension of input features hidden_dim: Dimension of output features chains: Base of each chain bias: whether using bias or not

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:42.237595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:41.284896Z digest=sha256:00615586c6e8846f2aafd9178ada5a10f7e3c688cb58258ad80a607a11589015

Pith citing papers

Observation e497e98a-88ab-4ef3-b3c2-e6f89884989e · inbound

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning cites this paper.

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning Chain-of-Model Learning for Language Model

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:02.904533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:37:02.506800Z digest=sha256:49c774dcb29a3d5c16c32e41acc53623985a52a2e5b58dbc5c1d8584c49633bc