Pith. sign in

Paper Citation Record · LEDGER

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch

As of 14 August 2026, this Paper Citation Record lists 100 of 165 outbound references and 4 inbound Pith citation observations for arXiv:2501.07124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07124 v3

Coverage vector

measured 100 of 165 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:54:02.331862Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:18:00.321479Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

100 of 165 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 353c996a-49a5-4e25-999a-9704cb14b243 · outbound

This paper cites 01-ai/yi: A series of large language models trained from scratch by developers @01-ai, 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch 01-ai/yi: A series of large language models trained from scratch by developers @01-ai, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.955542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.955542Z digest=sha256:ef0f38ff4e8ec8aff91a1916b3491a7233855b9c3bf085b87c473642ab93df3c

Observation 26c84eaa-e468-4105-8b3a-b46224bdf5a6 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.959421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.959421Z digest=sha256:3b1b47a94e3867ef2a2d5bb37604ed566685ba46399aff2a9e3e385eae57906b

Observation 5cbd8f9c-904d-4d6a-bf0e-3a628577094b · outbound

This paper cites Nemotron-4 340B Technical Report.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Nemotron-4 340B Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.963309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.963309Z digest=sha256:7391a5f17deb1c16e3697bfc831afbbd54f784a5dd7dfad36e19aa70fcabecca

Observation b16dbe74-6fee-486c-8ee9-a708f81ba64e · outbound

This paper cites SantaCoder: don't reach for the stars!.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch SantaCoder: don't reach for the stars!

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.967459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.967459Z digest=sha256:4af1444f5c9477c882a293c2d1e913e0d2c18ae525b1b0438a2f1ac25d9ab743

Observation 7b8d6cd3-4c44-4ac9-a8f4-69e8e4a6b565 · outbound

This paper cites Smollm - blazingly fast and remarkably powerful.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Smollm - blazingly fast and remarkably powerful

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.971970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.971970Z digest=sha256:1e7bff3b660932edadf5a2749d4762000ad68753a146bf0da12637d4a45f324b

Observation b9908567-1d0d-4b41-8e27-64a7bc2147e9 · outbound

This paper cites The falcon series of open language models, 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch The falcon series of open language models, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.975834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.975834Z digest=sha256:c12b7816f4087e60af017e0222c47afc44a20182737fc4f830188471c4831233

Observation 459c08ae-b3ec-4f36-9a60-a65d1a30a75e · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.979803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.979803Z digest=sha256:f60ae800d145b35bd808b372fc05e4122db999be7f67522860dd23b3c6a6f2c3

Observation 677cf518-4f43-443e-b81a-4ece58f79b7c · outbound

This paper cites GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch , 9 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch , 9 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.984191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.984191Z digest=sha256:9eadee76e3242d14183ad67790ca03e94df8fe52321319046473468aead1fcbf

Observation 5130e73c-5385-4d44-8658-a5a94aa78cb2 · outbound

This paper cites PaLM 2 Technical Report.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch PaLM 2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.988432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.988432Z digest=sha256:603592e6d3d38ddc05827a9303cde511e94c580910090ea3d95bde407f40aa1b

Observation 682bede8-499a-4b44-a522-9ba7dfe1c06d · outbound

This paper cites Model card for claude 3.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Model card for claude 3

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.992058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.992058Z digest=sha256:012a0b11ef36ee6505490d87514a5666bbefd91c87528d0be02fda6b75a78615

Observation 3bc24cd8-44e7-41a9-b8df-30c5b5e53710 · outbound

This paper cites Program Synthesis with Large Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Program Synthesis with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.995171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.995171Z digest=sha256:5a2390f9565b21b323806fa0cf551fd3cf05bd7ed4af5981ea14b884d5e3796a

Observation 4f63f515-d423-47fc-85fa-f112e1074f35 · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:01.999295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:01.999295Z digest=sha256:3c9344083d2e5264acb673906f9a569e2f06671895f2c66c4c94592423b0b5a4

Observation 37c5bd7a-7e0d-4beb-b9eb-bc1be24a34cb · outbound

This paper cites Qwen Technical Report.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Qwen Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.003077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.003077Z digest=sha256:1b80d0040da55cf50fe521870e5f23687e5621dcf5e507f62a269ca7efb1d1a1

Observation e68d8bd9-e96d-472e-bd0b-f254ccfbfccf · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.006684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.006684Z digest=sha256:0f83daa853b8c67952a3456c2e98c390720dc95e782ba8e6cbb000fe9c154757

Observation bdd10a75-959f-4984-9a84-51d3010e6953 · outbound

This paper cites Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.009921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.009921Z digest=sha256:6351167306564a23b4c203f6f137880c4462fa7412afc75ef0552ec8e2386e94

Observation e0b9ef25-a8b6-48d7-ba3c-c601cfbea8b4 · outbound

This paper cites Efficient Training of Language Models to Fill in the Middle.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Efficient Training of Language Models to Fill in the Middle

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.013010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.013010Z digest=sha256:189146da8db870e790543f9acbc2761df0e98b1256c8048234270b1665464059

Observation 733e0b5b-8126-49bf-a5d6-9ed6adf2cb99 · outbound

This paper cites Infinity instruct.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Infinity instruct

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.016393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.016393Z digest=sha256:e1fc7fa50296d9163b20599102fc300984eca21c9b0bc49c5b03625079ec8341

Observation d32ee3df-d104-47ef-a4d0-887de42d0615 · outbound

This paper cites Stable LM 2 1.6B Technical Report.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Stable LM 2 1.6B Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.019512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.019512Z digest=sha256:665d2053cb4210cf61fb10e10b5d30c3d682d84b559f0feb11f5dec78e45f9f9

Observation a929b4cf-f6fd-462e-8488-1ffe950eebfc · outbound

This paper cites A framework for the evaluation of code generation models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch A framework for the evaluation of code generation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.023089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.023089Z digest=sha256:757c688b7174ee5b686be645768bb745d6539f2f296b7f84b52d751f8e75f565

Observation de808395-7c4c-4c84-b40d-ad225cf6b9a2 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.026489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.026489Z digest=sha256:313e08dd3038dbde826a1edfee015104415474564c90e2c4636aa047172071ac

Observation 34a0f1f4-db22-4fa9-a52c-942427e07f4a · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.030313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.030313Z digest=sha256:0b5438ff4652f6ea4a5f5577d9fe45c5cad0adbf6cff619900f3f6ce6c38a33b

Observation 3ede848b-1766-4ace-88b5-a8b9c5e9e298 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.037382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.037382Z digest=sha256:73994e08b4705b0d669bbfcd2df3c69a19db136673f0c29badd9a161da70c055

Observation f9c5f72e-490a-4b04-84af-3cec111c53e3 · outbound

This paper cites Emergent and Predictable Memorization in Large Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Emergent and Predictable Memorization in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.041642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.041642Z digest=sha256:8e4141c9ac693f08b5b059542ade8345f1ebe03210af6700ae3e9da472e30934

Observation ad9d6489-289d-4997-ac5c-ec97a7927f3c · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Pythia: A suite for analyzing large language models across training and scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.044863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.044863Z digest=sha256:d25d9adbacc0314ecfa5eaef3227aa86fb5249551e7676e1c4b0f31c65fa9dca

Observation 62104c6c-a437-4c88-80fc-f7098987ac89 · outbound

This paper cites Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.048149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.048149Z digest=sha256:06eb3b206e2154beb3fb41f92d192132b4ce6a75d192bd9917c05946d731e81f

Observation 908d7370-691f-416b-bbad-3a5670b12da7 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Piqa: Reasoning about physical commonsense in natural language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.052615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.052615Z digest=sha256:8182314043649b3bd93f26e05a93fea5cbcd52dcc825e9d9d9ed49244386ad70

Observation 9f4df066-a5cc-4c95-b81d-681f5ea5b8e3 · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow , March 2021.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow , March 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.055732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.055732Z digest=sha256:288d4072d4db3b478f7528397cbafc0bb58361bb8c4679b94d788bb4dca7e864

Observation a11260a7-cfb1-423d-baa6-241e93d60615 · outbound

This paper cites Gpt-neox-20b: An open-source autoregressive language model, 2022.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Gpt-neox-20b: An open-source autoregressive language model, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.059993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.059993Z digest=sha256:60f82b48440c9edaff11afcf013da05abc725345b799c14fa9698b21e6804f30

Observation 3e2eaa45-bdb0-42fd-bae0-e1210afce069 · outbound

This paper cites Language models are few-shot learners.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Language models are few-shot learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.063408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.063408Z digest=sha256:f613af9ccd190f07f103cf09970bcf1711a988ed08960838e5c550d219822df5

Observation 2d2fe6f1-c841-45f2-a154-7985062c84be · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Evaluating Large Language Models Trained on Code

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.067418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.067418Z digest=sha256:bff0a8747f31bc959b0c0e00d786d7bd043824b2b769c43d052ed4b9f29fa3b6

Observation 2320637d-67ee-4f2e-b709-1c9e3d35ce82 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.071072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.071072Z digest=sha256:75dae388806961c5b2188e348a74d7083b54536da5802b6e95423c1ef7c5c12a

Observation 71e82385-988b-4d51-9e8a-bc6ca1b06298 · outbound

This paper cites Berkay Celik.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Berkay Celik

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.074985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.074985Z digest=sha256:e53c1551c4c3bc22205a57c81f41c5f7bdd5135d913b6bae4359453f3f2098fa

Observation 11fceb6e-415b-49f5-94fd-6366f32272ab · outbound

This paper cites Palm: Scaling language modeling with pathways.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Palm: Scaling language modeling with pathways

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.078843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.078843Z digest=sha256:50c6078bfa36572003d41b342fd9082e51d69e30d0bf69b7b206cb11726017b9

Observation dd6e64d5-4ce1-4f4d-abea-fe6a2cff3d28 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.082316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.082316Z digest=sha256:6d930acb76f716c362691a6b723f20e0e78d1082059a52cd143cac5ff7f7725b

Observation 745ddf9b-c294-4ab2-8a07-e663e61bec8d · outbound

This paper cites Claude 2.1 model card.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Claude 2.1 model card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.086091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.086091Z digest=sha256:4879221e376535bd712336a570287a3e6cefcc4496e775ff116dbd26de020b24

Observation c0f9f6ed-ae70-4d8b-93ae-3a1f3deef052 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Training Verifiers to Solve Math Word Problems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.089721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.089721Z digest=sha256:af19022159a8140f4b995df8123efc24bd70fccf46aeca8d7533f465a5b632ed

Observation f8072b34-19f6-4572-999c-fecd73c524d4 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.096331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.096331Z digest=sha256:dfa5745677ded2bfad282b1f2fe41c3828f0d4872c81f6e4a552ae8d0df71856

Observation 620c4cc6-e3c9-4eca-a133-013a26907290 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.100332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.100332Z digest=sha256:46d6a3c4aad2441e8b8daf7b4742f0dcea9d8ced253ad82286d38289e2d4c117

Observation 9e3c24f1-d64a-4c22-9cc9-8dcdbfa8247f · outbound

This paper cites Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.104556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.104556Z digest=sha256:f1716689eb1ce525b5d278756135a037e9da7afadb703be357875a60f9dcb94c

Observation 500e964a-2772-4318-b08f-e8986998d485 · outbound

This paper cites Mpt-7b: How we scaled to train the world's largest open-source model, 2024.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Mpt-7b: How we scaled to train the world's largest open-source model, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.108199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.108199Z digest=sha256:45aa8ad31ea66bc86ac31472e95876f5df6f4ecffd95f38f95810cc64e188702

Observation 3c526ebc-31c5-411a-950a-8a978bb560d9 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.111522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.111522Z digest=sha256:78a6ce5a905bb108ffbfb586d3d901f280152e5fb9eb6ef9a31974e0f8f2776d

Observation 17ce54dc-9bdf-478d-8e8b-159b4dfa5895 · outbound

This paper cites Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.114677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.114677Z digest=sha256:c384b1f09cb11565c4a520953eff2c8531daab6f59c5cf6326d7f66e8fc1c5cd

Observation 665d302d-37e6-4b56-992b-cbb4e5602f87 · outbound

This paper cites Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.118352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.118352Z digest=sha256:686980b09b7b7bae30d61b422209dd54ec33e6a1514531f1c73fd0292e59fcd9

Observation 1802e637-0315-493f-abe2-81c4158263af · outbound

This paper cites Measuring the carbon intensity of ai in cloud instances.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Measuring the carbon intensity of ai in cloud instances

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.122269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.122269Z digest=sha256:c0aeb1df851d81c621fa65a3b036184e807215ec96a73a9b2ef8606789861c97

Observation 2c49ece4-f296-4f06-ae2b-8a317b4b6d0a · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.129330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.129330Z digest=sha256:38b2114ae7b962f1419051f9ab4d600b976b5b50b7336b4fa7de80ba265179a8

Observation 3cc63d1a-799e-46e8-8f24-a53bb8ac485f · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.133111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.133111Z digest=sha256:fcdba1e09559699dbc4d82c7829e8489c10f345ba161dab015add897897d3499

Observation d7d70984-70d3-4e0e-840e-47c57aa7773f · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.136349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.136349Z digest=sha256:f7377d70d946129b478241becd7977f6ab09a15e375d3fbe369d30b8d2b7bacc

Observation b77448b8-6375-44ba-9d28-855f93286586 · outbound

This paper cites A framework for few-shot language model evaluation, 12 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch A framework for few-shot language model evaluation, 12 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.140349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.140349Z digest=sha256:962b215b3cdc358af164c0d28307a63f05e711dc3bd37ec164263697b77f29cd

Observation 4b5f0ed7-9e0a-47b2-b4fe-04e8dc1286ea · outbound

This paper cites Openllama: An open reproduction of llama, May 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Openllama: An open reproduction of llama, May 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.144639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.144639Z digest=sha256:dc03f598e654255aebab5cad0b7f871dc27bc2d6b0229e0eb6660083d4a8526e

Observation 9faaac25-56ca-4b78-88ce-6602d87306ab · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.147830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.147830Z digest=sha256:bbf649ea8baa35990839fe63ad78550f9056a43b7d3d49983d3f0fd6a0f626f7

Observation e02df997-f95e-4d92-97dd-3e27849233d9 · outbound

This paper cites The Llama 3 Herd of Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch The Llama 3 Herd of Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.151311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.151311Z digest=sha256:98df2215f86b5234c93d14e854ce165dbeaa77a5c772d7360d7181be095d64e8

Observation 41635d41-28b7-476e-80a8-269d0057f979 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch OLMo: Accelerating the Science of Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.155176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.155176Z digest=sha256:7997899197dbabcc1e3f0fd2e7c6619dddd2aa625e01cda3cff584f0f48cc35c

Observation 99c9f3ce-63b8-4508-bcbd-56b96cc00e00 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.159446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.159446Z digest=sha256:70685ecd9f51bcdd81f126c518183562a35f2dd9fb9664c735917c5841186397

Observation cbe8a428-2578-4ee3-9822-7cecafdaa0a4 · outbound

This paper cites Textbooks Are All You Need.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Textbooks Are All You Need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.163641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.163641Z digest=sha256:6cb39e9e791723aa0eb8e32b6f9eddf8eadb6b8d5851b93e1e863790de8b228d

Observation 377b48e5-89ed-4b0d-a99b-c6ab94cf3c84 · outbound

This paper cites Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.167170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.167170Z digest=sha256:975926ad960b506438a728c00271474e9bb8e0f2c224f581a1d6895ed232ae5e

Observation e37e348a-888a-4f97-ae71-b057a7d4ea76 · outbound

This paper cites Measuring massive multitask language understanding.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Measuring massive multitask language understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.171187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.171187Z digest=sha256:955dbf55830623827ffc2e576670481eaeb65d7802db0b46c949caa0dae09a6c

Observation f3d54395-f465-46cc-9e3b-3821f6a517ab · outbound

This paper cites Scaling laws and interpretability of learning from repeated data, 2022.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Scaling laws and interpretability of learning from repeated data, 2022

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.174107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.174107Z digest=sha256:4e16d4a9c1ce6a1cb92308fe1741a3584cae6ae7a67bb3254e8255c358083615

Observation 5325d773-ce83-463a-b940-e68cce82adcf · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.177541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.177541Z digest=sha256:ed4518df89a1e31d1d501505382994d201faf7f5cf06e72d7a8a50fa0505dcba

Observation 019bd9f7-cc7d-43c8-9bc3-a389d6834dee · outbound

This paper cites Mistral 7B.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Mistral 7B

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.180688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.180688Z digest=sha256:229ff2f521c1950ab551f8fa9b0cb4c80bdce0cc61214cb087b50a42440db8e3

Observation 04a9b181-53cc-4dbd-87d9-511039fadd5c · outbound

This paper cites Mixtral of Experts.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Mixtral of Experts

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.184639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.184639Z digest=sha256:65b033a311779d6fbef2ba8008c1bf06e9e0a6f9e3e77f3c0c33959526652c4f

Observation 15571009-fc40-49e2-b260-bef286dbf032 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.188363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.188363Z digest=sha256:f1d1485da2063cc65056d6fa2eabc9abe9a5aa3710cf14146d021c2ab48ef583

Observation dc8b8ae4-7d60-429a-bff1-b104b48a1384 · outbound

This paper cites Pubmedqa: A dataset for biomedical research question answering.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Pubmedqa: A dataset for biomedical research question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.192535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.192535Z digest=sha256:2a02db0d1ae63e59f762a143589cdbf782e910b11b901530fd255c4a1a4f87a4

Observation 8a727b01-f1b0-422f-a0c3-e0b97f204c1d · outbound

This paper cites ProsocialDialog: A Prosocial Backbone for Conversational Agents.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch ProsocialDialog: A Prosocial Backbone for Conversational Agents

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:54:03.247928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T20:54:02.195859Z digest=sha256:7c27f75afd449000ec9a7c14840c172643a0f3c13aafa6cdc120b454b9a66272

Observation b145c6c1-3727-4b12-90a0-e0f9b34b67e0 · outbound

This paper cites Reducing activation recomputation in large transformer models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Reducing activation recomputation in large transformer models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.199547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.199547Z digest=sha256:e1c5ad61a87df00444ab3d5f6f4df5ed7357f6dc7b60cb036bcef2e1d65924d1

Observation d80d1715-a930-418b-af5c-4fb7b296707a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Gonzalez, Hao Zhang, and Ion Stoica

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.203992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.203992Z digest=sha256:cd82fb3c7427e1e0fe4d6f7adca02f13db80cd53569340424c987c8e36a671bb

Observation 867df05e-ff7f-4362-96f3-e6604ffbbaa0 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.207771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.207771Z digest=sha256:1763901ff9e95c37b8baade1b2cf6a0020e5e264855acd6552ebff4d6a8f2390

Observation 79945661-1790-45bd-b89b-76ecfc9a2961 · outbound

This paper cites Amp: Automatically finding model parallel strategies with heterogeneity awareness.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Amp: Automatically finding model parallel strategies with heterogeneity awareness

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.211814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.211814Z digest=sha256:0300334584cab82357d57448b1819dc8100cf11ea150f9d3f2c1487fddfa0146

Observation 72d2d7cd-df33-489b-8151-edcb3e122410 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DataComp-LM: In search of the next generation of training sets for language models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.215712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.215712Z digest=sha256:120f67789f21059e33fc7d9bc1bf660200d7fdae80f0466414f58e5a3cd3db66

Observation cc166aed-8f57-46ec-9290-45d1c0aabcde · outbound

This paper cites StarCoder: may the source be with you!.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch StarCoder: may the source be with you!

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.219380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.219380Z digest=sha256:2cdf8b858a0a77181842a37ddf075b6564b9c890c70f0e21df95873baefcd7a2

Observation 76c1f094-4242-437c-848a-a1e86a42b946 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.223259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.223259Z digest=sha256:cd59dfe641f4a9d7627a32bdfe539e575367d6ac9845ef60c17dd2494dadefa6

Observation 0c0de2d0-6fa5-46c1-9f6c-e60b38fa38d7 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Textbooks Are All You Need II: phi-1.5 technical report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.226972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.226972Z digest=sha256:cbbd9f51a12520ac252f6d5e57c4527e12eb297625077622cd7b9fb772efa1f7

Observation ac1dd331-5d0f-4cb7-bf85-552a2c51d7f7 · outbound

This paper cites Against the achilles' heel: A survey on red teaming for generative models, 2024.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Against the achilles' heel: A survey on red teaming for generative models, 2024

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.230532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.230532Z digest=sha256:3be8fb31f92c38dd0950362c2098f42d81f65ba00cfc302d3de3fccd00cada67

Observation 09d746df-dcec-428f-9498-20a3734e6b20 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods, 2021.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Truthfulqa: Measuring how models mimic human falsehoods, 2021

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.233747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.233747Z digest=sha256:a0ed69a3843410009e606c62cb437838846ab35ab0b3e951247fa4a0bcf7dde3

Observation bfd16ff5-243f-48be-8b4e-a9f4127b063a · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.236804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.236804Z digest=sha256:1c10f0699a33a627ca3ca21cb5d1fe51fb92dc03f507cb4a6dcbd0665addea82

Observation e6119b63-645e-4fc5-aba1-bc75ed39954d · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.240581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.240581Z digest=sha256:a02fd42399dbfedb93f55f3f635c250856848a8da7b97dad692c6edba35625b1

Observation 0cd7980e-610b-49d6-befe-0a085fb9fe08 · outbound

This paper cites Goal-oriented prompt attack and safety evaluation for llms.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Goal-oriented prompt attack and safety evaluation for llms

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.244556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.244556Z digest=sha256:e87d67a114cb34f5686cf7ab8f121be6ca37e451779f7be3dde029c0287c4dd8

Observation e21b2f94-9be7-4401-b0f4-580752f4d2f2 · outbound

This paper cites Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.247757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.247757Z digest=sha256:c8db05def61077474fb29cfbd979ec99f2a4d5c1767003b745e83a5dea904b13

Observation 2ff9f191-945d-4bcf-b78b-9dc5f8da0586 · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Prompt Injection attack against LLM-integrated Applications

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.251901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.251901Z digest=sha256:99bf336e419c5aff57ef297d9903fcc42d45bf81ba8359bf380d13a66a46d835

Observation 542add97-136b-468a-a54b-75501854fe1a · outbound

This paper cites LLM360: Towards Fully Transparent Open-Source LLMs.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch LLM360: Towards Fully Transparent Open-Source LLMs

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.256160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.256160Z digest=sha256:b4549d4c723701259d428d6b81d69a9dcbca908fbc6a00b295b0972a13ccefff

Observation 72ba1396-22b3-4e23-b784-47c5fab262fb · outbound

This paper cites an unresolved cited work.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.259959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.259959Z digest=sha256:50357a039582a4df0eae46a181d775c2ae309f51d2dc87d40b99b1b31403107b

Observation 5fe49e8b-9b45-41dc-8283-4ae44f9daa5c · outbound

This paper cites an unresolved cited work.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.263440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.263440Z digest=sha256:acd047bf9a8592b66bd0ff1399421d64b6a2a0b4543375e5724bc58427a94484

Observation 3e7ff170-1ce7-4598-9f69-437e880b870b · outbound

This paper cites The Flan Collection: Designing Data and Methods for Effective Instruction Tuning.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.266755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.266755Z digest=sha256:d926b2f63a9eb2c19cd56d3b2c5276f23aa5c39991f371785ea8ba1acafd58e5

Observation b51fcd39-1874-432e-b5b8-36942066a195 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch StarCoder 2 and The Stack v2: The Next Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.270581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.270581Z digest=sha256:00f5d41c35795b251594614dcd92d75da30dfd0b5ebb0cf5c0c55003f0ccd3bf

Observation 3ea9b17c-4447-4e72-86c2-1e3b52be9c68 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.274300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.274300Z digest=sha256:2b6a541e09b3d3f311a40bcce79c7d938f0be1f349d113fb3c46f30ab14d9c4c

Observation 5a7f8115-75e3-4b52-898b-ec1748c49fa7 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, 2024.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Introducing meta llama 3: The most capable openly available llm to date, 2024

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.277701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.277701Z digest=sha256:7fd8f917659ea740b37a4238204616159d74c82f7af8c567086aa7fdbb297f06

Observation f846330e-948a-4a92-b447-9ce66143fda4 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.281307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.281307Z digest=sha256:932914b786aaef5c64a4e5d2124532fa6bf92e0013ad3ded36962d4d10c17455

Observation 353657b0-39cc-4b40-83df-16a9c7bab53d · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Crosslingual Generalization through Multitask Finetuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.285019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.285019Z digest=sha256:7dd0f9cec68adb40104c435066a575932a1ee701bd0b8f029b8f65160cb6c237

Observation 12aa9019-3dc4-4a0f-87c8-8a55e4c1961f · outbound

This paper cites OctoPack: Instruction Tuning Code Large Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch OctoPack: Instruction Tuning Code Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.288356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.288356Z digest=sha256:4b94bee12dcd34e0b8fa7b1ba43aff5a7e00f015ed6131cfca59e9f865dc1915

Observation bcd71929-54f8-4010-a38a-41c0d3be8ac2 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch OLMoE: Open Mixture-of-Experts Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.291740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.291740Z digest=sha256:9dd247c11d9b104bcd4dcdd07a7545239a622b930930684ade001f0f8dfa0173

Observation c43956bb-d0ef-41bb-9819-4b5c5f1fae0a · outbound

This paper cites an unresolved cited work.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.295140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.295140Z digest=sha256:7fab69bc368b8d5b54694832e9c1f369928706f40845ca2085835103705c6626

Observation feb5e4f5-1390-458c-ada9-a52aea10514a · outbound

This paper cites Pipedream: Generalized pipeline parallelism for dnn training.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Pipedream: Generalized pipeline parallelism for dnn training

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.298396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.298396Z digest=sha256:6ac43ce985649bc528b145f1a1e3e11f7f56d4aec5055bded478c752e4390f17

Observation 6412da5c-14a2-4153-ad18-0a6f1342978c · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.302336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.302336Z digest=sha256:aea885daac8994edbc51713d23cb96591d7389d6e320720bcbb3a6f5d06e2f82

Observation de3f0b55-32de-4602-9e3a-e9848c6574c9 · outbound

This paper cites Gpt-4 technical report, 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Gpt-4 technical report, 2023

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.306122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.306122Z digest=sha256:4eabcad3bdab20c6b6748c33ac70424d6fddaf023283cbdcc8b35e9d07bf88cc

Observation f67a740a-c36b-407b-a5f7-5bee299582d3 · outbound

This paper cites Grok-1: Explainable ai system.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Grok-1: Explainable ai system

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.310216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.310216Z digest=sha256:3b3f774f8454c2fbf0f4af67015a918ba5dec14b82b59e5b9f463aa2e6718047

Observation f7a866fc-febe-41f6-8dce-43297f6a59de · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.313497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.313497Z digest=sha256:76980fbc9e27bcecc769ccf28c5adbf35aafe2fae24eed8a8a308a8ba6ae1cdb

Observation 023d7c34-7e5c-4458-95b9-03a59bda5e8a · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.316876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.316876Z digest=sha256:5f1addabbe8f95f14cf966c6d5c7e2ce4ec9e07528b446230d4d6724703584a5

Observation a43284d0-6a4f-4d19-8bfb-f6b2e6f56e71 · outbound

This paper cites Nemotron-4 15B Technical Report.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Nemotron-4 15B Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.320492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.320492Z digest=sha256:38bc14f65efd994d1be0e0ed114a09ae32af03f7ae5cc965b907e7a3b442c6f1

Observation 2058f417-e1a3-47d0-9406-3be02f93d793 · outbound

This paper cites BBQ: A Hand-Built Bias Benchmark for Question Answering.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch BBQ: A Hand-Built Bias Benchmark for Question Answering

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.323853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.323853Z digest=sha256:5f4573008294fa8ce69463c3742f47e9ca6319380477820fafab76146165bca8

Observation 3d2c8ee2-1ab4-44e5-b939-561b88c95719 · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text, 2023.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Openwebmath: An open dataset of high-quality mathematical web text, 2023

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.327870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.327870Z digest=sha256:a01c00900ab3173f3902c9bae7826a6bd4caa7299da392395cb1fc8dd7080d28

Observation c7372d9d-0fc7-49d3-9f4e-1f76f1322290 · outbound

This paper cites Carbon Emissions and Large Neural Network Training.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch Carbon Emissions and Large Neural Network Training

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.331862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.331862Z digest=sha256:d986dbcccf14897d31228ea6d1e1954d06f1d6daf9522a9daa5ce7033e20e8df

Pith citing papers

Observation bca93d8b-50e0-40af-ae8a-631cc3c33fcc · inbound

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition cites this paper.

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:00.321479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:18:00.321479Z digest=sha256:4ae1718ba29da23861bc6220561e5d05c28f778e819a64b5dbf37bb355ce6b8c

Observation 44acbbac-0a63-49ac-a2ef-18634bc53a3d · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.436455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.436455Z digest=sha256:a42a995b9b8c9d08b49569f24dca83ff3c7813d6f4bcf130330adf0a457ead1d

Observation 3fceb114-49db-44b3-8a73-7bad8b49d944 · inbound

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models cites this paper.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:11:57.125399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:b459a026d6d3be6a2c056364b73bb89e1c221c4bd3a6c763ff2eed715b5ee12d

Observation 37d2254b-2d91-4067-b330-26975a3278ec · inbound

CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion cites this paper.

CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T09:38:44.568685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:38:44.568685Z digest=sha256:9d9960edc2383f950e4cfdd02b7281561cecf7e43fcaf89f1c2042f544fc45bf