Pith. sign in

Paper Citation Record · LEDGER

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2505.20225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20225 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:51.437970Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f7cb68c-c85a-429a-b05b-0f20b46cd02f · outbound

This paper cites Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:44.098280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:44.098280Z digest=sha256:5a3e8829fb746034ce01eb3f3c0a34fac81a9d56407adff1eea76e3b2415c570

Observation c9399a12-234f-4d58-953b-785a7351a57b · outbound

This paper cites Emergent and predictable memorization in large language models.Advances in Neural Information Processing Systems, 36:28072–28090, 2023.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Emergent and predictable memorization in large language models.Advances in Neural Information Processing Systems, 36:28072–28090, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:44.185028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:44.185028Z digest=sha256:0dcae30e213828f304dc3eac4cf4a183b09ff0252e3f760c1a7c211d4ca6ff3c

Observation b1066643-3fa7-4900-bf49-dd1df442b687 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.347994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:44.270403Z digest=sha256:c051ba2a40663ec40e6557d12157966b09a697a3020544b32d383e5c142ba2bb

Observation 19b5c372-8288-4763-b652-6a08202985e1 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models PIQA: Reasoning about physical commonsense in natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.131655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:44.386672Z digest=sha256:96373ff055a16ff20bd4cd79c9d3a4eb0bfbc180887fc386bdd10ca75b6218bd

Observation 4cc5e103-b3f9-402d-a872-6b5f11822cc6 · outbound

This paper cites On the representation collapse of sparse mixture of experts.Proc.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models On the representation collapse of sparse mixture of experts.Proc

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.991684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:44.465623Z digest=sha256:33978b745a69ebcb978d11ea35456f9f86f1795d1035aa6be4701bb1d8a28bc9

Observation 17867122-d6c0-4dd7-a013-08ec8d2efbcf · outbound

This paper cites Think you have solved question answering? Try ARC, the ai2 reasoning challenge.ArXiv preprint, 2018.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Think you have solved question answering? Try ARC, the ai2 reasoning challenge.ArXiv preprint, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.694982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:44.535555Z digest=sha256:366b759e6d3a0fa4407d865bb40ae588bb9ff6aacb661a48e9f6adac906e85cf

Observation 979caedc-952a-44ab-8c32-168840215b12 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.JMLR, 2022.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.JMLR, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:44.870448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:44.870448Z digest=sha256:a4546957c85660c977db338a13f63fe3e97aa5b6633bcd595ec233275a319651

Observation ce109ed3-19b0-4f5f-8f44-e4aa15571a9f · outbound

This paper cites A framework for few-shot language model evaluation, 2023.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models A framework for few-shot language model evaluation, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:46.280209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:46.280209Z digest=sha256:0eb723e88b2b555e9a65a96fb7b258ea3f0c3d92fc47b799ed21b4ebcc61cc49

Observation 0260bff8-ab34-40aa-a0e4-839c568c536f · outbound

This paper cites Fastmoe: A fast mixture-of-expert training system.ArXiv preprint, 2021.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Fastmoe: A fast mixture-of-expert training system.ArXiv preprint, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.459809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:46.500738Z digest=sha256:06283e31fae6bcc25b169f11e919c2bc130959004939b190fd60ce8ef992accd

Observation 837e542e-4add-4452-860f-13db535962b7 · outbound

This paper cites Measuring massive multitask language understanding.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Measuring massive multitask language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.294055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:47.956423Z digest=sha256:eb126c3924265a92b00c74213b35a6c5fd5ae0cba14be990a6ad61327e0053df

Observation 2a02d5d3-a089-44f8-a061-99e75156a5d4 · outbound

This paper cites Rae, and Laurent Sifre.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Rae, and Laurent Sifre

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.108386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:48.421407Z digest=sha256:83696648177c037e13699d152bac83463d56d904f0c43239731a45d8cee8336f

Observation 1ec5f9de-fb94-47fb-8161-ede4a9723c2c · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies.ArXiv preprint, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.955220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:48.559340Z digest=sha256:50d1e761455d693faa50ccfa500ca88c5cfe50c9069ee1cf43d56113ec8fd355

Observation 4d514ff9-0f8f-4d37-909c-3756780d7cc0 · outbound

This paper cites Demystifying Verbatim Memorization in Large Language Models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Demystifying Verbatim Memorization in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:48.631328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:48.631328Z digest=sha256:a031fd9f9c3cafec94e7f91d1322aff8df2d2c5002c36005f213d6fdcbc5e126

Observation 9cf0d8ca-ff43-48f7-8858-1af889e32b07 · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale.Proc.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Tutel: Adaptive mixture-of-experts at scale.Proc

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.778999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:48.696471Z digest=sha256:568f62078292d23cc2965212b6edd6e2ec04cf0fd4829e427377f4258fed8488

Observation 7836f992-70e3-44e5-96aa-b364a4855eee · outbound

This paper cites Mixtral of experts.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Mixtral of experts.ArXiv preprint, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.645820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:48.774085Z digest=sha256:f639869038ed7ee4234bd2aecdb28c8fbbf6ab75feba1dbe4d4b16e5ec09eff7

Observation 600cf498-58d5-4668-a673-ceed6b75624c · outbound

This paper cites Kingma and Jimmy Ba.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Kingma and Jimmy Ba

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:48.903883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:48.903883Z digest=sha256:e7c921dbf965adb2fab7350f6fcaccb5d56595f66b5ce226154d16143c9379a9

Observation 91fb215e-94c9-4a4a-ba68-68fac46da7b2 · outbound

This paper cites Gshard: Scaling giant models with condi- tional computation and automatic sharding.ArXiv preprint, 2020.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Gshard: Scaling giant models with condi- tional computation and automatic sharding.ArXiv preprint, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.494798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:48.998050Z digest=sha256:f348ea9b73141c19c9647e2209254dfe8a8530898d90ceefb3f4cb2e6f9e9155

Observation 57be5bee-4181-4a08-a047-9b862dc2c2d6 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Base layers: Simplifying training of large, sparse models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.322288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.095709Z digest=sha256:ea2f92c8bd26934425ad43d5d8d0d280f4b242eefd62f180b0cdee1b689a9950

Observation 32991eed-24c1-4f1d-8a80-38227fbea49e · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.097841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.225017Z digest=sha256:467e061234ef8c3474e5f36b68c3f9155e76c2e0d1452fffc9473480832bb421

Observation 205aaada-f018-4f97-8661-88b85581ae3c · outbound

This paper cites DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model.ArXiv preprint, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.915523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.299350Z digest=sha256:378415d499c5b66e409f66084f02017094a81a0b8af07a19b658fe498b47fd87

Observation 2488fdc3-327d-4a77-a490-b7155880a952 · outbound

This paper cites Deepseek-v3 technical report.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Deepseek-v3 technical report.ArXiv preprint, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.699463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.376132Z digest=sha256:28b159997929916e8336cec40898f65dd189fa8ac08a453073165380e2962b1a

Observation c893d78f-3ea9-4cd1-bfd2-30062ab246b4 · outbound

This paper cites Tending Towards Stability: Convergence Challenges in Small Language Models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Tending Towards Stability: Convergence Challenges in Small Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:02:51.956017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.501554Z digest=sha256:0757683853ff29b55145408d23ce4deaa40fe60cae053afb2ebb696c5f56590a

Observation 347390f4-ba5f-4991-9f2d-07015a526b31 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:49.622260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:49.622260Z digest=sha256:f3f3f2a4faa6bf7c5b90cfdd455480a73e06b9106499a21db9b494eb1754cc83

Observation 377c94ac-495e-4e96-84f2-1f3ca751b85a · outbound

This paper cites Can a suit of armor conduct electricity? A new dataset for open book question answering.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Can a suit of armor conduct electricity? A new dataset for open book question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.543552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.725393Z digest=sha256:1360f3adede45d774709c7854071304a74be77d3f0fac9775827a4750deb74f5

Observation 2e3005a9-6025-4c94-977a-647799d1facc · outbound

This paper cites OLMoE: Open mixture-of- experts language models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OLMoE: Open mixture-of- experts language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.352140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.831717Z digest=sha256:c43e58a466a5d4c519136166b44256e12a56f1045e7db87af7811157a43b6aae

Observation 42cb3bec-3ff6-43e2-9c9d-19449809682d · outbound

This paper cites Attributing mode collapse in the fine-tuning of large language models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Attributing mode collapse in the fine-tuning of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.222040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:49.925737Z digest=sha256:921154e4d42cc4334888286f69f27f1c24399cf8fc85a477794770a7dede2113

Observation 59ccef2d-e88b-4f1f-b0ff-247c8dcb8098 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.005119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.016518Z digest=sha256:262e41369cbf5e69d822656a79e5b34920ef019e9717e0154e634596de87bf89

Observation 9049d2e6-f188-4158-a940-595c4b82c435 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.845348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.132350Z digest=sha256:74b00dee9d868e41190949dad40a30bb243f9d5fadfd90e09f71f8ff3c4d9ee5

Observation ee80a897-dbc9-4f7e-8163-6394dfc5ff60 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.ArXiv preprint, 2017.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.ArXiv preprint, 2017

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.671642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.256613Z digest=sha256:165b188b803dc4b98810fc9a5440048ba9edda81a5e4dd7ab0fcb254bc5bb4f7

Observation 604aefab-5833-440f-b97d-1813b83ca0a6 · outbound

This paper cites JetMoE: Reaching llama2 performance with 0.1 m dollars.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models JetMoE: Reaching llama2 performance with 0.1 m dollars.ArXiv preprint, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.496149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.353681Z digest=sha256:f5daa03bf44529acf11e928e2c604e51f171d16dc00463a31c25417fe98f67a4

Observation 1fc7123e-646d-4393-8b86-6fab5e0cb37f · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism.ArXiv preprint, 2019.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Megatron-lm: Training multi-billion parameter language models using model parallelism.ArXiv preprint, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.348965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.455008Z digest=sha256:ab59c1d0f7ecf474f9a723d1975ef48af2d4f6459e2760cf3c4f5f3d0020fe75

Observation 3b62554e-adc6-4893-8488-a990ba9a1b9b · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.561305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.561305Z digest=sha256:f7ee376a840d375aaa9419ba46725ad77d9166216e43fc44fc3c979c4b77c785

Observation 59facc6e-3fac-4b6c-90c8-1093272969e9 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Free dolly: Introducing the world’s first truly open instruction-tuned llm

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.158449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.679920Z digest=sha256:8faa5d7371c0e6ca02f767c399c5001d6eeea05e9788c15c9f048c82541240f1

Observation b53b6c84-44e8-4434-9e21-a2ac9b6741f3 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.ArXiv preprint, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.961360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:50.778370Z digest=sha256:a6424e7db26d1b8011f2b44a68fb28411b3233e1aac3c6820a25ecee6d9d1931

Observation 675da49c-9505-4b43-86d1-2942a9db9fec · outbound

This paper cites LLM Circuit Analyses Are Consistent Across Training and Scale.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models LLM Circuit Analyses Are Consistent Across Training and Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.878275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.878275Z digest=sha256:c753a465d4df269abdaa39fb23f7bee6bec986ed3cdc69367f027d34aa195167

Observation 7797fdd4-71a3-47c6-ab3e-5ba497646968 · outbound

This paper cites OpenMoE: An early effort on open mixture-of-experts language models.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OpenMoE: An early effort on open mixture-of-experts language models.ArXiv preprint, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.797912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:51.005458Z digest=sha256:8c7ad5c15036ba9de8be746329e27507c0bea2692c738c21cedb956e60051c29

Observation 236e214e-2813-4c4f-a6d9-7e262475f5f7 · outbound

This paper cites M6-T: Exploring sparse expert models and beyond.arXiv preprint, 2021.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models M6-T: Exploring sparse expert models and beyond.arXiv preprint, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.594448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:51.127391Z digest=sha256:14f61f6dc2706677646ed1c1b2568dbe5612dd39ab5819295c97b7fc9c1d46a4

Observation 8fb563fd-cd57-47fa-8787-b1992250163a · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? InProc.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models HellaSwag: Can a machine really finish your sentence? InProc

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.401440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:51.236526Z digest=sha256:b3538bdd5dacc3dd59ab2ac9acde8d32cd31e1fad8335c36965b2dcaeb2e9c4f

Observation 7475c41a-9ede-430d-bf65-5afe0a0e104e · outbound

This paper cites OPT: Open pre-trained transformer language models.ArXiv preprint, 2022.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OPT: Open pre-trained transformer language models.ArXiv preprint, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.220119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:51.365700Z digest=sha256:b38065e99f024b6349b8a138d0eb290f2345c91c448a6b3b22af480cf2dbd27a

Observation 846f8683-413e-4ca0-842b-a049c33d399f · outbound

This paper cites ST-MoE: Designing stable and transferable sparse expert models.ArXiv preprint, 2022.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models ST-MoE: Designing stable and transferable sparse expert models.ArXiv preprint, 2022

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:02:51.728312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:02:51.437970Z digest=sha256:8fab2219738b72ae5f172e34199d07a1e40690199c91fa9c1c0a98798e382127

Pith citing papers

No inbound Pith citation observations are available.