Pith. sign in

Paper Citation Record · LEDGER

SpeLLM: Character-Level Multi-Head Decoding

As of 20 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2507.16323.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16323 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:19:17.012836Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact7
  • verified fuzzy5
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ebcf8dc-70f1-4854-b083-3987f5d9a141 · outbound

This paper cites A new algorithm for data compression.

SpeLLM: Character-Level Multi-Head Decoding A new algorithm for data compression

Reference 1

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:19:17.302442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.918735Z digest=sha256:ff0db30a93b8a6044f2c63cc7528be8a293bb69f72ddc3e95e734c99bb22bf12

Observation 186c0948-d305-433c-97e0-694897ad8467 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

SpeLLM: Character-Level Multi-Head Decoding Neural Machine Translation of Rare Words with Subword Units

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.922411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.922411Z digest=sha256:ba501e19deaecdda05176da21f01c89a5bb38066b43a4ca792acdc64ad44ea70

Observation 270fd262-2d24-4068-a12c-90dcd69305b1 · outbound

This paper cites The Llama 3 Herd of Models.

SpeLLM: Character-Level Multi-Head Decoding The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.925226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.925226Z digest=sha256:b94838a8d5b00f77e2ef00d43839b9f79fd198c121daca159e243b1e8a894685

Observation 7261426e-2a4c-4606-9029-5c0cc507b9ef · outbound

This paper cites Gemma 2: Improv- ing Open Language Models at a Practical Size.

SpeLLM: Character-Level Multi-Head Decoding Gemma 2: Improv- ing Open Language Models at a Practical Size

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.928045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.928045Z digest=sha256:092b5d357db4d282f3e8bf2865f4dbebff2679f8708ead3a324604462371f924

Observation 5e0db424-262a-4df7-a490-4d1d5b4a2c45 · outbound

This paper cites Do all languages cost the same? tokenization in the era of commercial language models.

SpeLLM: Character-Level Multi-Head Decoding Do all languages cost the same? tokenization in the era of commercial language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.931347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.931347Z digest=sha256:8c7971e9b543f8ef93e266d76d135f47cbbc23b6d6dcc6e25bddfd6cb564c202

Observation c8c8da7a-cc3c-4150-8d43-ae6a796e6743 · outbound

This paper cites Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models.

SpeLLM: Character-Level Multi-Head Decoding Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.934493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.934493Z digest=sha256:d26b29b2906f71e13b797d8946b4f69f45b00aa16bace2dbc46670ce1d27bdb1

Observation 5c3ac7fe-313c-4a4d-82a4-ae89033ee9eb · outbound

This paper cites Language model tokenizers introduce unfairness between languages.

SpeLLM: Character-Level Multi-Head Decoding Language model tokenizers introduce unfairness between languages

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:17.348538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.937949Z digest=sha256:d94538c078ea746e657722e2c62e3cbeb7c4a49e5f21312e9ac2267c6846e85f

Observation f7c98384-9b4d-475c-b113-1333c85bd11e · outbound

This paper cites MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers.

SpeLLM: Character-Level Multi-Head Decoding MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.940215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.940215Z digest=sha256:e904c7d393fff8a60b0954b5f7c7527caa92fe59b4a2313f171c664f1fd9521a

Observation 245cef04-c4a6-4041-b171-e7fca218e328 · outbound

This paper cites Byte Latent Transformer: Patches Scale Better Than Tokens.

SpeLLM: Character-Level Multi-Head Decoding Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.943166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.943166Z digest=sha256:67cb12da9686efab7a1865dbd0cd1fca6a5b7a0f31860efc2543f5ee796f47ee

Observation 0cb5616c-25cc-486e-aac7-191873976f6d · outbound

This paper cites Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens.

SpeLLM: Character-Level Multi-Head Decoding Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.186769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.945692Z digest=sha256:e50ffcea61a87f728b8e51ea2b085402971a44946a2ab003293e5ecb8b38a493

Observation 9192d1c9-37cf-4b15-b38a-64929256038e · outbound

This paper cites What do tokens know about their characters and how do they know it?.

SpeLLM: Character-Level Multi-Head Decoding What do tokens know about their characters and how do they know it?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.234460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.948956Z digest=sha256:a7af176f28b57c34f9fcc5e070341dbe01fa800d0b84959db6a346f2c70e139f

Observation 97039f1e-3c11-485f-ac59-9ddf20a39391 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SpeLLM: Character-Level Multi-Head Decoding Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.951711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.951711Z digest=sha256:580af1c1dd6c230bf083ff538879935395296f8b59b08e0ddac2b85f08041010

Observation 3c03c5fc-4e5b-4ccf-a9a2-469a1174ba39 · outbound

This paper cites Teaching Machines to Read and Comprehend.

SpeLLM: Character-Level Multi-Head Decoding Teaching Machines to Read and Comprehend

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.954054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.954054Z digest=sha256:95da62d6e8ada36b7b4a6be7ee0565c9f82e445a8d3c7dffa7fa3d6f5b1278b7

Observation 1326b273-1c82-4c7d-a75f-6deb4246d885 · outbound

This paper cites Headless Language Models: Learning without Predicting with Contrastive Weight Tying.

SpeLLM: Character-Level Multi-Head Decoding Headless Language Models: Learning without Predicting with Contrastive Weight Tying

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.162960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.956368Z digest=sha256:2703251b1cb3609ea29628711e534ef780bc614165efee0d89d62963a09ac0c2

Observation 7144c661-b181-4a4d-86bb-7b1d37f95d2a · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

SpeLLM: Character-Level Multi-Head Decoding The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.958966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.958966Z digest=sha256:889b6f1d717411722dc21a4efe7bb4d558589bbc0c2fab37781be570a8cd5b84

Observation 6da07750-1c17-4ac1-a0a1-5da691814834 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

SpeLLM: Character-Level Multi-Head Decoding BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.961408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.961408Z digest=sha256:df3b026dfd4889b11bd4dcce5e58fafaeaa69ac0f0be403dae211540c3dc0bcb

Observation 8a199f57-e1fd-403e-bf94-d32da6b3a2ba · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SpeLLM: Character-Level Multi-Head Decoding Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.964358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.964358Z digest=sha256:573072fff97e49c94407b80a47982329534257976ee9c0a872cf442758bf7f58

Observation a2572573-efdc-481f-aa99-3d907ebf5d58 · outbound

This paper cites Efficient vocabulary reduction for small language models.

SpeLLM: Character-Level Multi-Head Decoding Efficient vocabulary reduction for small language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:17.341045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.966931Z digest=sha256:6aeee2c4bb2239262295b4f6cbb6ac4429f959c45075d7123f3665f7604e03f8

Observation 167698e1-f1d5-43b3-919a-c46c3cced097 · outbound

This paper cites T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings.

SpeLLM: Character-Level Multi-Head Decoding T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.131959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.969137Z digest=sha256:d7d01654276f47c011c7eb0e0f486d638a811aa2cbc950d0a1a499d08c434a0d

Observation 04412551-b22f-4a0b-8434-2f666041544b · outbound

This paper cites FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling.

SpeLLM: Character-Level Multi-Head Decoding FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.971760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.971760Z digest=sha256:ae9bc0fd97b07dbbd165a9037d6c24d97e5a7e0dcede28bd17b3a3b36629cd4f

Observation 7fab0aa4-c7af-4c6a-abf5-cffd343c0568 · outbound

This paper cites LLM Vocabulary Compression for Low-Compute Environments.

SpeLLM: Character-Level Multi-Head Decoding LLM Vocabulary Compression for Low-Compute Environments

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.114501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.974995Z digest=sha256:42e097f085d2298bb506d467bf8a002ee7b164f61419dad743b388ef8282517d

Observation 3c850fcb-1bc5-4df0-9620-536487f792ae · outbound

This paper cites Fast V ocabulary Transfer for Language Model Compression.

SpeLLM: Character-Level Multi-Head Decoding Fast V ocabulary Transfer for Language Model Compression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.977758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.977758Z digest=sha256:e5e7aab741d75b74da4295f697fc8ed57a625ec422fb8bf7c644d86064e2f1b6

Observation 33cc0c51-4968-4207-9c29-89d452e39af1 · outbound

This paper cites Vocabulary-level Memory Efficiency for Language Model Fine-tuning.

SpeLLM: Character-Level Multi-Head Decoding Vocabulary-level Memory Efficiency for Language Model Fine-tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.099342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.980301Z digest=sha256:5d9691394fe053d4b46ab694e87d78c45b2c65b63e4907bc7113520a10ad8a29

Observation 29f02703-7b4f-4211-bcfe-ce376e4cd92f · outbound

This paper cites Efficient softmax approximation for GPUs.

SpeLLM: Character-Level Multi-Head Decoding Efficient softmax approximation for GPUs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.982992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.982992Z digest=sha256:6c7d9a2c608179df59a0d74ed9bafc680a1060d6dcf43f8d415cb7fff0f6e477

Observation 3c189951-7a1c-414b-81d8-caf044f6a757 · outbound

This paper cites Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers.

SpeLLM: Character-Level Multi-Head Decoding Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T15:19:17.081790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.986067Z digest=sha256:5fa7a0c3b10e9522afa7759e9eedde38d721b0b45948b5d5f4df66307c0fad3c

Observation 27ef0d46-9eba-4202-b173-f7a08e14c87e · outbound

This paper cites Baharav, Ryan Kang, Colin Sullivan, Mo Tiwari, Eric Luxenberg, David Tse, and Mert Pilanci.

SpeLLM: Character-Level Multi-Head Decoding Baharav, Ryan Kang, Colin Sullivan, Mo Tiwari, Eric Luxenberg, David Tse, and Mert Pilanci

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:17.333809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.989283Z digest=sha256:40c34ad987dacd34d021703926cd5ca8e4f8325dd9a69f5e5150100cdd956fd4

Observation 95158648-9b5a-422d-ae41-9f01732ae97f · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

SpeLLM: Character-Level Multi-Head Decoding Better & Faster Large Language Models via Multi-token Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.994757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.994757Z digest=sha256:59f510024155329811863e20bd37f163dd124f069adc5ed29807ced1dc6141e9

Observation df19124a-1058-4bb8-8244-87eca278ffcb · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

SpeLLM: Character-Level Multi-Head Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.997333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.997333Z digest=sha256:a4f3fe31df1804385c242549d0adc602c3aebc1628113b03823f54f9d9a3e51b

Observation ce70b2e4-cfa9-400a-86c1-3b9912224178 · outbound

This paper cites Improving Multilingual Models with Language-Clustered Vocabularies.

SpeLLM: Character-Level Multi-Head Decoding Improving Multilingual Models with Language-Clustered Vocabularies

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:17.056539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:17.000545Z digest=sha256:31a8f9763757c0bac8cf68b635baf8c5b18517981506127e3c2ba7c444c214fb

Observation 9535f0fe-a3e9-452d-9c20-9fa2f60d2380 · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte-to-byte models.

SpeLLM: Character-Level Multi-Head Decoding ByT5: Towards a token-free future with pre-trained byte-to-byte models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:17.002827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:17.002827Z digest=sha256:7986697b8daa8f1023cfad59e2270e32c429cfd3699447ffa6469b009e452477

Observation c058ad57-4d09-4a1e-b962-1e36c8def3ec · outbound

This paper cites Clark, Dan Garrette, Iulia Turc, and John Wieting.

SpeLLM: Character-Level Multi-Head Decoding Clark, Dan Garrette, Iulia Turc, and John Wieting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:17.005211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:17.005211Z digest=sha256:b29c3351e02d0f11e600d7a5663b60cf525a8cf48d97025d05bf0566c07816f2

Observation d340e217-b17a-47d7-9b39-0580177e4155 · outbound

This paper cites Decoupled Weight Decay Regularization.

SpeLLM: Character-Level Multi-Head Decoding Decoupled Weight Decay Regularization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:17.318960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:17.007608Z digest=sha256:b4fb8f10144949f666c4142fbb499648d633c955cc7317dfcf0ba0693655cd32

Observation e7cccba9-6784-417d-af24-624aab0e5b2a · outbound

This paper cites All models are trained on a single GPU: smaller models on Nvidia L40S, and larger models on Nvidia A100.

SpeLLM: Character-Level Multi-Head Decoding All models are trained on a single GPU: smaller models on Nvidia L40S, and larger models on Nvidia A100

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:19:17.309902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:17.012836Z digest=sha256:d7e6452081d845e5ad0aa74183b073dc4c85ced54349f50ec181a1527b60a299

Observation d179371d-a7c4-4021-bad6-9b35891e93ca · outbound

This paper cites Decoupled Weight Decay Regularization.

SpeLLM: Character-Level Multi-Head Decoding Decoupled Weight Decay Regularization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:17.010080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:17.010080Z digest=sha256:d5da056ba05fb5dea7fba4d4390bb87b54f7e46ff3fdb35aea09191d856236c6

Observation ea75a3c5-db19-4383-bcfe-3ed4949d5877 · outbound

This paper cites an unresolved cited work.

SpeLLM: Character-Level Multi-Head Decoding Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:19:17.326714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:19:16.992044Z digest=sha256:9027d7830affe54353ca0c89e50032f23cb68a88ce150630c6650de2e1a4705e

Pith citing papers

No inbound Pith citation observations are available.