Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

As of 13 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2501.10322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10322 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:16:20.792657Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T03:42:21.307552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:47:48.559145Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact3
  • verified fuzzy14
  • unresolved32
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3de18d85-dfa7-4811-aa6a-c90d5856cd2a · outbound

This paper cites write newline.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.525205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.525205Z digest=sha256:4b0de27738ad105b3ead9db405b76432b5059494a56d7ccda9740951ba0a3ef2

Observation cbeb6181-a3c8-489b-9edd-16cc8ea47f3f · outbound

This paper cites C har2 S ubword: Extending the subword embedding space using robust character compositionality.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C har2 S ubword: Extending the subword embedding space using robust character compositionality

Reference 2

Resolution
verified exact
doi, observed 2026-08-10T19:16:21.079048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.531884Z digest=sha256:597be2fb89555d2254c66ff3cee58631f9b6d364655a59db34547849ac71089e

Observation 933d638a-7d56-4add-b6e5-84d415df53d7 · outbound

This paper cites Character-level language modeling with deeper self-attention.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Character-level language modeling with deeper self-attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.537264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.537264Z digest=sha256:a78969d72b33e85d07e6f04aa04a362ae086ec7436598a803a2e21b1bac5d6fe

Observation bb07750b-a32c-4997-a03a-f78160021689 · outbound

This paper cites an unresolved cited work.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.543430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.543430Z digest=sha256:1288b85fcc515abd3c279e41b28171c84c5170beea2c87a433d5462b0b1cab7d

Observation 8ece48a5-8ffd-460b-bafa-2bb2cc5d30a1 · outbound

This paper cites Semantic parsing on F reebase from question-answer pairs.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Semantic parsing on F reebase from question-answer pairs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.604001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.548554Z digest=sha256:6fcb4185afb8433d598f43355a018f21ba22e8d6ab38ceea3f8c8c77e5a35310

Observation 1220c9c1-7765-4959-b42b-3d103c8aa7e6 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.585064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.554031Z digest=sha256:63e9bf1d0d930de3f16559ff90f98d4f3ae40e9b26038769b3f10e7bc4fc87c5

Observation 414fed55-78ec-4436-92e8-2150c3d30aa3 · outbound

This paper cites Occiglot fineweb v0.5, 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Occiglot fineweb v0.5, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.567965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.559649Z digest=sha256:b21278e973b7202aaea644b9142a76fa62d800141650314bdb1461077fd0ae00

Observation c0102fa2-af1d-4228-8664-4be39b1f8aa5 · outbound

This paper cites Bridging the Gap for Tokenizer-Free Language Models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Bridging the Gap for Tokenizer-Free Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.565344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.565344Z digest=sha256:5e165d23fc5c23c99fdc386a144d1efbe1e671daf58d1d10846a85ce267a4282

Observation 6fe2a8b4-c99a-43ee-a251-0ea9812ca2f9 · outbound

This paper cites B ool Q : Exploring the surprising difficulty of natural yes/no questions.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models B ool Q : Exploring the surprising difficulty of natural yes/no questions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.570745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.570745Z digest=sha256:a273e2bbe7daa0bab9978e98ef6ad2f3320b743288f2a541603160ed5613d8fc

Observation e54a72f7-d208-4f53-ba61-ca3ff6cbd064 · outbound

This paper cites Bowman, Holger Schwenk, and Veselin Stoyanov.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Bowman, Holger Schwenk, and Veselin Stoyanov

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.540576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.581145Z digest=sha256:efbc2a97f731e562f67efac83bf254adfb11c208dd23eb6321be042450280683

Observation d9904d11-5bd0-4444-b023-d8eae768ce5a · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.523486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.586757Z digest=sha256:32135e3a8f6f0f6690f3dff1c535c10dc2802ef8f47978f7fa00c53ce7a809d5

Observation 8a168cc1-d65a-4b9d-aa61-d38ce01d6f32 · outbound

This paper cites T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.592713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.592713Z digest=sha256:20e17a0d6b83ba06a3987f1777d407300c8237f36a91003b83285f66ab0d9ee6

Observation 782f88d9-861e-4520-b119-60c30e302157 · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models BERT: pre-training of deep bidirectional transformers for language understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.597892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.597892Z digest=sha256:363c7580ba07ac4d8702dec7a4a936e3956ad76048c498a78ec0305243a756a0

Observation 5da89344-4a69-4cd2-823b-bccc91664fb0 · outbound

This paper cites The Llama 3 Herd of Models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.602612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.602612Z digest=sha256:108be2863c8f57002f9ec57e76d7016ba834f5e542d929c7359107c505a3c7c8

Observation 1b555442-5b81-4784-aa2a-f7b709d46e53 · outbound

This paper cites C haracter BERT : Reconciling ELM o and BERT for word-level open-vocabulary representations from characters.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C haracter BERT : Reconciling ELM o and BERT for word-level open-vocabulary representations from characters

Reference 16

Resolution
verified exact
doi, observed 2026-08-10T19:16:21.033923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.607848Z digest=sha256:f2352ec0d8634ffceff9a370f6b3db02c720364422234203ce91cef876507d6d

Observation 73b683fb-60de-4cce-92ee-48bb9f33c1c8 · outbound

This paper cites A new algorithm for data compression.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models A new algorithm for data compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.612830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.612830Z digest=sha256:478e28f8f49e723654b47fbc59c29a7b6ead801b53e93718e225c07c03e33dce

Observation 656f9556-601f-446c-b152-fdca2744171e · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models A framework for few-shot language model evaluation, 07 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.617517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.617517Z digest=sha256:fa41df90827b6fceaec661834f10b6aff80bf660a238786ea1e397c5da9c849f

Observation b78784f5-7632-4cdf-ad9a-b779d1a3f5eb · outbound

This paper cites Better & faster large language models via multi-token prediction.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Better & faster large language models via multi-token prediction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.485778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.623085Z digest=sha256:13a452ccf1ff9f2e3b146862c968ccb1f05fdbe8160f4af5e460c4311fac2fde

Observation d9eca35e-c863-4917-81db-3df6e9cd72de · outbound

This paper cites Botvinick, Ian Simon, Hannah Sheahan, Neil Zeghidour, Jean - Baptiste Alayrac, Jo \ a o Carreira, and Jesse H.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Botvinick, Ian Simon, Hannah Sheahan, Neil Zeghidour, Jean - Baptiste Alayrac, Jo \ a o Carreira, and Jesse H

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.468811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.628241Z digest=sha256:b39147201e18e9bf6d23903b043550c8f618abf2e54261481cb518664780253b

Observation 6f10e283-7a25-418f-b2b8-b3659d8ce36a · outbound

This paper cites Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.633112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.633112Z digest=sha256:c83e958e0b0a6f0730822e9cc72c40d7e2ebe91ca4ea39e0b4877272ca0daa88

Observation 93bdbdc7-51e1-44f8-a080-af2eab71825a · outbound

This paper cites Measuring massive multitask language understanding.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.638405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.638405Z digest=sha256:64c6a723fd6a7363a5cc7ccb35262678da8763c1258f439a91c1a147dfad7433

Observation 88d0a68e-a3a3-4cf2-8383-2381420f2759 · outbound

This paper cites T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.643281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.643281Z digest=sha256:9ff71bcb19cb7fe40bb3503030728f7cd235670feb2fe346718e69eac7a238e3

Observation be429736-ed51-4512-9f19-57a2ea1d3b50 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.648279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.648279Z digest=sha256:e50145e3a83e00156df5b947c2bbab7924d770b3b38ea9a520c10b60ad6c6a6f

Observation 660921d5-e930-4439-9f8b-10154c24dde9 · outbound

This paper cites Hashimoto.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Hashimoto

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.441723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.653654Z digest=sha256:255d897e6c883fb533d4958f2f9d179fd4287c8c8b4148d51b892c82d673111f

Observation 8933bb48-3fbe-4f11-8f7c-4ef8681aafba · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Truthfulqa: Measuring how models mimic human falsehoods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.658426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.658426Z digest=sha256:a8a0e8c2d56e5e987e6ff06c9704ec4fb10b1702072abe89f2ce37d52344b4ea

Observation 7ecbde2d-a12a-4623-81e2-61804aaee209 · outbound

This paper cites Decoupled weight decay regularization.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Decoupled weight decay regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.663439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.663439Z digest=sha256:959eb4eea4625897430ac7c957b64b98851a8cf70f5372fda2028110ddcb6204

Observation 2fa6d751-0c1e-498a-89de-b04d24e64ae2 · outbound

This paper cites C har BERT : Character-aware pre-trained language model.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C har BERT : Character-aware pre-trained language model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.668608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.668608Z digest=sha256:94a11d14fceae2ee647870e4ddc585d9847846746775bc570f0cc6665c6ac869

Observation 96634558-8a12-4a1e-883d-2375d3d142b5 · outbound

This paper cites Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.673355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.673355Z digest=sha256:a03a738f5208c10eaf39473d4607215c00832a641e86dc0d1d4ae255a6550ed7

Observation fbac2ce3-d16b-4cb6-a7e4-e25169452ad4 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.678410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.678410Z digest=sha256:33c20359b4351458298bd25a34ce7e00e336c32714aa92b4c97b798af591efee

Observation a7114735-6bee-4cc9-a853-5bd59a842d07 · outbound

This paper cites Mmmlu, 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Mmmlu, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.413125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.684951Z digest=sha256:1fed39be258fb9de12184954a1267be7df5d3d49835037eb75157fa4541adce3

Observation fb4d1ed9-bada-4e93-8008-4a8e2c546bca · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 32

Resolution
malformed identifier
no resolver link, observed 2026-08-10T19:16:20.689907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.689907Z digest=sha256:15b11164ff9549bf48c297708d0333da9e6094fa0b81a65ff416bbe4e412063c

Observation 4920c0d1-870a-4343-9e07-a1ec82e2e7b6 · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text, 2023.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Openwebmath: An open dataset of high-quality mathematical web text, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.695245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.695245Z digest=sha256:cdbe83619e737077c5a01900ad21232d481ced7d2e8543cd730559db0a412a27

Observation 0f524630-137f-4cae-ad5c-b0dfdde3b122 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.700411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.700411Z digest=sha256:0187dc32db6c82a614dc3fe26f6098f92c85bb37186602b50c2e847cf8d5d386

Observation b122c748-f7d3-4765-b9f1-722e781c350d · outbound

This paper cites Language model tokenizers introduce unfairness between languages.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Language model tokenizers introduce unfairness between languages

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.384278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.705646Z digest=sha256:67e9a8fbb612a7d45738cabe628f16a2f73cc7f5cfac12f915f2af0a7da0889f

Observation 0087ec6c-2aa5-4ee4-9ba8-bd65f7aed6bc · outbound

This paper cites W i C : the word-in-context dataset for evaluating context-sensitive meaning representations.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models W i C : the word-in-context dataset for evaluating context-sensitive meaning representations

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-10T19:16:20.711185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.711185Z digest=sha256:0a7da6bd1013e2db7c689ebfb3786f1fe42e1f3a9a9ba2509d926c6f82d3ab82

Observation 336de9ed-8e9f-4f7a-8c1e-2ee9afbacf21 · outbound

This paper cites Germanbenchmark, 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Germanbenchmark, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.367231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.716273Z digest=sha256:aed3c551e4c107dbb397354a9156564f286c870e29a9e030df42cd937d6bfd9f

Observation 38759c8e-ccbc-435a-aa92-a26b40a9f5cc · outbound

This paper cites Efficiently scaling transformer inference.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Efficiently scaling transformer inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.351227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.721304Z digest=sha256:e38d35c44a6c10de8243de6fc3e8ce7789b493f8f4456abf8f8ae30fd3ecfdeb

Observation 169d0db1-22b1-4897-958f-6eeaf8fe2c05 · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Winogrande: an adversarial winograd schema challenge at scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.726054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.726054Z digest=sha256:ec3b207203e5ec08a060d5864505eee189d412dd19a0253864e0b23eac79914b

Observation aa5f9226-e877-49ac-b01d-6c3740db7e4c · outbound

This paper cites Neural machine translation of rare words with subword units.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Neural machine translation of rare words with subword units

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.730719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.730719Z digest=sha256:8463d93d692483a90091c851d38899e52423e0e1c66e52bc268d707468bfba44

Observation 20926a3b-0f62-45e6-a6f7-e97a45f61b93 · outbound

This paper cites SpaceByte: Towards Deleting Tokenization from Large Language Modeling.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models SpaceByte: Towards Deleting Tokenization from Large Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.735877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.735877Z digest=sha256:856bf9db087c3eb4e26eca124bd65b063aa5107604ab835769e3e80a84f527f8

Observation b4a657ca-7ccf-4c2b-b7eb-eefefbfce51c · outbound

This paper cites From characters to words: Hierarchical pre-trained language model for open-vocabulary language understanding.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models From characters to words: Hierarchical pre-trained language model for open-vocabulary language understanding

Reference 42

Resolution
verified exact
doi, observed 2026-08-10T19:16:20.888678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.741445Z digest=sha256:1e7d4e3ff284e0b7fb5fe544dd2edf76674bf631e6d3f4b11d9cda3c45d2be96

Observation 3afa7e17-d33e-4208-9623-f0e9d23d5e58 · outbound

This paper cites Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.334501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.746281Z digest=sha256:56cb30c968e6cf45b019095615b5ceab1aef5f3e8573ab5d7b50975bf1e9e0f3

Observation e6ffb928-35e1-4086-96e2-56637c7195a6 · outbound

This paper cites Learn your tokens: Word-pooled tokenization for language modeling.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Learn your tokens: Word-pooled tokenization for language modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.751455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.751455Z digest=sha256:9795a9aaf188d5f76fc72a26d6847521fa0a340a1ed98e40ae2449dce917acb1

Observation 878598da-bec5-4c1e-8e7e-89b674f45275 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.756271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.756271Z digest=sha256:eb95585e34e56ed83d0f4e15c3fd6d3e8c544449a4607ab9b92df7f95c8bc2bc

Observation f7d60f3f-f0d0-4a13-9bd3-73083189e0f4 · outbound

This paper cites Skywork: A more open bilingual foundation model, 2023.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Skywork: A more open bilingual foundation model, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.761511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.761511Z digest=sha256:b7a2017d19d1ee56ddac23abd681fc3d9195e75c39a67e2f217bbab64e0c983d

Observation 937b5811-b6f4-445d-ba3f-d3069313addb · outbound

This paper cites B y T 5: Towards a token-free future with pre-trained byte-to-byte models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models B y T 5: Towards a token-free future with pre-trained byte-to-byte models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.766254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.766254Z digest=sha256:90d36601ad930eb6757f786c51c0ca8af63dc7b0bc02f240568a0c6c50b99ec1

Observation f5ac978a-6974-4c1c-aac0-18890fd35edd · outbound

This paper cites MEGABYTE: predicting million-byte sequences with multiscale transformers.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models MEGABYTE: predicting million-byte sequences with multiscale transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.306409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.771088Z digest=sha256:667c18841a876073938ce947a08a19bf8a2a1f313c3baf1377ad8c601248cc88

Observation a9035346-a518-4497-9d04-4106f6631988 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.775743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.775743Z digest=sha256:400bd0bbd8e4f561c0433933e4a9462078468508296bc38ec2dbd493109856df

Observation d8770478-1b87-44cb-9c73-05f7bd567182 · outbound

This paper cites @esa (Ref.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models @esa (Ref

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.781482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.781482Z digest=sha256:00a9efb0f5a04b735b4e7ca7477b734a0e42a9001a0392c8ba5e5febc99a7faa

Observation f090c69b-7d44-49b4-9c00-dfd4e5a5f5d7 · outbound

This paper cites an unresolved cited work.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.786674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.786674Z digest=sha256:b95931c362c1bf84d1d0d0478f6878516556f59acc21daea9f4e0a7f578e8c8e

Observation 1ce16206-2216-4b9e-9bdd-b8a2ac8537d4 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.792657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.792657Z digest=sha256:5ce09ce9f352036da2a413185fa70813b1a7c1d65ce80143341b04abea660580

Pith citing papers

Observation e2b73c34-d868-473f-b1c9-35637e98a049 · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.879364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:2d5eef128bd7015901ecd78d0fab6b4ccf88420ee8c33af842876dae00f8a151

Observation c8705c67-22f6-4b94-8b48-f2e04376b670 · inbound

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES cites this paper.

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:47:48.584286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-11T03:42:21.307552Z digest=sha256:b31a482b7394a25428ea196f0eb2c19d68e97407c2456b689529544a851be125