Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2501.10322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10322 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:16:20.792657Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T03:42:21.307552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:47:48.559145Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact3
  • verified fuzzy14
  • unresolved32
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3de18d85-dfa7-4811-aa6a-c90d5856cd2a · outbound

This paper cites write newline.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.525205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.525205Z digest=sha256:70096e2a8fa6a9b2087a60ce4b09cca4c883f400e00a3a79317027136b13ce46

Observation cbeb6181-a3c8-489b-9edd-16cc8ea47f3f · outbound

This paper cites C har2 S ubword: Extending the subword embedding space using robust character compositionality.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C har2 S ubword: Extending the subword embedding space using robust character compositionality

Reference 2

Resolution
verified exact
doi, observed 2026-08-10T19:16:21.079048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.531884Z digest=sha256:b814cf3cbb38ee5b7170692a1b0cf026143c3f2313608681293138a04b44aa65

Observation 933d638a-7d56-4add-b6e5-84d415df53d7 · outbound

This paper cites Character-level language modeling with deeper self-attention.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Character-level language modeling with deeper self-attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.537264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.537264Z digest=sha256:b7f7d552e4516b541772107950a617ca8e7ba7e5951d32416dd7a35de549492b

Observation bb07750b-a32c-4997-a03a-f78160021689 · outbound

This paper cites an unresolved cited work.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.543430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.543430Z digest=sha256:c52de522cf3757daeada4aaabd81ac94c9f9cb5bf95d29dfd2d42ef72245143b

Observation 8ece48a5-8ffd-460b-bafa-2bb2cc5d30a1 · outbound

This paper cites Semantic parsing on F reebase from question-answer pairs.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Semantic parsing on F reebase from question-answer pairs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.604001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.548554Z digest=sha256:ee198158d2ff31849ce7ac946152c1af7ce2a0b02a7674a114510806c69dfd21

Observation 1220c9c1-7765-4959-b42b-3d103c8aa7e6 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.585064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.554031Z digest=sha256:84b11675168fb918e00a91f6b90404f1f78a5b0a2ed44ae69c234cd67cfab8c6

Observation 414fed55-78ec-4436-92e8-2150c3d30aa3 · outbound

This paper cites Occiglot fineweb v0.5, 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Occiglot fineweb v0.5, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.567965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.559649Z digest=sha256:3e943b1a3e784b4b69007762466a19e84a2cae4a4965f492499eaafe91631244

Observation c0102fa2-af1d-4228-8664-4be39b1f8aa5 · outbound

This paper cites Bridging the Gap for Tokenizer-Free Language Models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Bridging the Gap for Tokenizer-Free Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.565344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.565344Z digest=sha256:bbb0edbaa3c6ffa1d01b4f625a4dd52a5242b52770700ca9314067794188461d

Observation 6fe2a8b4-c99a-43ee-a251-0ea9812ca2f9 · outbound

This paper cites B ool Q : Exploring the surprising difficulty of natural yes/no questions.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models B ool Q : Exploring the surprising difficulty of natural yes/no questions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.570745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.570745Z digest=sha256:c3947de3d8404ea64bdac3cc63351751a580a504b9d97b53dc45d4e0c353808d

Observation e54a72f7-d208-4f53-ba61-ca3ff6cbd064 · outbound

This paper cites Bowman, Holger Schwenk, and Veselin Stoyanov.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Bowman, Holger Schwenk, and Veselin Stoyanov

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.540576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.581145Z digest=sha256:e64f34eb456d624fb7605f07af59caa301d6a6c1571d7c7b99ce0ea9339a9ef9

Observation d9904d11-5bd0-4444-b023-d8eae768ce5a · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.523486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.586757Z digest=sha256:855b48e1009e161f3a119dce3df0a087d380ff9917c14a26d831c280ad9096f2

Observation 8a168cc1-d65a-4b9d-aa61-d38ce01d6f32 · outbound

This paper cites T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.592713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.592713Z digest=sha256:c4e8b20b118990f90c65d17ae499d9e0d4d5df24b1006e61a0e8e64f5f49f38e

Observation 782f88d9-861e-4520-b119-60c30e302157 · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models BERT: pre-training of deep bidirectional transformers for language understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.597892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.597892Z digest=sha256:3006f49a8450355d376d93019a3f33afb4770d1a27f6cd130065a27ea9bcc0ae

Observation 5da89344-4a69-4cd2-823b-bccc91664fb0 · outbound

This paper cites The Llama 3 Herd of Models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.602612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.602612Z digest=sha256:1eb6c22ea75b7884796ba34fe5c6154e649ee0b24409da2abab10daa96e4c4f8

Observation 1b555442-5b81-4784-aa2a-f7b709d46e53 · outbound

This paper cites C haracter BERT : Reconciling ELM o and BERT for word-level open-vocabulary representations from characters.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C haracter BERT : Reconciling ELM o and BERT for word-level open-vocabulary representations from characters

Reference 16

Resolution
verified exact
doi, observed 2026-08-10T19:16:21.033923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.607848Z digest=sha256:5e17333d6e7b3c8cfe10311640b49504cd2d97521e056270e35ab9ffeb88a472

Observation 73b683fb-60de-4cce-92ee-48bb9f33c1c8 · outbound

This paper cites A new algorithm for data compression.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models A new algorithm for data compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.612830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.612830Z digest=sha256:2d5cfadf9b91381c66f809fea25e69cc886a17a829b64c88a4e9cfa6b004f9b6

Observation 656f9556-601f-446c-b152-fdca2744171e · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models A framework for few-shot language model evaluation, 07 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.617517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.617517Z digest=sha256:be2563ead16b88da0db0489dbc806cb76a136f946b259df9cc616939dbf966ba

Observation b78784f5-7632-4cdf-ad9a-b779d1a3f5eb · outbound

This paper cites Better & faster large language models via multi-token prediction.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Better & faster large language models via multi-token prediction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.485778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.623085Z digest=sha256:784855f05b1a1915a8a76951f37464040eeeea695ee076ea064a590b9456606f

Observation d9eca35e-c863-4917-81db-3df6e9cd72de · outbound

This paper cites Botvinick, Ian Simon, Hannah Sheahan, Neil Zeghidour, Jean - Baptiste Alayrac, Jo \ a o Carreira, and Jesse H.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Botvinick, Ian Simon, Hannah Sheahan, Neil Zeghidour, Jean - Baptiste Alayrac, Jo \ a o Carreira, and Jesse H

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.468811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.628241Z digest=sha256:4845fe1e1a3ac8391008925b2337979b71b9ccd7470cbaceee080d5987b102dc

Observation 6f10e283-7a25-418f-b2b8-b3659d8ce36a · outbound

This paper cites Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.633112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.633112Z digest=sha256:9380cb0232fc91c0732e5efe0017764dbf7b14edd33ba5db35765c8b23f16f7d

Observation 93bdbdc7-51e1-44f8-a080-af2eab71825a · outbound

This paper cites Measuring massive multitask language understanding.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.638405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.638405Z digest=sha256:c67d4f6de23274ea8e63b1e1769004bde86690e1f6694df43126cf007a741093

Observation 88d0a68e-a3a3-4cf2-8383-2381420f2759 · outbound

This paper cites T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.643281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.643281Z digest=sha256:9c4a7025b2c677fd82e84b77f2d16bb0fe07bbc624e642096372f5b4db130b49

Observation be429736-ed51-4512-9f19-57a2ea1d3b50 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.648279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.648279Z digest=sha256:2a975834396562cb2b76d90a204197def7437324617b260ba2aaa95922650952

Observation 660921d5-e930-4439-9f8b-10154c24dde9 · outbound

This paper cites Hashimoto.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Hashimoto

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.441723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.653654Z digest=sha256:b5085b3caa5cfe7f72fd9bcadbeeadc618b9c3f06d123b4dfe1f80d680ae3023

Observation 8933bb48-3fbe-4f11-8f7c-4ef8681aafba · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Truthfulqa: Measuring how models mimic human falsehoods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.658426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.658426Z digest=sha256:0e0fa0d2eb5385bbd48d06aaccecb8af3d52a56b3caf029f4bc5ceff3fcabaeb

Observation 7ecbde2d-a12a-4623-81e2-61804aaee209 · outbound

This paper cites Decoupled weight decay regularization.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Decoupled weight decay regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.663439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.663439Z digest=sha256:4bc8f48eaab77da1bfd803b2fb2a0be7b5a7152b70c27cd2e84745b206184df4

Observation 2fa6d751-0c1e-498a-89de-b04d24e64ae2 · outbound

This paper cites C har BERT : Character-aware pre-trained language model.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C har BERT : Character-aware pre-trained language model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.668608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.668608Z digest=sha256:d9c50bda1070be754600c9934a2d67aa4c24fa9d4892a158d5222c9e5fcbf16f

Observation 96634558-8a12-4a1e-883d-2375d3d142b5 · outbound

This paper cites Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.673355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.673355Z digest=sha256:6d8f512bde241f390ad20ae3687703b18bcc62e026987422cdfab515631c330f

Observation fbac2ce3-d16b-4cb6-a7e4-e25169452ad4 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.678410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.678410Z digest=sha256:f9859087990e1923ce83fa5286d90db2e9ae4647ec0888f0ec1d1ce8cad4308b

Observation a7114735-6bee-4cc9-a853-5bd59a842d07 · outbound

This paper cites Mmmlu, 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Mmmlu, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.413125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.684951Z digest=sha256:11061cc0f325a013d8156c1ef3dbf28b757bd8419c1aef236b0046e628a602f4

Observation fb4d1ed9-bada-4e93-8008-4a8e2c546bca · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 32

Resolution
malformed identifier
no resolver link, observed 2026-08-10T19:16:20.689907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.689907Z digest=sha256:8372f3bfbfe3a1db97aac192f486db09d52e916b498f788a68610c7798da1300

Observation 4920c0d1-870a-4343-9e07-a1ec82e2e7b6 · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text, 2023.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Openwebmath: An open dataset of high-quality mathematical web text, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.695245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.695245Z digest=sha256:03bb802c925dbaa90a3666e2217b2b09e155d317866438396d72ee80d07b393b

Observation 0f524630-137f-4cae-ad5c-b0dfdde3b122 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.700411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.700411Z digest=sha256:e8565bae4e2c2267564ea1a9afdfefe63ff57d7bdbdceb1b6bb4331b8961de00

Observation b122c748-f7d3-4765-b9f1-722e781c350d · outbound

This paper cites Language model tokenizers introduce unfairness between languages.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Language model tokenizers introduce unfairness between languages

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.384278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.705646Z digest=sha256:27cffd7e08e9d98ed03b69a2b7be9f63838ac451183d444888f7a20239cb78e2

Observation 0087ec6c-2aa5-4ee4-9ba8-bd65f7aed6bc · outbound

This paper cites W i C : the word-in-context dataset for evaluating context-sensitive meaning representations.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models W i C : the word-in-context dataset for evaluating context-sensitive meaning representations

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-10T19:16:20.711185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.711185Z digest=sha256:8c138024c9e4e985c1e209b462ca5f1767df636615abc42400996f410edc755c

Observation 336de9ed-8e9f-4f7a-8c1e-2ee9afbacf21 · outbound

This paper cites Germanbenchmark, 2024.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Germanbenchmark, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.367231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.716273Z digest=sha256:70a5a31210785f5ec2e27d36414ded6cad2113de34abd0d25e21efe844b256b3

Observation 38759c8e-ccbc-435a-aa92-a26b40a9f5cc · outbound

This paper cites Efficiently scaling transformer inference.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Efficiently scaling transformer inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.351227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.721304Z digest=sha256:8b77e992c9b5988c7efe3b0b084162860b7ab0b8e08e923f191de66246524af9

Observation 169d0db1-22b1-4897-958f-6eeaf8fe2c05 · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Winogrande: an adversarial winograd schema challenge at scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.726054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.726054Z digest=sha256:793d956be13f70f4e395cf07933d06ab189d16e891894b0af94bcab357f326e7

Observation aa5f9226-e877-49ac-b01d-6c3740db7e4c · outbound

This paper cites Neural machine translation of rare words with subword units.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Neural machine translation of rare words with subword units

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.730719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.730719Z digest=sha256:63b004e4a3de03cb50d6103ee673e951758b272804b58e312a44a22ac9fa3fb3

Observation 20926a3b-0f62-45e6-a6f7-e97a45f61b93 · outbound

This paper cites SpaceByte: Towards Deleting Tokenization from Large Language Modeling.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models SpaceByte: Towards Deleting Tokenization from Large Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.735877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.735877Z digest=sha256:4b3c867094086f11ec601f5b076d775e0d171847bab23db92e5b197e886cb15e

Observation b4a657ca-7ccf-4c2b-b7eb-eefefbfce51c · outbound

This paper cites From characters to words: Hierarchical pre-trained language model for open-vocabulary language understanding.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models From characters to words: Hierarchical pre-trained language model for open-vocabulary language understanding

Reference 42

Resolution
verified exact
doi, observed 2026-08-10T19:16:20.888678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.741445Z digest=sha256:081c5c6457b3d2e808317f79846138ab38392f2c2d45c8e504dfb499267311fa

Observation 3afa7e17-d33e-4208-9623-f0e9d23d5e58 · outbound

This paper cites Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.334501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.746281Z digest=sha256:f33378b27d9e7e25e087b432c7e24c5c54db670a69f98fd43845dfab307328b6

Observation e6ffb928-35e1-4086-96e2-56637c7195a6 · outbound

This paper cites Learn your tokens: Word-pooled tokenization for language modeling.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Learn your tokens: Word-pooled tokenization for language modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.751455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.751455Z digest=sha256:98300e49adeb7b4a64dfc0f7ce907aaf5cbbc0667fa607e048aa6f61a7136e78

Observation 878598da-bec5-4c1e-8e7e-89b674f45275 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.756271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.756271Z digest=sha256:49c513258f9d4378e46a45c4be907de3ca40c452f78a8791cf38fc242954ae60

Observation f7d60f3f-f0d0-4a13-9bd3-73083189e0f4 · outbound

This paper cites Skywork: A more open bilingual foundation model, 2023.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Skywork: A more open bilingual foundation model, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.761511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.761511Z digest=sha256:5059e5daf77020ccfc739955982d3a420d4126408632effdfc70db3c558b7713

Observation 937b5811-b6f4-445d-ba3f-d3069313addb · outbound

This paper cites B y T 5: Towards a token-free future with pre-trained byte-to-byte models.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models B y T 5: Towards a token-free future with pre-trained byte-to-byte models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.766254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.766254Z digest=sha256:780924a0e2a3cc9e9a7e3a51deee553675d2e8b4cd3540c376da2b5027e713e0

Observation f5ac978a-6974-4c1c-aac0-18890fd35edd · outbound

This paper cites MEGABYTE: predicting million-byte sequences with multiscale transformers.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models MEGABYTE: predicting million-byte sequences with multiscale transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:16:21.306409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-10T19:16:20.771088Z digest=sha256:6bb430a6d8c3e26c3178fb81160ee50f0454de4d91047a885a6970cbe11a0558

Observation a9035346-a518-4497-9d04-4106f6631988 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.775743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.775743Z digest=sha256:1d505db764001e61d6fcadafe740d9af23c97fdc56fddb95322e4fc461753898

Observation d8770478-1b87-44cb-9c73-05f7bd567182 · outbound

This paper cites @esa (Ref.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models @esa (Ref

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.781482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.781482Z digest=sha256:215b39feab68d83527c3e50dc89a4a1652e37ef6fac6362af234aee3d6c46fbe

Observation f090c69b-7d44-49b4-9c00-dfd4e5a5f5d7 · outbound

This paper cites an unresolved cited work.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.786674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.786674Z digest=sha256:21e1e44c914571931016b4bdabea15cf0908e6f7c553c3638b4b2534102ff42a

Observation 1ce16206-2216-4b9e-9bdd-b8a2ac8537d4 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T19:16:20.792657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:16:20.792657Z digest=sha256:1d59f4c39fbffe7b251ea593904c968518a10f0554a64067385223573880497f

Pith citing papers

Observation e2b73c34-d868-473f-b1c9-35637e98a049 · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.879364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:bf0a3540d585d5deab5a9b9a1d1cc0b6dd96693d583687b95c500f77bbda6391

Observation c8705c67-22f6-4b94-8b48-f2e04376b670 · inbound

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES cites this paper.

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:47:48.584286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T03:42:21.307552Z digest=sha256:f838e68f144f0a7e96e238597b121cc68c8a0cf75c4b84ba8108d1ae83649b2d