Pith. sign in

Paper Citation Record · LEDGER

Controllably Efficient Language Models

As of 10 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2511.05313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.05313 v2

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:34:03.638787Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved88
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f31ee432-0c57-4f01-b740-c4b3f637ab54 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Controllably Efficient Language Models gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:51.967485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:51.967485Z digest=sha256:64d1d805970b5b019576e6740fbd6e495ffa0b81618d49ac5a5f756eca4f497c

Observation 84bbd185-97ea-424b-ba7f-cd7235fec54e · outbound

This paper cites Zoology: Measuring and Improving Recall in Efficient Language Models.

Controllably Efficient Language Models Zoology: Measuring and Improving Recall in Efficient Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.100391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:52.100391Z digest=sha256:21a3681a83e05382bb94f9de85e2a80ac568fcca3795dd37f51d4cb9e9896862

Observation c3755546-f36b-4f49-a888-6856a99e84b3 · outbound

This paper cites Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes.

Controllably Efficient Language Models Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.278564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:52.278564Z digest=sha256:9465ac2ce7b4ab9071e4fc3450ce341e7310f33fd68eb4036b2a8e91220ac5ef

Observation b6a5d79a-bbc0-4398-9c2e-0d29f0220b10 · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

Controllably Efficient Language Models Simple linear attention language models balance the recall-throughput tradeoff

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.464870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:52.464870Z digest=sha256:44583bc00581db188c75a97f93058eed2c6a093df481df3cc4b74359d5b09f03

Observation 6e049d7e-d541-4901-a43f-5631b65a9810 · outbound

This paper cites Just read twice: closing the recall gap for recurrent language models.

Controllably Efficient Language Models Just read twice: closing the recall gap for recurrent language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.586902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:52.586902Z digest=sha256:37fd88d810530e8740cd8198c9ef1a7fab3f3952280adf650e37ac643b81952c

Observation 0d7a5402-de5f-4207-b442-21a3b8508d60 · outbound

This paper cites How to scale your model.

Controllably Efficient Language Models How to scale your model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.719531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:52.719531Z digest=sha256:4e067488cb250e301e405c9c08e6e58a0ade96d6f00e6ee43701fbf6030c2a18

Observation 73751852-6914-4cfa-88ea-b6650c79a247 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Controllably Efficient Language Models Neural Machine Translation by Jointly Learning to Align and Translate

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:52.869712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:52.869712Z digest=sha256:db618eb5ef0181e6fca5f32daa3515dc5bb62bec3a13dae1fb4e466891098e12

Observation d152bfad-8134-44c8-9b7a-f20422a98c28 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Controllably Efficient Language Models LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.021093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.021093Z digest=sha256:d568d0965859818e03bbbc88a97d1ad7bc039868ffa4418a55fda3dee5e10747

Observation 993f876d-5050-4296-8947-5fd8d7825eff · outbound

This paper cites Transformers need glasses! information over-squashing in language tasks.

Controllably Efficient Language Models Transformers need glasses! information over-squashing in language tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.144347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.144347Z digest=sha256:dc05eb319b1261233758be83cf8ec43aab510defe7f51c5619319e9365d79503

Observation 2276db91-4b47-41f6-8b8b-0d0dd4484280 · outbound

This paper cites Large Concept Models: Language Modeling in a Sentence Representation Space.

Controllably Efficient Language Models Large Concept Models: Language Modeling in a Sentence Representation Space

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.304599Z digest=sha256:5ba7eb317cbc12ea75edbc67edee4a54380062b2aa37addda684a387fea82969

Observation 8ca648cb-47da-487f-b18f-5029b19abe03 · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Controllably Efficient Language Models Titans: Learning to Memorize at Test Time

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.461340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.461340Z digest=sha256:6f840c93db4f9703fe7200f2bc7e8dc2283e2b069a03606e2a9e0417fdf2cd7f

Observation a642a052-83c7-4be8-a96a-fbfb33159988 · outbound

This paper cites Longformer: The Long-Document Transformer.

Controllably Efficient Language Models Longformer: The Long-Document Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.602718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.602718Z digest=sha256:1b6ba756cf32b2bf0082aadbd986d9010e7348218fd40ca3afb566651480ce4f

Observation 110e89bc-a0a2-47e3-bcd4-03bda494dd5d · outbound

This paper cites Flexivit: One model for all patch sizes.

Controllably Efficient Language Models Flexivit: One model for all patch sizes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.751932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.751932Z digest=sha256:06a2de8085948e97b4110b7f61cda600ff6186b8412005e473ef64ab7f15956f

Observation e4e77e54-04d4-4e2c-9459-fd07059fbf30 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Controllably Efficient Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:53.987182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:53.987182Z digest=sha256:cf0c150a471996a17e8c1dbbc6be0d2660ec934a14c8af7f6a53f1d81466b546

Observation 96683785-752a-412a-993f-85bec3c24e50 · outbound

This paper cites Adapting language models to compress contexts.

Controllably Efficient Language Models Adapting language models to compress contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.115968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.115968Z digest=sha256:a42bec2ffeb365ab5367a7f7e30910ca2250c30696756415556309bea0169c02

Observation 4c865657-ed21-45ef-9bed-d0a716a929ef · outbound

This paper cites Overcoming a Theoretical Limitation of Self-Attention.

Controllably Efficient Language Models Overcoming a Theoretical Limitation of Self-Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.260769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.260769Z digest=sha256:49fa233a8225118facfcdd1a97675e825b3a1be7ebc6348100d44f12d7e2b086

Observation 5df8a36e-2400-407b-9a9c-f2035357cf9e · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Controllably Efficient Language Models Generating Long Sequences with Sparse Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.352747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.352747Z digest=sha256:175b0e2c6ec477241a6fb454c838fad61ac789c9ecbb9b73128ddc2c5d824932

Observation 52d5c907-7a1f-4118-9de4-03bd4e0be61d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Controllably Efficient Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.444640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.444640Z digest=sha256:ea42afaa8e21787c8413be4b6585e1d850e64069c7d246106fa66caef1b8bf1f

Observation 28f59c38-b09a-4d5e-946b-f7cbf66cb2a9 · outbound

This paper cites Funnel-transformer: Filtering out sequential redundancy for efficient language processing.

Controllably Efficient Language Models Funnel-transformer: Filtering out sequential redundancy for efficient language processing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.546457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.546457Z digest=sha256:320b949babd01e1b9c7708a681c87f20b531bd33ab904309558c72b7efa0b0db

Observation 8a44282c-6d9b-40c0-bb9b-04968cb2d662 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Controllably Efficient Language Models Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.645983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.645983Z digest=sha256:3e907b306a211867257c58b06b1c2b915e5bad7a32cccb762839c7d7ad52a75e

Observation c7ec9ae0-c2de-465f-b26b-fa48d9f64456 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Controllably Efficient Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.729595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.729595Z digest=sha256:1fe0781f35c9e6449a52b87e0291e5af6836a603185d823cb615747a27e95301

Observation 3d3b6f76-9922-4c8b-8191-94d830645711 · outbound

This paper cites Matformer: Nested transformer for elastic inference.

Controllably Efficient Language Models Matformer: Nested transformer for elastic inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.834896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.834896Z digest=sha256:3c97c89367b9fb491b77a2156e1f4027a9d5f7ab25013abbb4f34ce2d42bce7c

Observation 563bfbff-9d26-4e4d-bcfc-01b4def9fef9 · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

Controllably Efficient Language Models Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:54.935011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:54.935011Z digest=sha256:91307af41b181634829d47a0a54d3d36ea02a3b4e5dbda533157ca736df0d2bf

Observation 2c36dbf2-848c-4c5f-9355-ebe795c7abd3 · outbound

This paper cites DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs.

Controllably Efficient Language Models DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.068944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.068944Z digest=sha256:84f35bd6eda7aee26860da5fbab86d8fbfe65c4c12af4d95185fecb4cf64201b

Observation 459b76c8-ac26-4782-8cc5-a6066ff01919 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Controllably Efficient Language Models Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.173974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.173974Z digest=sha256:a0243a42593ae9d3939e1fd96290d21f0bae810a5841a6e2321c56a2164acdf6

Observation 596865b9-53a5-47c0-afad-0e1bb137fd30 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Controllably Efficient Language Models Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.285309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.285309Z digest=sha256:69de9e66f87244b63935e9c9bfdeb8023d6b6abda413c2e417674b5e2c60625d

Observation 4f67ff52-2dfa-4347-bcf6-4eebe876007b · outbound

This paper cites Ai and memory wall.

Controllably Efficient Language Models Ai and memory wall

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.408794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.408794Z digest=sha256:3e493f749fde28fb10945d64ea039d79a137dda0457205d66ccdda376c7415b8

Observation c28ea875-8e7c-43e7-ac6a-c860e6cbb4b9 · outbound

This paper cites Multi-Token Attention.

Controllably Efficient Language Models Multi-Token Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.536316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.536316Z digest=sha256:21ea0ca58a8eee9dec7e035532e0564e621a3939ae41799246e1961e0cc19fa5

Observation dd420224-1cd1-4fab-b3ea-3054f9e4d8cf · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Controllably Efficient Language Models Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.673830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.673830Z digest=sha256:209e3a5404eed44ccbd9616261079fab20bfffadc88c747a0a28835a07416510

Observation 42e700d9-5079-4f05-8136-aee825a7039d · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Controllably Efficient Language Models Efficiently Modeling Long Sequences with Structured State Spaces

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.769204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.769204Z digest=sha256:14024811f32d7de2e6036e14938f90c1e1a5873263229d9af790950e71911f3b

Observation 1bebe3b0-619b-43f7-afb6-29ba8e79e53b · outbound

This paper cites Transformer in transformer.

Controllably Efficient Language Models Transformer in transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:55.903631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:55.903631Z digest=sha256:d4bade2c2edc5c06d945f12a042600774ff5ecc0aff33cb6133fd616c79267af

Observation 7c548c31-0ca0-4a5e-ac79-487c786081a7 · outbound

This paper cites Block transformer: Global-to-local language modeling for fast inference.

Controllably Efficient Language Models Block transformer: Global-to-local language modeling for fast inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.006822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.006822Z digest=sha256:791174fdf1d3e0e189c02d420ef7f1e71f8f8af647e9e7b9590d9dec81ae1e20

Observation 81e42c13-d365-4e0d-af5b-68def8d0eac3 · outbound

This paper cites functorch: Jax-like composable function transforms for pytorch.

Controllably Efficient Language Models functorch: Jax-like composable function transforms for pytorch

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.116543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.116543Z digest=sha256:c65fb16c5f6692cbad1332b49e568b7370f3bff91c6d7960b5170924a2bae692

Observation 4fdf4524-0a13-45db-94ad-9cd088e66d7b · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Controllably Efficient Language Models RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.240171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.240171Z digest=sha256:edd4d229292bd7c55d03ca05e5d0d7196fac91e33dd7277269ebae1c4f5ccb2b

Observation 5e9d73fa-dc57-4337-abf8-ef5c8045dbe2 · outbound

This paper cites Dynamic Chunking for End-to-End Hierarchical Sequence Modeling.

Controllably Efficient Language Models Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.393190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.393190Z digest=sha256:a2fdea4a7814121094d421a7e7b9eee485b1e6f7fdcc1deb11bbcf9dd56d435e

Observation 8103b078-7df7-4329-8404-fb9a9755987a · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Controllably Efficient Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.548151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.548151Z digest=sha256:659236f1261aebfc25c025dff3c80dcebec916c2f38ae809f6147fccb16beeee

Observation af4c7222-d26c-4ab5-a758-e926e265ecd7 · outbound

This paper cites Mistral 7B.

Controllably Efficient Language Models Mistral 7B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.627634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.627634Z digest=sha256:634711b522d84b50f95b69022690e6699b39d724c640688c2e234cbc720a0267

Observation 1fecb2ce-a985-41f8-841f-e30c937b7451 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Controllably Efficient Language Models TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.777094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.777094Z digest=sha256:1e23b9eb226ef6198be5cea82e345f5c6c34b7d2d701266ea33e53ec137db214

Observation 52f09885-9892-499e-9b8a-8781767292af · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Controllably Efficient Language Models Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.934583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.934583Z digest=sha256:db821b36f427f11b0082555f633f96f092aeda91a1d31e0207f1931b62948caa

Observation d6f56781-7e16-40b6-8dc6-0feb16c7262c · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024.

Controllably Efficient Language Models Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.061128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.061128Z digest=sha256:b51ef3950f30ddc4b532c63b6d08d613382e653b6ee97bcc731f340d3eeeadfe

Observation 1f175543-bfcd-481e-89cf-42e0491e105a · outbound

This paper cites Matryoshka representation learning.

Controllably Efficient Language Models Matryoshka representation learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.173889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.173889Z digest=sha256:a30258b31752b2352805675c895e33b80d96864acdfaad07045259b06de4cec8

Observation f5fccd5f-e1b4-4801-9feb-36533a92919e · outbound

This paper cites Inference-time hyper-scaling with kv cache compression.

Controllably Efficient Language Models Inference-time hyper-scaling with kv cache compression

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.348990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.348990Z digest=sha256:c568058dbe5b4f89dfeaa3713338a1b5638112674b7e1b4c71425cd271b0025b

Observation 38fe05fb-1178-4439-8c23-8595ac381ef6 · outbound

This paper cites A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention.

Controllably Efficient Language Models A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.502308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.502308Z digest=sha256:295be043c8f7860b1ddae7865edebe3347528d02acd460c4b2ea43497d75bee1

Observation be78e69b-d60a-44a1-b951-354d1b35ac41 · outbound

This paper cites InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU.

Controllably Efficient Language Models InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.667935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.667935Z digest=sha256:1b9c6a6b0097a370540281de384a0b6b58c54a2a7ed6ba06d081d051d88675b7

Observation b4737703-15f4-474d-aad4-7df565ad98be · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Controllably Efficient Language Models Snapkv: Llm knows what you are looking for before generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.776369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.776369Z digest=sha256:b883356ed90ac1142d81e6cf195bcaf54f8ea4670422e50f104eb6cf6329afd3

Observation 5c41e4df-f6d0-4eb8-b56e-a50412000701 · outbound

This paper cites Openceres: When open information extraction meets the semi-structured web.

Controllably Efficient Language Models Openceres: When open information extraction meets the semi-structured web

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.878890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.878890Z digest=sha256:b75c2364cbe62cae7ff12dd019f66a0e723defbaa043e361b469fd1cc0868573

Observation 40a3483b-0d90-4f0b-aff9-ce0ed6750d7f · outbound

This paper cites Decoupled Weight Decay Regularization.

Controllably Efficient Language Models Decoupled Weight Decay Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.019489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.019489Z digest=sha256:ff771c9d444a50441f1da3c93d2a3847ffdcb4d5b26c974c7f8a97326ecc3da6

Observation f0433471-6744-4cde-8e03-160a390f6dfa · outbound

This paper cites BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining.

Controllably Efficient Language Models BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.134824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.134824Z digest=sha256:545f62c05b57b671e620c45bb076eec9d9c8aefad106aee925b003d96a0a59d0

Observation 76bcc4f8-8028-46d7-94ba-93e5722eede8 · outbound

This paper cites Pointer Sentinel Mixture Models.

Controllably Efficient Language Models Pointer Sentinel Mixture Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.233366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.233366Z digest=sha256:deb795d765d0a97084a321cd4a1c52932b9f8097066e1449c3f0961ac6f6ecdb

Observation efee2872-8e7c-48b7-a3d9-2e04f4343eb6 · outbound

This paper cites Online normalizer calculation for softmax.

Controllably Efficient Language Models Online normalizer calculation for softmax

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.412745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.412745Z digest=sha256:4a9a84b39ec5666e975f64d1f57b84fc4fe91ac586baea63f6109856800f4547

Observation 8a0768db-e298-4849-bca3-9133dbcbff82 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Controllably Efficient Language Models OLMoE: Open Mixture-of-Experts Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.538138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.538138Z digest=sha256:994d75961c13b841d15f14b165cdc969cd974c45a579e03c639d8279f0918e6d

Observation 0eb3e795-1ba5-45fe-bed3-4fff3bc1ada8 · outbound

This paper cites Hierarchical Transformers Are More Efficient Language Models.

Controllably Efficient Language Models Hierarchical Transformers Are More Efficient Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.639384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.639384Z digest=sha256:fab636077781a17f612317897ee6eb8c13729dc71b6a8d574b3dd701774f1adb

Observation 4819815b-cf87-4e23-ac21-f99d3131a3e3 · outbound

This paper cites Efficient Transformers with Dynamic Token Pooling.

Controllably Efficient Language Models Efficient Transformers with Dynamic Token Pooling

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.766767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.766767Z digest=sha256:62905d48d8eb78f5e3763fd0239def373b7d2430d8ac2b31d2fe5e6a4ef6d896

Observation b6763228-711e-4bcd-8240-3fc1b58fd1d5 · outbound

This paper cites Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference.

Controllably Efficient Language Models Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.921784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.921784Z digest=sha256:0afcf149edfc4576fb2cbf6f1974239e5c703286609fb44f73835151a1fc8b7c

Observation 47d5ac34-31c3-46d8-aad6-586dfe9b5dea · outbound

This paper cites The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs.

Controllably Efficient Language Models The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.082317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.082317Z digest=sha256:3980d6354a275050a157890dcd6ae8fa818c6afffbf2fa103321cec1cf410657

Observation feceae77-34d1-45af-8e80-c03eabab59b4 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Controllably Efficient Language Models Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.230152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.230152Z digest=sha256:48e5ad204d66f93122f56c9559b95a461ccabcffb6610088b2d066bdd73bc42d

Observation 73da1e76-a640-4cee-9104-fe5a029dab56 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Controllably Efficient Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.340379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.340379Z digest=sha256:830a535a7f4d2027a2ed104bcf309624b8c759a5f696a32e88367c31973c30e0

Observation 0c02a50b-e20a-4b05-bfdb-e6b92c29115d · outbound

This paper cites Hierarchical transformers for long document classification.

Controllably Efficient Language Models Hierarchical transformers for long document classification

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.513639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.513639Z digest=sha256:a6cd51f4087b060b1e19ea7d7e34d2809cf7b1cea554e64624312d25f86a4d85

Observation 86e1cb58-a968-4965-b88e-9374c0b2110a · outbound

This paper cites Image transformer.

Controllably Efficient Language Models Image transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.679673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.679673Z digest=sha256:99639635416bdf63152c5ed7553092a5d68f6347e7da5e1d7b945c181f5470a5

Observation 10255ec1-712d-4dd7-8133-72da4f8dce19 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

Controllably Efficient Language Models The fineweb datasets: Decanting the web for the finest text data at scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.847950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.847950Z digest=sha256:a09c66a51c1f1ab85bc6132df3050b10a1f03648ad6a81d00b60c1ec3e375da3

Observation bb3bf95e-dfd0-4a3e-9f33-0ab58208b486 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Controllably Efficient Language Models RWKV: Reinventing RNNs for the Transformer Era

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:59.981988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:59.981988Z digest=sha256:39cb27691bf716bd53f03ceaa0c1dac46f4c8f31a8f3c4b5f779ab985f5e3d18

Observation 8221b64f-0408-4959-90f2-7191a66bc340 · outbound

This paper cites Hyena hierarchy: Towards larger convolutional language models.

Controllably Efficient Language Models Hyena hierarchy: Towards larger convolutional language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.183830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.183830Z digest=sha256:dd9c0aa450a6183a8cca0dcd8db9ff912fee3013d6a5bf8c5ed3e49b07fe19d9

Observation 47e1585d-1060-4773-af56-61e5cd5886f7 · outbound

This paper cites Qwen3-next: Towards ultimate training & inference efficiency, 2025.

Controllably Efficient Language Models Qwen3-next: Towards ultimate training & inference efficiency, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.306085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.306085Z digest=sha256:87431781470eed1194a814d981be1b2d7229ddb75a74e2e9d59b3255a4df08c5

Observation e2273f60-b621-4d92-b865-eaaa36e7f3d3 · outbound

This paper cites Rae, Anna Potapenko, Siddhant M.

Controllably Efficient Language Models Rae, Anna Potapenko, Siddhant M

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.412422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.412422Z digest=sha256:7f175201bc9ad4a7f5687fd0d0b0bfae32bd720d1fa8dd7989da01b7cdf7d84e

Observation 234d56c0-f7ff-4d5f-ba8e-d3301656faab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Controllably Efficient Language Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.537110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.537110Z digest=sha256:6cffd0fe8656185dd6ffe6418ff455b11d5ed0cdadfb383c490bab1c8d8d56d7

Observation 11fa7bf4-aa49-41ce-8d0a-f56ba3fe0caa · outbound

This paper cites Know What You Don't Know: Unanswerable Questions for SQuAD.

Controllably Efficient Language Models Know What You Don't Know: Unanswerable Questions for SQuAD

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.637989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.637989Z digest=sha256:bebabc8b5dcb329bd51531e40c06184f9b30d5a1f93fd2193fa853581806a9dd

Observation 6f30815b-300c-4c39-8b6d-3ad0958f855e · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Controllably Efficient Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.714474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.714474Z digest=sha256:8948024645c41f49f331f543389a72c472dbd25c9963f109e41575e49f9e1ea0

Observation f5dd56e3-022e-4f4d-bd56-fcb3fc0462cb · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Controllably Efficient Language Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.847597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.847597Z digest=sha256:ee047755ba8dddec0a5cee9151ba26e38a33a8003ad2762ebfdb355901d3b6ce

Observation 34ea441a-a600-4372-8f3f-69fec8920cae · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Controllably Efficient Language Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:00.954690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:00.954690Z digest=sha256:723bc95c3161b9d984f7e6aedb0c78b2f86daf526f46c2b512d8e3d6537f169e

Observation e5e016b5-ff1a-47ff-801f-87f1ce3c3045 · outbound

This paper cites Spacebyte: Towards deleting tokenization from large language modeling.

Controllably Efficient Language Models Spacebyte: Towards deleting tokenization from large language modeling

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.058268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.058268Z digest=sha256:5e2d8fdd99b22f6f10cfe3e73f0a630f14eed463589f7b669f6ea8bccc2060fe

Observation cf322614-469e-447d-826a-2ea1b609e2e5 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Controllably Efficient Language Models Retentive Network: A Successor to Transformer for Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.136033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.136033Z digest=sha256:85de50fe4e4ccddcc82b2333bc46d374d986e81e7a19c38216c71330702b3386

Observation 2f78c862-f3a3-48f8-8b14-791d1edf10ee · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Controllably Efficient Language Models Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.311105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.311105Z digest=sha256:81621cfe1f60f741ac87a1c44a6728568863c4b77abb83cbd597a37bf554cdbe

Observation 6f0f53df-7c87-4a38-b167-3698409b3afa · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Controllably Efficient Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.405536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.405536Z digest=sha256:6970484eb42fd27de55781e3daeefea4e3efb75072a6c2fd00b2772e812e095b

Observation 95725d55-15aa-43a7-992c-9d7f3df25158 · outbound

This paper cites Attention is all you need.

Controllably Efficient Language Models Attention is all you need

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.628861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.628861Z digest=sha256:095ee8b378130eb089d12c37658cd68e190625972436a7d58c8c9dc995aea708

Observation b4f79500-c358-4059-bb5f-98b746648751 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Controllably Efficient Language Models An Empirical Study of Mamba-based Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.701947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.701947Z digest=sha256:08eb1de60a7d8e6cf7c39b8891f0bafbf4477adfa083ae85c187c17971f6bfb7

Observation 701d657e-5446-47c6-8977-09b3b9e41f66 · outbound

This paper cites A Systematic Analysis of Hybrid Linear Attention.

Controllably Efficient Language Models A Systematic Analysis of Hybrid Linear Attention

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.832889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.832889Z digest=sha256:b26e4c584ddf211e4e22bd760b2e4a910dd8ed5f9df8eec0e9f36b0a3c512ab3

Observation f8d687bc-a3ce-47d4-ad52-4884142b9fc0 · outbound

This paper cites RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval.

Controllably Efficient Language Models RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:01.985027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:01.985027Z digest=sha256:57474c75100b89118ae32e6baa31aa730456178ec53dbbea2fcb44fc9f243ce2

Observation a566006c-e113-46b9-8acb-eb055fbf78d1 · outbound

This paper cites Qwen3 Technical Report.

Controllably Efficient Language Models Qwen3 Technical Report

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.066497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.066497Z digest=sha256:cfc37e758ace9a1ce77c64c62da3cf6af079d9d68ca8ebce8b2c3e6b25b48fc9

Observation 8fca704d-1836-4dbd-a532-7dfdb71489d4 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule.

Controllably Efficient Language Models Gated delta networks: Improving mamba2 with delta rule

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.202813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.202813Z digest=sha256:12ae8f8180536ff933f40246877b7819f69190dce604710d083218841c886247

Observation 4fe82a1b-2d77-4d51-90d9-a34c60dfdacd · outbound

This paper cites Long-context language modeling with parallel context encoding.

Controllably Efficient Language Models Long-context language modeling with parallel context encoding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.335542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.335542Z digest=sha256:c9cf580f4633955867f5bb2122d219bbb08f4f813ebcb769e2f04187322ae5e9

Observation aa9a0b41-f33c-4c96-b949-0730b1adcf4d · outbound

This paper cites Megabyte: Predicting million-byte sequences with multiscale transformers.

Controllably Efficient Language Models Megabyte: Predicting million-byte sequences with multiscale transformers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.500224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.500224Z digest=sha256:7c73750fcbf28f8ecfec4d8e4e1c4dd13b56c51f87f8533bd2f39eab2edd225e

Observation 6ad72cb6-fabc-4662-910c-855e4f3259d5 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Controllably Efficient Language Models Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.651609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.651609Z digest=sha256:9f49e42146da50f93dad0b2b793ddf8c9e6e38dc7641e1a143a465fb127f42f0

Observation d868f5dc-9d09-41bb-8828-910976dee525 · outbound

This paper cites Big bird: Transformers for longer sequences.

Controllably Efficient Language Models Big bird: Transformers for longer sequences

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.781218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.781218Z digest=sha256:d7b22308a6d2699c88f0e05f409fb8fcfc65a1f219c1451c0cd97e93f3a3ab1b

Observation 31f4c21a-6f7e-402e-8f27-c7649ca3f272 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Controllably Efficient Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:02.975875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:02.975875Z digest=sha256:0a46be5ec6564d32a909bd579299f73d1a8793bf83577955a8054169bdc8b692

Observation e5772a9c-62cc-4f72-a18a-c57d24330541 · outbound

This paper cites Preference learning made easy: Everything should be understood through win rate.

Controllably Efficient Language Models Preference learning made easy: Everything should be understood through win rate

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:03.105381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:03.105381Z digest=sha256:5b4a94813295d4b5617ef4a852af71942b23ad24353fd9abf0f9a8684edceb7d

Observation f831c6f1-5bb5-4093-9c27-55c64da39b18 · outbound

This paper cites write newline.

Controllably Efficient Language Models write newline

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:03.290389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:03.290389Z digest=sha256:b9fd3c0a6a75cfece5c9c4d4028efa5f1fbbd5102df2fb8fd3af6b33a893401a

Observation 18133eed-25a0-4780-94cc-1e1e61e08850 · outbound

This paper cites @esa (Ref.

Controllably Efficient Language Models @esa (Ref

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:03.396939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:03.396939Z digest=sha256:dbe42e221fd41233f0420056d0fc9a0ed02ca84e218edacf6265a7bc4dabef15

Observation 27f6c759-d020-4966-99be-262a53674c00 · outbound

This paper cites an unresolved cited work.

Controllably Efficient Language Models Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:03.499389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:03.499389Z digest=sha256:f55994747310c54a1a25cc1aa5b615fab62110f6948415e17605113837b69bc4

Observation 56c04b40-c098-48b7-9951-b87380d5e5a8 · outbound

This paper cites global” chunk embedding. An embedder first compresses each chunk independently, then these “local.

Controllably Efficient Language Models global” chunk embedding. An embedder first compresses each chunk independently, then these “local

Reference 89

Resolution
malformed identifier
no resolver link, observed 2026-08-03T23:34:03.638787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:34:03.638787Z digest=sha256:714ce293a347067bc43f36db19e37248c25c56a656a624fda424b314a3e9b9db

Pith citing papers

No inbound Pith citation observations are available.