Pith. sign in

Paper Citation Record · LEDGER

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2505.17852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17852 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:49.061779Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T15:25:21.814232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:26:33.682527Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9085359-cf6e-40c8-9f34-c1ce24afa187 · outbound

This paper cites Demystify Mamba in Vision: A Linear Attention Perspective.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Demystify Mamba in Vision: A Linear Attention Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.162107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.162107Z digest=sha256:7c7e649d96f298c294a20aed07956b33e40d92515008ab5153b3365a81a76ab4

Observation 452e639c-9a53-41c3-9be9-ab3616972d59 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.233553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.233553Z digest=sha256:fa8eea31e50e3acd4004666702b051cee06cae20b681de27ada7964b09375121

Observation b5fb0433-8b3e-4338-bd1a-a288115f0777 · outbound

This paper cites FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.516439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.516439Z digest=sha256:509a0fb6d644b3706e7695a266eb9efc9942d23bd8d6bd0e8590e02e285c42cf

Observation c591d36a-529e-4088-ba46-6240fde96dd9 · outbound

This paper cites Aditya Rawal and Risto Miikkulainen.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Aditya Rawal and Risto Miikkulainen

Reference 11

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:44:49.663860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:44:48.591763Z digest=sha256:c531f7bac96b63f9dec4e7eb98c5ca3b219e6c5b47ca31be5366f42bcf5155bd

Observation 9cc023b2-3ebc-42ce-bb83-c01f2f4541a5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.962796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.962796Z digest=sha256:736f6d82d1968120334935d6814302f8ba4eb7c0bf74a18e781ce1adc53ec4fe

Observation d0da4f7d-486b-4e53-84fc-1612d11b0d46 · outbound

This paper cites Unbiased Online Recurrent Optimization.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Unbiased Online Recurrent Optimization

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.774666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.774666Z digest=sha256:bf4a6e00d805045887cd08381ca22a18abbc8b8c0808c9f595d4dfa109d8539e

Observation d32bcea1-64c7-49e2-9041-ea2875a5bec2 · outbound

This paper cites Suraj Srinivas Malladi, Xiang Wei, Josip Djolonga, and Dale Schuurmans.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Suraj Srinivas Malladi, Xiang Wei, Josip Djolonga, and Dale Schuurmans

Reference 2005

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:44:50.012543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:44:48.329435Z digest=sha256:6b9477b314df209b5c539b69a7f802988579a4409f3b0dba3d969c225b0e3dff

Observation e5a19286-959e-4c74-8927-59b1935e1d25 · outbound

This paper cites doi: 10.1145/2908812.2908941.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization doi: 10.1145/2908812.2908941

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.689680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.689680Z digest=sha256:cb0ad92cb348ad5575b053a9cd9d7889b3b9c1fd7a577ec00ba3d51b4d658273

Observation 4e49812c-ecd8-4cb3-817d-fde0fbffd9fd · outbound

This paper cites Efficient Transformers: A Survey.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Efficient Transformers: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.847581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.847581Z digest=sha256:f39aabb90df63fa0cd24af08a728adbcd0876c87e32ca4c28ba2faec5b291632

Observation f3cf41fa-b619-4722-bf73-eaa3d9399cda · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization OPT: Open Pre-trained Transformer Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:49.061779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:49.061779Z digest=sha256:069d206906e8bf9f3d0ac6f6a10e5bccaf57f5e6d564a69a797866258d79c046

Observation b04da9fd-e410-415c-8f3e-82eb4e7ac224 · outbound

This paper cites Koopman-informed recurrent neural networks.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Koopman-informed recurrent neural networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.740119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.740119Z digest=sha256:b32958fe3593f3c9cd1eae8b6820b4e07e07b9ff1aa9c4eaab329834d60a7ef7

Observation ccf275d2-8c12-401d-a93b-52688f3993e9 · outbound

This paper cites Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:49.001777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:49.001777Z digest=sha256:e200fd58dfd2295c0f505e33484aa4a1bf236dfdc59c4e552ce864c7291878fa

Observation 8e3decb9-8c5a-4957-b82d-50682f695791 · outbound

This paper cites Textbooks Are All You Need.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Textbooks Are All You Need

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.421248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.421248Z digest=sha256:1e21f7a64f7c948d1b8217a926ee2fa68d23969004bd8124cd89cd6ac5ce570b

Observation ec1e4f59-413a-46e9-8bfc-2cae6704f589 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.613549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.613549Z digest=sha256:47d6fbde2b26e560569526a220aa784a5deb99ad30a2a1d19aea8dafa2cec8e4

Observation d958b103-bc00-4191-8d6c-a7e0973594d7 · outbound

This paper cites Alex Graves, Greg Wayne, and Ivo Danihelka.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Alex Graves, Greg Wayne, and Ivo Danihelka

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.890058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.890058Z digest=sha256:1ad210f7c37ec700a9641b4ea61e9f567c9e26176e8e15ad583ba9c79763f3a1

Pith citing papers

Observation 4e24ab96-e287-4206-8acc-317294dcf2a4 · inbound

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers cites this paper.

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.685182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:25:21.814232Z digest=sha256:a7b10391071ae0f424aeba37bfeb8d89a61636af3df603e100fa90b5ee77c2fe