Pith. sign in

Paper Citation Record · LEDGER

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization

As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2505.17852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17852 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:49.061779Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T15:25:21.814232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:26:33.682527Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9085359-cf6e-40c8-9f34-c1ce24afa187 · outbound

This paper cites Demystify Mamba in Vision: A Linear Attention Perspective.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Demystify Mamba in Vision: A Linear Attention Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.162107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.162107Z digest=sha256:8a3de4ffa97c4ab7b4ccee4b660077aa43474aab700f943371523204cf320f0e

Observation 452e639c-9a53-41c3-9be9-ab3616972d59 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.233553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.233553Z digest=sha256:4c2a2ec5959bd4809d753a774b5c8b220654f237d2fe546f7f037be246a390e1

Observation b5fb0433-8b3e-4338-bd1a-a288115f0777 · outbound

This paper cites FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.516439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.516439Z digest=sha256:e2fb9f5fe8d6cecc51200218b7d0683bdec24f33e4df109fc9f18fe4ce5ea26b

Observation c591d36a-529e-4088-ba46-6240fde96dd9 · outbound

This paper cites Aditya Rawal and Risto Miikkulainen.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Aditya Rawal and Risto Miikkulainen

Reference 11

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:44:49.663860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:44:48.591763Z digest=sha256:6835bd519474f1c261540c35f8f70a307a058e43047d7c242b323b0617b2be0f

Observation 9cc023b2-3ebc-42ce-bb83-c01f2f4541a5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.962796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.962796Z digest=sha256:971562421e94781c978de67a009aba819f80b98aae909a71fd0fbc38c272a8a7

Observation d0da4f7d-486b-4e53-84fc-1612d11b0d46 · outbound

This paper cites Unbiased Online Recurrent Optimization.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Unbiased Online Recurrent Optimization

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.774666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.774666Z digest=sha256:b58dd2c6cb5f663051d707ff5c41ce1ba0f89f06fb579767160f243ea0815556

Observation d32bcea1-64c7-49e2-9041-ea2875a5bec2 · outbound

This paper cites Suraj Srinivas Malladi, Xiang Wei, Josip Djolonga, and Dale Schuurmans.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Suraj Srinivas Malladi, Xiang Wei, Josip Djolonga, and Dale Schuurmans

Reference 2005

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:44:50.012543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:44:48.329435Z digest=sha256:eac744de71d0cb22cf660ab4486c6e4003073e020d399e66e36606565e5f6250

Observation e5a19286-959e-4c74-8927-59b1935e1d25 · outbound

This paper cites doi: 10.1145/2908812.2908941.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization doi: 10.1145/2908812.2908941

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.689680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.689680Z digest=sha256:5c39adf795fcb5f5ea075587fdba0797ddf63147f1209e654b6ba6465e646fcc

Observation 4e49812c-ecd8-4cb3-817d-fde0fbffd9fd · outbound

This paper cites Efficient Transformers: A Survey.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Efficient Transformers: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.847581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.847581Z digest=sha256:64d4ea37f92a540dc4488247bb70e4191f18da6a4c5296131b53cb5a337f4450

Observation f3cf41fa-b619-4722-bf73-eaa3d9399cda · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization OPT: Open Pre-trained Transformer Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:49.061779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:49.061779Z digest=sha256:5b5d57ba1165447c3fa2971d85ad4d6cdc976a5ac3772623f0f402e7636e9bea

Observation b04da9fd-e410-415c-8f3e-82eb4e7ac224 · outbound

This paper cites Koopman-informed recurrent neural networks.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Koopman-informed recurrent neural networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.740119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.740119Z digest=sha256:63077f48b6db8e91c576e8b54742edc645458fa0de4c15c44aa1ec3dea36b3d7

Observation ccf275d2-8c12-401d-a93b-52688f3993e9 · outbound

This paper cites Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:49.001777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:49.001777Z digest=sha256:7f4d83d7c53386ad1f7c4a18fc7e7e2e4398073a03c03c09eb7c3493f2542a1a

Observation 8e3decb9-8c5a-4957-b82d-50682f695791 · outbound

This paper cites Textbooks Are All You Need.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Textbooks Are All You Need

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.421248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.421248Z digest=sha256:5c79bb37c0ce7895850f005483cd53bacea990652f0ba76ee2c9b4aacc2c2782

Observation ec1e4f59-413a-46e9-8bfc-2cae6704f589 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.613549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.613549Z digest=sha256:dacc3c0fa7e97fa8447df57d68d0db1f0d80572c501095672428085fb09e2ff3

Observation d958b103-bc00-4191-8d6c-a7e0973594d7 · outbound

This paper cites Alex Graves, Greg Wayne, and Ivo Danihelka.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Alex Graves, Greg Wayne, and Ivo Danihelka

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.890058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.890058Z digest=sha256:9957f9182bd4883158d1d8bc458570bae33dd0f6e59a7827bf777c497dc6fa55

Pith citing papers

Observation 4e24ab96-e287-4206-8acc-317294dcf2a4 · inbound

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers cites this paper.

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.685182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:25:21.814232Z digest=sha256:a5203adef2c68117728d7afe54d23719442c904ad34e2434919e9193b7c4eeb9