Pith. sign in

Paper Citation Record · LEDGER

Towards Distributed Neural Architectures

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.22389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22389 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:14:20.461704Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0975920-8ccb-4fc5-bfeb-2264c07accdc · outbound

This paper cites GPT-4 Technical Report.

Towards Distributed Neural Architectures GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.370267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.370267Z digest=sha256:607c3bcbf584ce606b501bf0adc3df331e638e4b210e89a87a33822721a64a45

Observation 2482fa62-7024-41dd-992c-ebe0d7109451 · outbound

This paper cites Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei.

Towards Distributed Neural Architectures Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.796382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.796382Z digest=sha256:0b6fcf7aad42cf62d00b97d13b4e9a74a6c41904742b1aa3bd3cc44d7e755d74

Observation 280f67af-b0b0-4421-a7a3-b66c0271f621 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Towards Distributed Neural Architectures Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.944144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.944144Z digest=sha256:99914ea7a1b6ee92a6416c4aa5a50d926e43a6c49fac8a2bfd81bdd0512e6f32

Observation 1ca5aea9-79ea-4d95-a551-6c355c8775e5 · outbound

This paper cites This may be because we are not considering a setting with high sparsity.

Towards Distributed Neural Architectures This may be because we are not considering a setting with high sparsity

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.567025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:20.148222Z digest=sha256:7a718595dd35e02bdd45817fa5e7ff9e951894c326482f2622481911cd12b741

Observation 16ee2b9c-a491-44bb-b343-14df6bd663b4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Distributed Neural Architectures DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.357714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.357714Z digest=sha256:9b1ebaa1bbbc607914587f97ffec66ad2bfe7ecaa091f505c70517b605e87b9b

Observation b37c74e1-eadf-4001-8422-2a432370ff72 · outbound

This paper cites DARTS: Differentiable Architecture Search.

Towards Distributed Neural Architectures DARTS: Differentiable Architecture Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.416905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.416905Z digest=sha256:a167b0ede8750b52bb34902cb46dd4bdb26a7337501cab1df51e34991a017d07

Observation 45d82e4b-be64-4207-a54c-5c5bbec99369 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu.

Towards Distributed Neural Architectures Fineweb-edu: the finest collection of educational content, 2024.https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:22.111671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:19.468676Z digest=sha256:b3bc9b27a26760e0fa296e609215ad33e520f57dfcc9f5c2851ad54991987aa8

Observation 691fbcff-885e-4b63-9d50-80ac488b6cdc · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Towards Distributed Neural Architectures Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.596966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.596966Z digest=sha256:b6b264132322aae64d716428f4ca1819f2dd4abc93793918ebe588a79b9827c9

Observation d1811a32-701b-46bc-9a0f-716da38d2905 · outbound

This paper cites Deep Information Propagation.

Towards Distributed Neural Architectures Deep Information Propagation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.643965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.643965Z digest=sha256:ddb086455d4eb390eb24cc40cc65ca63b7b9f648ae9c0419475600b5ce37ec93

Observation b3c29bac-d4fe-4df1-bc35-eb61e82108aa · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Towards Distributed Neural Architectures Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.690467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.690467Z digest=sha256:473cd4a9f7b12cf19a688eed946ac3b914bcf5f4e388931ba73a310f1c61bad2

Observation 2a4618a6-8bc6-4de1-82ac-58aa2f2a9a07 · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,.

Towards Distributed Neural Architectures Dropout: a simple way to prevent neural networks from overfitting.The journal of machine learning research, 15(1):1929–1958,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:22.010935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:19.742140Z digest=sha256:fc5b8b045dd44921ae7528494ccb12404ec53d52e05b3c38d0347ef6eb54ada6

Observation 66a3de26-abb7-40ef-a9d7-45045dcee1bd · outbound

This paper cites torchtune: Pytorch’s finetuning library, April 2024.https//github.com/ pytorch/torchtune.

Towards Distributed Neural Architectures torchtune: Pytorch’s finetuning library, April 2024.https//github.com/ pytorch/torchtune

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.905199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:19.869553Z digest=sha256:4bd2d390388ec9845828fb681b298f98d28d6f997813e1676b7c99a65d84a12d

Observation 9c6aa4a5-2d58-45df-9390-c1ea56231cb4 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Towards Distributed Neural Architectures Neural Architecture Search with Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.952473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.952473Z digest=sha256:d348f21e1c0277ea46e0b4cd36eb5022e09a787dea6dc026698cfe105a2b9a5a

Observation 126d581e-3eb2-482f-8a24-b3638f6ea7a0 · outbound

This paper cites an unresolved cited work.

Towards Distributed Neural Architectures Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:14:21.820871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:20.040098Z digest=sha256:6895baf9a055be679eda731b7fa865bb8a555ec48a81c74958aba5988c6911d7

Observation ba1c9cb3-de61-4bc8-baaf-a7e2e568a5b4 · outbound

This paper cites B Module Usage and Load Balancing We plot the module usage distribution for all DNA models used in the main text in Fig.

Towards Distributed Neural Architectures B Module Usage and Load Balancing We plot the module usage distribution for all DNA models used in the main text in Fig

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.678283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:20.045284Z digest=sha256:f1fa36bd66a355d02c457d724f3dcce1c8128faf3cd624e4b8d2075337d8b2b9

Observation 6395a7bd-f507-47d1-b607-48576b07aec8 · outbound

This paper cites The random noise is per-pixel zero-mean, and has a linearly decaying variance, starting at 1 and ending at 0 by the end of the optimization procedure.

Towards Distributed Neural Architectures The random noise is per-pixel zero-mean, and has a linearly decaying variance, starting at 1 and ending at 0 by the end of the optimization procedure

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.382742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:20.256809Z digest=sha256:417346064743cafdb98560d22a55f22716b392b8d397028414c1ca88bb81a875

Observation bbb959d7-42ec-4192-bbe4-8f53e54d3e56 · outbound

This paper cites 3, we find that the patches following the same path in a randomly initialized model share much greater visual similarities.

Towards Distributed Neural Architectures 3, we find that the patches following the same path in a randomly initialized model share much greater visual similarities

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.216489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:20.374982Z digest=sha256:18ad0c7a86b67edef6111784c0cf9902effedc84be7da1c04d8e6a67e9256490

Observation 08fb4c44-9031-43fa-a905-51093eb1bc92 · outbound

This paper cites 9 is a zoomed-in version of those two figures.

Towards Distributed Neural Architectures 9 is a zoomed-in version of those two figures

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:21.112995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:20.461704Z digest=sha256:b080d973350286e3f66d3efc3d3886c47e297e72977bf310f67d0a548f68783b

Observation 88848d09-fc48-4bfe-9fbd-977fbd2abc37 · outbound

This paper cites Scaling Laws for Neural Language Models.

Towards Distributed Neural Architectures Scaling Laws for Neural Language Models

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.232811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.232811Z digest=sha256:0360a4aa7e0af3c4f95450e78134a9c04ca89beef468a650e7cec80d7f752d1c

Observation 5161ef2f-64d8-456b-9d59-5bfba7bd16c5 · outbound

This paper cites RACE: Large-scale ReAding comprehension dataset from examinations.

Towards Distributed Neural Architectures RACE: Large-scale ReAding comprehension dataset from examinations

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:14:22.229680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:14:19.291521Z digest=sha256:9541e96a4efb2bbb4795c74dbfa9d72a622e21841f83027616021e32d5788d0c

Observation e9f46559-0637-4498-a104-7320a4abcc83 · outbound

This paper cites LLM Pretraining with Continuous Concepts.

Towards Distributed Neural Architectures LLM Pretraining with Continuous Concepts

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.808871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.808871Z digest=sha256:ce00087c9350042b5af20f52f4fab8ccc62ae4776f121a0b564592fb299c51a3

Observation 99607442-268c-4eac-9c9c-65012a38d471 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Towards Distributed Neural Architectures Distilling the Knowledge in a Neural Network

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.192043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.192043Z digest=sha256:8b5f8e78360eab9fe05594a6922ee5d6fac46c72523d1c96f0f76fb63ca2a854

Observation 1d627122-77c7-40cc-9e44-a97146762004 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Towards Distributed Neural Architectures PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:14:19.516841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.516841Z digest=sha256:e4958c05032a8464590e9e3464704508695aa93a01aac6eabece96835927ff21

Observation b2f7d93b-bee0-48bd-bfb2-1abc51917bee · outbound

This paper cites Do language models use their depth efficiently? arXiv preprint arXiv:2505.13898,.

Towards Distributed Neural Architectures Do language models use their depth efficiently? arXiv preprint arXiv:2505.13898,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.721573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.721573Z digest=sha256:25117061f63ed8002215bfc2a59be0abb1b423ec71bf3b06df758dc329330bd7

Observation e1b911b4-b597-4b18-8395-38a785d1ba97 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Towards Distributed Neural Architectures Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.635699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.635699Z digest=sha256:10eb7a336891c28d6cc7fbb96d5059c3a2f1d3b34b84d463cfdfd76a09e3112f

Observation ef8483eb-53bc-4772-9926-b0885c75eb1e · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Towards Distributed Neural Architectures FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.547802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.547802Z digest=sha256:ebd9dcda72a7ea5f8c0541596128c2e9fb51868358493e1c28c51c18000ec735

Observation 9122c42c-cc6e-464e-84c7-8063b72564ae · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

Towards Distributed Neural Architectures LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.881708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.881708Z digest=sha256:f40c022ee0495089ea69e86c3bc636278c81929c1d59c53a4fc3fa9ab812c2a2

Observation 538acbe2-ac9c-413a-928f-44ade83f4014 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Distributed Neural Architectures The Llama 3 Herd of Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.055624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.055624Z digest=sha256:400813f457ce395ff0c548656e3e8582bcf75a5f84105330c94d924da8d41697

Observation 1e1d56b4-b56f-412e-8ea8-f24089f2cff0 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Towards Distributed Neural Architectures Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:18.461421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:18.461421Z digest=sha256:25e0ac4de92ffbb638099e57db3a8a0f1de323af0e4094b07e498a23a7e1d982

Observation a0a084e5-e6ef-4663-85d2-3565f95c0be8 · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Towards Distributed Neural Architectures The Unreasonable Ineffectiveness of the Deeper Layers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.114883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.114883Z digest=sha256:43f3b40ca3ed0110450f8a5c36756a1599198cc2da140dda68312f6ab0e16134

Observation 98349d3c-4cf6-4b1e-bfa9-44ccd5c843da · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Towards Distributed Neural Architectures What do Vision Transformers Learn? A Visual Exploration

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.008165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.008165Z digest=sha256:8007816e08ca46d9f6a3a3187683680293aaf238dd996aabebae08ca41983f7f

Pith citing papers

No inbound Pith citation observations are available.