Pith. sign in

Paper Citation Record · LEDGER

Reducing Transformer Depth on Demand with Structured Dropout

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:1909.11556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.11556 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:17.834796Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

273
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1ad9f5c-7c94-4fca-b410-26e8483eec12 · inbound

PyTorch Distributed: Experiences on Accelerating Data Parallel Training cites this paper.

PyTorch Distributed: Experiences on Accelerating Data Parallel Training Reducing Transformer Depth on Demand with Structured Dropout

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:15:26.852762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T19:15:26.764929Z digest=sha256:f2839719d88d82e127a3f7f8b33b8b644b7a642653271fb1010242d0080c198a

Observation 686627a5-f0e5-4e24-a436-285b6556240f · inbound

Eliciting Latent Predictions from Transformers with the Tuned Lens cites this paper.

Eliciting Latent Predictions from Transformers with the Tuned Lens Reducing Transformer Depth on Demand with Structured Dropout

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:54:37.594176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T16:54:37.382049Z digest=sha256:bd0638e68a0f91c426965793b5f920fdc9443aa852008739ea2b2ddd485d8fe0

Observation 2a778857-1a0e-42dd-9422-7548c2e75db6 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Reducing Transformer Depth on Demand with Structured Dropout

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.141062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:b1b9c73e05d7ac9ae65466286a3a777c745ce9fbfee685954d2cda7d9d8dda70

Observation 2ac5c55d-a864-4dac-b941-0025254eda76 · inbound

AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer cites this paper.

AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer Reducing Transformer Depth on Demand with Structured Dropout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:17.834796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:17.834796Z digest=sha256:ba068f99156ca10778e1b51f2982aef0776fd24b5815df91e3e2ad2c46723fb8

Observation 211c76de-123e-4481-a77a-df405572e788 · inbound

Position: The Future of Bayesian Prediction Is Prior-Fitted cites this paper.

Position: The Future of Bayesian Prediction Is Prior-Fitted Reducing Transformer Depth on Demand with Structured Dropout

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:36.354788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:36.354788Z digest=sha256:8e027ea74e8fd0e7a37b3cb184e7b82c082a61fc51d5eb9f53bb9914c6a9ab62

Observation 3223e1a1-e09b-4363-b4b1-f661be8cbb40 · inbound

Learning to Skip the Middle Layers of Transformers cites this paper.

Learning to Skip the Middle Layers of Transformers Reducing Transformer Depth on Demand with Structured Dropout

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:53.679247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:40:53.679247Z digest=sha256:751f8fb4194e73f3675a916b423533ce1e03c4a1e75cfe2e16d1264f3df6a4e5

Observation 282d4541-9bd4-46a8-b938-768d71b1ce3b · inbound

Towards Universal & Efficient Model Compression via Exponential Torque Pruning cites this paper.

Towards Universal & Efficient Model Compression via Exponential Torque Pruning Reducing Transformer Depth on Demand with Structured Dropout

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:42.899383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:42.899383Z digest=sha256:4c9d356d932e40347d13f9754ee62362c2e4b72402227b210b55c06218d1fd83

Observation 54e80b2d-9a4e-45cd-b894-d3ffe7f781a1 · inbound

Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs cites this paper.

Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs Reducing Transformer Depth on Demand with Structured Dropout

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:19.194766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:19.194766Z digest=sha256:dbe1f05fd09828aa50fb12e3ef8ed1d27cec3e8147cdc440b8e390fc8f5580c1

Observation ec931f3a-5d9b-4e1a-b351-2bc27cd39f80 · inbound

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling cites this paper.

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling Reducing Transformer Depth on Demand with Structured Dropout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.447217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.447217Z digest=sha256:17e663382d6d110321138998a88aadc8d2667cc5280fc898e637045b544cc38c

Observation 3776073e-b025-4559-a80c-f2d61610c801 · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study Reducing Transformer Depth on Demand with Structured Dropout

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:44.726762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:44.726762Z digest=sha256:22a5133c671bd082ce5ee0a7e19d5e8f691aebc83d990613fed453d7d6d2c0e2

Observation b220cf03-cacc-40ac-99fe-602821af9ace · inbound

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code cites this paper.

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code Reducing Transformer Depth on Demand with Structured Dropout

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:01:55.897855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:01:42.228190Z digest=sha256:dda3b2625675ff44406ded1b9af04cefe804679316e2914715d36d2e6464990b

Observation e4d35bd6-cf40-4be1-9924-1ff35b2f3be8 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models Reducing Transformer Depth on Demand with Structured Dropout

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:36:56.313975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:34:38.247099Z digest=sha256:eb230dd950353349fc099e3081740265330f06ecf28c66f250874165b941e702

Observation 8b34d409-72f4-470e-8e2e-b23b693ea384 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models Reducing Transformer Depth on Demand with Structured Dropout

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T00:02:58.946950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:02:58.946950Z digest=sha256:49506018e22102894e99edb4846c6c9443d215dcae2547851530f248f148bfab

Observation 1beebdf1-25db-475a-9bc5-eb1303dad0e5 · inbound

Harnessing Input-Adaptive Inference for Efficient VLN cites this paper.

Harnessing Input-Adaptive Inference for Efficient VLN Reducing Transformer Depth on Demand with Structured Dropout

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:22.237618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:22.237618Z digest=sha256:03b3c4f0fd9360b60a1228ee31244501269d919cc482c67c4a504d347a288cf0

Observation 483f8e97-ee2b-4ef9-8ae3-c0b17994d552 · inbound

VISP: Volatility Informed Stochastic Projection for Adaptive Regularization cites this paper.

VISP: Volatility Informed Stochastic Projection for Adaptive Regularization Reducing Transformer Depth on Demand with Structured Dropout

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:37.008912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:37.008912Z digest=sha256:e2a906151f5365a131d9558112f9113672403dab2e30aab5f0251518a44b966d

Observation 3f4346a0-e7d6-4651-ba8f-7c985a2c2f7b · inbound

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures cites this paper.

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:38:07.556104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T18:35:23.814733Z digest=sha256:dba2f7c93a7d0e22e518584ea724a90b3b2912b1fc3e154f165c50c46143cd38

Observation d1ac1b49-3e8e-4b59-a63e-c429923851db · inbound

Depth Adaptive Efficient Visual Autoregressive Modeling cites this paper.

Depth Adaptive Efficient Visual Autoregressive Modeling Reducing Transformer Depth on Demand with Structured Dropout

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.625427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:22:55.035749Z digest=sha256:27cafc8ce8b38b93158bb271f020337297edc5c1b33e8bf80308d4561d2fe52b

Observation c01b4b4a-0906-42c7-af38-7334de1511bf · inbound

Language models recognize dropout and Gaussian noise applied to their activations cites this paper.

Language models recognize dropout and Gaussian noise applied to their activations Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:02.186788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:38:12.882914Z digest=sha256:b945a083094b222ee76f9cd52b3a6f9c1ba15b948c463af346ad45d749531412

Observation 7351aec5-c8c2-4627-b739-ae141d653acf · inbound

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing cites this paper.

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing Reducing Transformer Depth on Demand with Structured Dropout

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:58:12.041938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:56:48.015363Z digest=sha256:6090b3134b068627edb0f5a1116620b773389d95cc038012b0eb8268cbc02f9d

Observation 67ff6be4-e9a7-4148-b7d1-cbc5aefda9fb · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency Reducing Transformer Depth on Demand with Structured Dropout

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:14.586000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:5eebc9561a23b12dcfef52872f6159a35c228c2a18c0d55baf8d58db79d4e7ff

Observation bf2da7a7-a8d6-447c-a3d8-8ceccf243a33 · inbound

SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations cites this paper.

SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations Reducing Transformer Depth on Demand with Structured Dropout

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:29.267468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:08:12.351139Z digest=sha256:9b9b440f14393c0328e792328c7fe485d3d6f7a4b7b6ca9923a447973d20be5e

Observation c13c2308-d5ae-45b4-981d-07a7b3f47ed0 · inbound

Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility cites this paper.

Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:09.011862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:33:52.457620Z digest=sha256:0a38aa5adddac700dc406285302bee20116161ad3d0242ada71794aa313c903d

Observation e9c90031-af5b-441b-9322-bd19b77d0272 · inbound

Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs cites this paper.

Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs Reducing Transformer Depth on Demand with Structured Dropout

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:58:04.030648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:50:10.564922Z digest=sha256:39107f245432012ce98008a7b2eab25e440dbfde6b0bdea5a69bca05a22fca25

Observation 972ad1df-c680-4ad2-807f-6c3adf3cb646 · inbound

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility cites this paper.

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility Reducing Transformer Depth on Demand with Structured Dropout

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:39:47.975021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:35:09.705532Z digest=sha256:e5ad283a7471aaa346383734ae03d4406ce5b914dab83eaccff4942df5bf1ad2

Observation a9ae38b3-1ea4-455c-add4-905c249bdbf4 · inbound

Skip a Layer or Loop It? Learning Program-of-Layers in LLMs cites this paper.

Skip a Layer or Loop It? Learning Program-of-Layers in LLMs Reducing Transformer Depth on Demand with Structured Dropout

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:57.370073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:56:34.435152Z digest=sha256:1bd20e087fefb645d801b3c68b94d2db779f134019be341d336c2ae49c4b5105

Observation 7cada08f-0e51-4b6a-b246-8e3b26a7c03d · inbound

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation cites this paper.

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation Reducing Transformer Depth on Demand with Structured Dropout

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:41:03.364124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:32:09.784716Z digest=sha256:7e410b60317ddc5c23673546202d86cda0ddf28536af191edf1c92b4dc84c7fd

Observation 557c46b0-6da5-4ea7-a4ed-dee2dfdb2bb1 · inbound

Tapered Language Models cites this paper.

Tapered Language Models Reducing Transformer Depth on Demand with Structured Dropout

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.954130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:11:20.341634Z digest=sha256:c421d20843a99a0e686be5b47ade34103cd7b2260dc457d8833762f288bfa184

Observation 23ae4396-631b-4c90-a1cd-57c22b8282ba · inbound

Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping cites this paper.

Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping Reducing Transformer Depth on Demand with Structured Dropout

Reference 34

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T07:34:21.633951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:29:23.786653Z digest=sha256:d7e352aa0c974492bc644da2428bd2b0a9683b314d86dd3d85467984d3e2e720

Observation ee248a96-9780-4b07-b3fe-04d7b3b56d60 · inbound

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study cites this paper.

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study Reducing Transformer Depth on Demand with Structured Dropout

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:20:38.521564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:20:38.521564Z digest=sha256:c705f6bdab41902d86fd752fc4d5ff2da934595894ff2a63d55ccde7489e5b85