Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:43:31.987951Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2411.18995.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:43:31.987951Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c7de7d48-ba92-4b66-8577-9f33223c599b · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Layer Normalization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7475193f-5158-467a-ace5-713f4e0b2708 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers MMDetection: Open MMLab Detection Toolbox and Benchmark
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10321d41-0599-4d38-80f6-c09ac586fcb8 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers CycleMLP: A MLP-like Architecture for Dense Prediction
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a28443-e4a1-4509-9530-e1271565b2c6 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Twins: Revisiting the design of spatial attention in vision transformers, 2021
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 38fee022-f3ce-46fc-9af3-cb39d6b99425 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d348127d-e7fc-4b8a-84af-5b9fa2414222 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Randaugment: Practical automated data augmen- tation with a reduced search space
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aaa21eac-edf0-464d-87df-3800442065ce · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Coatnet: Marrying convolution and attention for all data sizes
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2c4525cc-c989-4802-9513-0e2f03475a57 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Imagenet: A large-scale hierarchical image database
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9841c4be-b33b-4317-beb2-93eade38f88f · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d87a41-ce01-46e2-bcab-4827e813cb0d · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Multiscale Vision Transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd35d36f-b548-49eb-af81-00f068998dad · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Segnext: Rethinking convolutional attention design for semantic segmentation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e9170c0f-c16c-4480-8fb4-a70225783446 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Visual Attention Network
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b86f3a-ccc3-42c6-9343-2aaeb8f003ab · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Deep residual learning for image recognition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1b8d0802-2d77-4adb-93bf-6fad9df567e1 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mask r-cnn, 2018
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9a869c6b-0f24-41a6-b7d7-23d3e7b6be31 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Deep networks with stochastic depth
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bed20438-0074-4926-b6e5-a5a779e30582 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Arbitrary style transfer in real-time with adaptive instance normalization, 2017
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 29986528-00a0-4f41-93d0-31d13389fea1 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Batch renormalization: Towards reducing minibatch dependence in batch-normalized models, 2017
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e595f8bc-5db7-46cf-a31e-7b9e1280c0fe · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa7156f-bf36-4b4d-8030-57372fbb9e40 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Relational self-attention: What’s missing in attention for video understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f2835c07-1940-4809-b32c-329e8f100596 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Adam: A Method for Stochastic Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5722a313-f18f-455e-935c-0cd05c2e3472 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Panoptic feature pyramid networks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a9de85a0-e180-435f-98e7-8c8daa20cef3 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mvitv2: Improved multiscale vision transformers for classification and detection
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d7166ac2-e40f-425b-b744-885ac63c940e · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers AS-MLP: An Axial Shifted MLP Architecture for Vision
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca758f5-1059-4aae-9d56-02f41a15dc4a · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Microsoft coco: Common objects in context
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331bfe55-e673-4641-87c1-7b13291d4340 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal loss for dense object detection, 2018
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 50af3cc3-dbaa-4b87-ae15-d1ae37dfcb64 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5732a23a-d106-486c-afc5-4c97c07cb58a · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers A convnet for the 2020s
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11d4a1a3-db15-4655-913e-64721d343196 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Decoupled Weight Decay Regularization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3637f4e0-8dfe-4153-86c9-994b75885459 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers How do vision transformers work?, 2022
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 66a09598-7ecb-4a39-b708-bd37f83e4407 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Semantic image synthesis with spatially-adaptive nor- malization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d79b78b-fc4a-46b3-94eb-64f62fc40bc7 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Pytorch: An im- perative style, high-performance deep learning library
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f007f04f-8fa1-4923-a4a2-595b4b2d6a9c · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Acceleration of stochastic approximation by averaging
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ec5e319-f1ad-4b0e-8730-94dc35119152 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers What makes for good tokenizers in vision transformer? IEEE Transactions on Pattern Analysis and Machine Intelligence,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 85eb279b-be1e-4cd5-971b-1d3f423a896c · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Do vision trans- formers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a066e8f0-2382-4d92-9860-4fdfb750d206 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mobilenetv2: Inverted residuals and linear bottlenecks, 2019
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43ce6aa9-f109-4aee-8aa8-3c833791980f · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Grad-cam: Visual explanations from deep networks via gradient-based localization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 00b7bd86-6602-4cb8-b4ae-862074991779 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Powernorm: Rethinking batch normaliza- tion in transformers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b13d27e-365c-4453-a5ce-616fe2a4970d · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ae794c-f344-4b07-9303-94610bc313a6 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Evalnorm: Esti- mating batch normalization statistics for evaluation, 2019
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2dd01271-4d17-4d0b-a9c3-50a8174ddd71 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Rethinking the inception archi- tecture for computer vision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4e1c54-1194-4a2e-a0ac-16d053fd602c · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mlp-mixer: An all-mlp architecture for vision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64fd3689-1a74-44e6-9569-c61fbf8cb174 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Resmlp: Feedforward networks for image clas- sification with data-efficient training, 2021
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a2ea657f-b050-44e1-909f-ceca805a7b5d · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Training data-efficient image transformers & distillation through at- tention
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e3b1bac5-b829-415c-8c37-3cb2e6eb3a82 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Maxvit: Multi-axis vision transformer, 2022
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a68175a4-aef5-436b-a78d-6f4fad36d37e · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers In- stance normalization: The missing ingredient for fast styliza- tion, 2017
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 010ca6d3-2c3b-44b4-ae3d-80234adaec9c · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Attention is all you need
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ccb88f5-0263-4385-bd4d-493a491ee74b · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d642eae-4c45-4e06-ae32-ec959ce8413e · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Ri- former: Keep your vision backbone effective but removing token mixer
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 934a293f-64b3-41c7-9779-35f66fab10b0 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 58281994-c54f-402d-868d-32042a40b407 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers PVT v2: Improved baselines with pyramid vision transformer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 71d710a3-29b9-41b8-8a47-834aab5ee576 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Active Token Mixer
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aab8e676-b42c-4244-a709-ec76e82795de · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Con- vnext v2: Co-designing and scaling convnets with masked autoencoders
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b7f03e5f-06f8-4ac8-a329-0c203d9ea072 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Towards stabilizing batch statistics in backward propagation of batch normalization, 2020
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fcb1aa8c-c73d-4f60-a7e9-adfe1931fdef · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal Self-attention for Local-Global Interactions in Vision Transformers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e22678-5467-4c24-bf79-d83fd66aa814 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal modulation networks, 2022
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2b092c62-72a5-4398-a493-7d7e1cf01390 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Leveraging batch normalization for vision transformers
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 482c18dc-8972-4260-b143-2d85abeea79a · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers S2-mlp: Spatial-shift mlp architecture for vision
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d9b8806-7a7e-4cbd-bc12-361009d8e795 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Metaformer is actually what you need for vision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 377dfb90-3f52-4454-99f2-561b28605c99 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Metaformer baselines for vision, 2022
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 686a3d8f-9eaf-4655-bd66-a73e7f99c996 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Inceptionnext: When inception meets convnext, 2023
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb307237-2f1d-4ee6-822f-a8b74b031bf0 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Cutmix: Regu- larization strategy to train strong classifiers with localizable features
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd0aee0-f4f8-4da9-bff4-12b173bcb47f · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers mixup: Beyond Empirical Risk Minimization
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c849c23-b374-4e99-a96e-0a26c994c2d6 · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Random erasing data augmentation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation efa4ee6a-aa2f-4b95-886f-71b7b68a6f3a · outbound
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Semantic under- standing of scenes through the ade20k dataset
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.