Pith. sign in

Paper Citation Record · LEDGER

Scaling Vision Transformers to 22 Billion Parameters

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2302.05442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.05442 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:11:41.496782Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

118
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23729b9b-4725-47c5-9564-2da4060c3202 · inbound

PaLM-E: An Embodied Multimodal Language Model cites this paper.

PaLM-E: An Embodied Multimodal Language Model Scaling Vision Transformers to 22 Billion Parameters

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.833939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:bc6484e00a825bd2b78874b2db05a2ac152820668f2232a19e0546ba5b3f47cf

Observation e95e3818-21d5-4680-8c7c-363a6e442c85 · inbound

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action cites this paper.

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Scaling Vision Transformers to 22 Billion Parameters

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:17:58.737145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T01:17:58.678036Z digest=sha256:b3216ea9d955e0d5fb063fe533c52c0a88d726a37f58562e6f0076a12fec505b

Observation 2688c09a-5d90-4442-b90c-d2cfa5d9aac2 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Scaling Vision Transformers to 22 Billion Parameters

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.596114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:1bb499470f4b8f5426c6eedf25ea9a3628c14b6d8d723a579ba9a27cce22df9b

Observation 29ea4e0e-312f-47b3-9f67-73fb0221328b · inbound

FuXi-$\alpha$: Scaling Recommendation Model with Feature Interaction Enhanced Transformer cites this paper.

FuXi-$\alpha$: Scaling Recommendation Model with Feature Interaction Enhanced Transformer Scaling Vision Transformers to 22 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T10:11:41.496782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:11:41.496782Z digest=sha256:58cc1e5034c750d2cb3a91879492647332b980443f2f98c71f6db17b98f03839

Observation 9573a299-a857-4fa5-b700-90716e570d7b · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.174869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.174869Z digest=sha256:5d46ff5060c1ff708e5155e4c8734b39ba45d0d5e3911b218ca5d2c82bb512d1

Observation 2133d3de-4b97-4476-9336-66851138504e · inbound

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models cites this paper.

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models Scaling Vision Transformers to 22 Billion Parameters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:44.160662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:44.160662Z digest=sha256:66199d897cf5ab8e3ead6b127825baf034bbe0bd6d31275a8643be67ba426bdc

Observation 7d02eae8-43dc-4400-8256-542870d7463b · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models Scaling Vision Transformers to 22 Billion Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.412772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.412772Z digest=sha256:7b0d88bcafce1ebda674fdc900921518d5a3728906cfcfe2174e6d5c0f52d73f

Observation 1a4ac500-b51a-4f0b-8103-a72bd24ed971 · inbound

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models cites this paper.

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models Scaling Vision Transformers to 22 Billion Parameters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:59.436591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:59.436591Z digest=sha256:6713a05908174b224319ee65f1eadcfc67cdac36372ad54862abd21830f8ba1a

Observation d38b0cd0-40aa-4a9d-a794-0a3d11f0e51d · inbound

Scalable Object Detection in the Car Interior With Vision Foundation Models cites this paper.

Scalable Object Detection in the Car Interior With Vision Foundation Models Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:11:50.831055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T21:07:31.589011Z digest=sha256:5f555bd939ddca698e05614d40f2571f641437958de790859a5604461c35557d

Observation adddb5ec-da02-467e-b030-836e6a73746c · inbound

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems cites this paper.

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Scaling Vision Transformers to 22 Billion Parameters

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:47:53.644339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:47:28.248540Z digest=sha256:1c18ed26c491a3559a366269ac5a3db5ed1061ab13255b7f491e8f2632010c00

Observation 91153d93-4a5c-4ad7-9ca3-b01c9c58f090 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Scaling Vision Transformers to 22 Billion Parameters

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:50:14.890183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:bf8c9436679aca75642f88e0b4e23ea4353b0affe90740e95d932614f559ed00

Observation 5aa45d76-8595-4ff7-b135-eef1d3dfef7b · inbound

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use cites this paper.

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.576201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:44:58.483183Z digest=sha256:c11dd970c250b4ba4b21989762a2d600af239db5f3a2b57044d2050aecf98107

Observation 4da0c83c-c463-475b-8c3a-4ccd33e04f18 · inbound

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use cites this paper.

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:03:46.390770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T21:02:51.889450Z digest=sha256:615e4566275789cc30e0660a5763c5893403071078da1a8cca43de64b56d5968

Observation 331735e9-43e3-47b6-941d-60a359b6b60a · inbound

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use cites this paper.

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:31:22.762383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:29:30.950904Z digest=sha256:60ce79b6544e209a2c54713f4648cc4f4dde6b0c8f657e50dcc607d67c02ed89

Observation 638b18c9-252e-4496-99a5-a2168512e572 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Scaling Vision Transformers to 22 Billion Parameters

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:51fd518a998a47adb70bececf44d7e0512880c5652288a88858f00a9c16424ac

Observation 927da818-1663-4882-bffc-704ad517ecdf · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Scaling Vision Transformers to 22 Billion Parameters

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:42.020874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:cfa78e83800dba447fcf3eed10c26256a4122a167fbdca903d429f9041c0d9cf

Observation 3e286a78-b870-4567-b118-a7e7fdb8cb3f · inbound

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation cites this paper.

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation Scaling Vision Transformers to 22 Billion Parameters

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.871113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:17:56.029187Z digest=sha256:acbc4f84836c8a615ed34585f6af5f117495ac3904bff6126367870d9b32bbdc

Observation 5f9f321a-8d8f-4e16-8130-897384d7fcbb · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Scaling Vision Transformers to 22 Billion Parameters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.473340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:37:02.350175Z digest=sha256:5ccf0209bc002fa87af667e6806f012560790b375ded9c130eff7bb84c8cc8ae

Observation d3d629c8-08dc-4f13-bf64-2fd1098ce959 · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Scaling Vision Transformers to 22 Billion Parameters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:16.375380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T00:34:21.224797Z digest=sha256:12b72a856879c469950ffeb9010925af9801ba1bec0e7ada5a2508adb2dd4769

Observation 81530ff2-576f-4bfc-a440-8cc5e3d12828 · inbound

Scaling Laws for Neural-Network Quantum States cites this paper.

Scaling Laws for Neural-Network Quantum States Scaling Vision Transformers to 22 Billion Parameters

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.071592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:30:23.893472Z digest=sha256:97aedc80a3864941817ff5a62b42094cccc207b33dc382e15d3359bb29a5e2bd

Observation 3a7d2d8c-ea92-4925-b702-774037e804a7 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders Scaling Vision Transformers to 22 Billion Parameters

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.697712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:d613e9c9c772023c35d2dfabe7529ad707b54898ce23209cc2c000e24a6374da

Observation 694d21cb-4f0c-46e9-aa13-177c648bdd48 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Vision Transformers to 22 Billion Parameters

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.807239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:8264db879b2eadbb508b789f958b179272cadfa548a8070006f6420e4bcff11c

Observation 3244b7a8-17a5-4b57-bf6b-2bc2a95d06ae · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Vision Transformers to 22 Billion Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:05.666991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:05.666991Z digest=sha256:98a99e91549bbfe4dd795a52124a66b635dd54ff7ab196f9fb8250111350579f

Observation fc1682b5-aa53-4aed-9792-b9fe167a72bb · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Scaling Vision Transformers to 22 Billion Parameters

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T06:13:51.944737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:13:51.944737Z digest=sha256:86a8f78c77e76917c2684eef929da88ffebe8a673e4e602feaf7766a158f8e11

Observation 6d2c0ac0-df7d-4e71-92ba-38b944d8b4a7 · inbound

Predict before you train: Scaling Laws for particle physics foundation models cites this paper.

Predict before you train: Scaling Laws for particle physics foundation models Scaling Vision Transformers to 22 Billion Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-30T23:56:38.595543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:56:38.595543Z digest=sha256:4d8544871f0ba4a7088907d7a68a4cbb6431ae735231b15be11b688c14cfe268

Observation a09046c1-03fd-451d-91a8-f87589e196ff · inbound

Opt.Gear Technical Report cites this paper.

Opt.Gear Technical Report Scaling Vision Transformers to 22 Billion Parameters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:38.653520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:38.653520Z digest=sha256:31db227b8fcf109d4c8562c52e875d134ce6d6fca3485a8d16cc4a6e6bc4ff54

Observation 698ecde2-46b1-494a-b9a0-6406481878e0 · inbound

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse cites this paper.

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse Scaling Vision Transformers to 22 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T15:20:15.447616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:20:15.447616Z digest=sha256:3c8fa608378163a5ec2c92fe8169ced805bc3fb13bb73e6b448598f381cdbaff