Pith. sign in

Paper Citation Record · LEDGER

CoAtNet: Marrying Convolution and Attention for All Data Sizes

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2106.04803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.04803 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:19:33.862479Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.799655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e08dccc-3a69-42d3-b1e7-7942c5e5b7ba · inbound

MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer cites this paper.

MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:46:35.211861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T20:46:35.073600Z digest=sha256:2b2254ad265123f72eb8c01266b9d7b6b6954ac271be7ce50288315759fb5f34

Observation bd6b2b45-d640-41e6-972f-b6f4603c7867 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.485888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:d7018f1af0b7b6d5cdc0b522872580f69dae9b0df4a89e09ee217809bdb41979

Observation 07410519-75ec-4f48-99b6-0727dd33450b · inbound

DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection cites this paper.

DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T11:19:33.862479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:19:33.862479Z digest=sha256:af5ddab4c87461371075611e82bafb6ed272b988fb363d8e1487063391c50429

Observation 2c9c4fae-ac83-4d4e-8b8e-4d89912c42c4 · inbound

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification cites this paper.

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:45.974828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:45.974828Z digest=sha256:3ade6d6c26898d76505f1160c20834bc2b846764b7801fd6d12d2e0879a6b924

Observation 3c15361b-e1a1-4a4e-bef6-28156de9a622 · inbound

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents cites this paper.

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.801223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-05T11:32:36.356530Z digest=sha256:819a644e57fe8da2be70f570162fa600131abd1681bbf6dd0aa7c6456ecf9972

Observation 30f27239-2882-49e5-bec0-8f7fd7a4ed2c · inbound

Advancing Vision Transformer with Enhanced Spatial Priors cites this paper.

Advancing Vision Transformer with Enhanced Spatial Priors CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.664602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T05:22:21.264807Z digest=sha256:1dac638e2b058f73ace104ab8f88054bb7b2abdba14a0fa207fd909f11c59671