Pith. sign in

Paper Citation Record · LEDGER

Exploring Plain Vision Transformer Backbones for Object Detection

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2203.16527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.16527 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:12.363276Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T14:31:40.521481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6fd6883d-85e4-451d-8bb0-c0bd008ac1b1 · inbound

Adding Conditional Control to Text-to-Image Diffusion Models cites this paper.

Adding Conditional Control to Text-to-Image Diffusion Models Exploring Plain Vision Transformer Backbones for Object Detection

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:43:10.956509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:43:10.880338Z digest=sha256:1c17bfcf61c07e43af1ca2d31bafc24c3abbf9a71e55a59620749a4d83b917bd

Observation 6e598cf3-8252-4fa6-ba2f-fedc9c962aa8 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video Exploring Plain Vision Transformer Backbones for Object Detection

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:40:23.949043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:43f0adff84468e204e98ce7cad3dcabe22facd38c367c94f639fe2f04ce14d48

Observation def6d8b0-b2e9-49fd-9d5e-d0e419a94406 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Exploring Plain Vision Transformer Backbones for Object Detection

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.112794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:9c8f910105c410c0cbc7399135fe20fed07cd44f5d5ec1e23fa68d6f7e1eb496

Observation 597da9c4-4597-419d-8f5f-9c49c7be8822 · inbound

gen2seg: Generative Models Enable Generalizable Instance Segmentation cites this paper.

gen2seg: Generative Models Enable Generalizable Instance Segmentation Exploring Plain Vision Transformer Backbones for Object Detection

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.524150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T14:31:30.651144Z digest=sha256:40ac6f43caea063726a57d48625a04be27ea4c5f542011a51e640bba2ef5228e

Observation b3524013-2008-4470-84d9-193b96830b34 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Exploring Plain Vision Transformer Backbones for Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.363276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.363276Z digest=sha256:22a208318461397e4382ce1635d761d30871e0bc2cbcf5d2b9062244309c5c8b

Observation 8ab5abb8-4ca3-4a0d-ab6e-e39a49dfe3fa · inbound

Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings cites this paper.

Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings Exploring Plain Vision Transformer Backbones for Object Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:54.608439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:54.608439Z digest=sha256:2b14490da1142dda657580b6ca60be40c0f46e85532ea668c62064cd9598fcd0

Observation 9d5ec591-45c1-41cb-b6d8-244f741b30bb · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring Plain Vision Transformer Backbones for Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.678581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.678581Z digest=sha256:a716c5e5f6613e4c45f88f6894f3a45e2e6465a8a74f1829060633411c2b2dcd

Observation 22e4caec-79fe-4af0-9f69-fe861da24206 · inbound

SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images cites this paper.

SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images Exploring Plain Vision Transformer Backbones for Object Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:25.296451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:25.296451Z digest=sha256:48df622e81c09cccf692736abcbddd03ab511137f2f217cc49a0c57a7ca3d289

Observation d8ed63db-1e74-4e59-9763-5dfe625835d7 · inbound

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation cites this paper.

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation Exploring Plain Vision Transformer Backbones for Object Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:51:04.256800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T06:51:04.256800Z digest=sha256:78f923fc8804e55c737393004c61650ddf3746ad041f6d8f819f913187ef70eb

Observation 4d5017c6-564e-456b-ab00-699f500f4462 · inbound

FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening cites this paper.

FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening Exploring Plain Vision Transformer Backbones for Object Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:14:45.185998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:14:45.185998Z digest=sha256:c7e3abdafa747f626ac72647896de6e474f3a234456d7f04c4c18a28da819839