Pith. sign in

Paper Citation Record · LEDGER

Simple Open-Vocabulary Object Detection with Vision Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2205.06230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.06230 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:44:58.769327Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.224246Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4390f54e-8a0f-49fa-b181-681292ccbac0 · inbound

Learning to Detect and Segment for Open Vocabulary Object Detection cites this paper.

Learning to Detect and Segment for Open Vocabulary Object Detection Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-24T10:26:08.081039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T10:24:52.653023Z digest=sha256:1927adc48e1c2021cdce1679381626b3685016400a0e0c241cf678363c52d5ed

Observation 02396885-aa78-4762-a020-785b413ac54b · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.537117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:469f431bcd86f309e79df3f5456073e9c6873f593addb79e974505fdba64ad95

Observation 16406840-911c-49fa-937c-63d6a09e9e8b · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:15:14.467897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:a28d6957d53ebe307be5362e1010abf0de5859c1cbb1c50342e922be0f38fea0

Observation 6d977316-3239-4058-9c0e-ba4d79384867 · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.432656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:69063a7573091aa9daa3e0ccceca9d24e79bb9cb201d98aa2c868c8ec56dfe96

Observation b8595cab-c45a-41f0-aefc-813690d76004 · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:18.076777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:db50fa1ae3513927277e8fbdfb72b0acd072c15b82fa040ca749e683e095edc6

Observation e6ecb145-e739-4e74-aab4-a03d0c1ea565 · inbound

Towards Wearable Interfaces for Robotic Caregiving cites this paper.

Towards Wearable Interfaces for Robotic Caregiving Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:44:58.769327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:44:58.769327Z digest=sha256:181671f92ac7dbea576b8f973ca7cf9dc5692e9cdd43cad60b9c32eaedb2d279

Observation ba198215-cfb9-4245-bf64-39ee441b0aff · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.785000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.785000Z digest=sha256:b6ccd02e891b423d5fee6ce6479dfbac2a3c6082505d64f9f3ca068f27341855

Observation c44236f6-515c-496e-babc-54a3a18326fc · inbound

Instance Segmentation of Scene Sketches Using Natural Image Priors cites this paper.

Instance Segmentation of Scene Sketches Using Natural Image Priors Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T20:56:19.716820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:56:19.716820Z digest=sha256:62f151bc4924d522ff163014e25a6eda3f3f92d1baf1f13bb15e02fc4482b7d9

Observation 1f3e67f5-1ce7-4381-899d-3a31c5f6a8b8 · inbound

Text-guided Generation of Efficient Personalized Inspection Plans cites this paper.

Text-guided Generation of Efficient Personalized Inspection Plans Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:52.175515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:52.175515Z digest=sha256:85242485c4d142ad459c95c2f927611c6902e2171fd06834ee8f048c7a2124ed

Observation 9970ec12-3997-4701-9dba-82053fae0d7e · inbound

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models cites this paper.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.316835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.316835Z digest=sha256:faaacbe813d05b4f5b1453b7e8d674638796d6e20aa384c4b5ee39eb384bee2a

Observation 62e13818-c8a8-4a18-bca9-4b17d6e68fa4 · inbound

(Almost) Free Modality Stitching of Foundation Models cites this paper.

(Almost) Free Modality Stitching of Foundation Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:23.658360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:47:23.658360Z digest=sha256:4be188803135f08029178896d88198b3313fb33404c4e649cddcd24fcb5313e4

Observation 1107d941-b2ec-4791-b94d-1aab7698439f · inbound

Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting cites this paper.

Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:31:48.438137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:31:48.438137Z digest=sha256:c490a3e133b5dfdb0bd34f8e6f2c8a5c65bba3b2653ad75fdfd6b126153e450f

Observation bf7986f6-73c8-488f-a28d-b700b392df51 · inbound

On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking cites this paper.

On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:54:40.668517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:54:40.668517Z digest=sha256:e7e3c2267e6bbedcff4fe196669e4eb2439dd8512fcf6c84931858dfdcad79ed

Observation 1fcdf4d8-a6a9-4990-af59-072fef2a708d · inbound

TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization cites this paper.

TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:46.082206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:46.082206Z digest=sha256:5e55cb27277eb27fb243bd9418073fae4e2c14fbc4417e8cf2033b4fb3dfdddd

Observation 11d4df83-795a-45c9-97b3-1e29731f0613 · inbound

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection cites this paper.

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.698939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:19:31.834944Z digest=sha256:38760e0b294e0fb3563a40d9ef6b97ee780741e62b5c84e8c43322b26494c726

Observation d643fb03-7aea-4702-a6fe-1494012bb72a · inbound

When Negation Is a Geometry Problem in Vision-Language Models cites this paper.

When Negation Is a Geometry Problem in Vision-Language Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:49:51.019505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T07:45:21.028636Z digest=sha256:28f7cec478b18759402fba3107dabae21fbaa66bbf805dca8d4a3a5cae3b94fa

Observation 678c51e9-fb8b-468b-b31e-feacfa06adc4 · inbound

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings cites this paper.

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:26:21.800655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:58:28.380188Z digest=sha256:d842e15fbef3e849e01ac6569a8b48301da9cc7c3eff8cbe5abc8ccea7693472

Observation b756cb49-806c-4ca7-b550-3bd9cb039bb5 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.521733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:afb11f98f24c383c464e241fcd857a137e510bfdeee3e6e63b3db94117bfe99e

Observation 7fb8cad5-9bb0-4a2a-a481-d87715022c2e · inbound

LV-OSD: Language-Vision-Complementary Open-Set Object Detection cites this paper.

LV-OSD: Language-Vision-Complementary Open-Set Object Detection Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.950038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:44:59.042817Z digest=sha256:bfdaee48922dc90d770761fc40d34f69bc61f385dc392b8d10ecf473936174fc

Observation 1db1c5c9-aee2-495b-9fc8-37ad18d75aa1 · inbound

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification cites this paper.

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:30.393879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:53:37.521755Z digest=sha256:5a7a872f978fb85179894b54a9eb6005725bb9cbda4e43dc423ea97fe4ec097b

Observation e625aa90-c184-4d32-8cd7-9810d78e45cf · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.225566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:33:25.373918Z digest=sha256:5f4261e4018c18d6eb09c686136c4bd85868916681ff0c2ddcc4b6aadf5671a4

Observation 7d817ee3-a392-40d6-979f-ea2f0edb6a97 · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:50.015989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:42:36.605869Z digest=sha256:e1b77166579923b90458b70d3a6f9561d689afba3869e26ee877ace959726bd6

Observation e9114509-68cf-441d-943b-a20f343a5371 · inbound

Sequential Planning via Anchored Robotic Keypoints cites this paper.

Sequential Planning via Anchored Robotic Keypoints Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T05:04:20.301670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T04:59:12.363425Z digest=sha256:b00c273bbdb2584aba5cd7d0e3991e4f8d4f2cb51bf5453f9228803713905328