Pith. sign in

Paper Citation Record · LEDGER

Visual Textualization for Image Prompted Object Detection

As of 18 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.23785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23785 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.897852Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf10745e-c8e4-4e84-81dc-bf97e68189d9 · outbound

This paper cites Lung image database consor- tium: developing a resource for the medical imaging research community.

Visual Textualization for Image Prompted Object Detection Lung image database consor- tium: developing a resource for the medical imaging research community

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.347738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:02.947288Z digest=sha256:9f0b4a391a8baf99500c03dca9154da615a38677fb68c838db8a1cd8f4a32d15

Observation a8a31ae4-53dd-4930-936c-bc8d7851bbba · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Visual Textualization for Image Prompted Object Detection Exploring Visual Prompts for Adapting Large-Scale Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.071096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.071096Z digest=sha256:be9c49574e242dfe8e85bf9d794c5a805870129baa08c1a7031bb2774fe72cae

Observation 494e3f34-7810-4fca-94fa-730bd867c9cc · outbound

This paper cites Fs-detr: Few-shot detection transformer with prompting and without re-training.

Visual Textualization for Image Prompted Object Detection Fs-detr: Few-shot detection transformer with prompting and without re-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.332668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.162281Z digest=sha256:2279ef564fa7793bcbc95eb26e38f87a4e1a356cbe3656923dd02e3098ca7cf0

Observation c10b563c-4060-489f-b9b8-d1a0a9c4ea95 · outbound

This paper cites Apollo: Unified adapter and prompt learning for vision language models.

Visual Textualization for Image Prompted Object Detection Apollo: Unified adapter and prompt learning for vision language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.316861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.251458Z digest=sha256:0b2cad6ddcadcf250404287bc4c16ac5bbfd6d53e05074a20b61d8cc3dd7f1a4

Observation f4129d0f-d7d2-41ab-ab60-dc6c1fc394c2 · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022.

Visual Textualization for Image Prompted Object Detection Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.299990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.345641Z digest=sha256:be7ba8e4afaf3adc80a8cd396806a2c9d9b5aa9f66ae3e71f9658ef4ac9cfc4f

Observation b47ad70b-726a-4417-8822-148ba1e39462 · outbound

This paper cites s- adaptive decoupled prototype for few-shot object detection.

Visual Textualization for Image Prompted Object Detection s- adaptive decoupled prototype for few-shot object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.421737Z digest=sha256:506b237d7696a9b08d30083f642437154311d6449e9206ad30d7c7086f0ff012

Observation 02332cf4-9657-440c-b52a-6f5b0fca1064 · outbound

This paper cites Learning to prompt for open-vocabulary object detection with vision-language model.

Visual Textualization for Image Prompted Object Detection Learning to prompt for open-vocabulary object detection with vision-language model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.258469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.502861Z digest=sha256:692d51bb83ebec04d730d08ce992a49d0a8246f9c35739df2637bd318d6e5038

Observation fc7a5f35-5371-4a1f-8b5f-125bf38d7737 · outbound

This paper cites The Turking Test: Can Language Models Understand Instructions?.

Visual Textualization for Image Prompted Object Detection The Turking Test: Can Language Models Understand Instructions?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.644571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.644571Z digest=sha256:175f3ce216e52e36ed509da0d38a63ba5f6fa5f3200564ab404de84a822f36f6

Observation 5d4286b5-c630-43ea-91e4-2648e1275036 · outbound

This paper cites The pascal visual object classes (voc) challenge.

Visual Textualization for Image Prompted Object Detection The pascal visual object classes (voc) challenge

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.743469Z digest=sha256:1647c9cf0d2a177d522d9652cb26d48102ade6cf8f5718e63e021a6955d55626

Observation da991085-3e35-48f9-92af-c7c2e6ba7bef · outbound

This paper cites Few- shot object detection with attention-rpn and multi-relation detector.

Visual Textualization for Image Prompted Object Detection Few- shot object detection with attention-rpn and multi-relation detector

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.829699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.829699Z digest=sha256:60e6549840652b03442a39041416eccec6ca6db095c107bfaa2e23a03e8a88f6

Observation d5dc14cd-cb45-4eec-8789-e2f26e00a6e2 · outbound

This paper cites Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network.

Visual Textualization for Image Prompted Object Detection Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.205060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:03.904842Z digest=sha256:e07333d76db58040f6701d4bcafa94c1666a2e44bb2b36db46bb6b78ba681201

Observation 0c33a867-fc55-40ee-9428-acedea4ca5b9 · outbound

This paper cites Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images.

Visual Textualization for Image Prompted Object Detection Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.185000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.004247Z digest=sha256:ab64856baca1f3adbc0e3da8c6329312066784ccb86b4240ecac31ac6da75245

Observation 482ae711-fa47-4911-ab40-5dd7c88d34a5 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

Visual Textualization for Image Prompted Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.064646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.064646Z digest=sha256:ecb7618dc7143906b840fb72f8b24b69867fa6d4827e5344839b5ce6db330df8

Observation 966a79fa-1baa-48e3-bde5-e59e1b4c9e28 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Visual Textualization for Image Prompted Object Detection Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.138434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.138434Z digest=sha256:a9637bcb6488a162514b71d1ed9f5bc907b285a67a27918bbd8e8d141abe42df

Observation d248f254-2407-4685-a5a7-ce97d22d99b7 · outbound

This paper cites Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.166743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.250057Z digest=sha256:7d2c41dc8122987e9b1ea191ab3797f13b12c0beebab38953c923ab121b76bd0

Observation 2c1e0f48-be20-4d20-b198-254041fd36b0 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Visual Textualization for Image Prompted Object Detection Lvis: A dataset for large vocabulary instance segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.149088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.369673Z digest=sha256:6ec1ab8a57036f136f0d9e90a042f6d530f22cd8761c8c702db5b99c5a593b22

Observation 75988dd5-13ac-45a7-8487-7f4ad5410f54 · outbound

This paper cites Few-shot object detection with foundation models.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.135836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.489927Z digest=sha256:d60ab1e51fe1191316e0f515ecd9570614747c8c7361ef87078ad474958ff314

Observation e848ea3a-8a08-4488-8694-157cb8550f99 · outbound

This paper cites Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks.

Visual Textualization for Image Prompted Object Detection Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.119455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.567751Z digest=sha256:e2801b888ea779d6be2e99e7d74c31283be43dc822b8fa472f5df8d09b19faca

Observation 0eb56578-ed90-408c-b085-e6eee5a18312 · outbound

This paper cites Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting.

Visual Textualization for Image Prompted Object Detection Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.660592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.660592Z digest=sha256:adb893185aaf166516df71c0bcbc296e09313bcce80772ed71f641cb116b1cd7

Observation 0b083cbe-5052-4226-b5da-d752dfbf10b3 · outbound

This paper cites Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment.

Visual Textualization for Image Prompted Object Detection Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.100727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.727426Z digest=sha256:bae618a0f548d83d111d4a05707997472eb7a627909db38eb299bfd9d666a1c5

Observation 68627033-c1b9-4d89-9ce7-a487d2873efb · outbound

This paper cites Few-shot object detection with fully cross- transformer.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with fully cross- transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.084388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:04.978030Z digest=sha256:b27c72bdc5208eacb4018ef0ddea055fdf9c9852b8bafaed29068d33bf6706c3

Observation a3f2ac37-d0a0-4ebb-8c91-099a6c3c8f29 · outbound

This paper cites Few-shot object detection via variational feature aggregation.

Visual Textualization for Image Prompted Object Detection Few-shot object detection via variational feature aggregation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.947853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.120058Z digest=sha256:e0d8a88b3d9693a55d02c4e587565a2deb1a14f2c5a16df2ea81dfe36fbac1a1

Observation 02f4ff15-dfd7-4c74-ba03-bcff5b1ea601 · outbound

This paper cites Visual prompt tuning.

Visual Textualization for Image Prompted Object Detection Visual prompt tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.931188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.202072Z digest=sha256:9b39744cba2c9b2f1b3e8f7ad247f4cc4e7849ebe57b6532e4f2706da0d8a3ae

Observation d8e8e50f-2e4b-4c46-af8e-5097595a48d5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

Visual Textualization for Image Prompted Object Detection Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.307636Z digest=sha256:cbd42a1870747d4b9276621d4641702d50f83b7eff37600959679e68412239ab

Observation 8fb616bd-3570-45dc-8811-9f2bf38cb355 · outbound

This paper cites Maple: Multi- modal prompt learning.

Visual Textualization for Image Prompted Object Detection Maple: Multi- modal prompt learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.891918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.404038Z digest=sha256:dac8f7869c268cd6803e5dec4903b060444aa65ba13881e3c0370f62cf529a09

Observation 5b792927-2702-4896-8eec-86b34a410cca · outbound

This paper cites A dataset and a technique for generalized nuclear segmentation for computa- tional pathology.

Visual Textualization for Image Prompted Object Detection A dataset and a technique for generalized nuclear segmentation for computa- tional pathology

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.871760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.530990Z digest=sha256:24905a2d3535c8e8d9296163368dd6063ae3d4d49525293d92a47783b37a0b9e

Observation d737176b-afa7-4a34-9af7-dd51e16057d7 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

Visual Textualization for Image Prompted Object Detection F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.649074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.649074Z digest=sha256:ea9d81aee05ec5543e33b92f9764d842b53c14a75cd6f4fb433d07c873c025ea

Observation 7be32b2c-81c1-47a8-b5ac-15dbca57af3b · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

Visual Textualization for Image Prompted Object Detection Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.855850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.790125Z digest=sha256:7ea065b3bb3ee7312c4dafdbbecfc484a676e062f424c41d0461b056ba68d3a1

Observation 002258a5-4867-4e66-936b-65487eb0445c · outbound

This paper cites Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective.

Visual Textualization for Image Prompted Object Detection Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.841297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.875694Z digest=sha256:71ca2bfae342e8d4439d1a2d260f48c1997dc5f4c8eae753cee1b5f2186e2f93

Observation e4c5e956-6447-4a4a-9cc8-20fcf753f8c8 · outbound

This paper cites Grounded language- image pre-training.

Visual Textualization for Image Prompted Object Detection Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.820928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:05.969037Z digest=sha256:aa3bf64baebbd5d5651496025239a73bc452a94b430fbd1bb46bac803860e2fb

Observation 9edc8e46-7db7-4c90-9e95-dd493c8605c8 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Textualization for Image Prompted Object Detection Microsoft coco: Common objects in context

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.788346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.045211Z digest=sha256:cc79f6ef42fd60a8dd7470acdff784c55f74434840e9bd0f9f39724fe3941ecf

Observation b2801c4e-e958-4df3-9fca-ba4b3f05eab5 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Visual Textualization for Image Prompted Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.107791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.107791Z digest=sha256:b05002c0060d60848b40cb20db74b0b5500614388a2d26939bd909c28b1e25d7

Observation 01277d58-5a56-4099-b94e-764c9deffc45 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Visual Textualization for Image Prompted Object Detection Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.250780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.250780Z digest=sha256:de8a9b26e5b320deef6d2cd309bf4c7f7ead581b6c41331068187b56cd70fb93

Observation 739b0451-22e2-4f15-a6aa-48fd1e79a463 · outbound

This paper cites Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.756412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.324343Z digest=sha256:db765c9199009218ed610ac9642c3ebf08cb922b86c489bc1d8cd664d38bca7a

Observation c39a64d9-a64f-48d1-99da-e3a2469eb890 · outbound

This paper cites Image segmentation us- ing text and image prompts.

Visual Textualization for Image Prompted Object Detection Image segmentation us- ing text and image prompts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.737378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.388853Z digest=sha256:0fc946848771927f44a7f92be0c3974a50e9945f3cf7fcf7fe0df1bf785e9ff3

Observation f0c2cc35-7f82-4ae8-af0e-e8d608ccaf31 · outbound

This paper cites Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection.

Visual Textualization for Image Prompted Object Detection Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.722011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.440503Z digest=sha256:a8d3609ae6c8944cf4e488f11d09894beccf8acf98e482c9e19190ad7a037cba

Observation 5f89670f-611c-4991-91f4-bc0e8c45a54a · outbound

This paper cites Simple open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Simple open-vocabulary object detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.701594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.497243Z digest=sha256:beccddaa5cdc3d46acfd410d6a8b3c6cbc8d8264bfab2330f598254d46d8ef10

Observation e0af7dcd-d85c-45f1-b10d-0fdc4ff3770f · outbound

This paper cites Scal- ing open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scal- ing open-vocabulary object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.678121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.557963Z digest=sha256:eb3d31b3e80a4ad63c38f90dbf532273d6f0021ec123cf60da8a251f6fdc877d

Observation a907e535-54c4-4178-8f1b-7ead7169f35e · outbound

This paper cites Defrcn: Decoupled faster r-cnn for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Defrcn: Decoupled faster r-cnn for few-shot object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.659393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.624223Z digest=sha256:d1d7f29e6e159ea4ab311df7a001eaab6685a4919e54f26873b9bec0e9af4baa

Observation 60171d71-81d6-4b39-a81b-16cc0df5ae2a · outbound

This paper cites Language models are unsuper- vised multitask learners.

Visual Textualization for Image Prompted Object Detection Language models are unsuper- vised multitask learners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.638175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.665207Z digest=sha256:978684f0df8148baf1b1730bae925c45f7614b202780221c9e2a504d616e56fb

Observation 3de2fc5b-5b78-44e4-a450-e5c16d79a758 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Visual Textualization for Image Prompted Object Detection Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.611623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.744231Z digest=sha256:3fd6ee615892ad57104c3fe9781258d6964db65f93e7ef18d80fb1ea477060f7

Observation 6157ef5e-b882-49eb-a9d0-f5d86407f505 · outbound

This paper cites Adaptive multi-task learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Adaptive multi-task learning for few-shot object detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.588564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:06.869279Z digest=sha256:40300a37385ec061a04f87e81eb7f131da6789ad3edd24b2a7f3e14b30cb10df

Observation 9042d0fc-ffe6-4c22-afa7-39b41f8cacbf · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Visual Textualization for Image Prompted Object Detection Objects365: A large-scale, high-quality dataset for object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.980275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.980275Z digest=sha256:edbce4dd83b7dd8594b419138acb323c26562ac6ec2764a91c2ca5a40ea04a06

Observation 9b62f9e8-eb53-4952-bd87-25f58b822722 · outbound

This paper cites Few- shot adaptive faster r-cnn.

Visual Textualization for Image Prompted Object Detection Few- shot adaptive faster r-cnn

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.551444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.045297Z digest=sha256:887749b97fbd481b3f1878bbd5a61d0bb799ededd8545dad6e7d9ddb18f5a9cd

Observation 86f89577-7cb6-453c-810b-285054f4d9b3 · outbound

This paper cites Frustratingly Simple Few-Shot Object Detection.

Visual Textualization for Image Prompted Object Detection Frustratingly Simple Few-Shot Object Detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.148413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.148413Z digest=sha256:d79629c32e8b6538116ed86a6c37e0bfcf8e456f0a823e96d6415489faa26306

Observation 3e24ebcb-cddf-4047-90bd-acd227207b64 · outbound

This paper cites Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation.

Visual Textualization for Image Prompted Object Detection Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.530773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.207653Z digest=sha256:d27d2de97171011bdab6b4a95512eed9496bfe17308aa5bc3875e760aef8fe14

Observation dfa01c67-5c5f-4333-b46e-b41aa717ed20 · outbound

This paper cites Multi- scale positive sample refinement for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi- scale positive sample refinement for few-shot object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.516661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.267887Z digest=sha256:700699738a718ae396fa029695107f6c855c6c6c37218b3eb6439a1c3eaa97ad

Observation c7e19578-69fe-445c-abf4-20e2d9dc6474 · outbound

This paper cites Multi-faceted distillation of base-novel commonality for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi-faceted distillation of base-novel commonality for few-shot object detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.501311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.331054Z digest=sha256:2521bb6c14876dbd8c2f1952f38e120c054ce90f5b2c2fa83ceb55ba34d8439b

Observation 1767531d-f860-461e-b608-09d2122bdc64 · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Visual Textualization for Image Prompted Object Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.485161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.456827Z digest=sha256:a236df2bceaf187aecc7300160a7e8eaea4853e45b33ff8587dc40f20a447162

Observation 2feb1fc3-d790-4a12-aaf4-a92cecd2fceb · outbound

This paper cites Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.469861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.524136Z digest=sha256:96413ee161eebed3bffc08645bf6a81a2d4882777f23f790075f1a3613a3cbb3

Observation a057a633-5693-40a9-82c8-13bf30038f4f · outbound

This paper cites Multi-modal queried object detection in the wild.

Visual Textualization for Image Prompted Object Detection Multi-modal queried object detection in the wild

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.454659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.577668Z digest=sha256:23a7048d5c40821e51fdde5d0cdd5b1941cd940d58aedd40bdb03a497cb940ec

Observation 0c0f5449-7811-4090-ba43-158692c4fef9 · outbound

This paper cites DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations.

Visual Textualization for Image Prompted Object Detection DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.600258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.600258Z digest=sha256:c285474d621aedc046a8f7c36982c780f0f7ec6d9806195e3b2ffeda33686be7

Observation 73132bc0-6d32-4dd3-b3f0-56030f5c197d · outbound

This paper cites Meta r-cnn: Towards general solver for instance-level low-shot learning.

Visual Textualization for Image Prompted Object Detection Meta r-cnn: Towards general solver for instance-level low-shot learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.438524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.694527Z digest=sha256:756af6db562a8e67f73ad6dc7196c8fe5851787d0f484a249c69be2f201ea12e

Observation e0e62d53-abc9-42a2-90d8-3c4e5d8d3c68 · outbound

This paper cites Meta-detr: Image-level few-shot detection with inter-class correlation exploitation.

Visual Textualization for Image Prompted Object Detection Meta-detr: Image-level few-shot detection with inter-class correlation exploitation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.421162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:07.849351Z digest=sha256:52c55ee189c5bcc3d34b9dcf8c2058c2eb482a90058380c2752daf86033f2d67

Observation e6ad7f42-c487-4313-b2c1-295f968ec37e · outbound

This paper cites Detect Everything with Few Examples.

Visual Textualization for Image Prompted Object Detection Detect Everything with Few Examples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.017364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.017364Z digest=sha256:6a173eb68d730badaee03a78098318d5f2446e763b2552a1cd33a7753286df80

Observation 42ad3728-b973-43cf-af2f-2cb443f79a27 · outbound

This paper cites Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.399983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.184430Z digest=sha256:f1d91ad5b8a4eabfb13effae3e9e70904f66a2808656682d05d6b29c2cb9abf9

Observation a5d45fcf-aa5b-47a6-8f2a-2ad0baa3c275 · outbound

This paper cites Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.377603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.351578Z digest=sha256:2dd0b57eef8cde5846875f35924cb7f173980a4be12640aac1d83eec54bd15dc

Observation 08905e76-b956-41e3-930d-b50160e36131 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Visual Textualization for Image Prompted Object Detection Regionclip: Region-based language-image pretraining

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.355023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.516525Z digest=sha256:a26f9d27d4909085c497e92bd3c42c7fddcb265d8cae878508d09207eeb6aa96

Observation 5cd787c3-1d4a-4e53-82d9-7dfe83fd4d3c · outbound

This paper cites Conditional prompt learning for vision-language models.

Visual Textualization for Image Prompted Object Detection Conditional prompt learning for vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.336727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.683479Z digest=sha256:3d1263d857476002049dca567d8c38d91cfa2b1f4e146f5776a9b7719a738edd

Observation f5b47bc4-f299-46af-a8da-1e0ec522b2e6 · outbound

This paper cites Learning to prompt for vision-language models.

Visual Textualization for Image Prompted Object Detection Learning to prompt for vision-language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.322179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.801070Z digest=sha256:eb06ae695a2d67adbe5889f7086b694d53d8b95e5d114990c9abb30f5d6c98a0

Observation 0e5d6e64-5eae-4609-bc1c-f0d391d14a4d · outbound

This paper cites fully connected (fc) + ReLU.

Visual Textualization for Image Prompted Object Detection fully connected (fc) + ReLU

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.304934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.863444Z digest=sha256:68313e7053f381ae8d29468430958a6247ecc1bbdb2df2524aa660f6f8c76406

Observation 08b86e82-9e0c-4100-8457-b110e286e90a · outbound

This paper cites 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab.

Visual Textualization for Image Prompted Object Detection 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.283734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.869019Z digest=sha256:de73259151721b99b8f1dd7a1f6fc4ff56e072bb104960e126ee46175104a30a

Observation d24f22c8-c160-4757-98b8-9f389092278b · outbound

This paper cites an unresolved cited work.

Visual Textualization for Image Prompted Object Detection Unresolved cited work

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.267762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.873838Z digest=sha256:66cfdc5c785df788f8a1dda2bfe77e4d6634e98560dcc64e559e93901a65bd1d

Observation 089e77eb-2e4a-4a55-8010-4c8199a976db · outbound

This paper cites 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF.

Visual Textualization for Image Prompted Object Detection 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.247972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.880919Z digest=sha256:dfe6b0cfe334509828236007d1d78d066b920fc7f9669c112114baf19ce5c0de

Observation 254e521a-4c25-470c-9c16-6511e3952a71 · outbound

This paper cites BG blur" technique performs best. It high- lights the target object while preserving some background, unlike.

Visual Textualization for Image Prompted Object Detection BG blur" technique performs best. It high- lights the target object while preserving some background, unlike

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.227510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.889886Z digest=sha256:6b3830c932116d8a0e12b8ca7c69e7ceb1084ac825659965ae85a2d9e5707669

Observation a9ddc82e-4ae4-4035-b9c4-0517e3f24bc1 · outbound

This paper cites Base- line.

Visual Textualization for Image Prompted Object Detection Base- line

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.203949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:37:08.897852Z digest=sha256:3b7a7a4642f639de358677ccde54236a654fbc347926446dbe68955265c1207b

Pith citing papers

No inbound Pith citation observations are available.