Pith. sign in

Paper Citation Record · LEDGER

Visual Textualization for Image Prompted Object Detection

As of 21 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.23785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23785 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.897852Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf10745e-c8e4-4e84-81dc-bf97e68189d9 · outbound

This paper cites Lung image database consor- tium: developing a resource for the medical imaging research community.

Visual Textualization for Image Prompted Object Detection Lung image database consor- tium: developing a resource for the medical imaging research community

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.347738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:02.947288Z digest=sha256:bd6de3cb3fcb4fcf254fbcd19e5fe8d705a15517de16143a05ee0b7a44a8315d

Observation a8a31ae4-53dd-4930-936c-bc8d7851bbba · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Visual Textualization for Image Prompted Object Detection Exploring Visual Prompts for Adapting Large-Scale Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.071096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.071096Z digest=sha256:79243ff613792883fa43ed0bcfb8f13ee003ca70d74d222722a86377aa6e28a4

Observation 494e3f34-7810-4fca-94fa-730bd867c9cc · outbound

This paper cites Fs-detr: Few-shot detection transformer with prompting and without re-training.

Visual Textualization for Image Prompted Object Detection Fs-detr: Few-shot detection transformer with prompting and without re-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.332668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.162281Z digest=sha256:6de7f119c425439cbcd6de554a0c7c599b844cd714a6b72613a7083eb6d8e709

Observation c10b563c-4060-489f-b9b8-d1a0a9c4ea95 · outbound

This paper cites Apollo: Unified adapter and prompt learning for vision language models.

Visual Textualization for Image Prompted Object Detection Apollo: Unified adapter and prompt learning for vision language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.316861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.251458Z digest=sha256:ec9dbdee6019b7ce3e11a626cdc9ba3c116945e95c6ccbc8471d5d6bfda712c1

Observation f4129d0f-d7d2-41ab-ab60-dc6c1fc394c2 · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022.

Visual Textualization for Image Prompted Object Detection Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.299990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.345641Z digest=sha256:cb4dcef95bcc7967e95cc45fa5e667c287f80b87badb800da3f7a7e14f5c4e99

Observation b47ad70b-726a-4417-8822-148ba1e39462 · outbound

This paper cites s- adaptive decoupled prototype for few-shot object detection.

Visual Textualization for Image Prompted Object Detection s- adaptive decoupled prototype for few-shot object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.421737Z digest=sha256:72ed1bbdb049327847572b813b9a11fe74616ed7043527e5355cf0395ee15ad1

Observation 02332cf4-9657-440c-b52a-6f5b0fca1064 · outbound

This paper cites Learning to prompt for open-vocabulary object detection with vision-language model.

Visual Textualization for Image Prompted Object Detection Learning to prompt for open-vocabulary object detection with vision-language model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.258469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.502861Z digest=sha256:5913221ba5709db1392b1ff65219c0ac64759824eac33364717e229cd0307565

Observation fc7a5f35-5371-4a1f-8b5f-125bf38d7737 · outbound

This paper cites The Turking Test: Can Language Models Understand Instructions?.

Visual Textualization for Image Prompted Object Detection The Turking Test: Can Language Models Understand Instructions?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.644571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.644571Z digest=sha256:8648f1996774676694d9aae373bc848d078252cc51fc8ad0eef4973854e70e0a

Observation 5d4286b5-c630-43ea-91e4-2648e1275036 · outbound

This paper cites The pascal visual object classes (voc) challenge.

Visual Textualization for Image Prompted Object Detection The pascal visual object classes (voc) challenge

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.743469Z digest=sha256:ef5a640bdd15f0ea916995fde5b2e34fbe5fc0b6fe08bcbb01ba9dd5297401b2

Observation da991085-3e35-48f9-92af-c7c2e6ba7bef · outbound

This paper cites Few- shot object detection with attention-rpn and multi-relation detector.

Visual Textualization for Image Prompted Object Detection Few- shot object detection with attention-rpn and multi-relation detector

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.829699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.829699Z digest=sha256:60e6549840652b03442a39041416eccec6ca6db095c107bfaa2e23a03e8a88f6

Observation d5dc14cd-cb45-4eec-8789-e2f26e00a6e2 · outbound

This paper cites Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network.

Visual Textualization for Image Prompted Object Detection Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.205060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.904842Z digest=sha256:11c926a5383b7f329b31eb47c288dd362dd6ef522b153495d43716fa0141f70e

Observation 0c33a867-fc55-40ee-9428-acedea4ca5b9 · outbound

This paper cites Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images.

Visual Textualization for Image Prompted Object Detection Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.185000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.004247Z digest=sha256:28183b567520e64211708d90e67c0ecacafefeb1aa6321764aa0ee2d4af57461

Observation 482ae711-fa47-4911-ab40-5dd7c88d34a5 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

Visual Textualization for Image Prompted Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.064646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.064646Z digest=sha256:6810f876f083c530efa66c68aa72a38b9ae662f73339a509725ddef3059e9ec4

Observation 966a79fa-1baa-48e3-bde5-e59e1b4c9e28 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Visual Textualization for Image Prompted Object Detection Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.138434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.138434Z digest=sha256:a9637bcb6488a162514b71d1ed9f5bc907b285a67a27918bbd8e8d141abe42df

Observation d248f254-2407-4685-a5a7-ce97d22d99b7 · outbound

This paper cites Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.166743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.250057Z digest=sha256:305ee8e98b3cf2a7e276c8a1fe6236b1b72833ed3fb3feb5963f23d3beb1a4ae

Observation 2c1e0f48-be20-4d20-b198-254041fd36b0 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Visual Textualization for Image Prompted Object Detection Lvis: A dataset for large vocabulary instance segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.149088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.369673Z digest=sha256:7f1965642c35da76cd9aa2c7843918d4c9bea9ca3ff89e8f5a9ba3b47cfda58b

Observation 75988dd5-13ac-45a7-8487-7f4ad5410f54 · outbound

This paper cites Few-shot object detection with foundation models.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.135836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.489927Z digest=sha256:fe7253714f38d4adc143076f8c4e30e4f0e88905c68b541389456b31fd8a1e3f

Observation e848ea3a-8a08-4488-8694-157cb8550f99 · outbound

This paper cites Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks.

Visual Textualization for Image Prompted Object Detection Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.119455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.567751Z digest=sha256:6cbada52ae5aa46a23bef3abdc0508847c94964ae735c9ce2eb2be63f8f597f6

Observation 0eb56578-ed90-408c-b085-e6eee5a18312 · outbound

This paper cites Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting.

Visual Textualization for Image Prompted Object Detection Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.660592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.660592Z digest=sha256:adb893185aaf166516df71c0bcbc296e09313bcce80772ed71f641cb116b1cd7

Observation 0b083cbe-5052-4226-b5da-d752dfbf10b3 · outbound

This paper cites Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment.

Visual Textualization for Image Prompted Object Detection Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.100727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.727426Z digest=sha256:2ecd32362a9fc9586596ade4049786fdefc4c2aedd65340e51c6840b1c7d7d45

Observation 68627033-c1b9-4d89-9ce7-a487d2873efb · outbound

This paper cites Few-shot object detection with fully cross- transformer.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with fully cross- transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.084388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.978030Z digest=sha256:cbbe4084047513c5b21e85d2e3fb0646a74ae22c87ccc048824c902567407053

Observation a3f2ac37-d0a0-4ebb-8c91-099a6c3c8f29 · outbound

This paper cites Few-shot object detection via variational feature aggregation.

Visual Textualization for Image Prompted Object Detection Few-shot object detection via variational feature aggregation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.947853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.120058Z digest=sha256:1e4eb91a860e3a5a6a29ee4fc743bdcf4241261f8f58ac9903254eba32132e8b

Observation 02f4ff15-dfd7-4c74-ba03-bcff5b1ea601 · outbound

This paper cites Visual prompt tuning.

Visual Textualization for Image Prompted Object Detection Visual prompt tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.931188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.202072Z digest=sha256:a6581b51d46ec68ed742ffbb46abd9614bf95a8ea70f6f4f1bf7cb0ab2c342f2

Observation d8e8e50f-2e4b-4c46-af8e-5097595a48d5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

Visual Textualization for Image Prompted Object Detection Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.307636Z digest=sha256:51089d7763530d4f838bc05cfa9197bffd81e7a5b4601aec1fdbd00ccfca5541

Observation 8fb616bd-3570-45dc-8811-9f2bf38cb355 · outbound

This paper cites Maple: Multi- modal prompt learning.

Visual Textualization for Image Prompted Object Detection Maple: Multi- modal prompt learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.891918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.404038Z digest=sha256:fcff31d7431e83fdfd397bf120ecccff1a0eb7a7176abd0b94e3ec4dd3a79ec3

Observation 5b792927-2702-4896-8eec-86b34a410cca · outbound

This paper cites A dataset and a technique for generalized nuclear segmentation for computa- tional pathology.

Visual Textualization for Image Prompted Object Detection A dataset and a technique for generalized nuclear segmentation for computa- tional pathology

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.871760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.530990Z digest=sha256:6685584138d57cef0d5b633df8d07d3c0b1327b86322a466b62983cf2efd4f87

Observation d737176b-afa7-4a34-9af7-dd51e16057d7 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

Visual Textualization for Image Prompted Object Detection F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.649074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.649074Z digest=sha256:ea9d81aee05ec5543e33b92f9764d842b53c14a75cd6f4fb433d07c873c025ea

Observation 7be32b2c-81c1-47a8-b5ac-15dbca57af3b · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

Visual Textualization for Image Prompted Object Detection Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.855850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.790125Z digest=sha256:aa30ff852de9396196d574824b712d356ebab1f9226f81f4efc61701e94882dc

Observation 002258a5-4867-4e66-936b-65487eb0445c · outbound

This paper cites Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective.

Visual Textualization for Image Prompted Object Detection Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.841297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.875694Z digest=sha256:d4b757395c9718ff67c96e016a00eddd6a1bf5237024303c6d77f5fa5af9f187

Observation e4c5e956-6447-4a4a-9cc8-20fcf753f8c8 · outbound

This paper cites Grounded language- image pre-training.

Visual Textualization for Image Prompted Object Detection Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.820928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.969037Z digest=sha256:5da66351641eadd33c2565241be601a8bed5350c775bfc2d7d680594f38fc55d

Observation 9edc8e46-7db7-4c90-9e95-dd493c8605c8 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Textualization for Image Prompted Object Detection Microsoft coco: Common objects in context

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.788346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.045211Z digest=sha256:c4a71d3cc6d9ee3565f7e223ebd6478c73a0fab8add4074b27b26c4032f3f994

Observation b2801c4e-e958-4df3-9fca-ba4b3f05eab5 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Visual Textualization for Image Prompted Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.107791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.107791Z digest=sha256:c0fb4da380357f6e6b084ab265cc1773fc1b2b70b7054097bdedc81c4e6bb465

Observation 01277d58-5a56-4099-b94e-764c9deffc45 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Visual Textualization for Image Prompted Object Detection Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.250780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.250780Z digest=sha256:de8a9b26e5b320deef6d2cd309bf4c7f7ead581b6c41331068187b56cd70fb93

Observation 739b0451-22e2-4f15-a6aa-48fd1e79a463 · outbound

This paper cites Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.756412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.324343Z digest=sha256:5f49c62fae7d7fb6b8e9f148a38cb815f805cba8ed9e6c984602e64575d8ba53

Observation c39a64d9-a64f-48d1-99da-e3a2469eb890 · outbound

This paper cites Image segmentation us- ing text and image prompts.

Visual Textualization for Image Prompted Object Detection Image segmentation us- ing text and image prompts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.737378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.388853Z digest=sha256:a8edffbbc4f66cea04b1a0f80444dd1698fa23874049b3dd16f19a501ae01c49

Observation f0c2cc35-7f82-4ae8-af0e-e8d608ccaf31 · outbound

This paper cites Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection.

Visual Textualization for Image Prompted Object Detection Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.722011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.440503Z digest=sha256:dac160bf092e373f801242cb3a47a182c1c9f77542b35577ce4b448e77cb753a

Observation 5f89670f-611c-4991-91f4-bc0e8c45a54a · outbound

This paper cites Simple open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Simple open-vocabulary object detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.701594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.497243Z digest=sha256:0102ef1c6aafa6decaca07aaae7a1caa29a9e11a91c01cd7c7a1b8f48893709c

Observation e0af7dcd-d85c-45f1-b10d-0fdc4ff3770f · outbound

This paper cites Scal- ing open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scal- ing open-vocabulary object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.678121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.557963Z digest=sha256:4146b92b58ad8800a2cfc6df9f082a9d9090c31ee7d163718529ec9dc787e746

Observation a907e535-54c4-4178-8f1b-7ead7169f35e · outbound

This paper cites Defrcn: Decoupled faster r-cnn for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Defrcn: Decoupled faster r-cnn for few-shot object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.659393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.624223Z digest=sha256:c7e8eee4a49d03d3a830462c3b9939a0f36986413dac8bed3edc6919d1539c74

Observation 60171d71-81d6-4b39-a81b-16cc0df5ae2a · outbound

This paper cites Language models are unsuper- vised multitask learners.

Visual Textualization for Image Prompted Object Detection Language models are unsuper- vised multitask learners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.638175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.665207Z digest=sha256:b95e924a998d34440bb441db0455d99fb38f894a769875d4868e98ca74657f0e

Observation 3de2fc5b-5b78-44e4-a450-e5c16d79a758 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Visual Textualization for Image Prompted Object Detection Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.611623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.744231Z digest=sha256:f6c7f4fe9b20e267a207cdc05579223886f02b426ca08f80befda3f6e8e8dd5a

Observation 6157ef5e-b882-49eb-a9d0-f5d86407f505 · outbound

This paper cites Adaptive multi-task learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Adaptive multi-task learning for few-shot object detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.588564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.869279Z digest=sha256:35a2e8e57be1e541267a2de4c51d779cfb8620dcc57e4f5ddf9bfa2f81896287

Observation 9042d0fc-ffe6-4c22-afa7-39b41f8cacbf · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Visual Textualization for Image Prompted Object Detection Objects365: A large-scale, high-quality dataset for object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.980275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.980275Z digest=sha256:edbce4dd83b7dd8594b419138acb323c26562ac6ec2764a91c2ca5a40ea04a06

Observation 9b62f9e8-eb53-4952-bd87-25f58b822722 · outbound

This paper cites Few- shot adaptive faster r-cnn.

Visual Textualization for Image Prompted Object Detection Few- shot adaptive faster r-cnn

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.551444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.045297Z digest=sha256:56c9c366e24357a7ef8ba28b6c82762c7ca5529ec325f90825a2ff67e0446516

Observation 86f89577-7cb6-453c-810b-285054f4d9b3 · outbound

This paper cites Frustratingly Simple Few-Shot Object Detection.

Visual Textualization for Image Prompted Object Detection Frustratingly Simple Few-Shot Object Detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.148413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.148413Z digest=sha256:d79629c32e8b6538116ed86a6c37e0bfcf8e456f0a823e96d6415489faa26306

Observation 3e24ebcb-cddf-4047-90bd-acd227207b64 · outbound

This paper cites Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation.

Visual Textualization for Image Prompted Object Detection Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.530773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.207653Z digest=sha256:a46e8f9cdb9427861a4019cbb945f4fcf6eece84e6484f882f36e0a1b82a144c

Observation dfa01c67-5c5f-4333-b46e-b41aa717ed20 · outbound

This paper cites Multi- scale positive sample refinement for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi- scale positive sample refinement for few-shot object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.516661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.267887Z digest=sha256:9046386d7f0104094e83ded470f98079eaaac53742f48085698d8a97ac8afba5

Observation c7e19578-69fe-445c-abf4-20e2d9dc6474 · outbound

This paper cites Multi-faceted distillation of base-novel commonality for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi-faceted distillation of base-novel commonality for few-shot object detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.501311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.331054Z digest=sha256:0d7b8977265d0730270027885055ddc8addb137af01721101f6473d15e01cf02

Observation 1767531d-f860-461e-b608-09d2122bdc64 · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Visual Textualization for Image Prompted Object Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.485161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.456827Z digest=sha256:29aa4e153dfb7d6de857d6b904c71c774ac3844a29a8dbcdb30f6f4dc94cbd29

Observation 2feb1fc3-d790-4a12-aaf4-a92cecd2fceb · outbound

This paper cites Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.469861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.524136Z digest=sha256:9b369868b80fbc4d7a2eaa968cd1e2b1833494b8ccdbaa8a1a3320d9fced6230

Observation a057a633-5693-40a9-82c8-13bf30038f4f · outbound

This paper cites Multi-modal queried object detection in the wild.

Visual Textualization for Image Prompted Object Detection Multi-modal queried object detection in the wild

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.454659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.577668Z digest=sha256:dbe0b6aa26483ca762bd5e521cc1017a8245fc7f0fe5e2f260c695e27300cea4

Observation 0c0f5449-7811-4090-ba43-158692c4fef9 · outbound

This paper cites DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations.

Visual Textualization for Image Prompted Object Detection DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.600258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.600258Z digest=sha256:233ee38d425a52332bd5a3b5792ce4bc8ed6f86692757fe879f77c7de3f6b777

Observation 73132bc0-6d32-4dd3-b3f0-56030f5c197d · outbound

This paper cites Meta r-cnn: Towards general solver for instance-level low-shot learning.

Visual Textualization for Image Prompted Object Detection Meta r-cnn: Towards general solver for instance-level low-shot learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.438524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.694527Z digest=sha256:5e6647a8896810791aaf76328edf58208b9e8f4de1e0389843a7df93424b39cb

Observation e0e62d53-abc9-42a2-90d8-3c4e5d8d3c68 · outbound

This paper cites Meta-detr: Image-level few-shot detection with inter-class correlation exploitation.

Visual Textualization for Image Prompted Object Detection Meta-detr: Image-level few-shot detection with inter-class correlation exploitation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.421162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.849351Z digest=sha256:a0ef17be9163b20ebbe18f0df3538fdcc8bb5ee1310c256f76ecb1d2d53efb6a

Observation e6ad7f42-c487-4313-b2c1-295f968ec37e · outbound

This paper cites Detect Everything with Few Examples.

Visual Textualization for Image Prompted Object Detection Detect Everything with Few Examples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.017364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.017364Z digest=sha256:2387d20fb430d77739a9b7413151404b072376dfaf5787cc467990b23ddf8e86

Observation 42ad3728-b973-43cf-af2f-2cb443f79a27 · outbound

This paper cites Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.399983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.184430Z digest=sha256:26909436f870a028b9e0ca4e44d0d5ab38a2043c185d5591889d01720794aaf2

Observation a5d45fcf-aa5b-47a6-8f2a-2ad0baa3c275 · outbound

This paper cites Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.377603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.351578Z digest=sha256:2595edf373e24f2dd912019622639f7f58f8d9050f18c4f2216e05f56547119f

Observation 08905e76-b956-41e3-930d-b50160e36131 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Visual Textualization for Image Prompted Object Detection Regionclip: Region-based language-image pretraining

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.355023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.516525Z digest=sha256:a8d81244e2012acb3c3825b7df6796517f3857cfd2a215cf9bb3148302de639f

Observation 5cd787c3-1d4a-4e53-82d9-7dfe83fd4d3c · outbound

This paper cites Conditional prompt learning for vision-language models.

Visual Textualization for Image Prompted Object Detection Conditional prompt learning for vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.336727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.683479Z digest=sha256:210dd03412990f4727b767ef9ae55ae22f4eab47fea578c727873e51f5fbd5c2

Observation f5b47bc4-f299-46af-a8da-1e0ec522b2e6 · outbound

This paper cites Learning to prompt for vision-language models.

Visual Textualization for Image Prompted Object Detection Learning to prompt for vision-language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.322179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.801070Z digest=sha256:45e4ce792dc0fc6c86315bbdea7591bb4a3a6dd0ee248caa9c5a7c76b6512bb4

Observation 0e5d6e64-5eae-4609-bc1c-f0d391d14a4d · outbound

This paper cites fully connected (fc) + ReLU.

Visual Textualization for Image Prompted Object Detection fully connected (fc) + ReLU

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.304934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.863444Z digest=sha256:d0006c2500bf2df12e9c7d0a0993b6321ecafd874c69b629607071c8f37f43ed

Observation 08b86e82-9e0c-4100-8457-b110e286e90a · outbound

This paper cites 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab.

Visual Textualization for Image Prompted Object Detection 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.283734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.869019Z digest=sha256:70ccf7c8b5c4f845ea0f144ab9db31004403a65444c9af04ccb13dddda328a7d

Observation d24f22c8-c160-4757-98b8-9f389092278b · outbound

This paper cites an unresolved cited work.

Visual Textualization for Image Prompted Object Detection Unresolved cited work

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.267762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.873838Z digest=sha256:c54d47bc7e77c47812a33eef32a0680fe4e4fffa47e8f4f09d3927c7308cb308

Observation 089e77eb-2e4a-4a55-8010-4c8199a976db · outbound

This paper cites 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF.

Visual Textualization for Image Prompted Object Detection 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.247972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.880919Z digest=sha256:e2890fce13336296a60ca59613fefacd94856b45138c5b575efc24cc29a87551

Observation 254e521a-4c25-470c-9c16-6511e3952a71 · outbound

This paper cites BG blur" technique performs best. It high- lights the target object while preserving some background, unlike.

Visual Textualization for Image Prompted Object Detection BG blur" technique performs best. It high- lights the target object while preserving some background, unlike

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.227510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.889886Z digest=sha256:a69b41e5ce7def883984bb549f7acc8e288d7566f32016c6f9e5145ba29e355d

Observation a9ddc82e-4ae4-4035-b9c4-0517e3f24bc1 · outbound

This paper cites Base- line.

Visual Textualization for Image Prompted Object Detection Base- line

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.203949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.897852Z digest=sha256:8481be6372830998a74759e63c8e64cc0429fd1d8f0569eafa79b3959ae59396

Pith citing papers

No inbound Pith citation observations are available.