Pith. sign in

Paper Citation Record · LEDGER

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 6 inbound Pith citation observations for arXiv:2501.18954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18954 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:55:23.361176Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:49:13.593500Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:46:33.282729Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1573677-fb1b-4b81-b5fe-61fdc1448642 · outbound

This paper cites Nltk: the natural language toolkit.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Nltk: the natural language toolkit

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.933368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.151880Z digest=sha256:38913aa776b57e2a1d25c55de48da9754043588261e26f72bfa806d14288d5bb

Observation 47ec76e3-c514-4b0b-b9ee-a76b070087c2 · outbound

This paper cites End-to- end object detection with transformers.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models End-to- end object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.922207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.158157Z digest=sha256:63822867769edf816d4e7daa1e34e77916179c5ec24259819cee53d4b5b78d6d

Observation ffaa8454-edde-48f0-8aeb-cbebda13f315 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.161762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.161762Z digest=sha256:906f332f6951fe3ab1bb0da168ce73e4c2e23b7adc317a3c4298f9fde8b79924

Observation f7a62b63-cfb7-4405-af9c-40b0302fd257 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Sharegpt4v: Improving large multi-modal models with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.912272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.168129Z digest=sha256:1cf291c080b2a7a903022efb5c10416bb4c5b1e61f661533da6d4cc6a552d832

Observation 84399f18-64fb-43a7-ab53-df1d0b9172fb · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.171583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.171583Z digest=sha256:8d5025a5db8f133d675ae93a0a9891d293d56cef5d6af64ac32e9c0e1053628d

Observation 68873fa9-8add-4bbe-86b5-b98e519d44b3 · outbound

This paper cites Yolo-world: Real-time open- vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Yolo-world: Real-time open- vocabulary object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.898006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.175076Z digest=sha256:ba4622cd4902b74834564bb63466e4787424b1905d161c2e93ad0277e8dd193d

Observation 6e36856d-eaae-4521-94f2-71467bd0f696 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.178553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.178553Z digest=sha256:17240a73cd5ffbfef97c4d509582fab971a58845b7b29025c96fa86da71d059d

Observation 901397cd-38b0-4712-8301-fbbd736bb35e · outbound

This paper cites Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.181974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.181974Z digest=sha256:4aa72178b333ded41e08aab722b4c88e1ab0e15c401c85765eaed261dc68c464

Observation 44edde5d-1910-4bd9-818a-bb26ab22c5b4 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.185433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.185433Z digest=sha256:df93912595af40d8f838da375137e89442ee3de8733635b86c040742c31ba6a9

Observation e407b773-545a-425d-9a2a-bfe6decd759b · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Coarse-to-fine vision-language pre-training with fusion in the backbone

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.877896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.188567Z digest=sha256:93f1435730dd634f4b65b22070c481650bbe289278054326f6fbafe44c7686c9

Observation 175448ee-958f-4fe1-a951-929a89ed7b21 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.191659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.191659Z digest=sha256:c0a66161cd4229b9808959bf323a563d8867962b20f2442d48a669eb14ecb4b5

Observation da140a42-570e-4172-b2cc-e20ac0e05163 · outbound

This paper cites Asag: Building strong one-decoder-layer sparse detectors via adaptive sparse anchor generation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Asag: Building strong one-decoder-layer sparse detectors via adaptive sparse anchor generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.869046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.195329Z digest=sha256:abc6f290a44bb4ed3d62d39175b43fc78f751e4fab39a3eca1524dbc322edc00

Observation d8d29c7b-f8bd-4807-af7b-1f1079d04e04 · outbound

This paper cites Frozen-detr: Enhancing detr with image understanding from frozen foundation models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Frozen-detr: Enhancing detr with image understanding from frozen foundation models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.198600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.198600Z digest=sha256:f0939e54b1031666cc4fe0f4c3b2a6aa17b9247d7deb65670e0d8379bbb50f3d

Observation e2b73a3e-4a5c-4cb3-9251-8fc272f08d66 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Open-vocabulary object detection via vision and language knowledge distillation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.855042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.201596Z digest=sha256:c1c74e796e511325de8c5f27823cee33be3b9641b0db2867fef1b29520e9ab67

Observation 71a56cef-51ce-4b25-8fbd-b500139973c8 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Lvis: A dataset for large vocabulary instance segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.204637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.204637Z digest=sha256:456ac358eed5f8594406b3dc5cdb39f907a82eca3a1b0360089a7c3b6bc63197

Observation d42e2777-51d9-42b8-b4ad-eb92e7620146 · outbound

This paper cites Mask r-cnn.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mask r-cnn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.840719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.207869Z digest=sha256:ed271f259e7a1654354b48fbe423737e520415fa65a0b86d28a83bd2eb0ea1a9

Observation 0b63dcd3-1a56-409b-b130-18fe2204f11c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.830403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.210819Z digest=sha256:27e47f0714b48bf8630b1e5083aca260257e30e7227bb1bb8ac92789599f78d4

Observation d75afc4b-f51c-4786-af1e-61c751d28bff · outbound

This paper cites GPT-4o System Card.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.213818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.213818Z digest=sha256:fee830ed7e3b83f30a16b303ca79387666d7adff4c8345d0c3684ba49a01ae1e

Observation 928aa28c-d9e7-4b7e-9bd6-83bac3d42c6b · outbound

This paper cites T-rex2: Towards generic object detec- tion via text-visual prompt synergy.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models T-rex2: Towards generic object detec- tion via text-visual prompt synergy

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.820218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.217121Z digest=sha256:4b949299810a2ccc04a68eaeb262e0b7f8407abaa1f744e2774c2fa77e592e1b

Observation f7fff0fe-110f-4403-980e-52c9497ababb · outbound

This paper cites Mdetr- modulated detection for end-to-end multi-modal understand- ing.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mdetr- modulated detection for end-to-end multi-modal understand- ing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.809160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.220120Z digest=sha256:1a900cdde97bdab160d72515752f49d1c025304498811c9e175affc30f7b9f90

Observation 22f02348-d893-4d0b-9c76-0394a80040a3 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.798966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.222588Z digest=sha256:a824f5ab00d7cd1faad2f061ccbeca5fdbb67f6b3cdbaf3d8c6c612c66fcd045

Observation 36ab6bc2-80a5-4269-82c7-8b948a420e03 · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.225150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.225150Z digest=sha256:5812f36d7c2a6754cb4feaa9a05f9a67f04577cb9c4e56b55c7a80fc96354245

Observation 76bde895-cbad-43ba-996f-a18531bfe81c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.227783Z digest=sha256:da0f9305d20368b5382452fe139a855afc25f0983722df90e51ddb264e55d429

Observation 1178452f-d5c1-46a8-af4b-0188eb0c6efb · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.783598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.230754Z digest=sha256:a15c9eb8082a5a4a2a2630396c74345e81afae8999baf2663e31653e458eec1c

Observation b405fdf2-d4e0-4aac-adfd-102410a18cda · outbound

This paper cites Desco: Learning object recognition with rich language de- scriptions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Desco: Learning object recognition with rich language de- scriptions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.774357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.233206Z digest=sha256:7f4b1c7872ecd4f37ddeb5cf69050fc478b567d8920b5d12ba7bbd3bf03dc5e6

Observation 0778c048-4d9f-4952-b70b-342ad58ac51b · outbound

This paper cites Distilling detr with visual-linguistic knowledge for open-vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Distilling detr with visual-linguistic knowledge for open-vocabulary object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.235667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.235667Z digest=sha256:d603fb9a490c45a4f69e3ac8fcf1c5217368434d93ea0b12fd6580b53449406b

Observation 8d840ada-62b2-4d70-bf8e-129906e71278 · outbound

This paper cites Grounded language-image pre-training.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Grounded language-image pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.760243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.238489Z digest=sha256:a978444da0ad3f06cc836228f743297690e282aa19c7dadeed7a8deae5a9070d

Observation e0d13269-2714-4591-9ff8-b2fe067bdcd5 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Evaluating object hallucination in large vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.751470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.241190Z digest=sha256:4fe5b110d6fe8932132e2e6abc00893cff668259af17ba0fc0e8cea4a6527c71

Observation 9323eaa6-52d3-4fb5-a424-9c20843a5806 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.742548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.244124Z digest=sha256:5316214749febb400d7dd88e6806c8042775c316cc6e96f93366038598d59604

Observation b73b9807-e4e7-4d0d-afe8-9b52010fdc72 · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Generative region-language pretraining for open-ended object detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.733418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.247039Z digest=sha256:878ac40cbcc5a2d3aa7eda3fadd78af8306e08fcb620d945bbfb3a20b58bc1b4

Observation 6fe0c555-14ec-4412-900b-f6aebbac2119 · outbound

This paper cites Microsoft coco: Common objects in context.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Microsoft coco: Common objects in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.724447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.250070Z digest=sha256:5188f4002681364bffb28a7c6c8384523e81b33283a2f4fc333c0c6742852fe7

Observation 525b4851-844e-4af1-ab32-5fd0ae659897 · outbound

This paper cites Focal loss for dense object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Focal loss for dense object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.715499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.253093Z digest=sha256:3420eea888e9b65458cf4b922aac90caf37d6e40d7d02831800b6fd8e57c24cc

Observation 40bb858d-7e45-4f0d-9c33-322bdeddbc64 · outbound

This paper cites Gres: Gener- alized referring expression segmentation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Gres: Gener- alized referring expression segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.706794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.256227Z digest=sha256:d6898881cdf70ddbb14c710ac7a49448ded5b745718d16402c5f02265db50f42

Observation 263e4df8-4a1a-49c4-8417-fdfb524b940f · outbound

This paper cites Visual instruction tuning.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.698025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.259327Z digest=sha256:bbcb1a978ce277fedcc21221c9a787b2d4c301bafda0c3cf7f06c588b6dae344

Observation c0044068-c4d7-4333-92a1-f099cce853d3 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Improved baselines with visual instruction tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.262296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.262296Z digest=sha256:db1d9dc4e23f60b0456b28ffcf0f8045b1ac04e87f5cb9d734ede3053819ab11

Observation 2aab2b4b-8dc7-4437-93e1-880f1eea506a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.684444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.265451Z digest=sha256:a59b2edb9a55a6e35f4e656f20c970dd7e611963b3ec2a3061bab815eb3fd7eb

Observation d17e59ce-34b0-4672-953d-50571de4e5a7 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Swin transformer: Hierarchical vision transformer using shifted windows

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.675415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.268477Z digest=sha256:d300ac8f673cc7fef721b5d365dbf01e052aac1be3e202dd4fb5e9b67fdc637a

Observation 8682bee4-9018-49de-9b74-103dff20e1c1 · outbound

This paper cites Capdet: Unifying dense captioning and open-world detec- tion pretraining.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Capdet: Unifying dense captioning and open-world detec- tion pretraining

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.666285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.271366Z digest=sha256:6fc69aa5da8e5da5f9ea8c8906cb208f93a1992d3aec6608c9e96ff392f49fd3

Observation da99e344-8dd1-4402-94a4-0fa5a80f38b3 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Generation and comprehension of unambiguous object descriptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.657526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.274310Z digest=sha256:07b11f4638b1f2e65914f210f43271558f65117c0dd4eb0a086335c2cb61299b

Observation 1fd409d3-7413-4b59-8c7d-bf22e5ec03c1 · outbound

This paper cites Coco-o: A benchmark for object detectors under natural distribution shifts.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Coco-o: A benchmark for object detectors under natural distribution shifts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.648586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.277500Z digest=sha256:dcb6c64474b101fb503145282b9ba76962e196baf0bd4b3de9be9b431ff204a3

Observation 3bae787a-db2c-4633-98d6-3d71437859ce · outbound

This paper cites Scaling open-vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Scaling open-vocabulary object detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.639857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.280494Z digest=sha256:37de12ab2f1f105dc2fa5435447ab7e4b8a2584e03d33821a69781bc7a0e8e36

Observation 5b43bc4b-c505-4adf-913a-2cc5e42c6bd7 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.630943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.283447Z digest=sha256:dc33ec24af6d932809c55196921f4fd48f5b2c48a859a50380ce839bd5213543

Observation 9b1f626b-f0e1-49f3-8c6b-5cdc6950d1e0 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.286704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.286704Z digest=sha256:f154801e2f23409a2a21b2c218596d43bbfbcb2f5333039c11753bcea60f4512

Observation c459f17a-db0b-4374-b117-cae7cdcc54d4 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Objects365: A large-scale, high-quality dataset for object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.289705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.289705Z digest=sha256:bed1688ac399afd4a64273bd68823211926d2ea1911f9e81a1d12d0198003f0c

Observation a837e111-9987-4453-b261-63cc72631c9b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.292515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.292515Z digest=sha256:26aad28c7bf4122784b5b570e3d92ea40a1c96bdd826f5e8585e955a4ef580e6

Observation 685e67aa-bf45-4915-bfc3-5aa41164105c · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.295553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.295553Z digest=sha256:b0d5209a5d744cffde9a2ab5bf2be6347ab77fc119b232d2487f4f10417db0a4

Observation fab3f358-ac94-47e6-90ac-0b5a9222a54b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.298981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.298981Z digest=sha256:18d84bd174b54b98f932068bd28990d2a616a613ef366011e44beaa4f0aac6df

Observation 9c2b995e-ef0a-4846-ad97-705907b0062e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.302192Z digest=sha256:11022c64b8a6fb7869db70ee8196c8fe49860e9a918924b3a45738257d02cdd8

Observation 6ceadd55-2371-4dba-803b-7276702c7e60 · outbound

This paper cites OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.305127Z digest=sha256:bef52e9b003daaeff0fd1334762c6f4ef78e7846390fdda958aacce24ef906fb

Observation 3fe787cf-ccf6-4697-aa3f-ba1ea02fbf10 · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models V3det: Vast vocabulary visual detection dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.607473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.308354Z digest=sha256:b9dc12216d9424c1e9ff1d56c9ca3cf194d4a77662c0afbaf1c47d8f0bb45902

Observation 58a6c0de-f752-4992-b1d8-a769f3b7ce75 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.311214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.311214Z digest=sha256:35b3a6f5051a064abceccd93df5bcd0eae958e38786da4c49c62612b485119a3

Observation 8635b23c-e02a-49cd-8d03-bcc3efcfcb22 · outbound

This paper cites The all-seeing project v2: To- wards general relation comprehension of the open world.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models The all-seeing project v2: To- wards general relation comprehension of the open world

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.598662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.314542Z digest=sha256:e3938b405e755865c9389de4af668b76f42ffba261dfdfeb9e2e9688b09d5252

Observation 30fe3dfc-7868-4563-b9ae-62493519089c · outbound

This paper cites Grit: A gen- erative region-to-text transformer for object understanding.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Grit: A gen- erative region-to-text transformer for object understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.589693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.317329Z digest=sha256:34b5f90326bd568a20b9cb362cd6c11385da8b85ba602723569424433c7c2877

Observation 379c936c-c3ec-4e99-a626-0661e3828ddd · outbound

This paper cites Aligning bag of regions for open- vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Aligning bag of regions for open- vocabulary object detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.321030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.321030Z digest=sha256:7cead04e4c4ca9b4754c9c5fce295a78c151a89dc9f3dcd23eca83b8774f54bb

Observation 49e41f53-ee25-4ceb-af31-9f6821eaa9fe · outbound

This paper cites Qwen2 Technical Report.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Qwen2 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.323886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.323886Z digest=sha256:e95e8841ac511a344138dde5d3b3ed89fa74617bf38b35a7f71198abefe84b78

Observation e5bba39b-7a2a-4bf9-b6f4-4fd26499815e · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.576360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.326997Z digest=sha256:39b6af85a8bcc3fd6c0b24c45ffe0b469a3da500daa4d0ae3040392acba5e9c9

Observation a6401795-8ca0-42d3-97ec-24b156b955eb · outbound

This paper cites Detclipv2: Scal- able open-vocabulary object detection pre-training via word- region alignment.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detclipv2: Scal- able open-vocabulary object detection pre-training via word- region alignment

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.567971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.329849Z digest=sha256:f97dcd2551f5b63ed4fbcace63ac17f4883900049fa8d12ecede4a035986168e

Observation c740e57d-d35b-4a47-b774-49fe0eb0225c · outbound

This paper cites Detclipv3: To- wards versatile generative open-vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detclipv3: To- wards versatile generative open-vocabulary object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.559626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.332659Z digest=sha256:1abb84a2704a364adcb4843292b8a74c266c1f0ac91b39cb2a83ed0fd24ed119

Observation c8c4055f-d33d-4e3c-abf8-528817b67173 · outbound

This paper cites Modeling context in referring expres- sions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Modeling context in referring expres- sions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.551103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.335463Z digest=sha256:ef62e8bd2a66567299e037cb361ca64e2131f3ec3e389200f72897b595b4c752

Observation 186e798d-e220-45a8-9f39-441239d814b0 · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.542460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.338352Z digest=sha256:a8d2e5776f2c287f29c499ba45f14a18327ab4a3f4994b12373b07bfaea80a68

Observation 11374560-06eb-44b7-92b9-335ce53c0f5b · outbound

This paper cites Sigmoid loss for language image pre-training.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.533868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.341115Z digest=sha256:57eb366e726089371008a83aacd666009a0124d7860c11140ef8e076481675ea

Observation e2f4d276-e4cc-4a3c-a279-99e31ae5b901 · outbound

This paper cites Glipv2: Unifying localiza- tion and vision-language understanding.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Glipv2: Unifying localiza- tion and vision-language understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.524614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.343845Z digest=sha256:a65193b24453de9bc18a689049262c35f09cceeb66069731c8cb9b29cc3dee99

Observation 446fe260-9a89-40b8-8fed-ec76fc9d0af3 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Dino: Detr with improved denoising anchor boxes for end-to-end object detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.515459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.346548Z digest=sha256:1bdb1a12b225ddc719998853c546c15d9853a788ee8d449a3228321cad23e16f

Observation f8cf1f89-d93d-4b5e-a524-0f87991f34d2 · outbound

This paper cites Generating enhanced negatives for training language-based object detec- tors.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Generating enhanced negatives for training language-based object detec- tors

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.506604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.349188Z digest=sha256:f012f4f7b4b527c98b7f8f0a4fac8608f0fe9e7d4fcefd75ea06d61ccc8e2c9c

Observation 2a559df2-72fb-4529-82c1-3936457f60a1 · outbound

This paper cites An Open and Comprehensive Pipeline for Unified Object Grounding and Detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.351984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.351984Z digest=sha256:0fa83ca06467312a45d6f5f32a670369226d932ccd4be160d96acbfd0d594c86

Observation b9aa56ca-f211-4658-a001-dc00bb9d5df4 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detecting twenty-thousand classes using image-level supervision

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.497229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.355152Z digest=sha256:322f11fdc7455c06157188027e578610959d69fffef6fd43a2b347121739cd6f

Observation 7bade701-8eca-4ae2-9985-41543c2c0371 · outbound

This paper cites Mova: Adapting mixture of vision experts to multimodal context.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mova: Adapting mixture of vision experts to multimodal context

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.488169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.357983Z digest=sha256:8a4332541d667234c9dbdd4e02ad9e52b76c96a134e3937c3bb6d168224907e0

Observation e1d54f5f-1378-46b5-98c1-321d736e2443 · outbound

This paper cites the man” and “umbrella.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models the man” and “umbrella

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.478023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:55:23.361176Z digest=sha256:9cc80f2e55e0a17b09eaeb4f58f607f65a5c27d6e54ee72cf09fa16476ce0b13

Pith citing papers

Observation f63d6b4c-fd85-4bc0-ad99-27228d92063a · inbound

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion cites this paper.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.436018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.436018Z digest=sha256:51fe753a86c2ea0db26bea53affe377f33cf9d179bef54979716b8c0388112a7

Observation b9aa25c1-7b01-4448-b0df-06e09001edc1 · inbound

RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution cites this paper.

RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:35.846198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:33:35.846198Z digest=sha256:5cc4502c9cb8b5fcc3fda8ae140addf47c6dd5478708d565fc736a275f35dc43

Observation 03db883a-530e-44f5-af38-8974a28cb22c · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.920344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.920344Z digest=sha256:c0018d11031f565893aa9d531c0c853b51a6de85747c362df4d3cd4f4270a546

Observation 925712cc-f512-4634-8d6b-39ea59245400 · inbound

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation cites this paper.

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:33.284362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T09:40:04.685274Z digest=sha256:d5cacfc44ffa86cd8026ee8b91140c50e802acc5b3278072813c22ee9b981665

Observation 28f99f0d-5609-453e-9ea9-da3eb702dda7 · inbound

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts cites this paper.

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T14:44:05.247859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:44:05.247859Z digest=sha256:bed141bb4a37914d20f3a571243411d4b7f62e5be54543b58feb0034eceaae99

Observation 2170e321-15aa-4d70-867c-ce1a46fcb057 · inbound

Kitchen Robotic Manipulation utilizing Foundation Models cites this paper.

Kitchen Robotic Manipulation utilizing Foundation Models LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:49:13.593500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:49:13.593500Z digest=sha256:ad5ea51673cb7ba29a4a65c9f55cbd253b741658ee6721908922fa39fc1e8749