Pith. sign in

Paper Citation Record · LEDGER

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval

As of 16 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2412.18806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18806 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:32:44.459932Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved22
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec488dff-669b-41a2-bbff-64aa37573a2a · outbound

This paper cites Label-embedding for image classification.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Label-embedding for image classification

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.172563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.172563Z digest=sha256:f6633dd52fdf1825438751e3ebfed29e313e8f44f744a37df4d4355d4a4dd5ac

Observation c125ee8e-1fac-4ae5-93ba-7684d01cdf5d · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.176758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.176758Z digest=sha256:2bb12addeeb160c3b486a6570921c18d6dcde73b37f99916af85d3c1e0c515b2

Observation 9e82b47a-cec5-48bb-8e04-6d544622e4e8 · outbound

This paper cites Pseudo-labeling and confirmation bias in deep semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pseudo-labeling and confirmation bias in deep semi-supervised learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.184879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.184879Z digest=sha256:64de5df4682050fd271e5ae10c8546260b0d274a015a7d64c05384608273d6f8

Observation d26f7b55-1ea4-47bd-b0f5-230404ba5725 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Bridg- ing the gap between object and image-level representations for open-vocabulary detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.188636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.188636Z digest=sha256:a498a08204489d74c6e3ac6522aeb5ccd6228c6457a3cc4f0f271089f5636da6

Observation 41970f72-8636-4b34-8d16-633b727c8d03 · outbound

This paper cites Zero-shot object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Zero-shot object detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.411789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.192362Z digest=sha256:eb22b63851335b79b8ab75967ab2710e69e5a1cd13f0b1653b916aa197b78180

Observation 05dc9e84-02a1-4631-a38e-71b6cfa5fae3 · outbound

This paper cites Mixmatch: A holistic approach to semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Mixmatch: A holistic approach to semi-supervised learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.196412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.196412Z digest=sha256:b55b355f190cc26d2689e6841fa8dd944e5379ae4f799ebb7129df63659de871

Observation 17a5de76-97ca-4339-98f7-13b0d9685fbc · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.394749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.200071Z digest=sha256:e6b5e94301c55c065ccc9cd52fdbbf4c80e60865ebb59c4cb17af2e6a4cb3248

Observation 4e90383b-b9f4-4eec-a88a-bb0a7e0f2dbe · outbound

This paper cites X-detr: A versatile architecture for instance-wise vision- language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval X-detr: A versatile architecture for instance-wise vision- language tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.384161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.203649Z digest=sha256:c59c86f61a1e7e87a4766177387338e4f77b5644d52686bbf19565764e11e974

Observation f87f0292-183b-4672-acc3-4919ee74a76a · outbound

This paper cites End-to- end object detection with transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval End-to- end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.373408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.207011Z digest=sha256:711539d122012f9641e24f26a78fb253a711dff650c3ab54585c4086c6f37d86

Observation 56821bf1-e47b-455d-95fc-595fd94361ff · outbound

This paper cites Big self-supervised mod- els are strong semi-supervised learners.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Big self-supervised mod- els are strong semi-supervised learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.363471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.210464Z digest=sha256:429846e94131e6bd354aac1c9b573ddf82229b235a6e3e2836342258a6b8eb70

Observation 0a833bd5-93f0-4192-95a1-42d607bf4af8 · outbound

This paper cites UNITER: universal image-text representation learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval UNITER: universal image-text representation learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.352804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.214032Z digest=sha256:3cd1bb8e4eba3cebadcc4c239821238d74d28289797e6407f006919d689703e1

Observation 61919452-1d70-4fe3-9c12-8542943725a4 · outbound

This paper cites Proba- bilistic embeddings for cross-modal retrieval.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Proba- bilistic embeddings for cross-modal retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.341194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.217445Z digest=sha256:7777c17359fcd12e51b6de8b3396c768733a86006e2440d71b6fc000407918fd

Observation 903289dd-14d6-4b02-9e5f-95e5d290c92b · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.220762Z digest=sha256:71bac192d4855180c77fe4bfebd24a4f0c47644bcd57ce49693b854b888b0779

Observation 6ce47396-f825-43a8-a6cf-66eae25fd29f · outbound

This paper cites Finding beans in burgers: Deep semantic- visual embedding with localization.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Finding beans in burgers: Deep semantic- visual embedding with localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.318051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.224063Z digest=sha256:162325dff515105a2f145b7ee3e4aae0a5dde7aceec96a49156eb19031d2881d

Observation 953e852d-7825-4d4a-95a3-827dbc5c7084 · outbound

This paper cites Fleet, Jamie Ryan Kiros, and Sanja Fidler.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fleet, Jamie Ryan Kiros, and Sanja Fidler

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.306459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.227298Z digest=sha256:103ca52969bc7c26fae0311cb55389d3759f39a82ce4c597a5609ad671d934a5

Observation 06144a13-524b-496b-b7e6-1f4c216a364c · outbound

This paper cites De- vise: A deep visual-semantic embedding model.Advances in neural information processing systems (NeurIPS), 26, 2013.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval De- vise: A deep visual-semantic embedding model.Advances in neural information processing systems (NeurIPS), 26, 2013

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.294893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.230559Z digest=sha256:dee82ab642af9307ea6f09e2bd6b2f75ec60e56ed991ea4f14f77c34b7784d99

Observation 3313c946-2ae3-4911-999a-8236d4c16ea0 · outbound

This paper cites Understanding the diffi- culty of training deep feedforward neural networks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Understanding the diffi- culty of training deep feedforward neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.283263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.233871Z digest=sha256:d20ee1a1074d67993e6c623e587f9729b1262e8c16b3f0cb996bc39dc2693331

Observation 6863ac69-5e23-4c13-8d6e-3a8a42c5e877 · outbound

This paper cites Improving image-sentence embeddings using large weakly annotated photo collections.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Improving image-sentence embeddings using large weakly annotated photo collections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.272919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.237175Z digest=sha256:a5669b487ba0a0b14d91058777108f8ff82d17a2d6b29192d1e598eb753b32f0

Observation 719f7dc7-c08f-42a5-9630-e5623dc54cf3 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Open-vocabulary object detection via vision and language knowledge distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.262866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.240663Z digest=sha256:981c117642dcc7175ede94f9d6c9695508960e56770bbf7607d1c51452c5042d

Observation 143294bd-f34c-424d-8c02-172b4b6160b1 · outbound

This paper cites Girshick.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Girshick

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.251143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.243834Z digest=sha256:779ebc27360743f7d6482a97461fa43f380b16f17ff4ff547b65d0c67a86a0c0

Observation 875578ae-d670-4d24-a181-b266a7f3887d · outbound

This paper cites Generative multi-label zero-shot learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Generative multi-label zero-shot learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.240567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.247322Z digest=sha256:aba1c3a3ff2651e0cffddc7e3192751682e2669f8ee20dea114e23ad7cfbb173

Observation e3af98b8-c68a-4ff9-a994-01a437976edf · outbound

This paper cites Mean average precision map@k metric explained code.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Mean average precision map@k metric explained code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.229970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.250616Z digest=sha256:9095400b003657df6ee76b3b58cab6bdc17f2eb3c8b54778a30aaba338e44720

Observation eadb2934-3082-4975-8320-8af75c355b0a · outbound

This paper cites Instance-aware im- age and sentence matching with selective multimodal LSTM.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Instance-aware im- age and sentence matching with selective multimodal LSTM

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.219450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.253776Z digest=sha256:fe12ba0f1ffd9127dfa505b5758660845de81c1b4f31adf46ac3819103642654

Observation 997aa0cb-9f17-408b-b1a0-f8a892d9fe35 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.256855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.256855Z digest=sha256:fd9732ee3140ea6f4c0c86617d5ad54b43a0f52408415118330863d32a28d932

Observation 9ec5a05a-8bdd-4ae1-9026-fae029bf1a22 · outbound

This paper cites A shared multi-attention framework for multi-label zero-shot learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval A shared multi-attention framework for multi-label zero-shot learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.208593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.260734Z digest=sha256:b1c5f600e4fcc13d836adbf424cc12fab0518f331f2ce0ea5788190aae4e4b8a

Observation 3fa3cc9b-d828-48df-83e3-961b39cba4ff · outbound

This paper cites Saliency-guided attention network for image-sentence matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Saliency-guided attention network for image-sentence matching

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.197759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.264069Z digest=sha256:985d2356eb18b5d3410d979f1bfaaddd8c952eae069019dbb620b6e7b05cb0ca

Observation 686a888e-1eba-4f25-8004-13e084b54a3f · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.268216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.268216Z digest=sha256:79ecca4635c02cac1fa09c4561834f4c8020f2fbf9da13aa12bb2612f18312a0

Observation eafb4ef9-c241-417d-8886-47d9e74930e1 · outbound

This paper cites Billion- scale similarity search with GPUs.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Billion- scale similarity search with GPUs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.180726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.271625Z digest=sha256:244093fe3e5f92bea91bc582720f8f97b0df127d21881600876f59159b711f28

Observation d8e8ddab-b78e-4966-b790-3786abc2f5d7 · outbound

This paper cites Deep fragment embeddings for bidirectional image sentence map- ping.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Deep fragment embeddings for bidirectional image sentence map- ping

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.169751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.274901Z digest=sha256:386d81c70f61936df13d94bfc87383a56038689388372e97e23466c08a653c72

Observation 21fc5f81-024b-4a2c-b6a7-7b36f78de260 · outbound

This paper cites Kingma and Jimmy Ba.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Kingma and Jimmy Ba

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.278151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.278151Z digest=sha256:146bb844938acffd1db5540735af4db9095e851ae4fe6d98ab1cc3dc9d49a6b9

Observation 8a12dc4f-a81c-44cc-a008-49f6370ae8d7 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.281134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.281134Z digest=sha256:85675cadb8c464d2c719356d175e7a2d79c2e0f12c077e511dc2938721316cb0

Observation 95c20d47-1580-4854-a49b-3dc08b8d5f11 · outbound

This paper cites Shamma, Michael S.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Shamma, Michael S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.152272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.284372Z digest=sha256:e6f412aced08b52c7c4482d85e6351f112a15ce34c8c92a363f22bf10c9509f6

Observation 6df81d8e-eff7-40a4-a593-cdfc267d3d75 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.141789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.287351Z digest=sha256:014db66a0a74938bc6b001d2e6acff76f6f9efea830bf0f479dff18aea3b0b8d

Observation b3b67e33-dcd7-43e6-9c36-f586355d5870 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.131234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.290556Z digest=sha256:771fb93799883e1ee4d5b5043f4b2f924ea0e2a2eeae2e9b353bd097a46d8084

Observation 60c05cb3-e6c2-405b-af93-126efa21482b · outbound

This paper cites Temporal ensembling for semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Temporal ensembling for semi-supervised learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.119941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.293820Z digest=sha256:27cbdb5bb77d7a7804e83e418b4dc8a1ce5eb3dff1a55ef3cf69b7015d07c00d

Observation de4c5e1e-d551-435e-a233-d19e650eb4ce · outbound

This paper cites Pseudo-label: The simple and effi- cient semi-supervised learning method for deep neural net- works.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pseudo-label: The simple and effi- cient semi-supervised learning method for deep neural net- works

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.108387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.296862Z digest=sha256:dd1368f524909de0455ff5033851e3a338313ab46cc3eeda9d5cfab6d46110e2

Observation e3ead0e3-686a-4c35-89a1-81cc86ebb2b4 · outbound

This paper cites Stacked cross attention for image-text match- ing.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Stacked cross attention for image-text match- ing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.097553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.299959Z digest=sha256:1e1aaa23b93eaaffdb1c88200a22b5f19c9cef3907026bcdd19ec12b555221ba

Observation 9dc3596a-371b-4a4f-9ae1-751f4c6e4a23 · outbound

This paper cites Object- centric open-vocabulary image retrieval with aggregated fea- tures.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Object- centric open-vocabulary image retrieval with aggregated fea- tures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.086738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.303233Z digest=sha256:f59aec88431bc54486d04ea996ffedc90ecf1f67b679509b914944333e0944ed

Observation 146a311d-96ca-4ecb-873c-1483b1c4eb96 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.066222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.310751Z digest=sha256:8a759832c24ef38b16712cb383506862b366463027d2d14cef21a019c1976dc3

Observation e34b8b82-f596-4a95-a4ee-346ed7268fc4 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.314386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.314386Z digest=sha256:eca87d99ec0fe0b95dc2c55dafc2498ca996f7650c5a9312ee8e08c2d415f389

Observation 59b5766e-5fd1-4b2e-b8f9-6b2f750695a0 · outbound

This paper cites Selvaraju, Akhilesh Gotmare, Shafiq R.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Selvaraju, Akhilesh Gotmare, Shafiq R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.050149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.317593Z digest=sha256:dc4962af2ae4c5d682277a1159f92701621b4d51726bfc86e703c5333de94cb5

Observation fd915810-0a36-4df9-a6ea-3e382c155e70 · outbound

This paper cites Adapting CLIP For Phrase Localization Without Further Training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Adapting CLIP For Phrase Localization Without Further Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.321352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.321352Z digest=sha256:1496f1fdb0f00fa2f7e7a8e37e08253a95bf28eb365550a49691c51790939523

Observation d046fad0-f1d1-480c-8022-6a12fe3da4a4 · outbound

This paper cites Visual semantic reasoning for image-text matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Visual semantic reasoning for image-text matching

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.039224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.325446Z digest=sha256:9687de1d572339a81575177f6578b3cb38a4693943a820a9c8c06d3664450ced

Observation 73bd46c6-5fc3-4ab1-a26a-26cf35164df4 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.028368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.328831Z digest=sha256:4cffb6214316b79febb4050d1e552bfcd68fd64b3bfef4d8d293296a91025198

Observation a9d6e88f-414b-42f1-b0b8-507a899234a1 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.897271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.336016Z digest=sha256:00bea49545a656263d90d6c0d5d153951c9a5c895c241e288bad842a0fdd5a5c

Observation 2b349096-06b7-42cb-a4c6-9cb48cf6f698 · outbound

This paper cites OVIS: open-vocabulary visual instance search via visual-semantic aligned representation learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval OVIS: open-vocabulary visual instance search via visual-semantic aligned representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.887428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.340173Z digest=sha256:f2e7f07c5f2e6191c362f006371e8c55ee67f2bff1d55116cbf88ffa378a9184

Observation 16a27aa1-d413-4d77-8973-cf788d4a5f0c · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.877132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.343611Z digest=sha256:73f5454ac35048b5b8641e28fa4ab33d55714cb35a728c631f3feab3c135d604

Observation d94d9d91-28c5-4225-994f-2c6853b905e0 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Simple open-vocabulary object detection with vi- sion transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.867278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.347063Z digest=sha256:2bd8ae8cdd264f5565319b7bf81122abcf5dd38c52e00fdd3794f222b5def992

Observation 7588f2e9-508f-4bb2-a304-3e9c275e8002 · outbound

This paper cites Gritsenko, and Neil Houlsby.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Gritsenko, and Neil Houlsby

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.856527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.350738Z digest=sha256:c69db8759e2e6bbbb22546b70d1dcf789f33698f3cfc448fbbfc4a22cf527f6c

Observation 5bf0f9be-7dff-4b7f-8e14-5aa812923dba · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:44.845186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.354036Z digest=sha256:5d89dcb639f9d6eadab0bddb2766d6d6fc3d090205e5f4c1395e939f5e426c42

Observation 09142528-ce7a-4f0b-ac31-6b5639a54405 · outbound

This paper cites Zero-Shot Learning by Convex Combination of Semantic Embeddings.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Zero-Shot Learning by Convex Combination of Semantic Embeddings

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.357178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.357178Z digest=sha256:a908ef55b5137ba794a7baab29548eff500d4b4e6c0a42c0f8c5b485414cd90e

Observation 1fc4f01e-9860-440f-ae21-fb0089197ab5 · outbound

This paper cites Meta pseudo labels.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Meta pseudo labels

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.834732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.361301Z digest=sha256:2528308516bedc6330b6c4a1ff8c2c2abcb2905a68ce98a3bd3287e392bd76ec

Observation 3fa9b5d8-4ad4-4319-8180-845013a1315b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.823825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.364047Z digest=sha256:40af764116a3faa65d5721abe9c461456f38aab5c4c7a9ed07971d84d30f562d

Observation 8279678d-be3a-4a0f-907b-8b02f10f7380 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Learning transferable visual models from natural language supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.812922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.366883Z digest=sha256:3e003f20cf5007e60a620f9416be236cc345965dea2b3b70c8b30679ba8c7b2b

Observation 32507526-2ecc-4074-846f-32b1062fe148 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Denseclip: Language-guided dense prediction with context- aware prompting

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.802015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.369964Z digest=sha256:05e0f982727997e45acc3203a00b721725f8d6dbf88dd983d54883d096b22ff6

Observation 344bb416-94ed-4c4b-afe8-07253d2c8211 · outbound

This paper cites Regularization with stochastic transformations and perturba- tions for deep semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Regularization with stochastic transformations and perturba- tions for deep semi-supervised learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.372916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.372916Z digest=sha256:97b3a1e0634443241c5591d69c61e37a42c5da13dd759034df5cde5338940940

Observation f9390e8f-d0e3-4351-a377-a724965289cc · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Objects365: A large-scale, high-quality dataset for object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.783806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.376152Z digest=sha256:4105f2542dc09d33a657bcfdd7b1d0cd5f4c00981bd5c619b4510027b9e58541

Observation 709102e8-1cd6-4c97-8075-a6903002902b · outbound

This paper cites Fixmatch: Simplifying semi-supervised learning with consistency and confidence.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fixmatch: Simplifying semi-supervised learning with consistency and confidence

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.379094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.379094Z digest=sha256:960effbcf5257e568babbf3c8f3d1a56404e0cb87c3be4b4468c5d359afea96d

Observation 0a9d0bf0-888d-426d-89a1-2fc076229288 · outbound

This paper cites A Simple Semi-Supervised Learning Framework for Object Detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval A Simple Semi-Supervised Learning Framework for Object Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.381983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.381983Z digest=sha256:8ff3c337d1948dbc0edf434998b8c28f5b615c473f179784708a68e5446fac2a

Observation ea46336a-b1d4-4448-9650-8cdf05a0cf7b · outbound

This paper cites Dualcoop: Fast adaptation to multi-label recognition with limited annotations.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Dualcoop: Fast adaptation to multi-label recognition with limited annotations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.765945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.385620Z digest=sha256:1c04d5e1fdfb1696f908541ac3b631131b69bd8c993d958e8f6680d399aa2b06

Observation 4b43bf7a-9756-4073-aead-773778dcfe73 · outbound

This paper cites LXMERT: learning cross- modality encoder representations from transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval LXMERT: learning cross- modality encoder representations from transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.755377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.388942Z digest=sha256:49659acc8a427778764452ce07ad64dc6868458700368d08b769614d7d0ca455

Observation c7ad6b8d-16d5-47a5-a1e2-e70438683180 · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval GIT: A generative image-to-text transformer for vision and language

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.744179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.392434Z digest=sha256:b6a02d5b79a6551e83c77f378fdbe4771c6d0928bf6ff84630d158919950a20a

Observation ef0f87d3-b8a2-405e-a4c5-f16b7fdc7551 · outbound

This paper cites Object-aware dis- tillation pyramid for open-vocabulary object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Object-aware dis- tillation pyramid for open-vocabulary object detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.733306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.395836Z digest=sha256:4ac3cee839bcb0e662b3cebb3fe03bcfb4be951660660b532121a7901b550a80

Observation a81d8b9c-e971-4e63-a19f-8ba16a6494e7 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Simvlm: Simple visual language model pretraining with weak supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.399222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.399222Z digest=sha256:19ce236233a6c415b90b0bd8dd21802d460bd9fec8e6b11d401021a13c745a56

Observation 49996732-0701-4f2d-9d60-70578c39dc16 · outbound

This paper cites Aligning bag of regions for open- vocabulary object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Aligning bag of regions for open- vocabulary object detection

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.716100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.402977Z digest=sha256:11996ea6c36d099da304bb68de56ccf97c4df615bf6331ff248f937c6e0ff43f

Observation 498f067e-e35f-4230-8853-ecb759bcd7db · outbound

This paper cites CLIPSelf: Vision transformer distills itself for open-vocabulary dense predic- tion.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CLIPSelf: Vision transformer distills itself for open-vocabulary dense predic- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.706543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.406238Z digest=sha256:ab7cbfae7306736e1336554ff9454332930a03fe4cec23920974164ed382a1cf

Observation 81b8bdef-5eb0-46d9-9bba-cabc7ba1b374 · outbound

This paper cites CLIM: contrastive language-image mosaic for region representation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CLIM: contrastive language-image mosaic for region representation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.696334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.409651Z digest=sha256:8f4237ef821f03d20ca79e8748d3aab8c83919d6e6d6106b69f4bb6105900348

Observation 580337d6-a7cf-4fc9-b58b-182f5997df2c · outbound

This paper cites CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.686378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.412748Z digest=sha256:4dd13d619ed0a68fc4d3c23fd125d9918dafdd3a67f6a7198d5a50525fd61a96

Observation 700ef036-57bb-4c77-ab99-8e0878c5917f · outbound

This paper cites Unsupervised data augmentation for consistency training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unsupervised data augmentation for consistency training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.676599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.415985Z digest=sha256:acc719f020ac90405ba990bbaa590f8f9cc70b31af37576b958a5030a7ed861c

Observation b744b86f-6194-4436-b81d-dd6704429373 · outbound

This paper cites Self-training with noisy student improves imagenet classification.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Self-training with noisy student improves imagenet classification

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.666904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.419323Z digest=sha256:c19e2b8b6aa83ea65e60112c3af7a6ba1a7d25c45c6b5a378a802310d5541452

Observation 048912da-d9df-44e4-b709-f355f2b41bd4 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.422649Z digest=sha256:77283df43c357051ec231695c16823c3372e4d471c6a1c389e08406186d685e9

Observation bd62beab-7379-4888-8130-214531072f82 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Coca: Contrastive captioners are image-text foundation models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.647759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.426262Z digest=sha256:bee7296a9007dadcc2d9a462ef4b3bbc058b3dfaf85b1afe96db3569a242de14

Observation b4d89a9c-6988-460f-b56c-92e848b601fd · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Florence: A New Foundation Model for Computer Vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.429626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.429626Z digest=sha256:59c93c10be61fad3e146ac8dc6c1a37b272a87cfcb1f9f264e2c9f5be94822a4

Observation d208642c-8131-4507-af4f-51998798543d · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Lit: Zero-shot transfer with locked-image text tuning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.637917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.433395Z digest=sha256:130ee02ebfe532bb6c480fae80f0bbab857a4d4b6a8f0a9caf44b7bd20a3724b

Observation ef404a21-badb-4213-b54b-cac7d18684fa · outbound

This paper cites Fast zero- shot image tagging.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fast zero- shot image tagging

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.627313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.436720Z digest=sha256:dbdaff6763ce7b5ee1300ce12304b37c859fa082160135302ad52dbdae5b9e1c

Observation ee36940d-78f8-4565-a92f-1b331c27776f · outbound

This paper cites Exploiting unlabeled data with vision and language models for object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Exploiting unlabeled data with vision and language models for object detection

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.616505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.439922Z digest=sha256:15eacf4ae486df552fe4631007ac0ded36d374f90a107fb0fd0563aa2fa04ad2

Observation e2e330fd-2ab1-4fbe-94ce-3da4a3e6a41e · outbound

This paper cites Regionclip: Region-based language-image pretraining.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Regionclip: Region-based language-image pretraining

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.605565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.443115Z digest=sha256:1a6e485ce0da7a69e31f27517293a7b5e363889f90a3438016a1591c41da3ebc

Observation b62ee730-0a93-42a7-87f4-f0bb529cb6f6 · outbound

This paper cites Extract free dense labels from CLIP.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Extract free dense labels from CLIP

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.594224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.446293Z digest=sha256:32f254082645f9deb7222764fb4da80bd151eb9146c35811175374b724136fc0

Observation eb1e4021-4501-4b31-ad38-015d563c7db8 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Detecting twenty-thousand classes using image-level supervision

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.583536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.449513Z digest=sha256:3000e7c17fd58c8d9ad9e57b8a68dcb6afe630c4692bcfff9212b66ce506ec48

Observation 95111b72-017c-45d8-bace-6ee30db8b557 · outbound

This paper cites Semi-supervised learning literature survey.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Semi-supervised learning literature survey

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.573034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.452756Z digest=sha256:c144df40d2089891738302fd5a949bc53a4a9c78858a16e79283aa0f5154a64f

Observation 351aa591-742d-4d2d-89d0-32ec6b46c002 · outbound

This paper cites Rethinking pre- training and self-training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Rethinking pre- training and self-training

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.562509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.456090Z digest=sha256:031ca73e1ff4ecc5dda51a6d4f8c52d3fc5b5ce1631783e2d30decb18fc984ec

Observation e59c07a9-5278-4433-89ee-a27e483a05aa · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 85

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:32:44.550453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.459932Z digest=sha256:5a717386dec9a59e364b648b8982056c7194ec0b479681a2050cc11b5428522c

Observation bf8ebb87-d2c6-425d-a6f2-0de489193ef3 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 137

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T04:32:44.906987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.332198Z digest=sha256:ddeabbd479b1c245afd89feafdded4cef895bd3c864944807b90d9dbe513f980

Observation c7121743-9217-492c-aba3-3e4c4fe05e1a · outbound

This paper cites 1, 2, 3, 4, 6, 8, 13.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval 1, 2, 3, 4, 6, 8, 13

Reference 608

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.076156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.307277Z digest=sha256:81117da8de8e9586acb34eb84e02a373a3e915c893834c8f148d6bf9a70a78b9

Observation 17f61c72-0b65-4ef8-9c9d-831b1e11f4cd · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 6086

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.180725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.180725Z digest=sha256:a0c2fb11d5cad2a3703906a47f0c48f6bac478065b3665b1a2518bb03c9a6c70

Pith citing papers

No inbound Pith citation observations are available.