Pith. sign in

Paper Citation Record · LEDGER

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval

As of 17 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2412.18806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18806 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:32:44.459932Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved22
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec488dff-669b-41a2-bbff-64aa37573a2a · outbound

This paper cites Label-embedding for image classification.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Label-embedding for image classification

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.172563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.172563Z digest=sha256:f0ac82a00bb0530051aea5fa7de2324651d7c96c23662c55e27c76b871282b90

Observation c125ee8e-1fac-4ae5-93ba-7684d01cdf5d · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.176758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.176758Z digest=sha256:61c270bab01726cdf464507ccc6a3ca6cb311b37dfa7708c658c85d018dff481

Observation 9e82b47a-cec5-48bb-8e04-6d544622e4e8 · outbound

This paper cites Pseudo-labeling and confirmation bias in deep semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pseudo-labeling and confirmation bias in deep semi-supervised learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.184879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.184879Z digest=sha256:5d38c25af03c4315556e24bc1de8039b6e517b84c49057ddce438fcf0c0b3643

Observation d26f7b55-1ea4-47bd-b0f5-230404ba5725 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Bridg- ing the gap between object and image-level representations for open-vocabulary detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.188636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.188636Z digest=sha256:912025d5601abd5aa3af82ac985df17ec2e7629bdde76f6fa558e3c9c62cb41c

Observation 41970f72-8636-4b34-8d16-633b727c8d03 · outbound

This paper cites Zero-shot object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Zero-shot object detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.411789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.192362Z digest=sha256:02d66d75ea0c39f020c24d277ee6490429cb98a6b42816bd320487e381b8f9b0

Observation 05dc9e84-02a1-4631-a38e-71b6cfa5fae3 · outbound

This paper cites Mixmatch: A holistic approach to semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Mixmatch: A holistic approach to semi-supervised learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.196412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.196412Z digest=sha256:3763c04cf6545058b6ebba9e542c6e115dc2f250ec98eb259d136f390ef63069

Observation 17a5de76-97ca-4339-98f7-13b0d9685fbc · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.394749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.200071Z digest=sha256:58f385eaea88ea1906d809be21abbc91e972746ff8cb8bfb15cca8f1cb0287a3

Observation 4e90383b-b9f4-4eec-a88a-bb0a7e0f2dbe · outbound

This paper cites X-detr: A versatile architecture for instance-wise vision- language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval X-detr: A versatile architecture for instance-wise vision- language tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.384161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.203649Z digest=sha256:2be9ca422141823771b7413c5256e3f025a9c8e145e094d797b39fddbcf90aac

Observation f87f0292-183b-4672-acc3-4919ee74a76a · outbound

This paper cites End-to- end object detection with transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval End-to- end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.373408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.207011Z digest=sha256:2854e55f0dfc258a63342c427a6896c5b5dd35d6d8d9039fae9a6ba901679b11

Observation 56821bf1-e47b-455d-95fc-595fd94361ff · outbound

This paper cites Big self-supervised mod- els are strong semi-supervised learners.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Big self-supervised mod- els are strong semi-supervised learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.363471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.210464Z digest=sha256:42148790df71a6f34e47c5e2d67b7c2093f80148de75a5a08c4d82c2b4efc2fe

Observation 0a833bd5-93f0-4192-95a1-42d607bf4af8 · outbound

This paper cites UNITER: universal image-text representation learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval UNITER: universal image-text representation learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.352804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.214032Z digest=sha256:363b7f8254166fbbf222f1a9f4b6049721d71951637141814a0e0b87ed1adbd7

Observation 61919452-1d70-4fe3-9c12-8542943725a4 · outbound

This paper cites Proba- bilistic embeddings for cross-modal retrieval.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Proba- bilistic embeddings for cross-modal retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.341194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.217445Z digest=sha256:a2fccecf8540f58ce262a604d26d03629475920414f1ba16f5e688586ec84d5e

Observation 903289dd-14d6-4b02-9e5f-95e5d290c92b · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.220762Z digest=sha256:d82ce7339ee23f62fc676bcb5e98a01b4e4abca93a8cd12eae76024d9db6ad51

Observation 6ce47396-f825-43a8-a6cf-66eae25fd29f · outbound

This paper cites Finding beans in burgers: Deep semantic- visual embedding with localization.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Finding beans in burgers: Deep semantic- visual embedding with localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.318051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.224063Z digest=sha256:1ffc27121a24c7b478f36e7efe12f22ce2ed7ba1c07ec560834330aada5474ba

Observation 953e852d-7825-4d4a-95a3-827dbc5c7084 · outbound

This paper cites Fleet, Jamie Ryan Kiros, and Sanja Fidler.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fleet, Jamie Ryan Kiros, and Sanja Fidler

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.306459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.227298Z digest=sha256:bf5ce776154bae247517c2267b4d9c354d7b89a7938cb892b42eb8c0f3d73176

Observation 06144a13-524b-496b-b7e6-1f4c216a364c · outbound

This paper cites De- vise: A deep visual-semantic embedding model.Advances in neural information processing systems (NeurIPS), 26, 2013.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval De- vise: A deep visual-semantic embedding model.Advances in neural information processing systems (NeurIPS), 26, 2013

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.294893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.230559Z digest=sha256:a73765c3e58bc940ce23439982cc07941fe197478e28d87426e895ca1cf79d04

Observation 3313c946-2ae3-4911-999a-8236d4c16ea0 · outbound

This paper cites Understanding the diffi- culty of training deep feedforward neural networks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Understanding the diffi- culty of training deep feedforward neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.283263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.233871Z digest=sha256:83535e2c30d8e545b87716d9e919cf2046258e786a448a6f0a683810eacbf3e2

Observation 6863ac69-5e23-4c13-8d6e-3a8a42c5e877 · outbound

This paper cites Improving image-sentence embeddings using large weakly annotated photo collections.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Improving image-sentence embeddings using large weakly annotated photo collections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.272919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.237175Z digest=sha256:8e005a14b950235dae99db8e53fb2039f677dd25e9b2cd36f85854eb7f015a91

Observation 719f7dc7-c08f-42a5-9630-e5623dc54cf3 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Open-vocabulary object detection via vision and language knowledge distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.262866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.240663Z digest=sha256:65d78228a8861c9bc5134cd6b53a29ec2155de3fa494ee54cba531861efbf0c8

Observation 143294bd-f34c-424d-8c02-172b4b6160b1 · outbound

This paper cites Girshick.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Girshick

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.251143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.243834Z digest=sha256:dfb4c8c3c2742f309834ac799676dc6359eb759c760c3aaf26a75b813f73c882

Observation 875578ae-d670-4d24-a181-b266a7f3887d · outbound

This paper cites Generative multi-label zero-shot learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Generative multi-label zero-shot learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.240567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.247322Z digest=sha256:636f58e52fff97cb9bded6d32974b4a986ede627be19b3d80ea145490326cd28

Observation e3af98b8-c68a-4ff9-a994-01a437976edf · outbound

This paper cites Mean average precision map@k metric explained code.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Mean average precision map@k metric explained code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.229970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.250616Z digest=sha256:4605842f3d3e1e2739a61eede8f1101d61da71bdc8cb0b354c1d8bc4da0a9563

Observation eadb2934-3082-4975-8320-8af75c355b0a · outbound

This paper cites Instance-aware im- age and sentence matching with selective multimodal LSTM.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Instance-aware im- age and sentence matching with selective multimodal LSTM

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.219450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.253776Z digest=sha256:539d1992050f350f9c057b34d0db506d0758fe525999ce9fe97fe4284eb7cc72

Observation 997aa0cb-9f17-408b-b1a0-f8a892d9fe35 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.256855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.256855Z digest=sha256:6f6a6ed59e57e505f234a3363a30e1bc4f215e58fa714a19ebd5f811440f7584

Observation 9ec5a05a-8bdd-4ae1-9026-fae029bf1a22 · outbound

This paper cites A shared multi-attention framework for multi-label zero-shot learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval A shared multi-attention framework for multi-label zero-shot learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.208593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.260734Z digest=sha256:b5ccc3b8be66701407281f52ceeb127045056f0e8de15d172899b0b18c6c7d6e

Observation 3fa3cc9b-d828-48df-83e3-961b39cba4ff · outbound

This paper cites Saliency-guided attention network for image-sentence matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Saliency-guided attention network for image-sentence matching

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.197759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.264069Z digest=sha256:fd7a5f28ffb4e5b8f9fc075e30a3eff92598b9eeea92ad004950363515db9327

Observation 686a888e-1eba-4f25-8004-13e084b54a3f · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.268216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.268216Z digest=sha256:2db21f0a3764bbcbc6097e9f30caa516f2e1d1203fa772c6e4a99ded3eb8e6d2

Observation eafb4ef9-c241-417d-8886-47d9e74930e1 · outbound

This paper cites Billion- scale similarity search with GPUs.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Billion- scale similarity search with GPUs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.180726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.271625Z digest=sha256:a6a48d8d2b1708bde0dfd8839daf256a3d4aca7a01c31930d5230805159549b3

Observation d8e8ddab-b78e-4966-b790-3786abc2f5d7 · outbound

This paper cites Deep fragment embeddings for bidirectional image sentence map- ping.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Deep fragment embeddings for bidirectional image sentence map- ping

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.169751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.274901Z digest=sha256:9e0a563f49a1dc37ee6d71ca346a7c0cc56ca4acc997cc9f76c9e031a2de94ae

Observation 21fc5f81-024b-4a2c-b6a7-7b36f78de260 · outbound

This paper cites Kingma and Jimmy Ba.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Kingma and Jimmy Ba

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.278151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.278151Z digest=sha256:8121ad24d26c6f3375d478b1e7d0b3a71ac1a3871b50cdff8c2225e4272d702a

Observation 8a12dc4f-a81c-44cc-a008-49f6370ae8d7 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.281134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.281134Z digest=sha256:ace405bbf32f95b0f845405562816a17f7d4e20138af1b32c83bd58143bce063

Observation 95c20d47-1580-4854-a49b-3dc08b8d5f11 · outbound

This paper cites Shamma, Michael S.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Shamma, Michael S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.152272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.284372Z digest=sha256:537aa28749563e7b0bc6f197f0438aeee91c794607f41e2fe8cb07d52389d402

Observation 6df81d8e-eff7-40a4-a593-cdfc267d3d75 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.141789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.287351Z digest=sha256:4d5d6dddf274bfae3a169c8316e881e7a1e6b17097fcd2c8e547406d8ce565dc

Observation b3b67e33-dcd7-43e6-9c36-f586355d5870 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.131234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.290556Z digest=sha256:a40e117f0ce76fdb1bfe0e07bd66e4280834cacea4461eeafcdb33d41f66bbef

Observation 60c05cb3-e6c2-405b-af93-126efa21482b · outbound

This paper cites Temporal ensembling for semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Temporal ensembling for semi-supervised learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.119941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.293820Z digest=sha256:ec2c6537719adc03aa9efc3f34906e8949ef1d37317251e58494e46285750cad

Observation de4c5e1e-d551-435e-a233-d19e650eb4ce · outbound

This paper cites Pseudo-label: The simple and effi- cient semi-supervised learning method for deep neural net- works.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pseudo-label: The simple and effi- cient semi-supervised learning method for deep neural net- works

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.108387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.296862Z digest=sha256:dddcc08cdd5aaf2e1708961314b11ecc78a21aafd7933520f166f8cb1d62c191

Observation e3ead0e3-686a-4c35-89a1-81cc86ebb2b4 · outbound

This paper cites Stacked cross attention for image-text match- ing.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Stacked cross attention for image-text match- ing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.097553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.299959Z digest=sha256:a7cb691c765dbd680ff98c26a008e033eb1ce6b99cd10fe308f9eac43ccd20fb

Observation 9dc3596a-371b-4a4f-9ae1-751f4c6e4a23 · outbound

This paper cites Object- centric open-vocabulary image retrieval with aggregated fea- tures.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Object- centric open-vocabulary image retrieval with aggregated fea- tures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.086738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.303233Z digest=sha256:050bbbe61e1b4ceebd270a2b085488edea5e90a84c8c5f562f0194d3f8387c1b

Observation 146a311d-96ca-4ecb-873c-1483b1c4eb96 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.066222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.310751Z digest=sha256:1190004a8695ff504bcbcd5b26097fb339ed24fc71d2c38e1e56f4a8b65ebc9f

Observation e34b8b82-f596-4a95-a4ee-346ed7268fc4 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.314386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.314386Z digest=sha256:633e8514eefc715ee6d173182ce507bbaeb13d6cda8334e9ce079f3d3834089b

Observation 59b5766e-5fd1-4b2e-b8f9-6b2f750695a0 · outbound

This paper cites Selvaraju, Akhilesh Gotmare, Shafiq R.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Selvaraju, Akhilesh Gotmare, Shafiq R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.050149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.317593Z digest=sha256:55fb093686e850f81dd17db74bbbace372669a90bf48933913f247039722c5a4

Observation fd915810-0a36-4df9-a6ea-3e382c155e70 · outbound

This paper cites Adapting CLIP For Phrase Localization Without Further Training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Adapting CLIP For Phrase Localization Without Further Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.321352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.321352Z digest=sha256:d71c26fb792b86c36e9d2ae4b606ddcf166ca34672a1dc0097c24e04bb089284

Observation d046fad0-f1d1-480c-8022-6a12fe3da4a4 · outbound

This paper cites Visual semantic reasoning for image-text matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Visual semantic reasoning for image-text matching

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.039224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.325446Z digest=sha256:cbc23e084efa7443eb7cb521c91659bdca739fa26823948de819b67006185337

Observation 73bd46c6-5fc3-4ab1-a26a-26cf35164df4 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.028368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.328831Z digest=sha256:fbd68cf824a432886e44aa6d844340a71d2a11b642ca5234d55e2d6b1155bf98

Observation a9d6e88f-414b-42f1-b0b8-507a899234a1 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.897271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.336016Z digest=sha256:2c5c7e0d1953cc0785ca7c9b0e3ab2d934d4bc1745e3fab3b66aa34c35d4a3d8

Observation 2b349096-06b7-42cb-a4c6-9cb48cf6f698 · outbound

This paper cites OVIS: open-vocabulary visual instance search via visual-semantic aligned representation learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval OVIS: open-vocabulary visual instance search via visual-semantic aligned representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.887428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.340173Z digest=sha256:ee097270b031e395f24bfe32fd23a8c52b53b3e6fa68605e7e592a136de9fb8d

Observation 16a27aa1-d413-4d77-8973-cf788d4a5f0c · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.877132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.343611Z digest=sha256:38ab1d06ed0b6fc6f6255e64392801368ab403b264085d04dac44b90a560b4f2

Observation d94d9d91-28c5-4225-994f-2c6853b905e0 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Simple open-vocabulary object detection with vi- sion transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.867278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.347063Z digest=sha256:670c31783cc411cc11be57e078ec0312146fd66d7c4daa509a929e37924b8689

Observation 7588f2e9-508f-4bb2-a304-3e9c275e8002 · outbound

This paper cites Gritsenko, and Neil Houlsby.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Gritsenko, and Neil Houlsby

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.856527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.350738Z digest=sha256:919642c081792d1e0237a490fe526a24d90b65e388182aab992c8061ce55c3a3

Observation 5bf0f9be-7dff-4b7f-8e14-5aa812923dba · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:44.845186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.354036Z digest=sha256:8b9584f600f8521df4e5fa9669c9d3268102d0ad62dbf76b4f6e4138327064b5

Observation 09142528-ce7a-4f0b-ac31-6b5639a54405 · outbound

This paper cites Zero-Shot Learning by Convex Combination of Semantic Embeddings.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Zero-Shot Learning by Convex Combination of Semantic Embeddings

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.357178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.357178Z digest=sha256:2e981213ef983acab899b6c40b676e305039319d98ab55487e9e7ddd87f0eea8

Observation 1fc4f01e-9860-440f-ae21-fb0089197ab5 · outbound

This paper cites Meta pseudo labels.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Meta pseudo labels

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.834732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.361301Z digest=sha256:2c5cc45663375489359e71e0d104f4ab754cce4d23315493c7345a61282a4c6a

Observation 3fa9b5d8-4ad4-4319-8180-845013a1315b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.823825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.364047Z digest=sha256:ff966a6a2f6e44392d9017925bfd091f3d20071f0a6243a20e7785c8a2c2f52a

Observation 8279678d-be3a-4a0f-907b-8b02f10f7380 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Learning transferable visual models from natural language supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.812922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.366883Z digest=sha256:ca0088fb770c8137313989f55099239f6841e77f47b26d4bd0e952c0eb1e65ee

Observation 32507526-2ecc-4074-846f-32b1062fe148 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Denseclip: Language-guided dense prediction with context- aware prompting

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.802015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.369964Z digest=sha256:b931a159e621e5fc3b74b53267e581cb9ca7864991dc717f67c2ec0886e714e0

Observation 344bb416-94ed-4c4b-afe8-07253d2c8211 · outbound

This paper cites Regularization with stochastic transformations and perturba- tions for deep semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Regularization with stochastic transformations and perturba- tions for deep semi-supervised learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.372916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.372916Z digest=sha256:c0f16d1e4e9ff885ddc7d61a1ab0684d830c0f16add83f5343beb05ff17aaced

Observation f9390e8f-d0e3-4351-a377-a724965289cc · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Objects365: A large-scale, high-quality dataset for object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.783806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.376152Z digest=sha256:2c9ba1bc51b38f341b59c6020c88d7b4e06c49104e2f11d987f91f0702d42cc0

Observation 709102e8-1cd6-4c97-8075-a6903002902b · outbound

This paper cites Fixmatch: Simplifying semi-supervised learning with consistency and confidence.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fixmatch: Simplifying semi-supervised learning with consistency and confidence

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.379094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.379094Z digest=sha256:9c1d11c24e03189e0d553a60dec2f968831fb7bef4d44831ff6ec9701d486d38

Observation 0a9d0bf0-888d-426d-89a1-2fc076229288 · outbound

This paper cites A Simple Semi-Supervised Learning Framework for Object Detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval A Simple Semi-Supervised Learning Framework for Object Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.381983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.381983Z digest=sha256:ddab583b00ad88905a4548bb73298e8a8a232dd2a4a9a5e21ab90f6d07ab919d

Observation ea46336a-b1d4-4448-9650-8cdf05a0cf7b · outbound

This paper cites Dualcoop: Fast adaptation to multi-label recognition with limited annotations.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Dualcoop: Fast adaptation to multi-label recognition with limited annotations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.765945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.385620Z digest=sha256:8d0797a3590e768e3f0ccf474dcedb087f8d2252c6adbed16d60ad1e853e8d48

Observation 4b43bf7a-9756-4073-aead-773778dcfe73 · outbound

This paper cites LXMERT: learning cross- modality encoder representations from transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval LXMERT: learning cross- modality encoder representations from transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.755377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.388942Z digest=sha256:82cd0658ada919f63f8db423024ca85f7df923e18810028b2a67baa5c486e74c

Observation c7ad6b8d-16d5-47a5-a1e2-e70438683180 · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval GIT: A generative image-to-text transformer for vision and language

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.744179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.392434Z digest=sha256:4da346ab91c439ed891a6e0915048a4478de03c224e727cb0e7d1c8d6ac1a227

Observation ef0f87d3-b8a2-405e-a4c5-f16b7fdc7551 · outbound

This paper cites Object-aware dis- tillation pyramid for open-vocabulary object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Object-aware dis- tillation pyramid for open-vocabulary object detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.733306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.395836Z digest=sha256:d564da31cd98a5cf3e6d58d9fde212a7edd4a80218204ed304f8a15fd09698c1

Observation a81d8b9c-e971-4e63-a19f-8ba16a6494e7 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Simvlm: Simple visual language model pretraining with weak supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.399222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.399222Z digest=sha256:17e5a492afb00195328fad24c1ed7cba2d33b4506da9c7819e6f92d7bc3e97ca

Observation 49996732-0701-4f2d-9d60-70578c39dc16 · outbound

This paper cites Aligning bag of regions for open- vocabulary object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Aligning bag of regions for open- vocabulary object detection

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.716100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.402977Z digest=sha256:1c3724c604c3b64cb29c03f4f5f424650fe4a161161f29fef6fddfb0499ce6f3

Observation 498f067e-e35f-4230-8853-ecb759bcd7db · outbound

This paper cites CLIPSelf: Vision transformer distills itself for open-vocabulary dense predic- tion.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CLIPSelf: Vision transformer distills itself for open-vocabulary dense predic- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.706543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.406238Z digest=sha256:8c93e9306bf9b37ad43049d5fe5882553aa939596efeb37cb4dc871a69ed9c00

Observation 81b8bdef-5eb0-46d9-9bba-cabc7ba1b374 · outbound

This paper cites CLIM: contrastive language-image mosaic for region representation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CLIM: contrastive language-image mosaic for region representation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.696334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.409651Z digest=sha256:d1fafa2b43ec8f62443a1a2c208cd7529e91bdccabe5b8db55d96b64b46ba893

Observation 580337d6-a7cf-4fc9-b58b-182f5997df2c · outbound

This paper cites CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.686378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.412748Z digest=sha256:b42a7e4a1953b9061197beb9d06545cbc6dc357f2ada6881db2a187222caef87

Observation 700ef036-57bb-4c77-ab99-8e0878c5917f · outbound

This paper cites Unsupervised data augmentation for consistency training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unsupervised data augmentation for consistency training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.676599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.415985Z digest=sha256:1914efec0f04c5b90b4b304bc6be24f0680a308531178ea53b62663abcb6956f

Observation b744b86f-6194-4436-b81d-dd6704429373 · outbound

This paper cites Self-training with noisy student improves imagenet classification.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Self-training with noisy student improves imagenet classification

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.666904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.419323Z digest=sha256:98296f5f3d8d794b785104ff1e43ed99ec5567f0f2bf3eeceb250ab28449981e

Observation 048912da-d9df-44e4-b709-f355f2b41bd4 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.422649Z digest=sha256:7f0b71cff5bedd15e91d9fb65a4a49133028303f4c778a4af9b58e7b5496cd5f

Observation bd62beab-7379-4888-8130-214531072f82 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Coca: Contrastive captioners are image-text foundation models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.647759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.426262Z digest=sha256:8b7744c0c9a00ff94aedbcfd9fc0e09e220ff334187bd967f21934611774a6b8

Observation b4d89a9c-6988-460f-b56c-92e848b601fd · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Florence: A New Foundation Model for Computer Vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.429626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.429626Z digest=sha256:3e1f14855822d9223712c375de0f71104a3197d31369f0a221f35320ee51f06d

Observation d208642c-8131-4507-af4f-51998798543d · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Lit: Zero-shot transfer with locked-image text tuning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.637917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.433395Z digest=sha256:0518b469f00671e80963a600c887ba4658482f1da3f53d8fdbb7cbb3cf54c07e

Observation ef404a21-badb-4213-b54b-cac7d18684fa · outbound

This paper cites Fast zero- shot image tagging.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fast zero- shot image tagging

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.627313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.436720Z digest=sha256:200f7ce7ebe73827e8dc692e34a3c3ca08936c93ab689bb93800878cc10cae5c

Observation ee36940d-78f8-4565-a92f-1b331c27776f · outbound

This paper cites Exploiting unlabeled data with vision and language models for object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Exploiting unlabeled data with vision and language models for object detection

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.616505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.439922Z digest=sha256:94198229f2f5b88ab773ed2eba4ef5db72d80577db69e256c0af21c77aad918b

Observation e2e330fd-2ab1-4fbe-94ce-3da4a3e6a41e · outbound

This paper cites Regionclip: Region-based language-image pretraining.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Regionclip: Region-based language-image pretraining

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.605565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.443115Z digest=sha256:0629794b725e24a3d8093b7f87c516a87a6bdc327944f98517b1feceb7d71be3

Observation b62ee730-0a93-42a7-87f4-f0bb529cb6f6 · outbound

This paper cites Extract free dense labels from CLIP.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Extract free dense labels from CLIP

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.594224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.446293Z digest=sha256:4ad568af89e2f3ccc45bfabb2b5ca163b71b6cea43a997d38bdd41dbd3684a70

Observation eb1e4021-4501-4b31-ad38-015d563c7db8 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Detecting twenty-thousand classes using image-level supervision

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.583536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.449513Z digest=sha256:7b717575b32911a5828e7720b7637e22f6b6ec383a76aaf491dea167da275d82

Observation 95111b72-017c-45d8-bace-6ee30db8b557 · outbound

This paper cites Semi-supervised learning literature survey.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Semi-supervised learning literature survey

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.573034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.452756Z digest=sha256:284d3e909545d7aae91dd211564a4292690db479b7671f0b8d1f8980ed17da07

Observation 351aa591-742d-4d2d-89d0-32ec6b46c002 · outbound

This paper cites Rethinking pre- training and self-training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Rethinking pre- training and self-training

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.562509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.456090Z digest=sha256:f707528e1cb09abd5ffd747a4ebe0f9c911487b3e444c493ef9e015afbbf8337

Observation e59c07a9-5278-4433-89ee-a27e483a05aa · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 85

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:32:44.550453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.459932Z digest=sha256:c6c5836c733bdd7c6fcefad68a44fea3d90414c6cb56d2de6ffb2bbb1a750250

Observation bf8ebb87-d2c6-425d-a6f2-0de489193ef3 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 137

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T04:32:44.906987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.332198Z digest=sha256:446be16cbbcf71c29173296d6a8f796d7e822f5f2378f5b7e2a4504b2ee1ceae

Observation c7121743-9217-492c-aba3-3e4c4fe05e1a · outbound

This paper cites 1, 2, 3, 4, 6, 8, 13.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval 1, 2, 3, 4, 6, 8, 13

Reference 608

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.076156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:32:44.307277Z digest=sha256:b0fd10ec74f8feabb482f8143b049210f906cb147eec9a2ff150400a24d77424

Observation 17f61c72-0b65-4ef8-9c9d-831b1e11f4cd · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 6086

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.180725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.180725Z digest=sha256:7a9cf3fd0dbb57f57ee6aeb4f69a641c4ea0ae4d0fa7d3bda9bb4cfef0356a30

Pith citing papers

No inbound Pith citation observations are available.