Pith. sign in

Paper Citation Record · LEDGER

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

As of 15 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.24025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24025 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:43.126490Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7fa3421-7e9e-4cd0-ad68-5d7b237a90b2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.156625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.156625Z digest=sha256:61fc6c14acac0101c0381fb14209e7e0c7a70d28978234209c9411c172426625

Observation b4ecf9c9-bc7b-4d08-8601-fb38be14cdfb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.261189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.261189Z digest=sha256:b5884d7af57d44049f0e6bed505123c8c16199f201892b68b5db6e1cde47d260

Observation bb8a39e6-f4c2-4445-97de-695c4cf08d6d · outbound

This paper cites The Llama 3 Herd of Models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.390073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.390073Z digest=sha256:f77d924d63aaf88fe9d81e08bb543014f5f2be29be9ad45805b52aa37bb236da

Observation 4eb85102-92d4-4566-b328-cb88248aed08 · outbound

This paper cites Qwen Technical Report.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.493697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.493697Z digest=sha256:e7b2f6ed123e4444782b062c040e0f8702ffbafc05687317fa9d40faf3c9e7e6

Observation d29c3809-a7bd-40d4-a555-ef565faded41 · outbound

This paper cites Qwen2.5 Technical Report.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Qwen2.5 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.629796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.629796Z digest=sha256:d94013383b5c652a895a036d8c505f8133cb65789a8d32c560db4e08ee61c41c

Observation 33e50349-0dd1-4cca-b99c-0339394e082a · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.746391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.746391Z digest=sha256:934191a11f85822eb39b97c0cbff28ba1fecd82afb8bd8b460f129273c15829e

Observation a067b816-b85d-497d-8084-7155d5c6def3 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.859878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.859878Z digest=sha256:dee92b2c6c97d7671f386f2d8a6d81a08469f6220a390bd140e4f85c4a9846f7

Observation cb988395-0cfc-42b3-9717-fa3fc90fd8fd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.902531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.902531Z digest=sha256:b69c39b914e7a6781fb7ba1a1b37fe7bb51f0268f97a33222dfd3f285fe17eaf

Observation e41cc792-baac-424c-a1fc-20f44a7e3636 · outbound

This paper cites DeepSeek-V3 Technical Report.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek-V3 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.970665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.970665Z digest=sha256:4e0e97f9b98cdaa14a648c3647f5122de7c9b9929ab2695ea9f5dac38c3d613b

Observation 0050adac-68c7-4a78-ad03-b15c7ce21f64 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.038626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.038626Z digest=sha256:f03b53040cc4dffa46e9bbff2e33206e673f050c0e530cd71daa100d84923d20

Observation b9a6c6ab-b501-420e-8a4f-927aaaa92e05 · outbound

This paper cites Segment anything.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Segment anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.094947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.094947Z digest=sha256:3ebc6b112f8f20c1d4dd00ad3d68e67adc9c1b52c7d3addf98c55e807d9dab02

Observation e0991557-f1d0-44cd-899e-fa00cafaf8d3 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models SAM 2: Segment Anything in Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.167283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.167283Z digest=sha256:89b0edad9b47a00588eb129f3e0da4c01616a0d5029ae4e263996cca9b4c96eb

Observation 5c9bfed1-f09b-4e01-8ef0-5f4d300276f4 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.233737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.233737Z digest=sha256:2c950577f46de8380d5a6a2a362b99bc861cd05734fcbc8fee208f17f71a465f

Observation 66406ad9-3fa3-4b00-9c68-4f52956ca41d · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.298980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.298980Z digest=sha256:c6016c19e3eae569e5d1727aea51005fe6ce5b18da7105c6dfdf4d7e0e9b6c15

Observation 7f08e0fa-4254-4761-b8a1-c7ffb6ac8e98 · outbound

This paper cites DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.392644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.392644Z digest=sha256:55adacd597004c87e18aefe09ffae4cd2e28af85e9c14b1b58499fcacc3c6bb0

Observation c05c271e-7514-4b73-874f-6de89913f62b · outbound

This paper cites Visual in-context prompting.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual in-context prompting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:46.051518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:39.501161Z digest=sha256:82d91162437ceb8756e25a1a481478a37bc9a7cb2b6234c147ba0e80e864a26a

Observation 9e1ad738-9420-425b-a973-25c6ab61b52d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.596806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.596806Z digest=sha256:fcf2540ea0779211f4a82761c73b2f3935a851813f4a65d8d342319b75402fb9

Observation 704b60cf-de98-437a-8e4d-430091d07f43 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.690023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.690023Z digest=sha256:845065a1412cb1235641680d66e351b4c87df5ad2792ffd64d932f8b99adc3c2

Observation 662ab28b-6bef-4529-b622-eed32a677db4 · outbound

This paper cites Microsoft coco: Common objects in context.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Microsoft coco: Common objects in context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.743576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.743576Z digest=sha256:3d6c82a85479704344a0c53009c6155264becb43195bb756cf27a1f02c17aee7

Observation a4030498-18b2-404f-801d-ebf925e00178 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Objects365: A large-scale, high-quality dataset for object detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.812688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.812688Z digest=sha256:c7c4430e26947e069e97bb13933a9467e180bb52bf7d2c1a57525ea0894344e6

Observation ae087a71-c295-45c0-b2f2-1221703ca4cd · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.891022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.891022Z digest=sha256:b9cb3f32e4577f2f06350e5bd7ef9d4cd328c5ffc66067119806f605562b8d05

Observation 5f9ef87f-503d-4289-b3e4-5f2af612fbb1 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models A simple framework for contrastive learning of visual representations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.964947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.964947Z digest=sha256:9627fc24ae44412f0fe2c49e3710326f35d16f9c34fad979e4a540a424616f84

Observation b477893c-00b4-4249-93b3-d4d253260570 · outbound

This paper cites G-simclr: Self-supervised contrastive learning with guided projection via pseudo labelling.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models G-simclr: Self-supervised contrastive learning with guided projection via pseudo labelling

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.911668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:40.061380Z digest=sha256:fb8a2c8ec476234eadb59a8f45cbe64728ab0c4454d61e991de92149c9ff808b

Observation d06069fb-48f5-4f1e-ab3f-1f9a3a56e64b · outbound

This paper cites T-Rex: Counting by Visual Prompting.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models T-Rex: Counting by Visual Prompting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.179541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.179541Z digest=sha256:605c1df303d78956f63678ed9cb20be2db0f027d650391e2867da83bc5fb7261

Observation 73714e8e-8a42-401a-9e20-194a605f915e · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.840820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:40.294747Z digest=sha256:82ce77ead5c11afbbba07b4f2f2569907862222a856d5dd1226a95f1813417fd

Observation 9f2710a2-c5df-456b-9f9d-9ec81e47460f · outbound

This paper cites Cp-detr: Concept prompt guide detr toward stronger universal object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Cp-detr: Concept prompt guide detr toward stronger universal object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.729040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:40.365687Z digest=sha256:156da29dc383bcaf6d7fb80d3159e43bd9afecedc054b16dcbda8ce29b2a18c5

Observation 4a7b5899-47f8-49a3-81a4-9ccbbd8385d4 · outbound

This paper cites AutoVP: An Automated Visual Prompting Framework and Benchmark.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models AutoVP: An Automated Visual Prompting Framework and Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.486017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.486017Z digest=sha256:9020002d4df76389fadfacd32b3e95fc9d39574fbd39b78e6d8479168bdfbffb

Observation a2913642-ecaf-4e01-92af-966cba330f9e · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.605490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.605490Z digest=sha256:f1c19890aaeaa1899f1e919277a94ef8eecaccabb27fd6f18ea63f0c1b50022b

Observation d1b9627f-cdc5-49b7-b38f-e8d6f1209201 · outbound

This paper cites Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.601553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:40.702290Z digest=sha256:4b6d89b1e44a5ef7912833aaa99c074f247799555c7cc57fcdb927477cce0ff8

Observation af1f007e-283e-4536-b106-f94e0420e065 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.833021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.833021Z digest=sha256:78c590e595c77dbceac378e151250c9a884cce6436e4b53e065ce03e6d3365eb

Observation 2ef018e4-7d7d-4683-a8c9-2f0605bf3126 · outbound

This paper cites Prompt engineering for zero-shot and few-shot defect detection and classification using a visual-language pretrained model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Prompt engineering for zero-shot and few-shot defect detection and classification using a visual-language pretrained model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.422370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:40.955884Z digest=sha256:076c4af1fd194446013e2e8e994384e8bd28eccc394abd6e3d0ee44837323aec

Observation 8a1757fb-0847-4a57-b776-379e6a3ef43b · outbound

This paper cites Promptcharm: Text-to-image generation through multi-modal prompting and refinement.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Promptcharm: Text-to-image generation through multi-modal prompting and refinement

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.316318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:41.022941Z digest=sha256:eed40fbe73d77357e49a693ebd6201874ea58b60071e9c25a643b62a5a52e1c1

Observation f161193d-fc2f-4a2c-aea5-812bfbab78d8 · outbound

This paper cites Prompting industrial anomaly segment with large vision-language models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Prompting industrial anomaly segment with large vision-language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.217572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:41.126318Z digest=sha256:ebdafb89bf1652abdbe0a577012039d5426df531b91574f6fadf91a354b14f8f

Observation a40d6278-33db-482d-a494-cb1094d91504 · outbound

This paper cites Visual Prompting in Multimodal Large Language Models: A Survey.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual Prompting in Multimodal Large Language Models: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.257873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.257873Z digest=sha256:fca3bc954581872ee60cf374ace255a4a66120996da4f69a70ae165648de96bf

Observation 2b52cc8c-8fb5-46f8-8fb8-5630793c7e8c · outbound

This paper cites Understanding and improving visual prompting: A label-mapping perspective.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Understanding and improving visual prompting: A label-mapping perspective

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.134917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:41.378888Z digest=sha256:fb3e3da5ff1dd0531da1b5aec77b20d40ea7c7a63c29923f8765be8b0ef3a28f

Observation 011313bf-cbfa-42f9-ab5a-d3ee850c0512 · outbound

This paper cites Vp3d: Unleashing 2d visual prompt for text-to-3d generation.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Vp3d: Unleashing 2d visual prompt for text-to-3d generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.952477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:41.458839Z digest=sha256:a54a92dd77233b04e28d18c2b457070075e5846af9509cca595eefceb0855e42

Observation b35e5bdd-f5c5-40c4-907c-40339020a79f · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.549563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.549563Z digest=sha256:7a3c7b11ef9073bca3c8e5652ded323085bb840e9bf32263d679c28664d54ac9

Observation 1efe85d8-c42e-4f8a-a403-3af50ae9eedf · outbound

This paper cites An Open and Comprehensive Pipeline for Unified Object Grounding and Detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.692788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.692788Z digest=sha256:9c7128c87db8c3dc9ec4c4670e9c6d4e5a3240dbed3df1888937afa3ad9b0fa2

Observation ee2b72e2-a38b-4cb5-8332-77c425c04928 · outbound

This paper cites Learning to prompt for open- vocabulary object detection with vision-language model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Learning to prompt for open- vocabulary object detection with vision-language model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.738724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:41.798426Z digest=sha256:2d88761e0bbf8e003d95c3f2762bd29d5456c49e51bbeeb59b7631d02d24e7a9

Observation 3660d5f6-6942-4fb8-b902-dd8dce926aff · outbound

This paper cites Proximal Policy Optimization Algorithms.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.874245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.874245Z digest=sha256:08418cbad0be6648f3ef8ba7c9296c1236da803bb1706ddb8e8e27f0495ce8e2

Observation d96daeee-c489-4fa0-b0b0-53c03e340c5f · outbound

This paper cites Training language models to follow instructions with human feedback.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Training language models to follow instructions with human feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.985819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.985819Z digest=sha256:1057f1a8911a0df2841fd173f668d0fbb3d37cdda504010622e7f8d280ddac5f

Observation 187d8867-3352-42a4-8935-d9094727e887 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.119299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.119299Z digest=sha256:90590f9481ad3c12433e1f659be8fc3ca927b03c85a405079526721a9d217432

Observation dd638852-dbbf-4862-ae9d-ad64d68c96b5 · outbound

This paper cites Learning transferable visual models from natural language supervision.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.236006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.236006Z digest=sha256:16efdf5d6fa915e266c81127fa6d30c7d33de769f09a21ae81b2867e2ee49562

Observation 5d3c849b-b1d9-4e16-a256-ce0158fff4bb · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Florence: A New Foundation Model for Computer Vision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.302476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.302476Z digest=sha256:c4731b8c2352cb01785314eab33895bb2f7667e9369b98c7337e5914d4a96bea

Observation be4cbdb3-a54b-421d-bb99-570a88d2ebf7 · outbound

This paper cites Grounded language-image pre-training.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Grounded language-image pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.362469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.362469Z digest=sha256:ae48dfb429f59fb1c806bcc613bfa81c0e645d18b5cf1f4264b07e5e239f4ee1

Observation 3f019d7f-0cb2-411a-8a18-fd4273d538d0 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.418637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.418637Z digest=sha256:ffd67dc5047b36c5f2cc01e3188905ae6dbe8e81ee5335195f756bb505d62a13

Observation 4e5b519a-8bbb-4317-ac31-80c6cddfb1cc · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Swin transformer: Hierarchical vision transformer using shifted windows

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.491065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.491065Z digest=sha256:078e5811748d4a4e6bdb5cba62ac663e017880f8c0155cb7e7c84e8f12684ab1

Observation 3e583b6b-35c1-4fc9-a44a-a84b6bbd4f35 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.554071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.554071Z digest=sha256:8da0886c34ff6ee6e326c420f92958998c1f9c973393270824007b581c2c10d1

Observation 0325b4b6-57f1-4fb1-b1e2-5125694fc982 · outbound

This paper cites Yolo-world: Real- time open-vocabulary object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Yolo-world: Real- time open-vocabulary object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.581367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:42.632729Z digest=sha256:52d34ca12f282e789b033bc2536b9940b1140ea5efe6611b31e686d032cd099b

Observation beab8506-1c34-412b-8baa-9a6374c7213c · outbound

This paper cites End-to-end object detection with transformers.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models End-to-end object detection with transformers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.490030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:42.690710Z digest=sha256:52d7286627ec89d8f04e516f6a0ba99f23f2e8a4e719c6c7c664a3db8a83eb0d

Observation ce2b969e-c6cb-44f4-9c79-594a261c9621 · outbound

This paper cites Dn-detr: Accelerate detr training by introducing query denoising.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Dn-detr: Accelerate detr training by introducing query denoising

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.346485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:42.749116Z digest=sha256:920678d54929198c8be78de5414da456184b45cf9431cb0880c1d324f65d130e

Observation bab7b3c6-915c-4f98-ab1c-60e8a3c80c80 · outbound

This paper cites DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.812679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.812679Z digest=sha256:2b1890b062a6f78365b74f56d17611959620acda8c0cc68b9c8ee46ffdac3ba8

Observation e6b29265-bd35-49d5-9570-0c7116e77fe6 · outbound

This paper cites Open-vocabulary detr with conditional matching.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open-vocabulary detr with conditional matching

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.203328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:42.850234Z digest=sha256:e8a7df5b8f09c0c71a3c7475f67f9efb4c7c47069dd576b454e0512782d94a3a

Observation 8ee742f9-e387-402d-9859-1ce46891a8d3 · outbound

This paper cites Open-vocabulary object detection using captions.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open-vocabulary object detection using captions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:43.935458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:42.909379Z digest=sha256:9a2ba15eafe702c20421e81e130e2f7cb8ddccad95f7c25c7e33b6dc7851ad19

Observation 6c86923d-74a5-40fd-8a62-21503d1aa8c2 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.949925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.949925Z digest=sha256:aad2ed7b82015962e5a189d69adeb092e7587d571924cfdb75d6456bbe836b17

Observation 0534f1ae-8262-4bb0-b7cb-0844e041a973 · outbound

This paper cites Open- vocabulary panoptic segmentation with text-to-image diffusion models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open- vocabulary panoptic segmentation with text-to-image diffusion models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:43.786937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:43.017255Z digest=sha256:e8a53b76848608adc96b7c6de90c59a17ec38d1ccf7ca65242bdd04c66bf6bd4

Observation 12105216-6a32-4a2a-8378-39d2edc7625f · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:43.064774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:43.064774Z digest=sha256:a60030c97461944a1fe632176ccfa9ce1a9faf755902c8f35ae98a907da8e902

Observation 7126fe7f-fbd2-4ecc-9a7a-d89e7ebb504e · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Lvis: A dataset for large vocabulary instance segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:43.605224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:43:43.126490Z digest=sha256:b7f6fd811e1e535ad905ed9fc9c6572ba2c22e7187e36365ffb02fa666eb538d

Pith citing papers

No inbound Pith citation observations are available.