Pith. sign in

Paper Citation Record · LEDGER

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.24025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24025 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:43:43.126490Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7fa3421-7e9e-4cd0-ad68-5d7b237a90b2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.156625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.156625Z digest=sha256:89d1e84f6b699890e265846f029447c461f70d14dcc935568cd5e15f1312a1c5

Observation b4ecf9c9-bc7b-4d08-8601-fb38be14cdfb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.261189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.261189Z digest=sha256:85a4ef8ae850887cc8bab4d5ba91ac199a2f8d5a4b03d040523c9bc9089376fc

Observation bb8a39e6-f4c2-4445-97de-695c4cf08d6d · outbound

This paper cites The Llama 3 Herd of Models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.390073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.390073Z digest=sha256:9a61e2e5a50f6013112fd75cbaffd6a42ec3735a6e3ed6906e802fee6baac414

Observation 4eb85102-92d4-4566-b328-cb88248aed08 · outbound

This paper cites Qwen Technical Report.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.493697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.493697Z digest=sha256:3be9d0f04802c0960d8009541309c73c7ed7a1b8f481eee264aa5e5128e948b5

Observation d29c3809-a7bd-40d4-a555-ef565faded41 · outbound

This paper cites Qwen2.5 Technical Report.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Qwen2.5 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.629796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.629796Z digest=sha256:18c7b36282640b0a68206d402341587629ad5bdc274cf634fb72ec102e8ad88b

Observation 33e50349-0dd1-4cca-b99c-0339394e082a · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.746391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.746391Z digest=sha256:110d6d889f8671c5a55ee0571b5d4d3cbde76359e260e043c8b2b7a16907e343

Observation a067b816-b85d-497d-8084-7155d5c6def3 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.859878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.859878Z digest=sha256:15ea93415756988b67d58e0397e421896bcdaa9a2acec41721473a46588a742d

Observation cb988395-0cfc-42b3-9717-fa3fc90fd8fd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.902531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.902531Z digest=sha256:b743ba7522ac15d4a58abf1501ddeebd8dcba33b64d9b48fd4ece386a734ca40

Observation e41cc792-baac-424c-a1fc-20f44a7e3636 · outbound

This paper cites DeepSeek-V3 Technical Report.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DeepSeek-V3 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:38.970665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:38.970665Z digest=sha256:e14089d0ec7d3f25dda1789bf75fdb25a5b10e89a70e584b9aa7de649808b4eb

Observation 0050adac-68c7-4a78-ad03-b15c7ce21f64 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.038626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.038626Z digest=sha256:2c81ddcf3b56cdd027e5a7ebc478dd93135704296070d148945380ba0c216cb4

Observation b9a6c6ab-b501-420e-8a4f-927aaaa92e05 · outbound

This paper cites Segment anything.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Segment anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.094947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.094947Z digest=sha256:b3201f892159a7cfa12232dd00f391373a71d1f67b04f00bedb1704342c9c3b5

Observation e0991557-f1d0-44cd-899e-fa00cafaf8d3 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models SAM 2: Segment Anything in Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.167283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.167283Z digest=sha256:67e3ca3d4a5f69a78225b6b4296ec1c279a77daccc81e48eba6a089a79d3059e

Observation 5c9bfed1-f09b-4e01-8ef0-5f4d300276f4 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.233737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.233737Z digest=sha256:0982e66454996902f0df38f79d8bd5b4f715fc07db4abd447f5856a7bb2d6527

Observation 66406ad9-3fa3-4b00-9c68-4f52956ca41d · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.298980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.298980Z digest=sha256:9e22630e689337b37846d8b6aeeec432af1368df477e2b3346ff6f5778662044

Observation 7f08e0fa-4254-4761-b8a1-c7ffb6ac8e98 · outbound

This paper cites DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.392644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.392644Z digest=sha256:b4080926bea58b7b1ad4c040b040e8416e77ae3c13b41139d6300747cbc7ad1f

Observation c05c271e-7514-4b73-874f-6de89913f62b · outbound

This paper cites Visual in-context prompting.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual in-context prompting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:46.051518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:39.501161Z digest=sha256:279fe4432e0726a98e85ef2c53c2b859708ae892035f24ddd417f5a634042163

Observation 9e1ad738-9420-425b-a973-25c6ab61b52d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.596806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.596806Z digest=sha256:0f75a7df5c24e34f3d5975313ca34bfb14065c455df3060a935a818a7874585c

Observation 704b60cf-de98-437a-8e4d-430091d07f43 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.690023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.690023Z digest=sha256:a7fc3c70ead85e151a991733235948432fb52e5c3942148e57428ea6e6970f10

Observation 662ab28b-6bef-4529-b622-eed32a677db4 · outbound

This paper cites Microsoft coco: Common objects in context.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Microsoft coco: Common objects in context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.743576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.743576Z digest=sha256:8e3904fc6eb31066b3323f14b9f64380bce2d55aeb11e34eb9c2cf27988e07f8

Observation a4030498-18b2-404f-801d-ebf925e00178 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Objects365: A large-scale, high-quality dataset for object detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.812688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.812688Z digest=sha256:785c1c3b4cbb0d56216d19941d868380f378a014b050a3933da5900864ce8f55

Observation ae087a71-c295-45c0-b2f2-1221703ca4cd · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.891022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.891022Z digest=sha256:72d488ff1a66ffaf6c4f1e2399318d837106ba29c99481e52c8071b1ff5fc8f7

Observation 5f9ef87f-503d-4289-b3e4-5f2af612fbb1 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models A simple framework for contrastive learning of visual representations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:39.964947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:39.964947Z digest=sha256:1dab143e2b821f08cd8b5ef4f9fd36656d4f59ca077219fe761263a64295c1b3

Observation b477893c-00b4-4249-93b3-d4d253260570 · outbound

This paper cites G-simclr: Self-supervised contrastive learning with guided projection via pseudo labelling.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models G-simclr: Self-supervised contrastive learning with guided projection via pseudo labelling

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.911668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:40.061380Z digest=sha256:1bbe6c2f3fa1f13df0ec9743ec96b2e71a48d2fb351e030672a0e742899065e9

Observation d06069fb-48f5-4f1e-ab3f-1f9a3a56e64b · outbound

This paper cites T-Rex: Counting by Visual Prompting.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models T-Rex: Counting by Visual Prompting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.179541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.179541Z digest=sha256:395d63118e2372ed4d8d907817a240c951333af520f84a13d423afc653deb64f

Observation 73714e8e-8a42-401a-9e20-194a605f915e · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.840820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:40.294747Z digest=sha256:282e43c336be828f5533016febef1d29cf986c8c66c48cf69670dc0b130081c3

Observation 9f2710a2-c5df-456b-9f9d-9ec81e47460f · outbound

This paper cites Cp-detr: Concept prompt guide detr toward stronger universal object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Cp-detr: Concept prompt guide detr toward stronger universal object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.729040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:40.365687Z digest=sha256:3de8dc4351a46a6a74309323570dd225f7a0b8da78d7a3a5ee0d772172d0a1b1

Observation 4a7b5899-47f8-49a3-81a4-9ccbbd8385d4 · outbound

This paper cites AutoVP: An Automated Visual Prompting Framework and Benchmark.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models AutoVP: An Automated Visual Prompting Framework and Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.486017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.486017Z digest=sha256:13c71af00ebac395d2755d98facad6324fcce519e93138a3f44f2c42d5f64985

Observation a2913642-ecaf-4e01-92af-966cba330f9e · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.605490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.605490Z digest=sha256:6c3237a5f8793d0a5f276a2e2971325aa63f0eccecfebb8d4d9c708575732a26

Observation d1b9627f-cdc5-49b7-b38f-e8d6f1209201 · outbound

This paper cites Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.601553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:40.702290Z digest=sha256:64bce86d712131fb23dbdba6ad543dc6baf07fbcc807ceb4f4da47193b8de787

Observation af1f007e-283e-4536-b106-f94e0420e065 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.833021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.833021Z digest=sha256:f79d764eef4f3405ca8223a6f646f77749e0d007e61649c8d31cc5443b057559

Observation 2ef018e4-7d7d-4683-a8c9-2f0605bf3126 · outbound

This paper cites Prompt engineering for zero-shot and few-shot defect detection and classification using a visual-language pretrained model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Prompt engineering for zero-shot and few-shot defect detection and classification using a visual-language pretrained model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.422370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:40.955884Z digest=sha256:1e520bfd5ef5c04444649b0949c209522e3febc41ad4f392986b3fdf009d0725

Observation 8a1757fb-0847-4a57-b776-379e6a3ef43b · outbound

This paper cites Promptcharm: Text-to-image generation through multi-modal prompting and refinement.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Promptcharm: Text-to-image generation through multi-modal prompting and refinement

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.316318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:41.022941Z digest=sha256:a3ec986f143b4a33f0527767273cf6d0b7dca5e43790c57028e1b7462e396cb2

Observation f161193d-fc2f-4a2c-aea5-812bfbab78d8 · outbound

This paper cites Prompting industrial anomaly segment with large vision-language models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Prompting industrial anomaly segment with large vision-language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.217572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:41.126318Z digest=sha256:5eb1029ce5b52e15dd0800a08b28580e6a933200fe894a063960a8ca9db836cb

Observation a40d6278-33db-482d-a494-cb1094d91504 · outbound

This paper cites Visual Prompting in Multimodal Large Language Models: A Survey.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual Prompting in Multimodal Large Language Models: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.257873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.257873Z digest=sha256:7f2e2c46cff6f8cae78bb1b39bc8cfb512c0fcb30b842cb8462f024de0be9752

Observation 2b52cc8c-8fb5-46f8-8fb8-5630793c7e8c · outbound

This paper cites Understanding and improving visual prompting: A label-mapping perspective.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Understanding and improving visual prompting: A label-mapping perspective

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:45.134917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:41.378888Z digest=sha256:59f65c1d8aef4f50446b4482f557ce0b3ae33b0436e97d3ee24129c3c7f5820a

Observation 011313bf-cbfa-42f9-ab5a-d3ee850c0512 · outbound

This paper cites Vp3d: Unleashing 2d visual prompt for text-to-3d generation.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Vp3d: Unleashing 2d visual prompt for text-to-3d generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.952477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:41.458839Z digest=sha256:f66e27be84a8dbbf4034c35b3551362639614f99bf68ca105d53dc5c2909d168

Observation b35e5bdd-f5c5-40c4-907c-40339020a79f · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.549563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.549563Z digest=sha256:9fc27180b020bb051d6770917e965075a4c741950b5b2dcaebf9df848efc03dd

Observation 1efe85d8-c42e-4f8a-a403-3af50ae9eedf · outbound

This paper cites An Open and Comprehensive Pipeline for Unified Object Grounding and Detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.692788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.692788Z digest=sha256:e4dd2f98cb176f09a851b9ea298dbca126b6f2bbed1fb43d8cc215e5aa634bf8

Observation ee2b72e2-a38b-4cb5-8332-77c425c04928 · outbound

This paper cites Learning to prompt for open- vocabulary object detection with vision-language model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Learning to prompt for open- vocabulary object detection with vision-language model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.738724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:41.798426Z digest=sha256:53481fdb97b339b063061fce418d5c296fcfda426f469a8afda4b8fe7f0f7c2d

Observation 3660d5f6-6942-4fb8-b902-dd8dce926aff · outbound

This paper cites Proximal Policy Optimization Algorithms.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.874245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.874245Z digest=sha256:e9737ab107918e8b75fd84f77a31f0c472792c66687ba6d255a7635a3fb8e516

Observation d96daeee-c489-4fa0-b0b0-53c03e340c5f · outbound

This paper cites Training language models to follow instructions with human feedback.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Training language models to follow instructions with human feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.985819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.985819Z digest=sha256:b710dd911b822c01f98014d925550ad15efa1ba4314df3aeb094299b4f6fe3e8

Observation 187d8867-3352-42a4-8935-d9094727e887 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.119299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.119299Z digest=sha256:fc243ebe667e7b3d18e461b6b5e9bbe4b2e6c53e8d7bd9bb4e9a38ac668d7c8e

Observation dd638852-dbbf-4862-ae9d-ad64d68c96b5 · outbound

This paper cites Learning transferable visual models from natural language supervision.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.236006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.236006Z digest=sha256:488168988f54e9d25f7d217ffd0f657b4d122020b3cee4f7a18f03bf2851d257

Observation 5d3c849b-b1d9-4e16-a256-ce0158fff4bb · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Florence: A New Foundation Model for Computer Vision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.302476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.302476Z digest=sha256:1144f71a2bd4aa7bb7d3ede11aa35e6a55b34a21ed07041017ccc8d2d55d62fd

Observation be4cbdb3-a54b-421d-bb99-570a88d2ebf7 · outbound

This paper cites Grounded language-image pre-training.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Grounded language-image pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.362469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.362469Z digest=sha256:de368e2b3e27a19e268600ee39d7a42fdf3e9ad2f470eacb429911f955fc7134

Observation 3f019d7f-0cb2-411a-8a18-fd4273d538d0 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.418637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.418637Z digest=sha256:be87e26a3cd09924d5f1adeb7598931ffd3f71be7392e4b34147ae602b340915

Observation 4e5b519a-8bbb-4317-ac31-80c6cddfb1cc · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Swin transformer: Hierarchical vision transformer using shifted windows

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.491065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.491065Z digest=sha256:a07bc25982ba07f9ecd78e40d178c8e626967e793ed347e3330abcd033459127

Observation 3e583b6b-35c1-4fc9-a44a-a84b6bbd4f35 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.554071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.554071Z digest=sha256:cc570030ade9411910c15c2242c6bdd5bab41625dd52d814b9a4a18dc148f6fc

Observation 0325b4b6-57f1-4fb1-b1e2-5125694fc982 · outbound

This paper cites Yolo-world: Real- time open-vocabulary object detection.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Yolo-world: Real- time open-vocabulary object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.581367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:42.632729Z digest=sha256:f9446201d783179fd2e69519d97750ff96dcdbc76a5046aca96937b706ad4dad

Observation beab8506-1c34-412b-8baa-9a6374c7213c · outbound

This paper cites End-to-end object detection with transformers.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models End-to-end object detection with transformers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.490030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:42.690710Z digest=sha256:765bac154f681cae02c45bfb1f9f934959d1632f76a8c97d606f30ff692bbad4

Observation ce2b969e-c6cb-44f4-9c79-594a261c9621 · outbound

This paper cites Dn-detr: Accelerate detr training by introducing query denoising.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Dn-detr: Accelerate detr training by introducing query denoising

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.346485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:42.749116Z digest=sha256:f5c5f578729425cf7943cfbbf86a1764ce964e0a0bf96b7689fdb5e0e80548aa

Observation bab7b3c6-915c-4f98-ab1c-60e8a3c80c80 · outbound

This paper cites DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.812679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.812679Z digest=sha256:a57016351202e2e1e583fd2f641aebf22c899d79c0975422f25b5aa1f8d57796

Observation e6b29265-bd35-49d5-9570-0c7116e77fe6 · outbound

This paper cites Open-vocabulary detr with conditional matching.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open-vocabulary detr with conditional matching

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:44.203328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:42.850234Z digest=sha256:474a08b5b7aadbe70a0de7a3972703424fe267072dc3492406d1d7c3cfea80c1

Observation 8ee742f9-e387-402d-9859-1ce46891a8d3 · outbound

This paper cites Open-vocabulary object detection using captions.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open-vocabulary object detection using captions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:43.935458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:42.909379Z digest=sha256:85a399523ee468cccbabdd255f4c3a7505b66c8e824890e88e603d37925afc23

Observation 6c86923d-74a5-40fd-8a62-21503d1aa8c2 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:42.949925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:42.949925Z digest=sha256:51028b87e08a5a6c12a7be23193adce59e46e573d2ec79ecfa3a4e72f8533066

Observation 0534f1ae-8262-4bb0-b7cb-0844e041a973 · outbound

This paper cites Open- vocabulary panoptic segmentation with text-to-image diffusion models.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Open- vocabulary panoptic segmentation with text-to-image diffusion models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:43.786937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:43.017255Z digest=sha256:e3bdada887cb5e49b16a8245f4bff7e8feb52fc83c0c0fe1c2303d62e1b9943d

Observation 12105216-6a32-4a2a-8378-39d2edc7625f · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:43.064774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:43.064774Z digest=sha256:31297585ab13d77de17c940e2228e0a295e2ae81e16f708390dfcb70ff7c1c50

Observation 7126fe7f-fbd2-4ecc-9a7a-d89e7ebb504e · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Lvis: A dataset for large vocabulary instance segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:43:43.605224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:43:43.126490Z digest=sha256:ee7d33bf8109e364c74dd1f43eacb6168de4f47f71ef1a1c2592c0d43d3d532a

Pith citing papers

No inbound Pith citation observations are available.