Pith. sign in

Paper Citation Record · LEDGER

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 7 inbound Pith citation observations for arXiv:2506.04277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04277 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:08:47.365596Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:10:17.325958Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:47:30.631926Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5f23e84-ba59-44de-bd41-5be8ded3fc81 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Hallucination of Multimodal Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:40.402070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:40.402070Z digest=sha256:0199b761731d48a04fe0c16e5f19a22945676df006bdc8af05653f004e6d97d1

Observation ba55ffd2-d882-44ca-85b4-7107ee658769 · outbound

This paper cites Reframe Anything: LLM Agent for Open World Video Reframing.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Reframe Anything: LLM Agent for Open World Video Reframing

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:08:48.258624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.474413Z digest=sha256:aa50901a59054c7f1c93b94444b43b444f676153c5b65177a43ffb09ffe72b96

Observation af604303-f121-4e89-bf79-d0a64c592fa4 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:40.556277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:40.556277Z digest=sha256:961ff3658dc20bb6938c07c5979d726a698be1b2d9e51ca6ad1537fd27d3e416

Observation c4bbb67e-1a05-4f59-a731-b4446a15aba3 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.729787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.681671Z digest=sha256:3959a3cbf68fe516104407b25f20fdbe1222689d3388d5f7c7411c3b52e5df7c

Observation e46b9ce0-3507-42ed-9fa0-5445fb7a13ac · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.514876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.787454Z digest=sha256:be5bcad39a7e92cfdd5b928359032ec95a87d4bb5f451dd52811fd61fdd9640c

Observation 2ec218c6-b22e-4362-881e-4b7d6aff0071 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.264144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.984131Z digest=sha256:d52cd7470b1e92901a07640df29f5974dbb7e8cf8e23d2b47caba09e1142cccf

Observation 69a6d90d-2fa3-43ed-9d7c-6d693f9152a7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.110594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.110594Z digest=sha256:689cd4f6f2ee333e5254ed8ab154e4a2bb936fc544de261686ac2a1d4a93ee90

Observation f75fec81-34f7-44c8-b523-8cb03ab78cd9 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.046756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:41.266042Z digest=sha256:025af108509a0f12c93aaceca75fa86499f4801f430fc4226bf41e3c59d85efa

Observation 332913ec-3213-4238-b852-00678b9d6229 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross B.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:08:53.818075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:41.417931Z digest=sha256:17726583ea1c3305f2d987a43f5604a566c680793608dc132f1c6aa64610a041

Observation 6ca3f72d-c73e-40b3-a3bc-0b126bcfcb13 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Gonzalez, Hao Zhang, and Ion Stoica

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.567535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.567535Z digest=sha256:bc8734788d0adf632ca8b0b8b1d511e523bb8ba48bdc104f0f832daadfb0252d

Observation b7ba99d1-bc73-4c31-8fa0-5837d3cb84a8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.796732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.796732Z digest=sha256:138a37fb0292e850f39920bbb7c3b27f5f6db70d5a5e683b035baddf218b1f0d

Observation 86bf2cb7-94da-403f-8ed4-05db4472013b · outbound

This paper cites VEU-Bench: Towards Comprehensive Understanding of Video Editing.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought VEU-Bench: Towards Comprehensive Understanding of Video Editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.965687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.965687Z digest=sha256:de4685fc5f18ed6ee5bd3b021164ed2ff15d941673bc9f97eb9e70c12a288f3d

Observation 8aa6eeb3-2245-4d99-8f1c-4cef689c819a · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.177990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.177990Z digest=sha256:7d6a61c42605e383aaf2d92b489c4cff5700d9c1a714fdf9b57e0f30998e9dce

Observation e32e10fc-095e-4f50-a471-72f556a82b9a · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.515400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:42.377897Z digest=sha256:962436d02be4d04f2b59f07a4ca8c59e7d554a007f08b5f5a4c89de87b90527b

Observation a1aa0c70-a29b-4b96-b5e4-a09febec9bff · outbound

This paper cites Visual Instruction Tuning.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.538130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.538130Z digest=sha256:bd3c9483b4db74f5c2cdcdf8d7a9db1b91cdd31c28338dc3b355452213e99d9c

Observation 0b696a43-dd1a-4874-b45b-28cd24fca905 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.665200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.665200Z digest=sha256:024bcfaaf080f0a1d2277b16ffbfc9c475ce3cec81079efe7e8d3859974c4555

Observation 02ce7b08-8488-445d-88ce-899b1600cbcb · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.819240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.819240Z digest=sha256:c0e7d10222a03d86e2a2d0b6b710ebae82b380a6fbe5f6675f6205a405570595

Observation d78528e8-4679-4c96-9299-0e685a999408 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.235178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.047540Z digest=sha256:914a7a5b3d84c2c85e53ebc1de5ee44b36ff8ee074e562f5bdbdd8f028a0ee55

Observation 38c7cbb9-2171-4374-95fe-3453f32c5ea8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.008973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.249635Z digest=sha256:a3412de3276df48465fd5565f08727e73a0b438d3740ed9b43b75c6a65dcf139

Observation 1086a62e-8c4b-4569-b9c5-8c0f98f29315 · outbound

This paper cites GPT-4 Technical Report.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:43.409280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:43.409280Z digest=sha256:5c689754ddbac7b542efe4f8bf2f97aa0cf1d530452c7e8f7fb3e5b07d000a48

Observation 07bd95a2-0b61-41ef-8f6f-ef55e92d43b4 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:43.637508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:43.637508Z digest=sha256:bfdfd9f6cbfa6e5d1f4e9a049f8ca73da0fe00c59c5756bf9e96f92fcff85dc9

Observation 59c95392-210b-4abb-a5dc-0851f2c3d032 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.759769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.844069Z digest=sha256:3e644a7f13210be91c6c687372d6684f66cb333dc4467cb3af152a54327309bf

Observation c8cc211c-6e79-45ad-921a-179e08be9277 · outbound

This paper cites TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.046424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.046424Z digest=sha256:73f9b411b2c7aeebec6f0c0e503da761231bf55d53dc26a974a284cd5d7b9bde

Observation d361bb9e-de9f-4f66-8159-23c9defc968d · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.450294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.100873Z digest=sha256:d29848bafefe14cde0077924503b7427d4b5b78fdcb6648d824db71c0d75c785

Observation 666dbb4a-c7d0-4f11-875b-1b49aa22f766 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.190797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.190797Z digest=sha256:ad3b1b16ad6a22eaa801fbd4683c099c880b10552107c52556d803954d13902c

Observation bc1cc798-6871-4981-ba66-7bedd1d77ff8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.265401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.257827Z digest=sha256:9bd59d8984726a137a9a10b63172dfaf3b9c8c0f0c38a708f5b9bbfda9add5f6

Observation 28d2689c-4b13-477e-b09d-a028170a9f29 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.329476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.329476Z digest=sha256:b68bc28c053eb1900a36ab1be2ffc111e249ca28e93075f5a1f29d46b2c71215

Observation f553a1ed-58f5-4e64-ade8-c85ee0a09b8d · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.024786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.414812Z digest=sha256:11c774675463ca0e199f315478e3d87356bea3638995f2fd947180453e71b754

Observation 37da8636-f985-4409-96fb-2f399c848057 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.516389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.516389Z digest=sha256:a50666ce8f07d4b8f679aee13c9cc5361d1c8a4f632a626a7cc3ca3c3d5ae2b0

Observation 52cb0a5b-b851-4a36-955c-2efaf3e4ed51 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.644702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.644702Z digest=sha256:7ca405de5871db337900bcfa3d7e584e11915bad20868a578eb51cdeb5717a33

Observation 1892add0-d767-490e-97af-c7f2addd3e02 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.782181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.764286Z digest=sha256:5d731331de6d19b091b2c964d320253caeb02f673f19d94e2e3a9a72890fc449

Observation e36f32f9-6776-45a0-ae8e-a42c8fc583af · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.852833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.852833Z digest=sha256:6c7d73e2b5a799463d370ab26e165ef5c98de7d97758d40e0e0516a185c1a519

Observation 65624622-d2fc-4af3-a3c5-ff048ec6b058 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.565953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.994367Z digest=sha256:d22c400cf99057f2709c1f9bc898d9c53cb2b44013aca8349c9e392ca2c90f6c

Observation 07a56731-bf64-4a31-adb2-3fa73e7968cd · outbound

This paper cites Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.066926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.066926Z digest=sha256:6f07de84191c62dda92fef4a3fa4feacedc89f893449552411fffc3e3ced1c38

Observation 6d5f953e-df5d-464b-891d-3f3984db2244 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.197404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.197404Z digest=sha256:b0dff90d7a22022540496f5fac4e9af54c8c72197176931d02bfc9efbfa9daf6

Observation 9aa33499-ec22-44ca-8371-05a812bd4fe2 · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.273322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.273322Z digest=sha256:cbbb1a31a3b8501892e60aa8dd28df37f852f18539e99f378ca425f63a4b7a0b

Observation b54b7879-cd9e-4278-a8fd-3b4d1c1f13f1 · outbound

This paper cites DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.339504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.339504Z digest=sha256:da222679dc664695381d609757db5e3fd16b954c157ff07afd5361ff48a63cd5

Observation e62bb179-dca9-4451-af74-35c82c258d88 · outbound

This paper cites Number it: Temporal Grounding Videos like Flipping Manga.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Number it: Temporal Grounding Videos like Flipping Manga

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.397760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.397760Z digest=sha256:5087089c3b8741939467cbf9971f3aa495a1d5f37b460d930cd7dcd321fd5490

Observation 0cdca2a2-1491-42a1-84a2-f137d4b60ea8 · outbound

This paper cites Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:08:47.792006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.459689Z digest=sha256:0cc850fc94542b180ee055d3ed76eb81a4e204d7a28b6f63e77b3a36bcf5b280

Observation 51cb02f5-cad9-44dd-bbc9-01845a6243e2 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.301911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.523531Z digest=sha256:16f839413fc37436ee360117bc7cda45c269fbb1a4ecaca85fd937ff7cc652fa

Observation a2e3dec0-6585-4e5c-8a44-a33bd8395fcc · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.086971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.602039Z digest=sha256:a502dfb4c0665b13b653fb1faab1cda949fb8b8e6f33de549949b7e8d6b87d2a

Observation 7d3edb90-745d-4e70-ad13-949b3b486d8d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.712931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.712931Z digest=sha256:8211065e9323da1d307e4eb9fd41708dbf3fee59e9947c5aced6bfb979bbfe48

Observation a5ccdc74-a977-4c00-bfe6-279859639c39 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.838207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.811259Z digest=sha256:8df754e0df3c03a1528bb62a38de0e70132182a5d57522bd1eb18849cc31bc28

Observation 6e8370fa-e4b9-496a-83dd-5abf89a01f1a · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.591962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.874839Z digest=sha256:9add84d082a45451ced0bcae8d396251a313f80b2f8f73781dddd98205790d77

Observation 74d27c5d-3e66-4841-823a-d5a35c83a16f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.368405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.955480Z digest=sha256:06b1e8d68566282b6f2c9716d0b3e074f020a56e144d4b0ccaf9824de2c0ba25

Observation f6460ecc-2516-46b6-8951-e023a26d0931 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.024434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.024434Z digest=sha256:2a2780720d36145b0bcbe6fb5389e581cf6871e63beb9b2c965b5777ecd13851

Observation ce95bdd1-bec3-4e60-a458-dfa9cc0aaec8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.131375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.080861Z digest=sha256:832064d07b714e91b5fb8d36a913df41a130f6674b82882914535c6d0ae3c73c

Observation 1af5bb6a-a3d7-4a4d-b806-8c0e4ecfae47 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.859512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.216434Z digest=sha256:5ed09b0a9395df4f266bda6438f98b946407d101cc7ca0e567dc5a511081317c

Observation 7f8f49d4-b716-4bba-9456-e474910fbc0c · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.617597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.357132Z digest=sha256:198dc8302dd3f8a305321201b2c14bb888df00f0ed100c71c2f840f0de1fa8d9

Observation 30762b84-bfb6-40ed-9b56-44f745255e11 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.407536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.504809Z digest=sha256:f4b765414083f10f47dff0ef0940bf94228561a951a7364d84cd5f3d7e2237dd

Observation b0b74032-a9cf-4e96-b770-10ae30cdb709 · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.629692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.629692Z digest=sha256:cf483702c76830e17f3f1cc5d65047d58f5f720ced8cc0c824814cd0a7db15ea

Observation ad0ceb4c-2c7f-4b19-9b02-a574aa1b35a1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.785107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.785107Z digest=sha256:dbc93b175eac83bdafad0e964cf8f01e317c9308ea617f12a0314bee7fb4e774

Observation 30acb137-6da1-4085-9efb-e2691c641a7f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.214546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.880493Z digest=sha256:270ece4f71eeecc01f82bb1dc654dab98a09f9e6e67d6966b11ce7bc6e840489

Observation 3f162f72-5613-42d6-a8d9-85c338a50d2f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:48.965251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:47.023203Z digest=sha256:7215269ffdfc48f077f20b13bf720f8c4ef8a0b3fc584708b754a717aeafdf8c

Observation 584af511-db2a-4278-891a-eaf7fb72764c · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:48.633512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:08:47.113216Z digest=sha256:0643a9b2c431ea8c9f616638f1e527092a3ee849d16b46b44b827f2093ba505d

Observation 45e7b5e3-81bc-4183-babd-a1581b8e9909 · outbound

This paper cites online" 'onlinestring :=.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought online" 'onlinestring :=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:47.214335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:47.214335Z digest=sha256:529a0a950c17d905b7d0717a8edd00348fd114fff037ce9b97bc003e7e611b9e

Observation b80bce5c-b602-4db9-862e-9dbc50ccd111 · outbound

This paper cites write newline.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:47.365596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:47.365596Z digest=sha256:cda9300aa3af8a58e4e7eea758b01b99de77f8b8888cee2183e71eb2353d0645

Pith citing papers

Observation 0ffd6ff2-f54f-4d8a-9d37-5055ffd9f570 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.519580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:0f395a4dfef5f33bdf4c63ca840012c44dda4ceedf69f38e0a060d172840cfed

Observation 49641d1b-4775-40c6-a75b-a7368f48152b · inbound

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding cites this paper.

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:17.325958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:17.325958Z digest=sha256:1fb264a5228559678437f1f182ec11828073e9b8635ed52469fda4ad5eaeefc1

Observation f7e37821-8e46-41e3-ba15-3eaf6e4791dc · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.528679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:626eecaed0289f2c491e3fb1be608becd24c878a48ffe14583fdd67b0b5072bb

Observation f33bb049-8095-4067-98f4-d6b7bc5feac5 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.098104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:284a293cf647b1ca552559936215db9860074818a7c16f8bea8b3d698f0db0f8

Observation f5f8ffce-3e0a-4b13-a6cb-a408093d7064 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:07.204123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:07.204123Z digest=sha256:68ba96ed674d4af6eb66b899e408f8f8f273a7f0e4465c1ebdceba82e4ac4aa1

Observation 97b80294-61cb-429e-b5e1-b3d601a20faa · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.431402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:cf95d9042753493be6f360955bcd2a39403b402d0159e38c4546bdf775cfad0e

Observation 5e176c91-e057-457a-99bc-1fa61be974c2 · inbound

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning cites this paper.

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.633469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T17:01:13.745646Z digest=sha256:1a5b78af77c7627eafa3a86db84e235731ae811c76e36fa1211968524f7eaa0b