Pith. sign in

Paper Citation Record · LEDGER

Reasoning Segmentation for Images and Videos: A Survey

As of 7 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 5 inbound Pith citation observations for arXiv:2505.18816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18816 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:18.789405Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:06:22.967237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.361578Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2535e581-f34c-49aa-a4cd-90ff023a2e52 · outbound

This paper cites UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model.

Reasoning Segmentation for Images and Videos: A Survey UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:27:20.276725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:11.018222Z digest=sha256:1b71e0075935fde7be28133c0f3bfff5a74e060aef02fa443f1a80287a09b256

Observation 5104d3b1-8a63-48c0-bd05-b019af82bf66 · outbound

This paper cites ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation.

Reasoning Segmentation for Images and Videos: A Survey ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.072412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.072412Z digest=sha256:eaeb23922a32c665b4ead73409455654e0baa503f07576c46835adfda55cd298

Observation 5628c86b-4642-451e-80ca-aa9b60782db4 · outbound

This paper cites Burst: A benchmark for unifying object recognition, 39 segmentation and tracking in video, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Burst: A benchmark for unifying object recognition, 39 segmentation and tracking in video, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.134478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.134478Z digest=sha256:9767ebcdfd0e420565a3dbba96718cc831763e856ed4c1eff4c40e5a84ad3ab4

Observation 1596bf16-837f-4dc4-bc2f-653d81fd4952 · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Reasoning Segmentation for Images and Videos: A Survey One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.218146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.218146Z digest=sha256:9d3de62eb2e7a54b893200e550fadb9381944089f88cb9dfaeb9856d2f6df671

Observation 98bc1725-92aa-476a-a66b-6da70b393a0b · outbound

This paper cites Cores: Orchestrating the dance of reasoning and segmentation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Cores: Orchestrating the dance of reasoning and segmentation, in: European Conference on Computer Vision, Springer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.273826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.273826Z digest=sha256:3ff0e077247b12ccefd5cf3bba193733c345713c5f35dd2ad3665fc41219c293

Observation f2799ce2-0ec9-4498-aadb-2f0dd8d18212 · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Xmem++: Production-level video segmentation from few annotated frames, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.329047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.329047Z digest=sha256:62970f562fa168ea3ed0c61e1cd751824d20a4ff1442c74cc66964d3646e1306

Observation 3e7199cf-1e6b-407e-b00d-f1c96ac9424c · outbound

This paper cites Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.431803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.431803Z digest=sha256:278990024c08d486803534381735749eeb1dabbcd0851ec7a2553d9912b18063

Observation 54b9c109-2a0d-4591-b8e9-1cb2d4b7aada · outbound

This paper cites Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.489319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.489319Z digest=sha256:18709513a5719b48d29c50b7d25d6c0d693732f93f68ffefc969f704124bc9bc

Observation f732ecc8-2281-48dc-9fa4-6ee936b81096 · outbound

This paper cites Pixel-Level Reasoning Segmentation via Multi-turn Conversations.

Reasoning Segmentation for Images and Videos: A Survey Pixel-Level Reasoning Segmentation via Multi-turn Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.587733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.587733Z digest=sha256:a8c493207df54c57f513620e6c510dfb66b4e491e1a9cd6e38a9ec5001c4f165

Observation a2672107-2874-4dd1-9b28-4eb09fc36747 · outbound

This paper cites End-to-end object detection with transformers, in: European conference on computer vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey End-to-end object detection with transformers, in: European conference on computer vision, Springer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.653276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.653276Z digest=sha256:6803182d0d2a3bd281bd71cf81f3900de2ec08dd2c1643521db3dd821d0b9a0d

Observation 9c17b8da-d672-4c2d-8dcf-003f75f2aa62 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.754067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.754067Z digest=sha256:9c9a2504cc13897c3e585b6b8812fc5521d38ed57a459dd2a89e015c557bd73f

Observation 0cf067c6-6e1d-49e2-b69c-d469484d8642 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.868491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.868491Z digest=sha256:df6d4340df7e1b393e53f3f875696fb462786527f0fa4a6ad3339099cdc0d510

Observation 0023a925-9097-4578-b971-51f60d688532 · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring ex- pression segmentation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Sam4mllm: Enhance multi-modal large language model for referring ex- pression segmentation, in: European Conference on Computer Vision, Springer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.942872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.942872Z digest=sha256:a037aaf677e6968443a952b0e6493b01405152c75f3d1b8b22a70d8ff9470017

Observation 4250dca0-38ea-4da1-bb92-80a94ba49762 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Masked-attention mask transformer for universal image segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.987613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.987613Z digest=sha256:9fedc7cdca8361a6972eab4afa6176578d7fd9319973618fc0d49068cf9cfdb0

Observation 1ed77509-9ee5-4d18-adb0-fd733b49b0dc · outbound

This paper cites Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model, in: European Confer- ence on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model, in: European Confer- ence on Computer Vision, Springer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.062712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.062712Z digest=sha256:7f646ad1ead1552de1eca6f6d9fe5e0322613881496bb79ec7e0009f2afcd2a0

Observation 47fd545a-b396-4222-bb23-690101e7ba9a · outbound

This paper cites Vocabulary-free image classification.

Reasoning Segmentation for Images and Videos: A Survey Vocabulary-free image classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.140729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.140729Z digest=sha256:819a89652b71f30b28140753d02145f060fb1052d41c5df40d79b58a261749eb

Observation cf1d2c9b-d29f-4510-bfab-118019625441 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.225657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.225657Z digest=sha256:8452c6f4925e753a3011807184da1013905a50859f928e4ea970132d96c31d40

Observation e90944a5-7290-4d20-87f3-5469fca853cc · outbound

This paper cites Tao: A large-scale benchmark for tracking any object, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, Springer.

Reasoning Segmentation for Images and Videos: A Survey Tao: A large-scale benchmark for tracking any object, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, Springer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.305165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.305165Z digest=sha256:0efcaea0c1d172b7ca7171897f9cbccb235e9ee6f579589cfdc0ea73a5ffb889

Observation 82302091-de9e-46e9-bae2-48fc6dbb0626 · outbound

This paper cites Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level.

Reasoning Segmentation for Images and Videos: A Survey Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.401764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.401764Z digest=sha256:3bed1602d7587d507479b9d9cba032d2e29e65d9c74948986c30714c13ac5b44

Observation a4b13401-579b-4aad-ac1a-9744ffa96064 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Mevis: A large-scale benchmark for video segmentation with motion expressions, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.500258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.500258Z digest=sha256:abd8213544860c4ab2b2b5ec8a75ea914b157b6b5da329215545b90087b73baf

Observation 783e111d-b744-4e99-b275-4ac4768809cb · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Mose: A new dataset for video object segmentation in complex scenes, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.592565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.592565Z digest=sha256:c683ca1f5f0c014ad7e27e7cde7496da5b4bfd014330b4062e82e50643678ce6

Observation 579de55c-4d1f-4f10-b152-c97f38c742c9 · outbound

This paper cites Panoptic Segmentation: A Review.

Reasoning Segmentation for Images and Videos: A Survey Panoptic Segmentation: A Review

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:27:20.231845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:12.659787Z digest=sha256:67d3396f0967917f8fe9fd5ff50af14825bbfed59744602b3be691838bf4cecb

Observation bcf025c8-490b-4b28-b496-4f577043241f · outbound

This paper cites Oops! predicting unintentional action in video, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Oops! predicting unintentional action in video, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.718042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.718042Z digest=sha256:0eb9bca000d5f425ce2b136a2f481206707326bc2b81b94fda6cae312553aaa3

Observation 56a0ce34-9da4-4754-8128-2335acf07537 · outbound

This paper cites A Survey for Foundation Models in Autonomous Driving.

Reasoning Segmentation for Images and Videos: A Survey A Survey for Foundation Models in Autonomous Driving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.818555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.818555Z digest=sha256:e86934a3f1caabce3a492f3227b71366961505fec1dc24a0469f0624fc4dc33a

Observation abe7de4f-a2cf-4111-ac9e-f8fe5fd37ed8 · outbound

This paper cites Part- aware panoptic segmentation, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Part- aware panoptic segmentation, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pp

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.927968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.927968Z digest=sha256:25be33ab7f13234283a93ed1d9633d0e5cfa04461178e9f0bbb2d7ee389d8afb

Observation 770e0673-4881-4b20-abde-e9f044696652 · outbound

This paper cites The Devil is in Temporal Token: High Quality Video Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.039735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.039735Z digest=sha256:6297dd7145ad28533be7ef0f8c934ebb419e0ff7d52049605ba5b22951260d57

Observation d3443753-da5e-48db-b968-d39eaebd1a60 · outbound

This paper cites The iapr tc-12 benchmark: A new evaluation resource for visual information systems, in: International workshop ontoImage, pp.

Reasoning Segmentation for Images and Videos: A Survey The iapr tc-12 benchmark: A new evaluation resource for visual information systems, in: International workshop ontoImage, pp

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.120490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.120490Z digest=sha256:55429efb3dad44fde6430fcb434e84911d394f2ded44bc84e76a335009a971c1

Observation b0496995-8774-403f-a4dd-cde1f9e7a271 · outbound

This paper cites Lvis: A dataset for large vocabu- lary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lvis: A dataset for large vocabu- lary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.222976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.222976Z digest=sha256:e790278eedd46093716ef67c31523989a7d98606f4495eb89cb6109f400dbf80

Observation 3b31db6e-4004-49a0-9ed4-73dddcf0523d · outbound

This paper cites A survey on instance segmentation: state of the art.

Reasoning Segmentation for Images and Videos: A Survey A survey on instance segmentation: state of the art

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.329467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.329467Z digest=sha256:e0ac920a1bbe9bc317abfcf931d6bb52915f93dcf93e0b77817e485c1ec0c8a2

Observation 99f2b6af-40a8-4263-8878-38917cb669dd · outbound

This paper cites A brief survey on semantic segmentation with deep learning.

Reasoning Segmentation for Images and Videos: A Survey A brief survey on semantic segmentation with deep learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.453671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.453671Z digest=sha256:5db4c9fb4ddb3d138fafbd55fa67f4178d26395fd39e0cb1c6eb4a61f9a83389

Observation c2a94a13-f9d0-4335-aa48-30e7e1af3941 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey LoRA: Low-Rank Adaptation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.560190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.560190Z digest=sha256:e7e30254402c1d098c1b96847c50cacf14e1c6df2bdaf18eb5698a4236883082

Observation 8ff5aa0d-221e-4702-92d0-32ff4c5ba8a4 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.682259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.682259Z digest=sha256:4afabd5466f054e8bf1cd33536f406ffccc4d4290d43808f7d3546eccdb79ee9

Observation d3d0aa40-7f13-4279-b0aa-637e14b53d26 · outbound

This paper cites MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.729538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.729538Z digest=sha256:6ed16ef44aecf414552e9540dcdcb192ef258fe39f2226adcc91d2321ba497e5

Observation 253e3da3-74e8-4901-9f12-5f9800cc4d08 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp.

Reasoning Segmentation for Images and Videos: A Survey Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.814778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.814778Z digest=sha256:586fd25b8556c9b1ca6ff608736a264406e7d1e942997d1299a220c3ed24e8dc

Observation e889dbfc-6c15-40ee-b6c1-ad95c4e89f14 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.880321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.880321Z digest=sha256:337db890e37a0fad2d47f50b040e827cad8d73009a159452a8c66ec807bd8426

Observation 21a0058f-f297-4cf8-8f18-d9b02ba6c83c · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.033054Z digest=sha256:73ede327c12a5d14f16da2deb99a6e09b18c9130d250c6a18e8f7e834f4b86a8

Observation 9bdf61cd-b914-4ae6-9b78-ede9ac9bb40a · outbound

This paper cites Segment Anything.

Reasoning Segmentation for Images and Videos: A Survey Segment Anything

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.085972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.085972Z digest=sha256:d3bb70c64d779b5a0b77770a9e127669a0e74337acb5f20165be63513243ceb3

Observation 35d9f209-ff49-4b67-aded-df2dbded454b · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annota- tions.

Reasoning Segmentation for Images and Videos: A Survey Visual genome: Connecting language and vision using crowdsourced dense image annota- tions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.115774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.115774Z digest=sha256:6343d7be093eb9e2fc83aec8886bd06594a2435be0d7cb5e2c3e25ff43c6525c

Observation 07e0de33-8137-44ee-ba4c-eb01420f4378 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Reasoning Segmentation for Images and Videos: A Survey The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.277333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.277333Z digest=sha256:44e0a9b2e71ddcc42e7590cecc4f402eb354fc2ef46d271f0fc29e7e201b3e61

Observation c4177e72-78d3-4531-aaf8-79acb0d3d0d0 · outbound

This paper cites Lisa: Reasoning segmentation via large language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lisa: Reasoning segmentation via large language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.380393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.380393Z digest=sha256:039fc13e305b4ad2074d33335fa6da79dc325e7908b635b031e1c5bd08e12c52

Observation 31641866-bf98-47f1-9203-f901a609f30f · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els, in: International conference on machine learning, PMLR.

Reasoning Segmentation for Images and Videos: A Survey Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els, in: International conference on machine learning, PMLR

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.526975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.526975Z digest=sha256:1b341b29c359296015c27bc4a8dfdebd1ac29ad79e60375ef795d193dbb21158

Observation a8b427f1-981a-4b88-9931-54c2b2ae607d · outbound

This paper cites SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model.

Reasoning Segmentation for Images and Videos: A Survey SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.624387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.624387Z digest=sha256:16a07f4e0bca1586bef974c8961b7debe30726b0aa27a4d079e0875be3ab2a37

Observation b28ed6f8-e70a-43ca-aa56-40198daea904 · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Robust referring video object segmentation with cyclic structural consensus, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.796318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.796318Z digest=sha256:f4e22cc4d6608574a8f43981973a17724a8e7ef44dca44150ee4e0dfca8a4b99

Observation ebc5c99f-5a2a-488f-b500-da0ab41c0c1b · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Llama-vid: An image is worth 2 tokens in large language models, in: European Conference on Computer Vision, Springer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.946198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.946198Z digest=sha256:fb82edf5ca4ba879d98a527d857ae97f1847beaedcd1832a10a11268f11e2de0

Observation 3b393a59-01d6-491b-9e1f-1e25d2744064 · outbound

This paper cites Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Confer- ence, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer.

Reasoning Segmentation for Images and Videos: A Survey Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Confer- ence, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.661879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:15.084904Z digest=sha256:2aa4c9c3dc7fa83aca61d544337990e422dd2e6646d4b2a62d8feae4374db7f9

Observation 9b0163e8-9b0b-48eb-ba6a-16065bb2e529 · outbound

This paper cites Gres: Generalized referring expression segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Gres: Generalized referring expression segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.653225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:15.241954Z digest=sha256:6bf6158fb12c79631a0c8fc2a2b2c03ea4c47754c0dba7303a24262cf59554bb

Observation 09a281c4-b093-437f-84cf-ce514a7497c9 · outbound

This paper cites Visual instruction tuning.

Reasoning Segmentation for Images and Videos: A Survey Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.643861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:15.352155Z digest=sha256:6ece105c995462426041890f40427efecc8eae8863b6b80660a3d1306bf0201c

Observation 17fd4db2-02f2-476e-a215-7c8d40dae819 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Reasoning Segmentation for Images and Videos: A Survey Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:15.482347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:15.482347Z digest=sha256:dffbdf764c825bca3df807eaa590ac7ddea84796e4b50e8dc1c95d1c57fc430f

Observation af282039-22b4-42fe-8e33-9d88f65ff8df · outbound

This paper cites Ground abstract structure concepts of scaffolding systems for automatic compliance check- ing based on reasoning segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ground abstract structure concepts of scaffolding systems for automatic compliance check- ing based on reasoning segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.635044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:15.625060Z digest=sha256:e24277209d0b2c6907a14ba22fc8f79345819928732c26ccb40f9ef34d16c99b

Observation e6c3dddf-c8cf-49f0-8e4d-304427d957a7 · outbound

This paper cites Video anomaly detection and explanation via large language models.

Reasoning Segmentation for Images and Videos: A Survey Video anomaly detection and explanation via large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.627054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:15.763488Z digest=sha256:dfcbc5f67a2fa95ef861cfd95c82051a3d124ece4f560ab61959678702db94e9

Observation dc0c66bc-dbca-49f9-9a4f-6239e505fd84 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Reasoning Segmentation for Images and Videos: A Survey Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:15.874035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:15.874035Z digest=sha256:d8704616451ea40dff81979cb823e90929c65e4a03cdb09dc86492b6ff7659d6

Observation 41bdf037-fc82-43e5-b1cd-8f4d9b98bd02 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.618602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.015561Z digest=sha256:9f608d148bb1358f647298a1e952127ffe3911620973cdaa1ac8f96370834988

Observation 4552696c-2c40-4fcf-8ca9-645ee71e25db · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Large-scale video panoptic segmentation in the wild: A benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.603088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.115719Z digest=sha256:20b829fd0d732c6b69ca89d1aa9d4012abfb4af841feb154e59e38538c4805f1

Observation d7785866-13af-4b67-a815-1f83c89251f6 · outbound

This paper cites Image segmentation using deep learning: A survey.

Reasoning Segmentation for Images and Videos: A Survey Image segmentation using deep learning: A survey

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.594634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.159160Z digest=sha256:033c1138509c4b7e52e5fa6c59287a5e2f61dcce8d653214a16f9f56befa7a07

Observation e0d0e917-465d-46e7-9e8d-6def35c33a95 · outbound

This paper cites GPT-4 Technical Report.

Reasoning Segmentation for Images and Videos: A Survey GPT-4 Technical Report

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.586466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.209495Z digest=sha256:759bd82a0f5db37b950708a33b06b2bed30461a6850d543d39facfe0939312ad

Observation 1d4d6a45-bf66-432c-8b22-6c4f1b2ab48f · outbound

This paper cites GPT-4V(ision) System Card.

Reasoning Segmentation for Images and Videos: A Survey GPT-4V(ision) System Card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.578848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.250907Z digest=sha256:4ef2112c673678e450efdb02164456c49e836400a4811ea18a34f11560fdc77b

Observation 53191d63-744b-4e11-9acd-8cc39efae377 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Reasoning Segmentation for Images and Videos: A Survey DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.276213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.276213Z digest=sha256:a175d33cf04b0032cd6c095053f636af7f61a97cb051aabb12a630ca29abc591

Observation 9e9a4395-0340-4ad8-92db-1bdbc82aa975 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.570111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.318324Z digest=sha256:1ffce0616c29787a6c7f77a6170caa3ad52c5f01cc5b5627e4c630a1ccbf6729

Observation 264110e8-d488-4fd7-ba05-e7207fa041de · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Reasoning Segmentation for Images and Videos: A Survey The 2017 DAVIS Challenge on Video Object Segmentation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.379401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.379401Z digest=sha256:3bccdd3ce7be80e89610d26b29c9250081eda4c4b920c28f563df850d4c63539

Observation 184032c2-86a5-4be1-8391-b84572ca2e99 · outbound

This paper cites Occluded video instance segmentation: A benchmark.

Reasoning Segmentation for Images and Videos: A Survey Occluded video instance segmentation: A benchmark

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.560685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.420028Z digest=sha256:de7a1042c57291858c05e79519c7c6b5f69b1b8d66c59acde6cecb9651b2144e

Observation 316d3aa7-825b-4253-8944-3f4469407fdb · outbound

This paper cites Reasoning to attend: Try to understand how< seg> token works.

Reasoning Segmentation for Images and Videos: A Survey Reasoning to attend: Try to understand how< seg> token works

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.459345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.459345Z digest=sha256:1f724b54ffdab1fcaff63890c82c4d1d19604f85276a25e28bc595d9154efaf3

Observation 1a67f527-dab8-4bf5-bee0-62f351497b25 · outbound

This paper cites Learning trans- ferable visual models from natural language supervision, in: International conference on machine learning, PMLR.

Reasoning Segmentation for Images and Videos: A Survey Learning trans- ferable visual models from natural language supervision, in: International conference on machine learning, PMLR

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.552768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.502970Z digest=sha256:d3329515e91ff743b8638c9dc338c6432ee1ffefa42bff811314bf5653a58b89

Observation e994d17d-0097-423f-9f4c-bbda7b8edf20 · outbound

This paper cites Paco: Parts and attributes of common objects, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Paco: Parts and attributes of common objects, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.544550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.551416Z digest=sha256:be1dacf7e08a13cc56eeff782b46acbad57770d4065a6cc6b29eb2470ce923a4

Observation f5f67165-0da6-4ff5-9b25-611289af9cd7 · outbound

This paper cites Glamm: Pixelground- inglargemultimodalmodel, in: ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Glamm: Pixelground- inglargemultimodalmodel, in: ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.536167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.605622Z digest=sha256:be07fd244680ebc9b72bac20ad3ea1701f4f63bc24c859799254cb133b85ddc1

Observation d5e00767-6e8e-4f2e-bc96-94c6afd01c47 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Reasoning Segmentation for Images and Videos: A Survey SAM 2: Segment Anything in Images and Videos

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.663543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.663543Z digest=sha256:021cfda0450c04674f99dd7f343d84019b9d38c5ebb315426088502eb9a389cb

Observation 13e78c09-8f72-4f36-a0a7-99b403efb154 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Reasoning Segmentation for Images and Videos: A Survey Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.717849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.717849Z digest=sha256:ad0c07408eb15d8ff0c3bf681a6a0fb605b6aeb6c56b88c88542029c9b1a3f3b

Observation 018d04b0-a9d8-4d8e-9aa5-5213379dccff · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Pixellm: Pixel reasoning with large multimodal model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.527449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.755628Z digest=sha256:8d5ca6c584c2c140b4a07f39023b29a186177856cd66a55f3eebae1138f88a62

Observation b7865683-1aa7-4cf3-8385-19d6aa57d54a · outbound

This paper cites Object Hallucination in Image Captioning.

Reasoning Segmentation for Images and Videos: A Survey Object Hallucination in Image Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.801705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.801705Z digest=sha256:8ede7901da228724dfdc8f8ec2ab097a4c77f4536f971322f00a2c1db431e96c

Observation f488d60a-fcf2-408d-908e-3ea4f8cb5ed5 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.518963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.840520Z digest=sha256:7f0f4556f4cc1d5d84076d201b8b34adfaf0599f38583b3f5fc39cefc5a82d37

Observation 8f6a796f-b9c3-43eb-afa1-2f3b46f58aae · outbound

This paper cites Position: Foundation Models Need Digital Twin Representations.

Reasoning Segmentation for Images and Videos: A Survey Position: Foundation Models Need Digital Twin Representations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.888010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.888010Z digest=sha256:6e0058c02ef161ee586eb7a88de8cf3f9af1553abb981a6f44d161fb28970fe8

Observation 6154626e-f598-42ac-b4cf-248e91f05f2b · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.510072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:16.937489Z digest=sha256:ab2ee2b7ba5868add68c91df3c6457ed8d1888d185b7762be57604f2a2c2f073

Observation ab7aff1f-a7d5-44b9-82d4-6c5c1e4bc53a · outbound

This paper cites RVTBench: A Benchmark for Visual Reasoning Tasks.

Reasoning Segmentation for Images and Videos: A Survey RVTBench: A Benchmark for Visual Reasoning Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.976460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.976460Z digest=sha256:4a35505fd8d192dad06cc4cc6e25d778bb94acb72893bb56197ddb2bcdf0c1b3

Observation 6260a082-1e7a-49cf-8066-b91dd9288ce4 · outbound

This paper cites Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins.

Reasoning Segmentation for Images and Videos: A Survey Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.030764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.030764Z digest=sha256:a6146f8605218924eca95d65aa941a025f113988465ffb3e924edad73ecbd7d4

Observation 92574bd9-3902-4306-b588-39bf8b41ed14 · outbound

This paper cites Online Reasoning Video Segmentation with Just-in-Time Digital Twins.

Reasoning Segmentation for Images and Videos: A Survey Online Reasoning Video Segmentation with Just-in-Time Digital Twins

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.086609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.086609Z digest=sha256:2facb32e5600253b47deea489bf63e7a315c57fa5f8213821f3b6a54174c1b53

Observation d201cc74-11a5-4ff9-b09c-545f0a9bf9ca · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.501488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.146232Z digest=sha256:3e594da1cbe9c6843a25866aa3dad20c803084adb8978f356bb1d40ee42bb029

Observation e75943ca-b73f-4740-b246-86923e47b367 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Reasoning Segmentation for Images and Videos: A Survey EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.198173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.198173Z digest=sha256:7b2598155e96e43251c36315cb0f62d9e61ea4736a4911fd494687b6640249eb

Observation ed7bddb8-2d12-4292-9d58-85e2c9f5c624 · outbound

This paper cites Growcut: Interactive multi-label nd image segmentation by cellular automata, in: proc.

Reasoning Segmentation for Images and Videos: A Survey Growcut: Interactive multi-label nd image segmentation by cellular automata, in: proc

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.493778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.223479Z digest=sha256:ce0e9aa1f23020bc8995e6c49a2da3622bfe86d65ed92b099c14c7c600d9e99e

Observation 02c0e07d-8a62-4754-b277-3d52456ef1a1 · outbound

This paper cites Reducing the annotation effort for video object segmentation datasets, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Reducing the annotation effort for video object segmentation datasets, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.484452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.274833Z digest=sha256:fc15f807dc62ea7b29641fdba3d531bbea9e9d8ec9d05b3b66dc7020dc45ce69

Observation 7b617589-b580-4d20-a999-c99295fe8c3a · outbound

This paper cites Prima: Multi-image vision-language models for reasoning segmentation.

Reasoning Segmentation for Images and Videos: A Survey Prima: Multi-image vision-language models for reasoning segmentation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.319384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.319384Z digest=sha256:2fe2a924af87c9d70374d19d14f849c154af9fcceafdcdd995b84bd0b582cce0

Observation a5194841-dd28-44a5-bad6-18c571d4eda9 · outbound

This paper cites Ov-vis: Open-vocabulary video instance segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ov-vis: Open-vocabulary video instance segmentation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.475048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.370790Z digest=sha256:493f7c239c6396f686ad33163ac7cec05c126d828e2b323a6313d162f2fe9bc7

Observation 4f1441ef-eeb7-4181-9c0a-80a193c00a37 · outbound

This paper cites Towards open-vocabulary video instance segmentation, in: proceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Towards open-vocabulary video instance segmentation, in: proceedings of the IEEE/CVF international conference on computer vision, pp

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.465799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.422485Z digest=sha256:01e3419724c9593c95ea6839397d3351072700c518e16ded0cee6eca152344e6

Observation 77313945-d7f1-4d7f-8ddb-150254cd2b55 · outbound

This paper cites Llm-seg: Bridging image segmentation and large language model reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Llm-seg: Bridging image segmentation and large language model reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.457907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.475549Z digest=sha256:f3b6dc8faf61e701b373c9b199f1ad21067785bce677b70f5fd284ede65ad5d7

Observation 135e0f7e-ac33-4b4b-822b-7803a27d02ac · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation, in: Proceed- ings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Unidentified video objects: A benchmark for dense, open-world segmentation, in: Proceed- ings of the IEEE/CVF international conference on computer vision, pp

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.450167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.515490Z digest=sha256:a5d0fdfef39c8a477ecbb737d37cb2eda28d93afb9e9ee70fe9d0c178eefdbb6

Observation 69049e09-7e92-4305-a151-604abd22817f · outbound

This paper cites SegLLM: Multi-round Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey SegLLM: Multi-round Reasoning Segmentation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.599286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.599286Z digest=sha256:7e484d8d0f2bfbdd419207a9c14275a8beabd1973285c19714444eb299792ee1

Observation c58eabf6-7f01-4e6e-92c4-d9cbc0d0ad04 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Reasoning Segmentation for Images and Videos: A Survey LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.662473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.662473Z digest=sha256:86abd58a421baef3b82d4360aaeec0e1e34e2271fe82c8613540008acf61d8de

Observation 84d1e245-ea60-4124-8039-1caec776250d · outbound

This paper cites InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.717014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.717014Z digest=sha256:b88b05295ca2179aa46ba9b9256da2e3cae78bedf750398d0904840b614b1c0d

Observation 87f917ea-32ca-4d10-a3e9-72157c73ce39 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reasoning Segmentation for Images and Videos: A Survey Chain-of-thought prompting elicits reasoning in large language models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.757478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.757478Z digest=sha256:253a0a2c51880337050db529935ab199a2c6a9783aba4c63e21280c4b59084c1

Observation f4eeaab8-288b-41c1-a4a6-1bde2a30ff2b · outbound

This paper cites Ov- parts: Towards open-vocabulary part segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ov- parts: Towards open-vocabulary part segmentation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.437125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.799215Z digest=sha256:50098499d65bbc8b03632733603bbc5b04afdd73a8802f7a89637357d2df9544

Observation d94a26f9-d9b2-495b-a66b-3aec75517bdd · outbound

This paper cites Phrasecut: Language- based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Phrasecut: Language- based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.428859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:17.895761Z digest=sha256:e15edf6c5107bdad0b0d63dce61714dd95fd5fe8d6c9b26630eca2091d3a04c9

Observation 482a4945-58ae-4189-81d3-31cdb93c7b47 · outbound

This paper cites See say and segment: Teaching lmms to overcome false premises, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey See say and segment: Teaching lmms to overcome false premises, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.420604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.002307Z digest=sha256:481c939c66271b142ef27eae161dc5e7921e72797e2cd5541217ad6d22512d80

Observation 1682cf05-b69a-4cd6-a225-5c8f435a83c4 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Gsva: Generalized segmentation via multimodal large language models, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.411628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.077397Z digest=sha256:8599affa24b6a353def3760ca2486d6c341127e7f4029db412ad2af3703e059b

Observation 2aeeff76-74e6-4d7e-a14a-eace992817c2 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp.

Reasoning Segmentation for Images and Videos: A Survey Youtube-vos: Sequence-to-sequence video object segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.402971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.152778Z digest=sha256:9605fd468799ef528eb3e16f95125da72bc1f2146db94a1c889ae36812be9cae

Observation a8374649-c868-43fb-a8d0-65a273149fd8 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Visa: Reasoning video object segmentation via large language models, in: European Conference on Computer Vision, Springer

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.395243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.244911Z digest=sha256:3c66faaf39d0a5968972c652cd4c3999fb50178b29bee718c8d2b8d193d09425

Observation a40796a9-b8ff-4ef7-8d6d-11208ac46186 · outbound

This paper cites Panop- tic scene graph generation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Panop- tic scene graph generation, in: European Conference on Computer Vision, Springer

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.386965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.337726Z digest=sha256:e85d284810d4dc9260b295f52362dfdad94a40130c8aadb92c6662776ed23743

Observation 0f8c2825-a571-46e7-ab1b-7d7dde838826 · outbound

This paper cites Video instance segmentation, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Video instance segmentation, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.379641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.419486Z digest=sha256:ead859f8870d21f7a135031228c404643c27bdd37109f19e45cb0a1d6bf929b1

Observation 5eadc24a-072c-40ed-a1b0-bfaa947003b7 · outbound

This paper cites Depth Anything V2.

Reasoning Segmentation for Images and Videos: A Survey Depth Anything V2

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:18.497202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:18.497202Z digest=sha256:6ca01693f2ceede9a85071d4f5f564e9f9fe3acf4f082635741aa9da80abf97e

Observation 2d846f45-b74d-4917-bef0-fe69edc5c59d · outbound

This paper cites An improved baseline for reasoning segmentation with large language model.

Reasoning Segmentation for Images and Videos: A Survey An improved baseline for reasoning segmentation with large language model

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.369615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.567677Z digest=sha256:beab5903f39098814efceeeed944f187c4e8cf6441efbcc57067bd7d7db8136c

Observation 61a7ef6b-27f3-4b50-8a02-cc88f72dbcf5 · outbound

This paper cites Empowering Segmentation Ability to Multi-modal Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey Empowering Segmentation Ability to Multi-modal Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:18.655378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:18.655378Z digest=sha256:4e1b772702407ba3776f7ab65d351a2dcd009b052e03a506b21e99ebadb4bcb3

Observation add5e5f7-086d-41d7-be11-5af856289345 · outbound

This paper cites Follow the rules: reasoning for video anomaly detection with large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Follow the rules: reasoning for video anomaly detection with large language models, in: European Conference on Computer Vision, Springer

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.360394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.716411Z digest=sha256:fdc51f0d6abc42825a90f1de1038feece8652ff68e93b711d616c91a6e6a6953

Observation a8ad9b85-c90d-4b59-b68f-c088e2d7806d · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmenta- tion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lavt: Language-aware vision transformer for referring image segmenta- tion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.350859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:27:18.789405Z digest=sha256:0c27f95865b1ef34fa6edbad5272c9d29f76dff44da02ec935887c8af6c95f98

Pith citing papers

Observation 6507a37d-5882-44da-844d-3acb78c39a8e · inbound

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction cites this paper.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Reasoning Segmentation for Images and Videos: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.967237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.967237Z digest=sha256:10ecc596b133c5f602e9f620154c71d778c158dbd3810e4c16f0d9c96adcafec

Observation 8dba12a3-d13f-4f62-8e1e-4ff91d86e7a5 · inbound

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality cites this paper.

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality Reasoning Segmentation for Images and Videos: A Survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:00.225117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:54:29.027805Z digest=sha256:f4f45237553c358e392ba35d191d56d941fc5615eaea64ed62f7f50c30e8f92b

Observation 99dbdaa5-f54f-43cb-989b-8cf35a603dd1 · inbound

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation cites this paper.

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation Reasoning Segmentation for Images and Videos: A Survey

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.156308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:40:16.265708Z digest=sha256:214a6a62b5a7a6136321750520f839d85447f3bbd6ad42ba14e30c5fd0e57e75

Observation faaf51be-153a-4a61-9ba3-1c7b12e6a4a4 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Reasoning Segmentation for Images and Videos: A Survey

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.363471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:77d3dd2bc93e28323c579d64be897fd6f925ea969d754a264125aa122a0cd9a8

Observation b5da47c9-260f-41ab-9dba-72e58d18e44a · inbound

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation cites this paper.

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation Reasoning Segmentation for Images and Videos: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T13:38:03.546834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T13:38:03.546834Z digest=sha256:9fe6c78df42890e93649771f079873b33fc65129706f16f69742a878ce8aa82f