Pith. sign in

Paper Citation Record · LEDGER

Reasoning Segmentation for Images and Videos: A Survey

As of 18 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 5 inbound Pith citation observations for arXiv:2505.18816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18816 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:18.789405Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:06:22.967237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.361578Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2535e581-f34c-49aa-a4cd-90ff023a2e52 · outbound

This paper cites UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model.

Reasoning Segmentation for Images and Videos: A Survey UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:27:20.276725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:11.018222Z digest=sha256:0a7b2d226220a226486c006791e4da46d2ae95f51f70931b9fcad4d3feeac723

Observation 5104d3b1-8a63-48c0-bd05-b019af82bf66 · outbound

This paper cites ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation.

Reasoning Segmentation for Images and Videos: A Survey ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.072412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.072412Z digest=sha256:28c59d44c3cb02476937a3727caa7277e16dacddfed1ee61c34fd37752d1c1fc

Observation 5628c86b-4642-451e-80ca-aa9b60782db4 · outbound

This paper cites Burst: A benchmark for unifying object recognition, 39 segmentation and tracking in video, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Burst: A benchmark for unifying object recognition, 39 segmentation and tracking in video, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.134478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.134478Z digest=sha256:fb06d2517d610321bb596c8ae4e6afa31b5ccaa7fb4bbbf45e146d613c524ef3

Observation 1596bf16-837f-4dc4-bc2f-653d81fd4952 · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Reasoning Segmentation for Images and Videos: A Survey One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.218146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.218146Z digest=sha256:c1d0651a8269021f17d6006b3feb480024907ed274f50b0a8dc79ca2d8e58f62

Observation 98bc1725-92aa-476a-a66b-6da70b393a0b · outbound

This paper cites Cores: Orchestrating the dance of reasoning and segmentation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Cores: Orchestrating the dance of reasoning and segmentation, in: European Conference on Computer Vision, Springer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.273826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.273826Z digest=sha256:56693ae28d5e0e1c3f00269e38605bcddc0dd2dc0c081a88c528112f16c018d4

Observation f2799ce2-0ec9-4498-aadb-2f0dd8d18212 · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Xmem++: Production-level video segmentation from few annotated frames, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.329047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.329047Z digest=sha256:5c082a2a485491417a8d1eae1fda7d0c0b412e4f440dcb5fb061b15e9f88bcbd

Observation 3e7199cf-1e6b-407e-b00d-f1c96ac9424c · outbound

This paper cites Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.431803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.431803Z digest=sha256:7d3dd2731fcfa4ce112359dc6edafecb5c20dbd0d6e576110608861d99e94d0b

Observation 54b9c109-2a0d-4591-b8e9-1cb2d4b7aada · outbound

This paper cites Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.489319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.489319Z digest=sha256:cd61ecbdb8acdc007202c5fbb9ea16f6621046a2b8339def6c89201b6b16032c

Observation f732ecc8-2281-48dc-9fa4-6ee936b81096 · outbound

This paper cites Pixel-Level Reasoning Segmentation via Multi-turn Conversations.

Reasoning Segmentation for Images and Videos: A Survey Pixel-Level Reasoning Segmentation via Multi-turn Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.587733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.587733Z digest=sha256:5c7e1996ab3bf29c813ec216a5e2e8a8dc3ca1bfcdffe62b109d80dacbcec538

Observation a2672107-2874-4dd1-9b28-4eb09fc36747 · outbound

This paper cites End-to-end object detection with transformers, in: European conference on computer vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey End-to-end object detection with transformers, in: European conference on computer vision, Springer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.653276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.653276Z digest=sha256:5f28b85c0b1e662cb511539a841672a5d63fbed778004f5f8e9eecdb977be9e2

Observation 9c17b8da-d672-4c2d-8dcf-003f75f2aa62 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.754067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.754067Z digest=sha256:b432e106e6b48e5001414cfc83357a908b3fd57430b378e0c83329c64589716a

Observation 0cf067c6-6e1d-49e2-b69c-d469484d8642 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.868491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.868491Z digest=sha256:c49c96543219c9875e9460c0e5667eb6cf0253d531c3f951cddf3f5e7d6c43f5

Observation 0023a925-9097-4578-b971-51f60d688532 · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring ex- pression segmentation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Sam4mllm: Enhance multi-modal large language model for referring ex- pression segmentation, in: European Conference on Computer Vision, Springer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.942872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.942872Z digest=sha256:03f3b5335c8f2791e353a2be5b3c6e11d24c707d5e836d0c60d8863dbc5bfcbd

Observation 4250dca0-38ea-4da1-bb92-80a94ba49762 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Masked-attention mask transformer for universal image segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.987613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.987613Z digest=sha256:0c88621e948edcf3993ba1251850408aafee1a1b548391140060f9cac3fea528

Observation 1ed77509-9ee5-4d18-adb0-fd733b49b0dc · outbound

This paper cites Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model, in: European Confer- ence on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model, in: European Confer- ence on Computer Vision, Springer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.062712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.062712Z digest=sha256:b23b4f2ebb5501b79d6699b555c261031d3595c9920c7f9dfef9fe8c4b59c467

Observation 47fd545a-b396-4222-bb23-690101e7ba9a · outbound

This paper cites Vocabulary-free image classification.

Reasoning Segmentation for Images and Videos: A Survey Vocabulary-free image classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.140729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.140729Z digest=sha256:cac405f010c342dae7dca9cf09eb7d19240971403c1bfe3588ad80c40c9ec5eb

Observation cf1d2c9b-d29f-4510-bfab-118019625441 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.225657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.225657Z digest=sha256:b2f9095f50d21696ecdff4c960652981a575c04830aa0e2fb2ebeb98bc5fd914

Observation e90944a5-7290-4d20-87f3-5469fca853cc · outbound

This paper cites Tao: A large-scale benchmark for tracking any object, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, Springer.

Reasoning Segmentation for Images and Videos: A Survey Tao: A large-scale benchmark for tracking any object, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, Springer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.305165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.305165Z digest=sha256:fc4b52b1a3c8c8ea1cc3c6f427184e95d3cb8dd1ec64a0c07f751114731493ca

Observation 82302091-de9e-46e9-bae2-48fc6dbb0626 · outbound

This paper cites Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level.

Reasoning Segmentation for Images and Videos: A Survey Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.401764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.401764Z digest=sha256:1bfb4f3791a24829c4fc65689f1e99a55425ba556b6ee90fc69b6114be48bb2a

Observation a4b13401-579b-4aad-ac1a-9744ffa96064 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Mevis: A large-scale benchmark for video segmentation with motion expressions, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.500258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.500258Z digest=sha256:bd8027db15b4b24e8027206b96506d7eed1ebf624e300348a140386b85936624

Observation 783e111d-b744-4e99-b275-4ac4768809cb · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Mose: A new dataset for video object segmentation in complex scenes, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.592565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.592565Z digest=sha256:16525bf71a6e2c22ebb7cf7c99e7526cc5a6052309d48f8d6921b8d9115bb7ec

Observation 579de55c-4d1f-4f10-b152-c97f38c742c9 · outbound

This paper cites Panoptic Segmentation: A Review.

Reasoning Segmentation for Images and Videos: A Survey Panoptic Segmentation: A Review

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:27:20.231845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:12.659787Z digest=sha256:295e8e77b80dddb8082a4f267e412671b42f2487a9bf0edf34790359f8fd64a1

Observation bcf025c8-490b-4b28-b496-4f577043241f · outbound

This paper cites Oops! predicting unintentional action in video, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Oops! predicting unintentional action in video, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.718042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.718042Z digest=sha256:0e8573e5cf0f9dfa8770f5ddfa5cbbc2101851bf16c93e8ea9658659ec6f0b32

Observation 56a0ce34-9da4-4754-8128-2335acf07537 · outbound

This paper cites A Survey for Foundation Models in Autonomous Driving.

Reasoning Segmentation for Images and Videos: A Survey A Survey for Foundation Models in Autonomous Driving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.818555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.818555Z digest=sha256:e1329f38076781b415a5a44025f3eac17d937f14c603759fb50783008b6c2a7b

Observation abe7de4f-a2cf-4111-ac9e-f8fe5fd37ed8 · outbound

This paper cites Part- aware panoptic segmentation, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Part- aware panoptic segmentation, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pp

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.927968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.927968Z digest=sha256:e5122695b329586070f0f7241d6c2d58e58de8fa1349fbaa84556d731d673da9

Observation 770e0673-4881-4b20-abde-e9f044696652 · outbound

This paper cites The Devil is in Temporal Token: High Quality Video Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.039735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.039735Z digest=sha256:07e291ac4c9b2dbd96bcfe9875a104e820ea7f28524132427304213ff0861707

Observation d3443753-da5e-48db-b968-d39eaebd1a60 · outbound

This paper cites The iapr tc-12 benchmark: A new evaluation resource for visual information systems, in: International workshop ontoImage, pp.

Reasoning Segmentation for Images and Videos: A Survey The iapr tc-12 benchmark: A new evaluation resource for visual information systems, in: International workshop ontoImage, pp

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.120490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.120490Z digest=sha256:8e308ce96be30d3bb9e4e18e89c0cfa0099f60f45cdd434cf1896e1cf6fcdf40

Observation b0496995-8774-403f-a4dd-cde1f9e7a271 · outbound

This paper cites Lvis: A dataset for large vocabu- lary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lvis: A dataset for large vocabu- lary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.222976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.222976Z digest=sha256:5c31a75934329b416721794818cd135b9ec40d7b9b900b0d8131c954e06be214

Observation 3b31db6e-4004-49a0-9ed4-73dddcf0523d · outbound

This paper cites A survey on instance segmentation: state of the art.

Reasoning Segmentation for Images and Videos: A Survey A survey on instance segmentation: state of the art

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.329467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.329467Z digest=sha256:53e476f5b6b7147aaab1cf010c0cf9f176675b93f1370d3e90f020caa7e92072

Observation 99f2b6af-40a8-4263-8878-38917cb669dd · outbound

This paper cites A brief survey on semantic segmentation with deep learning.

Reasoning Segmentation for Images and Videos: A Survey A brief survey on semantic segmentation with deep learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.453671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.453671Z digest=sha256:dbcf606f3a723925ff72d982c75cbd2e3e6e384f9d8c3126c234a908624b9340

Observation c2a94a13-f9d0-4335-aa48-30e7e1af3941 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey LoRA: Low-Rank Adaptation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.560190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.560190Z digest=sha256:79d8fde8aa7fe359620fcd472436575f8c34dc3aeba1369dbb29a6c63c4fa1d6

Observation 8ff5aa0d-221e-4702-92d0-32ff4c5ba8a4 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.682259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.682259Z digest=sha256:d4ee955e28682f2f2c58c9304adc92040e95bd6e4899a864ed61622897333899

Observation d3d0aa40-7f13-4279-b0aa-637e14b53d26 · outbound

This paper cites MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.729538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.729538Z digest=sha256:2809c26b8581ec48df1588166b163c5842ebccab6106835d6d3db35d134e96cc

Observation 253e3da3-74e8-4901-9f12-5f9800cc4d08 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp.

Reasoning Segmentation for Images and Videos: A Survey Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.814778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.814778Z digest=sha256:b5d285df73f25528e3a813c088044ea4a8b99ae1721067b925f37e0357f1a008

Observation e889dbfc-6c15-40ee-b6c1-ad95c4e89f14 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.880321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.880321Z digest=sha256:faf132b3c7ad7d589f2ebb0cd57d577d068b6a4c7c2889681e0ec8391121af0a

Observation 21a0058f-f297-4cf8-8f18-d9b02ba6c83c · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.033054Z digest=sha256:aab1e106f36bd0a724876650c275da1752354845968299a4feaed9bb922810dc

Observation 9bdf61cd-b914-4ae6-9b78-ede9ac9bb40a · outbound

This paper cites Segment Anything.

Reasoning Segmentation for Images and Videos: A Survey Segment Anything

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.085972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.085972Z digest=sha256:59bd2ee655682251c9dffa69ff21740591276bc13cc1b3828d39304c97bf2f56

Observation 35d9f209-ff49-4b67-aded-df2dbded454b · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annota- tions.

Reasoning Segmentation for Images and Videos: A Survey Visual genome: Connecting language and vision using crowdsourced dense image annota- tions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.115774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.115774Z digest=sha256:2b9d7bc66694d092e6498f90da101b7ffcd338ca3ed52ed8458b19b06257950c

Observation 07e0de33-8137-44ee-ba4c-eb01420f4378 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Reasoning Segmentation for Images and Videos: A Survey The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.277333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.277333Z digest=sha256:f668fb41ef2c6e3896f880a3b1df5b7b14d752a95d5fea48b77fd9515b4c3051

Observation c4177e72-78d3-4531-aaf8-79acb0d3d0d0 · outbound

This paper cites Lisa: Reasoning segmentation via large language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lisa: Reasoning segmentation via large language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.380393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.380393Z digest=sha256:cbb9bb77e1f58ddc91a1623c1b76032643ee19c0609c14736c80fa9c9d76c99e

Observation 31641866-bf98-47f1-9203-f901a609f30f · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els, in: International conference on machine learning, PMLR.

Reasoning Segmentation for Images and Videos: A Survey Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els, in: International conference on machine learning, PMLR

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.526975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.526975Z digest=sha256:e44f35ba47e6fc6c9eb5d62a86afa758b19c05274e261028057a5a552f118db1

Observation a8b427f1-981a-4b88-9931-54c2b2ae607d · outbound

This paper cites SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model.

Reasoning Segmentation for Images and Videos: A Survey SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.624387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.624387Z digest=sha256:d6f59afafeaaec0e4141eb8c377daffe344b5cc5dcdf2436b6501a717a1c4dfb

Observation b28ed6f8-e70a-43ca-aa56-40198daea904 · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Robust referring video object segmentation with cyclic structural consensus, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.796318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.796318Z digest=sha256:2f028d6296dc22ca66e61f67f967a8fe31d3835ff13b531ad9caf134ee11dc47

Observation ebc5c99f-5a2a-488f-b500-da0ab41c0c1b · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Llama-vid: An image is worth 2 tokens in large language models, in: European Conference on Computer Vision, Springer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.946198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.946198Z digest=sha256:0a068d6dc7f132198fa6cb1939d1ed3a0390e9ff3e6997e552d357c8a0151d5e

Observation 3b393a59-01d6-491b-9e1f-1e25d2744064 · outbound

This paper cites Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Confer- ence, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer.

Reasoning Segmentation for Images and Videos: A Survey Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Confer- ence, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.661879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:15.084904Z digest=sha256:b6de1dbd57dd39e53d02fe52f2745e434ec27380838ee25d77129b2672adbd45

Observation 9b0163e8-9b0b-48eb-ba6a-16065bb2e529 · outbound

This paper cites Gres: Generalized referring expression segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Gres: Generalized referring expression segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.653225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:15.241954Z digest=sha256:ec34dddfcb4423cdd7f3fedba1b9c05db2653b59d4e7c24a033429ff11e84f77

Observation 09a281c4-b093-437f-84cf-ce514a7497c9 · outbound

This paper cites Visual instruction tuning.

Reasoning Segmentation for Images and Videos: A Survey Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.643861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:15.352155Z digest=sha256:7606aa101091a94f010aadd46cb2b7604d1e847f25628ea3025aff09936d6f0f

Observation 17fd4db2-02f2-476e-a215-7c8d40dae819 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Reasoning Segmentation for Images and Videos: A Survey Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:15.482347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:15.482347Z digest=sha256:4443d1d8bdf5901ee444181f36165e7b816a05cb53c33276ee257cb201625ad2

Observation af282039-22b4-42fe-8e33-9d88f65ff8df · outbound

This paper cites Ground abstract structure concepts of scaffolding systems for automatic compliance check- ing based on reasoning segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ground abstract structure concepts of scaffolding systems for automatic compliance check- ing based on reasoning segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.635044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:15.625060Z digest=sha256:28160a0507cce5a88dddc9ed055859325c206dfc8bbaf05f3442ca889f9c68b3

Observation e6c3dddf-c8cf-49f0-8e4d-304427d957a7 · outbound

This paper cites Video anomaly detection and explanation via large language models.

Reasoning Segmentation for Images and Videos: A Survey Video anomaly detection and explanation via large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.627054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:15.763488Z digest=sha256:b1a42f0abb5844eb8d6656eaa3d14b3ac341ab32f78ac9acbc985ad22e1c3f38

Observation dc0c66bc-dbca-49f9-9a4f-6239e505fd84 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Reasoning Segmentation for Images and Videos: A Survey Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:15.874035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:15.874035Z digest=sha256:a1915c9835c605c447054110acf6d393b76f488031c31f24f2ab174269bcb6e8

Observation 41bdf037-fc82-43e5-b1cd-8f4d9b98bd02 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.618602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.015561Z digest=sha256:d735c96623521f7d08275e3901238889327e38245ee6247daae0567792289484

Observation 4552696c-2c40-4fcf-8ca9-645ee71e25db · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Large-scale video panoptic segmentation in the wild: A benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.603088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.115719Z digest=sha256:f4c34561a8abba6ee2cba1f47e45cab410e4e925d0ce7198370e695138299374

Observation d7785866-13af-4b67-a815-1f83c89251f6 · outbound

This paper cites Image segmentation using deep learning: A survey.

Reasoning Segmentation for Images and Videos: A Survey Image segmentation using deep learning: A survey

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.594634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.159160Z digest=sha256:8ef8463c9d38c53d33fb98163d827dfd070eb6f31774b80e19759870de51db26

Observation e0d0e917-465d-46e7-9e8d-6def35c33a95 · outbound

This paper cites GPT-4 Technical Report.

Reasoning Segmentation for Images and Videos: A Survey GPT-4 Technical Report

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.586466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.209495Z digest=sha256:55e635e70dd05a5715610c137efd2d4c4530e5c01f3f96c66eb919476bfd3834

Observation 1d4d6a45-bf66-432c-8b22-6c4f1b2ab48f · outbound

This paper cites GPT-4V(ision) System Card.

Reasoning Segmentation for Images and Videos: A Survey GPT-4V(ision) System Card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.578848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.250907Z digest=sha256:75a8d806809fcad2b1c3bfe940a90756c373f3598496e452e31bc85f65e0513e

Observation 53191d63-744b-4e11-9acd-8cc39efae377 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Reasoning Segmentation for Images and Videos: A Survey DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.276213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.276213Z digest=sha256:01523dfb9eb0254c81cdaf5a200aab4e54224f4310692352a19553dd33f095c7

Observation 9e9a4395-0340-4ad8-92db-1bdbc82aa975 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.570111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.318324Z digest=sha256:336cdb524f1e047b8da2b8f107e47605f93447b4fb3774f7627561e05d7c00f1

Observation 264110e8-d488-4fd7-ba05-e7207fa041de · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Reasoning Segmentation for Images and Videos: A Survey The 2017 DAVIS Challenge on Video Object Segmentation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.379401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.379401Z digest=sha256:1fa3dfa6935dd84cf8d06113a372b54f38a44a44daf56de8efc80ac568fc5775

Observation 184032c2-86a5-4be1-8391-b84572ca2e99 · outbound

This paper cites Occluded video instance segmentation: A benchmark.

Reasoning Segmentation for Images and Videos: A Survey Occluded video instance segmentation: A benchmark

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.560685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.420028Z digest=sha256:94021ca8e4db11a953cf7cdf0a9d1bd0f44527f34248bcdac7ac123272c95e25

Observation 316d3aa7-825b-4253-8944-3f4469407fdb · outbound

This paper cites Reasoning to attend: Try to understand how< seg> token works.

Reasoning Segmentation for Images and Videos: A Survey Reasoning to attend: Try to understand how< seg> token works

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.459345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.459345Z digest=sha256:f0377cbd14495eca23335a4a9f6d1f7c6c805c60a8ee289ba9011577fadc20c1

Observation 1a67f527-dab8-4bf5-bee0-62f351497b25 · outbound

This paper cites Learning trans- ferable visual models from natural language supervision, in: International conference on machine learning, PMLR.

Reasoning Segmentation for Images and Videos: A Survey Learning trans- ferable visual models from natural language supervision, in: International conference on machine learning, PMLR

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.552768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.502970Z digest=sha256:e0158d9dfd6e87c32d7e29f7d2d15e16050efebfe61cd6d65edf6853b01e5443

Observation e994d17d-0097-423f-9f4c-bbda7b8edf20 · outbound

This paper cites Paco: Parts and attributes of common objects, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Paco: Parts and attributes of common objects, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.544550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.551416Z digest=sha256:db5d431ec18266de8cb21c9063d39e0e80346587ff85bd8e18444b7b8c5e6738

Observation f5f67165-0da6-4ff5-9b25-611289af9cd7 · outbound

This paper cites Glamm: Pixelground- inglargemultimodalmodel, in: ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Glamm: Pixelground- inglargemultimodalmodel, in: ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.536167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.605622Z digest=sha256:ac79100f6be527212f3f80fa92ad5dc4203ad5865a2031300f44848ef0ecdb1d

Observation d5e00767-6e8e-4f2e-bc96-94c6afd01c47 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Reasoning Segmentation for Images and Videos: A Survey SAM 2: Segment Anything in Images and Videos

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.663543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.663543Z digest=sha256:33b2ae032db6d8078140b5bddab56607f23a6cd0c171af9c7e63dc963959a33a

Observation 13e78c09-8f72-4f36-a0a7-99b403efb154 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Reasoning Segmentation for Images and Videos: A Survey Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.717849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.717849Z digest=sha256:48204fcf0f26690793755845d7db841b814e79912050c53ce29358f2c1d726aa

Observation 018d04b0-a9d8-4d8e-9aa5-5213379dccff · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Pixellm: Pixel reasoning with large multimodal model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.527449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.755628Z digest=sha256:5840516d5c933884840f3c207db82e6b55d1800350f3e98b32e6a0315aaee735

Observation b7865683-1aa7-4cf3-8385-19d6aa57d54a · outbound

This paper cites Object Hallucination in Image Captioning.

Reasoning Segmentation for Images and Videos: A Survey Object Hallucination in Image Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.801705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.801705Z digest=sha256:99556e12b868722efc2a8d5556e51f9f917dbc49050df30c2f21714e36015d3d

Observation f488d60a-fcf2-408d-908e-3ea4f8cb5ed5 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.518963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.840520Z digest=sha256:759034aa778e637ebcae0b385793a6dc7587744379039036eab39f76197771b5

Observation 8f6a796f-b9c3-43eb-afa1-2f3b46f58aae · outbound

This paper cites Position: Foundation Models Need Digital Twin Representations.

Reasoning Segmentation for Images and Videos: A Survey Position: Foundation Models Need Digital Twin Representations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.888010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.888010Z digest=sha256:23ef1ca04ca8a0189a49c7584d60d7ec30870e7cf0834420fe8cb6a6e0e27251

Observation 6154626e-f598-42ac-b4cf-248e91f05f2b · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.510072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:16.937489Z digest=sha256:546c5d173aa003d8af87ee05753875f7da5358c2c1f589a5316483f35415f903

Observation ab7aff1f-a7d5-44b9-82d4-6c5c1e4bc53a · outbound

This paper cites RVTBench: A Benchmark for Visual Reasoning Tasks.

Reasoning Segmentation for Images and Videos: A Survey RVTBench: A Benchmark for Visual Reasoning Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.976460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.976460Z digest=sha256:578cbc80da5f6bbfc2f4e4ebc57abbad524b57ae3999570784c4c11e454f707b

Observation 6260a082-1e7a-49cf-8066-b91dd9288ce4 · outbound

This paper cites Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins.

Reasoning Segmentation for Images and Videos: A Survey Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.030764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.030764Z digest=sha256:8a208c40aa5422fa0d9d93f3c9da268d8f86b235b450e774b7129eee82ae9694

Observation 92574bd9-3902-4306-b588-39bf8b41ed14 · outbound

This paper cites Online Reasoning Video Segmentation with Just-in-Time Digital Twins.

Reasoning Segmentation for Images and Videos: A Survey Online Reasoning Video Segmentation with Just-in-Time Digital Twins

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.086609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.086609Z digest=sha256:ac14a8186c741fb61eae77b69a67297dfb72812aecc51b0f8548c5b6e0c12422

Observation d201cc74-11a5-4ff9-b09c-545f0a9bf9ca · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.501488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.146232Z digest=sha256:1a705f1d98499711b4cc579b60ee8e8d5962d30f1d03abde12df2781f6c09b97

Observation e75943ca-b73f-4740-b246-86923e47b367 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Reasoning Segmentation for Images and Videos: A Survey EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.198173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.198173Z digest=sha256:2948a5a5adc48dd2c4c8f0ec09e90639b7c3c7c778677941cb771a760d8a85f5

Observation ed7bddb8-2d12-4292-9d58-85e2c9f5c624 · outbound

This paper cites Growcut: Interactive multi-label nd image segmentation by cellular automata, in: proc.

Reasoning Segmentation for Images and Videos: A Survey Growcut: Interactive multi-label nd image segmentation by cellular automata, in: proc

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.493778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.223479Z digest=sha256:9c927c4370ee0c2e1b96231c4907c89bf5eb36410f0e94499175c3de42037f3f

Observation 02c0e07d-8a62-4754-b277-3d52456ef1a1 · outbound

This paper cites Reducing the annotation effort for video object segmentation datasets, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Reducing the annotation effort for video object segmentation datasets, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.484452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.274833Z digest=sha256:0257fb3a34e00f90db5b4d7efe3700ffd12b6e56927c457b6a1c0c00a8af50c8

Observation 7b617589-b580-4d20-a999-c99295fe8c3a · outbound

This paper cites Prima: Multi-image vision-language models for reasoning segmentation.

Reasoning Segmentation for Images and Videos: A Survey Prima: Multi-image vision-language models for reasoning segmentation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.319384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.319384Z digest=sha256:fcdfbb1db617d9fd723d8674bc5deaa4f50a2c09e39d5036db0b449a69591e4d

Observation a5194841-dd28-44a5-bad6-18c571d4eda9 · outbound

This paper cites Ov-vis: Open-vocabulary video instance segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ov-vis: Open-vocabulary video instance segmentation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.475048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.370790Z digest=sha256:773f393cded8baeb1cb28a8c8b5fcb0d98e5bf8e8c0c8b5f9c62f1e94a2b510e

Observation 4f1441ef-eeb7-4181-9c0a-80a193c00a37 · outbound

This paper cites Towards open-vocabulary video instance segmentation, in: proceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Towards open-vocabulary video instance segmentation, in: proceedings of the IEEE/CVF international conference on computer vision, pp

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.465799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.422485Z digest=sha256:eeb3c1376d737ed69cf1423df1149b2c38f3de67ab8d08e9b9239de62fb1a9a7

Observation 77313945-d7f1-4d7f-8ddb-150254cd2b55 · outbound

This paper cites Llm-seg: Bridging image segmentation and large language model reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Llm-seg: Bridging image segmentation and large language model reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.457907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.475549Z digest=sha256:a9377cf4bab9b00ccc40e8ec77214651cdc3f6ceb326625376813e53dd6d561f

Observation 135e0f7e-ac33-4b4b-822b-7803a27d02ac · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation, in: Proceed- ings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Unidentified video objects: A benchmark for dense, open-world segmentation, in: Proceed- ings of the IEEE/CVF international conference on computer vision, pp

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.450167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.515490Z digest=sha256:8bec30a2cd1c678cff7e53ae9b38001ec912e14a9bde8ae4790715008fb54959

Observation 69049e09-7e92-4305-a151-604abd22817f · outbound

This paper cites SegLLM: Multi-round Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey SegLLM: Multi-round Reasoning Segmentation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.599286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.599286Z digest=sha256:acb35f7c22d47778c4991ae0ad4b9d31e1c6df143faf87a00d3e896e7a34f37c

Observation c58eabf6-7f01-4e6e-92c4-d9cbc0d0ad04 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Reasoning Segmentation for Images and Videos: A Survey LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.662473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.662473Z digest=sha256:9a699c81d67d4473122cc89fa8a4a076844d9e5f0da285b906b26bfa8c30785c

Observation 84d1e245-ea60-4124-8039-1caec776250d · outbound

This paper cites InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.717014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.717014Z digest=sha256:0877af9c8c2202af5760db51683b708f8c697dd2765c8c4825bb73c797638319

Observation 87f917ea-32ca-4d10-a3e9-72157c73ce39 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reasoning Segmentation for Images and Videos: A Survey Chain-of-thought prompting elicits reasoning in large language models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.757478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.757478Z digest=sha256:f278c30719f879d3e3a3952492ea99d9ab96fedba77f412d47f547ca5d297b41

Observation f4eeaab8-288b-41c1-a4a6-1bde2a30ff2b · outbound

This paper cites Ov- parts: Towards open-vocabulary part segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ov- parts: Towards open-vocabulary part segmentation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.437125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.799215Z digest=sha256:c60a2eabe49f7ae61048bd8292b277642a332170cae053803f9daefe69d12ab5

Observation d94a26f9-d9b2-495b-a66b-3aec75517bdd · outbound

This paper cites Phrasecut: Language- based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Phrasecut: Language- based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.428859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:17.895761Z digest=sha256:7fa13d76f3f134ee5aaf03af52120c99408d96504e77b5ca754c95ae1c8615f8

Observation 482a4945-58ae-4189-81d3-31cdb93c7b47 · outbound

This paper cites See say and segment: Teaching lmms to overcome false premises, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey See say and segment: Teaching lmms to overcome false premises, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.420604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.002307Z digest=sha256:8781c8bf110dce9fb7ebe253a493e4250ea373783abf30f72c02754651a81c37

Observation 1682cf05-b69a-4cd6-a225-5c8f435a83c4 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Gsva: Generalized segmentation via multimodal large language models, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.411628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.077397Z digest=sha256:7738abc3bcc87375e8c291cc7bb6251eba5a28af7b5c952efe7a071fd006ce9d

Observation 2aeeff76-74e6-4d7e-a14a-eace992817c2 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp.

Reasoning Segmentation for Images and Videos: A Survey Youtube-vos: Sequence-to-sequence video object segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.402971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.152778Z digest=sha256:bc2d5d71165399a8be364cc0ce99d967d2148044938b7a76bcd6d6f36f41459f

Observation a8374649-c868-43fb-a8d0-65a273149fd8 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Visa: Reasoning video object segmentation via large language models, in: European Conference on Computer Vision, Springer

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.395243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.244911Z digest=sha256:a680208b20884038edcbaf65cd985269a2e9998f78bafb6f4201aa1de7fb9fcc

Observation a40796a9-b8ff-4ef7-8d6d-11208ac46186 · outbound

This paper cites Panop- tic scene graph generation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Panop- tic scene graph generation, in: European Conference on Computer Vision, Springer

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.386965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.337726Z digest=sha256:035fe0689f46e78a69dbb1fd6ad8fa265c59ec260308e83854aa4cf38779dff5

Observation 0f8c2825-a571-46e7-ab1b-7d7dde838826 · outbound

This paper cites Video instance segmentation, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Video instance segmentation, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.379641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.419486Z digest=sha256:e261adf855f5b12e711a83d010ad6ab7b6e4fe3ae93af78e870e9d5f8df6809c

Observation 5eadc24a-072c-40ed-a1b0-bfaa947003b7 · outbound

This paper cites Depth Anything V2.

Reasoning Segmentation for Images and Videos: A Survey Depth Anything V2

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:18.497202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:18.497202Z digest=sha256:323bbb855157fbd8fdb0c75dac4ffa612ceac507b8837f898e564330b1194807

Observation 2d846f45-b74d-4917-bef0-fe69edc5c59d · outbound

This paper cites An improved baseline for reasoning segmentation with large language model.

Reasoning Segmentation for Images and Videos: A Survey An improved baseline for reasoning segmentation with large language model

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.369615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.567677Z digest=sha256:5c14d245aa86b06b10c59c2d0f61ceda859206def2abb02fde921b181d5a4b0d

Observation 61a7ef6b-27f3-4b50-8a02-cc88f72dbcf5 · outbound

This paper cites Empowering Segmentation Ability to Multi-modal Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey Empowering Segmentation Ability to Multi-modal Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:18.655378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:18.655378Z digest=sha256:71a0f2b9efe8fb2507980e1d19f9848ffe6f69bb203d15bc4bdb2c1aeea5db1d

Observation add5e5f7-086d-41d7-be11-5af856289345 · outbound

This paper cites Follow the rules: reasoning for video anomaly detection with large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Follow the rules: reasoning for video anomaly detection with large language models, in: European Conference on Computer Vision, Springer

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.360394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.716411Z digest=sha256:7bd67649832c19a8f933c3117fb9d2563a21383a81c5431676e3ce8b3d8095d4

Observation a8ad9b85-c90d-4b59-b68f-c088e2d7806d · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmenta- tion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lavt: Language-aware vision transformer for referring image segmenta- tion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.350859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:27:18.789405Z digest=sha256:ff0fb2fa34869b3432407f9d89da4c8f4d52ae1449a706b011afaad1ca323513

Pith citing papers

Observation 6507a37d-5882-44da-844d-3acb78c39a8e · inbound

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction cites this paper.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Reasoning Segmentation for Images and Videos: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.967237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.967237Z digest=sha256:5985fe4993d27f673428005b34e1134bb6187a62dd75468564e2009dceb5edf3

Observation 8dba12a3-d13f-4f62-8e1e-4ff91d86e7a5 · inbound

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality cites this paper.

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality Reasoning Segmentation for Images and Videos: A Survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:00.225117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:54:29.027805Z digest=sha256:85a7f14bfff55ef856513088ea457a606ea25a77b1ef86ab6d860addfa69f6d1

Observation 99dbdaa5-f54f-43cb-989b-8cf35a603dd1 · inbound

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation cites this paper.

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation Reasoning Segmentation for Images and Videos: A Survey

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.156308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:40:16.265708Z digest=sha256:81d4825372999f8bf84792a479c5afaf13063b7e544904e791d27608565ee373

Observation faaf51be-153a-4a61-9ba3-1c7b12e6a4a4 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Reasoning Segmentation for Images and Videos: A Survey

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.363471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:f85731cf81852a040608ee413775591b129d8f7df014fa639412d8739f317355

Observation b5da47c9-260f-41ab-9dba-72e58d18e44a · inbound

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation cites this paper.

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation Reasoning Segmentation for Images and Videos: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T13:38:03.546834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T13:38:03.546834Z digest=sha256:9a6545675e5c5b1966b259b2dd2a45487ce6ed52f554216f3d62e56f84e56866