Pith. sign in

Paper Citation Record · LEDGER

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

As of 20 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2607.08688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08688 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T03:04:57.109924Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact4
  • verified fuzzy55
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 828e8133-13c3-4c05-9f41-374665f742f2 · outbound

This paper cites MOSE: A new dataset for video object segmentation in complex scenes.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation MOSE: A new dataset for video object segmentation in complex scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.962222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:6a4ce515e7adca0751f69faa54556ea09f734e2cc24615afe6c65f241ab748ab

Observation 50b89bc8-651a-4e15-9721-e9947a33758d · outbound

This paper cites MOSEv2: A more challenging dataset for video object segmentation in complex scenes.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-10T03:06:43.590309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:c5ce63a4765af41a1cb99f2fec459be71aa8aa20e5fa647465157dba99d776ca

Observation 9f1a42d9-39c3-4071-8276-5937315ea7cd · outbound

This paper cites Genie: A generalizable navigation system for in-the-wild environments.IEEE Robotics and Automation Letters, 2025.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Genie: A generalizable navigation system for in-the-wild environments.IEEE Robotics and Automation Letters, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.881974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:4c29cd6e8baf08f797a41b08e112378e447f321eb1e415162eab5e8ea860a28d

Observation 53afb98e-4a04-4271-8290-a538370892b3 · outbound

This paper cites Video object segmentation-based visual servo control and object depth estimation on a mobile robot.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Video object segmentation-based visual servo control and object depth estimation on a mobile robot

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.865260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:bbb118302204ad221511693ddc90bc0179012c201be1edbcf31a13e4e48770ad

Observation 3f9daccc-03a4-45c7-b1ff-df87d9306c23 · outbound

This paper cites Video object segmentation using space-time memory networks.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Video object segmentation using space-time memory networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.858864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:256db9b92d12c1eb2f478854dacfc7ea7f8aacc0d8b1c787936d99973ed76114

Observation c3a1d15b-29ff-4cd1-b967-0031e7aa4989 · outbound

This paper cites Rethinking space-time networks with improved memory coverage for efficient video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Rethinking space-time networks with improved memory coverage for efficient video object segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.941432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:ff5eb2f648d2d1f1e49ac7c97fc209a38c0e2ac3be4441c7b3809b19f321dcdb

Observation 35db4574-bd18-4991-8187-1479b52b351e · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.933424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:a9c5292c5c6ac1d2b3db34f3b0464c17f8e2667e82900645c477e160dc4207f1

Observation f6689e78-df52-4fa8-9ba6-5a156a931f38 · outbound

This paper cites Putting the object back into video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Putting the object back into video object segmentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.937573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:0b5aa995b352e14705514c7872bd6a29957304c1e95aa7415e07add0e04e4a9e

Observation a737784b-b10e-4c63-b788-e7ee57da023c · outbound

This paper cites Sam 2: Segment anything in images and videos.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Sam 2: Segment anything in images and videos

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.873184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:09c0bebeeda71626c8b688d61845e51f04e0fe63c42ad2eb67da2f998d806103

Observation 7fc1cf10-d5ca-4221-b3b8-44f58f2c6361 · outbound

This paper cites Sam2long: Enhancing sam 2 for long video segmentation with a training-free memory tree.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Sam2long: Enhancing sam 2 for long video segmentation with a training-free memory tree

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.892369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:4c35f0f7e299741033cded7698da0a699bdfaa9b5ddacea0de0f1c50f0426b68

Observation ab96b837-5e1b-4046-b5a1-46212837888a · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.938679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:0484d6d74b7380cc0085e9febd0467b4a5495eabe8ffb0cd1e7639a4ec184fb0

Observation a5e4d703-6beb-4763-b65f-13b6920a71cd · outbound

This paper cites MeViS: A multi-modal dataset for referring motion expression video segmentation.IEEE TPAMI, 47(12):11400–11416, 2025.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation MeViS: A multi-modal dataset for referring motion expression video segmentation.IEEE TPAMI, 47(12):11400–11416, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.936860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:4cb955e928e075cea9af0fe6f11415be1e7b4505e1bade66561da18e8b883a0a

Observation df158df1-66b3-4c15-8e00-63a9bffddb3a · outbound

This paper cites SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.593613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:e01f1cec076ff3b0ffe6965c11b21391c3f8aedad96d743209b3cdc4dafa6e17

Observation 7a7c2a76-4956-44fc-a2a1-3f2a261c5cab · outbound

This paper cites A distractor-aware memory for visual object tracking with sam2.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation A distractor-aware memory for visual object tracking with sam2

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.960246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:246b92e37a057101b35a6dba21e35cc9ab34217af02024a011e36fe9fb2a6b8a

Observation c46bdff7-3ec0-4f3c-83af-982324227488 · outbound

This paper cites Modular interactive video object segmentation: Interaction- to-mask, propagation and difference-aware fusion.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Modular interactive video object segmentation: Interaction- to-mask, propagation and difference-aware fusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.924366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:da37b7210e4868ed2d1eed205c1ae1c3ffed39935f3bdb727baef13e39b4cf97

Observation 358eb817-b68c-4835-83f2-ef9a107719ef · outbound

This paper cites Learning position and target consistency for memory-based video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Learning position and target consistency for memory-based video object segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.939435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:47422bb2c00999aafa4ac362b81b85cd031578d98d25e8341bbb57a3713de078

Observation 888d7a63-9384-4efe-aa14-e7e8079386e1 · outbound

This paper cites Efficient regional memory network for video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Efficient regional memory network for video object segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.954348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:23ad14680b4f286dd211722095c99e224b2834c118e251717a3f545d69fd92bf

Observation 161a26fc-b068-4cea-86b0-3bcea8907cc3 · outbound

This paper cites Swiftnet: Real-timevideoobjectsegmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Swiftnet: Real-timevideoobjectsegmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.932494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:15421a7ccb298df1203a26c03b1c4c82fedd507b3b54b04227e9810d799c533e

Observation 0a58ad69-d8b4-421f-b708-fefc99e3e5e4 · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Xmem++: Production-level video segmentation from few annotated frames

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.930626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:d7a8f4d31a9499ca35bbdb6b892952dc4d912685541733f6623f00e41d274f12

Observation 298dfe0f-b1aa-4bcf-9521-54b7e69cdce6 · outbound

This paper cites Tracking anything with decoupled video segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Tracking anything with decoupled video segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.913875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:91f5509d7efa378450e20d7f74239af3ce884afe19bcff354ab2411cf7222239

Observation 4eaf65a2-4ca2-4791-b4ac-8231477cf861 · outbound

This paper cites Lvos: A benchmark for large-scale long-term video object segmentation.IEEE TPAMI, 2025.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Lvos: A benchmark for large-scale long-term video object segmentation.IEEE TPAMI, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.919650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:00a5d1c4eafeef033bdec92e868e01732b3ae5d62841104f969c0b0a2966049f

Observation 3407bac9-c7c8-4b40-8a9e-f29a430064ce · outbound

This paper cites Hierarchical memory matching network for video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Hierarchical memory matching network for video object segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.913547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:97e6c59cce673e8d38cc6f59a209e293788b3324d97c4fe8f5ca7bc4bf50ddc2

Observation d5eac60c-6812-4860-928c-c81c9979056c · outbound

This paper cites Per-clip video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Per-clip video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.926261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:c503860419d488049c65b7fbc534260499c8634cf5a328daae7e2554e207980d

Observation 199e1015-8f1c-4b6b-adb9-90bae8869fba · outbound

This paper cites Transformer-based visual segmentation: A survey.IEEE TPAMI, 2024.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Transformer-based visual segmentation: A survey.IEEE TPAMI, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.928397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:280ed085465ba17717709966b675318c074d16535d8c572083083630aefabfbb

Observation e5b54164-1eb4-49dd-9545-d3effec09adf · outbound

This paper cites Towards open vocabulary learning: A survey.IEEE TPAMI, 2024.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Towards open vocabulary learning: A survey.IEEE TPAMI, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.905492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:3b65c889df69a1b260cd313a10d3ef55ca56d48224dd79cbe71a892fffff712e

Observation c3bb0124-60d0-4d1d-99e3-60b2b6e7d6cb · outbound

This paper cites VLT: Vision-language transformer and query generation for referring segmentation.IEEE TPAMI, 45(6):7900–7916, 2023.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation VLT: Vision-language transformer and query generation for referring segmentation.IEEE TPAMI, 45(6):7900–7916, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.934741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:f84e4f890a41be24f59dc118722b47807fabaf1e3194a9ada6e9621983b111e5

Observation 0eadb7b5-b2fe-4b00-9347-8776de0fbdb3 · outbound

This paper cites Multimodal referring segmentation: A survey.IJCV, 2026.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Multimodal referring segmentation: A survey.IJCV, 2026

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.903960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:d8a229e2c2853de50c2af62f8122bcbcc16911cf93e6af627f887fb05c8e0eb1

Observation b6a719ee-99cb-4528-88fe-cb3b270ccb94 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.NeurIPS, 34:17864–17875, 2021.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Per-pixel classification is not all you need for semantic segmentation.NeurIPS, 34:17864–17875, 2021

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.886050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:4bd72f5420a165486cdd690379258104267df8a7490211dc0a1d2608f357c27c

Observation fc179b54-0c28-428f-b06d-7b46a2b8a0a2 · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Scaling open-vocabulary image segmentation with image-level labels

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.917518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:5bc62a30ce084fae9917c508e0a85a98b9e859118ebde39ec7b65d99fd857072

Observation 725c7da1-da0b-4bc2-a376-bc8604f24d88 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Masked-attention mask transformer for universal image segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.921831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:22b8630d1033c78ef15bda287ad386c76dcb4ff5addc18b407946a534df624d6

Observation 58714759-d678-417e-8d25-9ab3673e352e · outbound

This paper cites Fastinst: A simple query-based model for real-time instance segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Fastinst: A simple query-based model for real-time instance segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.964130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:d69d0c66a95ff19f934a2fc3497c729bbf1b210801cef5763f5f07d1bf056e7c

Observation 4b04e013-8eb5-47c6-bbe4-d99598a44ca1 · outbound

This paper cites A survey on 3d gaussian splatting in segmentation, editing and generation.IEEE TPAMI, 2026.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation A survey on 3d gaussian splatting in segmentation, editing and generation.IEEE TPAMI, 2026

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.907542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:990560f40178dfba69a53b83e761d135b2bb0c498327d7b1199f1d4fde7a71dc

Observation e2844507-dded-4b42-bda8-e9846dacb814 · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.919825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:a3ab0e55d4ab63a05fe07b20231fea72ece3382e00d54e5dc8bc7b2201bc8e4a

Observation c81c906c-a51b-426b-ac39-48c4852c9b50 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Oneformer: One transformer to rule universal image segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.905656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:3c21c8221a436eae85d982ae3e125a5e4dc28d38db20facd2093057ac249fd4a

Observation 5a9ff550-7f68-42c7-a856-1a2a6a16ee27 · outbound

This paper cites Segment anything.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Segment anything

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.909684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:27d88fd51ac272c5fa9d58a63b20f77878c9f36fff1beac27fe5a5e63fc91082

Observation d9eb9627-9c59-4736-9ee1-bb79eb3c7b9a · outbound

This paper cites Segment anything in high quality.NeurIPS, 36:29914–29934, 2023.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Segment anything in high quality.NeurIPS, 36:29914–29934, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.899560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:e03eea7aa33e35af6e0da2036a887eb8cf7ac771e31b73f2266bd09675e0c508

Observation a2c2218b-556c-4bdd-91d4-c672bdeb00e0 · outbound

This paper cites Entitysam: Segment everything in video.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Entitysam: Segment everything in video

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.915786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:bcc0979adb06e0510d4df3c35257ed3c612b5758a3031b920c27601c0f74531e

Observation 9184c21b-0b06-4727-ade7-d8a3d7260127 · outbound

This paper cites GRES: Generalized referring expression segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation GRES: Generalized referring expression segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.926061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:262f52713d89107f54548c99bd8f5a94735af5c1b7f2f72f8de49ad437da1e0a

Observation 528dbb0a-9b39-4475-aa3d-9547e8dcf970 · outbound

This paper cites Vision-language transformer and query generation for referring segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Vision-language transformer and query generation for referring segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.877720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:9273efe54b99c435d78d24abd2479e044427718d82c78b487918fc1fb4736b43

Observation f39d8cb2-5a0c-4ccc-a3f7-7c6a1181246e · outbound

This paper cites Continual learning for image segmentation with dynamic query.IEEE Transactions on Circuits and Systems for Video Technology, 34 (6):4874–4886, 2023.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Continual learning for image segmentation with dynamic query.IEEE Transactions on Circuits and Systems for Video Technology, 34 (6):4874–4886, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.869386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:8d04e01a076b29a6cea9261645df35251f8837f7150e295a8a05de37c6e9d4b5

Observation 2d56c772-f1b1-4dc9-a720-126e5fb52226 · outbound

This paper cites Rethinking query-based transformer for continual image segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Rethinking query-based transformer for continual image segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.871299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:117352885da7b72d4ecf83127150b29d13667b2380bbb44fa8b6295eebdc1607

Observation 40c27ab0-2a6e-47a7-a907-2191408491a9 · outbound

This paper cites Primitivenet: decomposing the global constraints for referring segmentation.Visual Intelligence, 2(1):16, 2024.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Primitivenet: decomposing the global constraints for referring segmentation.Visual Intelligence, 2(1):16, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.893859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:36a48c9b98331fb489ff59bdfa037b90b6db8d67d09b683466e57e99e01d26ce

Observation c23fbff6-1f85-46ab-97e7-6a407f459a16 · outbound

This paper cites GREx: Generalized referring expression segmentation, comprehension, and generation.IJCV, 2026.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation GREx: Generalized referring expression segmentation, comprehension, and generation.IJCV, 2026

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.940636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:711de5eac05cf94276a778c02fe853e25cddf5367f8f63c3f4d4732d6047cf01

Observation 6d9ccf74-0757-4304-b706-49656dde6509 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Fully convolutional networks for semantic segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.897559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:0c72667f884ebc97957075ce9f0e350402bf30cc86efc9dd98dbebf326c20d7c

Observation f3acd39c-a09a-4082-8c6d-d3fa14e6d4db · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.IEEE TPAMI, 40 (4):834–848, 2017.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs.IEEE TPAMI, 40 (4):834–848, 2017

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.921990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:25c7b6196fb90d87e834b1202296bea1787a4c43e26d481c866f74b1fbb6db12

Observation 2768820e-3894-4f02-847c-a553ae81758f · outbound

This paper cites Mask r-cnn.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Mask r-cnn

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.965959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:8ce68df88fc850b0ba40636074920aa5e015fb0f5e55a1e871fa781dbc5a118a

Observation f380a21f-ad29-45a0-8f66-da1da372ddda · outbound

This paper cites Boundary-preserving mask r-cnn.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Boundary-preserving mask r-cnn

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.948469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:1bcf49ff284ef17a5b95ef3a0ae525399da91868fde6f182b54dbd090c85e7a2

Observation 1a457714-7e6c-4925-996d-6a14c7b32e49 · outbound

This paper cites Attention is all you need.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Attention is all you need

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.950617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:a22f41f3253e0ab16cc77ae17713913e125fbb706c93707905c99f621bd02097

Observation aacbb5a3-64a2-4e90-bd87-ffdedc95b691 · outbound

This paper cites Mask2Former for Video Instance Segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Mask2Former for Video Instance Segmentation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.602664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:d06f1ee5d063ff6f81db7b7d3ebe58ae0ea1c779c9a4b65f8a13ed5fbd44bd71

Observation 431a1c91-eb72-410d-929f-93f2236e182f · outbound

This paper cites Deep residual learning for image recognition.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Deep residual learning for image recognition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.944572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:073647a8ddb341500bcc7ef958fbec107a3ec0f2465850dc9e7110773129a0c9

Observation 6c91dd84-4627-4991-9a2c-2245878ef10e · outbound

This paper cites Lvos: A benchmark for long-term video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Lvos: A benchmark for long-term video object segmentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.942564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:bbe70e134246659fe19fca489632be73e16a8d4eb4ef5dda971f922d7a0e5618

Observation 269cdef4-8cf0-46aa-ae11-58c1b05fd85a · outbound

This paper cites Associating objects with transformers for video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Associating objects with transformers for video object segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.952449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:12bd6ddbb9ecb13436ecce49566168f5166176df6eb8ace9ba7005bb1323b4fc

Observation 60e55d6b-71be-4a68-8055-c51090c1c37d · outbound

This paper cites Decoupling features in hierarchical propagation for video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Decoupling features in hierarchical propagation for video object segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.956487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:5676f3f9bf8889a1dbcd9f67909e1763350f18be0ffbe92c75f3f0fa04596169

Observation 58aa81f1-2b53-49f7-99b7-a5a1df4f439f · outbound

This paper cites Recurrent dynamic embedding for video object segmentation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Recurrent dynamic embedding for video object segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.958457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:6d1c177531333be2054a11a749e9213eb8c0eb39114cd1116bf040e3dd98885f

Observation 323c6fae-5e33-47f6-9d5c-cb8cadaa007a · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Lisa: Reasoning segmentation via large language model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.967847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:5554bd307af36fd60c85d6315d9aa72e251792029a828cb6efad259d92a6ca90

Observation 18829729-5dd8-44c4-9eab-9bddecfc242f · outbound

This paper cites Adamcot: Rethinking cross-lingual factual reasoning through adaptive multilingual chain-of-thought.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Adamcot: Rethinking cross-lingual factual reasoning through adaptive multilingual chain-of-thought

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.896537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:15f5d9c000d79bad62c086d942bb8ac3e67780a18c55ca83cc7ab3ba1554d047

Observation 1791961e-de55-4ebb-8c53-3ee26c2ee274 · outbound

This paper cites Ccl-xcot: An efficient cross-lingual knowledge transfer method for mitigating hallucination generation.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Ccl-xcot: An efficient cross-lingual knowledge transfer method for mitigating hallucination generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.894233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:00ae538ac4f1f7b0f585856360e3cf7ff1e7ac1c2e44438419d9f6b0e437bfe4

Observation 416097cd-bf9d-4a1a-944f-ca0b51896f57 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation One token to seg them all: Language instructed reasoning segmentation in videos

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T03:06:43.946521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:7da023caf2d06e4328d37d115c2eb1a7258aafd1551de9854cf25ea6fa71c8e8

Observation ecbc4ea4-6c23-49b2-8f71-8b41ecf7c2cb · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.599834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:8997f77d4915256ad3e297a19f1f8bb20964e8c9397f13c143387bb15478bc7b

Observation 0177f891-ccea-405f-ad11-45cb2e1a67e0 · outbound

This paper cites InACM MM, pages 2880–2889.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation InACM MM, pages 2880–2889

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T03:06:43.596828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:6597669cafb3c15ddd4d8661f5d5401fe95df6a9ded32419d7194ab677b80ce8

Pith citing papers

No inbound Pith citation observations are available.