Pith. sign in

Paper Citation Record · LEDGER

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost

As of 16 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.01304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01304 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:49.988237Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11a488a3-9d79-45bd-84bf-3beb1059ecc5 · outbound

This paper cites Xmem++: Production-level video segmenta- tion from few annotated frames.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Xmem++: Production-level video segmenta- tion from few annotated frames

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.738085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.181841Z digest=sha256:9bf0500379adf37be499d3aefb2b94baa13a243f550537babe14b1d40053add5

Observation 366b346b-2fb1-42a3-9b20-82c3210325bf · outbound

This paper cites One- shot video object segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost One- shot video object segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.726052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.234531Z digest=sha256:a4aeccc0682ff99155febca7e99717214fbe0b9d21bc7d8eb5a1383199f0bc18

Observation 11561c84-e797-4fbf-950d-36de8a62d960 · outbound

This paper cites Rsprompter: Learning to prompt for remote sensing instance segmenta- tion based on visual foundation model.IEEE Transactions on Geoscience and Remote Sensing, 2024.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Rsprompter: Learning to prompt for remote sensing instance segmenta- tion based on visual foundation model.IEEE Transactions on Geoscience and Remote Sensing, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.714060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.287685Z digest=sha256:1d7c513d44527a64bca49c2471aa37cd7e358106f413a8667becbac27026627e

Observation ba02cf08-f5f4-4072-95c3-ddd96b478355 · outbound

This paper cites Adaptformer: Adapting vision transformers for scalable visual recognition.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Adaptformer: Adapting vision transformers for scalable visual recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.703563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.368695Z digest=sha256:9459f3b848a78dd1b9a643785b0c8629f9dc77802f6f89ba916b4c5d2309ec7c

Observation 4a334688-76cb-41b5-8674-3a3d106dde9e · outbound

This paper cites SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:38.439660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:38.439660Z digest=sha256:08a837fcb7a53f85a213fdf4b6890b5e02771f24a185d619842064cfc63aae68

Observation db998cdb-5490-40f5-97fe-37adf7115483 · outbound

This paper cites 0.1% data makes segment anything slim.NeurIPS,.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost 0.1% data makes segment anything slim.NeurIPS,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.692814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.533843Z digest=sha256:3d697b0d8423096b0cae35b73b6a889fc7424307c75d0edbaafa5f932a5cf369

Observation 4c34bce8-2e12-47db-abee-d3958b50182d · outbound

This paper cites Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.682205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.585479Z digest=sha256:07823b3ad6b60d47b9d91b309bebb2e2d64d96749263bb6547d64d94938f231a

Observation 4e577a05-e56e-4644-8cdd-8399b2e354d6 · outbound

This paper cites Modular interactive video object segmentation: Interaction-to-mask, propagation and difference-aware fusion.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Modular interactive video object segmentation: Interaction-to-mask, propagation and difference-aware fusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.669131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.657606Z digest=sha256:6bb24bea3dc5483ac399031628f3bf93a35a5063e68fd08b3354f72b8cf3bbe9

Observation 7e4ddc88-dcfd-4521-8ca4-3f5f80beee02 · outbound

This paper cites Rethink- ing space-time networks with improved memory coverage for efficient video object segmentation.NeurIPS, 2021.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Rethink- ing space-time networks with improved memory coverage for efficient video object segmentation.NeurIPS, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.656205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.735290Z digest=sha256:43fd1982b1ecbf304155ba5ccf5616d80d8cc2636257dcf4379f5de5340bb604

Observation f07074bf-67a3-4cd5-8cd0-9247113ec8ae · outbound

This paper cites Tracking anything with de- coupled video segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Tracking anything with de- coupled video segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:38.821648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:38.821648Z digest=sha256:05b421c04aa90352815fd0df69d3a17373e3b5d1e19768d8051b7e735ed6b797

Observation 68e2cdb2-f115-42c5-b0f3-2c108187e7a2 · outbound

This paper cites Putting the object back into video object segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Putting the object back into video object segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.636635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:38.900297Z digest=sha256:40f18f2e7c04906e30cec98b88854af29d4b3c8266ebc19aab1ffa7ec5f0594e

Observation 9b1cc4fd-526c-497a-b21e-a3484ac88c49 · outbound

This paper cites Segment and Track Anything.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment and Track Anything

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:38.984260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:38.984260Z digest=sha256:62fbeb9e28c041580c6eed7a786476b9088036a4b93f0639dbbc47d2c2981372

Observation a33b334d-66b1-4181-a184-740ff8cbc34e · outbound

This paper cites ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:39.082937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:39.082937Z digest=sha256:ca2522c2c98ef4ade23ad6de1448f0deb166815edbc18af174d964476bb4b065

Observation 409ef4a8-5aeb-4347-a106-dceda66c567b · outbound

This paper cites Learning the what and how of annotation in video object segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Learning the what and how of annotation in video object segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.625812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:39.212385Z digest=sha256:fe17c4c2ac23e14b7c07c2082b9fbf83357b2332c0e5bd8bc1dc6f975fdc2b1a

Observation 35c29cac-6a12-481e-84a0-ac4ac2ca8add · outbound

This paper cites Segment anything model 2: an application to 2D and 3D medical images.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment anything model 2: an application to 2D and 3D medical images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:39.281260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:39.281260Z digest=sha256:4f75e3f67eafb7183b8918c40c56a34eaff11ae511c1ceb81cbab9287db782b5

Observation 642f041d-7bee-48c2-94c4-6b3bac3f550a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:39.420425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:39.420425Z digest=sha256:64a52ccd6aa676d5cd2565bf69a41113fc367c0380444122a2d6be1c692aa14d

Observation 8c099b13-52de-4c6f-820c-5c523858aadf · outbound

This paper cites Interactive video object segmentation using global and local transfer modules.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Interactive video object segmentation using global and local transfer modules

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.615372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:39.603848Z digest=sha256:5dc80a5b1af8799c46e9034ce3ab9aa2a3374886a0713325c1dcaeaf32b3404c

Observation 58d1537e-41f7-47b0-a316-1aeda0409a70 · outbound

This paper cites A Neuromorphic Dataset for Object Segmentation in Indoor Cluttered Environment.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost A Neuromorphic Dataset for Object Segmentation in Indoor Cluttered Environment

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:50.520855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:39.861021Z digest=sha256:a68be48f5580bc305b9c96ed00d37289a2cf3667b9245283e46f9ba68146de14

Observation d6944856-87ef-495f-a04c-90eb2ac2d4c0 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Prompting visual-language models for efficient video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.604426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:40.192127Z digest=sha256:b9024e92809cbe465ca5893760e30cd68da19682655b0ce62c19939beaafbfe0

Observation 3c3594b6-e9d5-4f84-86ef-35f0455f93e5 · outbound

This paper cites Segment anything in high qual- ity.NeurIPS, 2023.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment anything in high qual- ity.NeurIPS, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.593770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:40.888261Z digest=sha256:08fffaee58ab449f1b855a611dbd4ce637d78ba5bb61edc7466e46dc1290a9e8

Observation 50853442-a2d1-4ad6-8d6a-4f62cb46ffe0 · outbound

This paper cites Segment any- thing.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment any- thing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.584383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:42.580715Z digest=sha256:e83af1eb131422ee1f25363a57546a368f97d24c231f5d6421d764f5d20fa66e

Observation 06580f31-8f5e-41a2-a50f-92f1890d0de8 · outbound

This paper cites Frozen clip models are efficient video learners.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Frozen clip models are efficient video learners

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.574762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:44.078694Z digest=sha256:d724514befd236616d4c659f3f61a70d49666e33c9022674b5bfe594baefa410

Observation 3d97282a-0322-42fa-a8d3-8595e413b810 · outbound

This paper cites Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:44.860444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:44.860444Z digest=sha256:75ef6578250ea03ca717ee1c734ab37812a85c486caee3269aa16079bea6de5c

Observation a27319da-c951-430b-95ce-aa2214fc1cdd · outbound

This paper cites Decoupled Weight Decay Regularization.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:45.975938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:45.975938Z digest=sha256:ef5af1a50d41b684ff6c8a0826b05ec8fa44a5acd6cbcc4f64929c72e05ca317

Observation 46538667-b3ba-4a45-9504-1aea96602a29 · outbound

This paper cites Segment anything in medical images.Nature Communications, 2024.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment anything in medical images.Nature Communications, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.562809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:46.692481Z digest=sha256:4eb7e89623487e3e935df22b8d1811ff35955750e20d79f401f46bf42a86fd43

Observation 296753c5-d4e9-4a3e-9d0c-e1f19ce65fb1 · outbound

This paper cites Video object segmentation without temporal information.IEEE TPAMI, 2018.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Video object segmentation without temporal information.IEEE TPAMI, 2018

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.545539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:46.789541Z digest=sha256:28cab8e8e7da4cf09b5a226189e14a2936521ea68ffb493fd43c44951b09d11f

Observation 97ffd79f-007d-4c78-84ff-5e994a6b19fc · outbound

This paper cites Segment anything model for medical image analysis: an experimental study.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment anything model for medical image analysis: an experimental study

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.517771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:46.896233Z digest=sha256:374a113af75e3032eff52fa7fd4fff6c147a5dd48635913ac5f66fbcbc74ecf4

Observation cf47f950-2c8b-4f84-9960-b8b403de068f · outbound

This paper cites Fast video object segmentation by reference- guided mask propagation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Fast video object segmentation by reference- guided mask propagation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.363558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.024429Z digest=sha256:cbee902c4f691264ac6b8ab277005b8a33cb15040454d9faf39cc08332e15064

Observation ade024f6-0491-46a6-86ce-d70d8e35fd4f · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.NeurIPS, 2022.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost St-adapter: Parameter-efficient image-to-video transfer learning.NeurIPS, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.187973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.074668Z digest=sha256:3453b7dd2670dd6f0a95ac2c31bab89aa5f50229b4ec8bced62644fd8e2fcb8c

Observation fce007a0-bd8a-485a-9603-4ae670e231c8 · outbound

This paper cites Dual- path adaptation from image to video transformers.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Dual- path adaptation from image to video transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:53.020964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.194040Z digest=sha256:13108a20671f05fdd3d19e46ce62363b6dbda5130f0e9ce24019dd304811dac6

Observation 1c66845a-b417-4fe3-be90-b52433a7d20f · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Pytorch: An imperative style, high-performance deep learning library

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:47.273487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:47.273487Z digest=sha256:cd88d80899754b4f1d2465aa02e352d924679c289d65d9aedcda77229df22cc3

Observation 54246d1f-937a-4143-8f44-ca248c0a50a6 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost The 2017 DAVIS Challenge on Video Object Segmentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:47.338977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:47.338977Z digest=sha256:e1c678cbd98c54c152820a7a5d3a5ec1e4652882686e2c14e914dba08c51957b

Observation 5ba348c6-d502-4d91-869a-c166418a4c9c · outbound

This paper cites Disentangling spatial and temporal learning for efficient image-to-video transfer learning.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Disentangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:52.962159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.428687Z digest=sha256:d177b3cd05035d64ce39768d8984ca3fa6e503d2957bd913f2ffe2a09cfc141f

Observation 4cdc9513-ab3e-4027-a34e-7f0167211ff8 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:47.507661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:47.507661Z digest=sha256:2acf04413d8b1459b9e3f506b40ca02a9da79afeef16b34349ff30be4d5587bc

Observation 54567cfc-c79c-4116-b667-2137cf7a270a · outbound

This paper cites Segment Anything Meets Point Tracking.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Segment Anything Meets Point Tracking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:47.613125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:47.613125Z digest=sha256:fec261e6360e603e48cfb6c3fce644e1e1c5f71a7812515f7667d73d49a8fed5

Observation 32d05b03-b271-4e21-ad9f-ba434b00d35e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost SAM 2: Segment Anything in Images and Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:47.689544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:47.689544Z digest=sha256:e7860381dd3d81368c28e4870414bd8a7d9815e4318afe52a4661e5c41504bf1

Observation b46505a9-ab4f-4412-a50f-b4457b8763be · outbound

This paper cites Seg- ment anything, from space? InWACV, 2024.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Seg- ment anything, from space? InWACV, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:52.634283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.882759Z digest=sha256:8fbb10ef5596fc2a586b5d3670d50fa4b7c2cc465c6789c66ca0dd1ceaa1e904

Observation d2d55a5b-8207-4979-97ea-d2284c541ef0 · outbound

This paper cites Learning fast and robust target models for video object segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Learning fast and robust target models for video object segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:52.441859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.977450Z digest=sha256:dc8bf6a8fb89b0bbc98d9f5873831c08ddc755f51dc3d6dd376ffcdbe0352197

Observation 6367c98a-127d-4bb1-8f02-dc5c67f8f2ba · outbound

This paper cites Interactive 3D Medical Image Segmentation with SAM 2.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Interactive 3D Medical Image Segmentation with SAM 2

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:48.066282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:48.066282Z digest=sha256:1ce1befcd1ca772c129c6d2f50aa13209802d50a0d33c825cf284905f24ab8e6

Observation 0af786a1-4ab5-4cc8-94af-fd5a658cad6c · outbound

This paper cites TinySAM: Pushing the Envelope for Efficient Segment Anything Model.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost TinySAM: Pushing the Envelope for Efficient Segment Anything Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:48.139202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:48.139202Z digest=sha256:599ca52170017d5c9a431f4dcab496e36892f98cf8f03a71e9f7a0860533ceb7

Observation 145770c0-a7d4-449f-ad6e-75da32188b50 · outbound

This paper cites Towards open-vocabulary video instance segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Towards open-vocabulary video instance segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:52.416992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.219346Z digest=sha256:f9245c41cd89ead3cb7957f3a47d76578d97e9ba3305cdc49ea3a24fc3c2e599

Observation a28ac603-12cc-45d1-b433-35d72d35297e · outbound

This paper cites F3net: Fu- sion, feedback and focus for salient object detection.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost F3net: Fu- sion, feedback and focus for salient object detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:52.133490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.404865Z digest=sha256:2425d3556b748e99d1c6a4a1719a332b41437a4128c71fecef56fa1729dc671c

Observation 8fa47d21-827d-415a-8d7b-019097e79d86 · outbound

This paper cites Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:48.476606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:48.476606Z digest=sha256:51dccd63994bfe7687c527190e456586b2e55b5f99a28da5a32815f5ccc3b403

Observation f66ab37f-aead-4a37-b22e-914904281cd9 · outbound

This paper cites Scalable video object segmentation with simplified frame- work.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Scalable video object segmentation with simplified frame- work

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.963083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.565837Z digest=sha256:e0dcb0cae5633b33b3e6c58ca08a8e4537bced016aaabe31d88d9e41c0f0af88

Observation 3cf8432a-bead-4b98-9543-340adc1bda27 · outbound

This paper cites Cat-sam: Con- ditional tuning for few-shot adaptation of segment anything model.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Cat-sam: Con- ditional tuning for few-shot adaptation of segment anything model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.886849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.671961Z digest=sha256:2699ead9af2c3a46f109e0ab3a6f7147856a3e98adefa69bcc81060d780dee2a

Observation 2d21e214-7128-4169-94a9-69c759c75ee2 · outbound

This paper cites Sam2-unet: Segment anything 2 makes strong encoder for natural and medical image segmentation.arXiv:2408.08870,.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Sam2-unet: Segment anything 2 makes strong encoder for natural and medical image segmentation.arXiv:2408.08870,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:48.758340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:48.758340Z digest=sha256:c9ad921ac9c21c901350d51df6f2f51f79eab1f2899f1bf98e53b32a30eb0c88

Observation f6ec5bef-6694-4bca-b495-7d1b0097f919 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.871711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.859621Z digest=sha256:f1bea62357f4ba7876a1911cdb436e166ac8257750b8030b55c3a894bd0b22ec

Observation 66dd7019-f941-415a-b79a-1d07b6480998 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Youtube-vos: Sequence-to-sequence video object segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.629829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.966528Z digest=sha256:de6ccd2792aa70fc882d0689cb798e1a8d94f69d2726bc88c88ff147376583de

Observation 9ac3aeba-a833-4a41-8e9e-70ccdfed6daa · outbound

This paper cites Track Anything: Segment Anything Meets Videos.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Track Anything: Segment Anything Meets Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.056362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.056362Z digest=sha256:9f9eca3fd3ed85513f6edb8104b4e618029d2dc5934a5be3e0310ecabf4c8836

Observation 960e79a4-35db-4e07-83ec-aca882a801dc · outbound

This paper cites Efficient video object segmen- tation via network modulation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Efficient video object segmen- tation via network modulation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.325468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:49.155216Z digest=sha256:5e653da3b3d8475a24f0c456d0776add2a2580c6c4109a95c81ee54d389c5563

Observation 71beb037-fafc-43c3-8e60-8a94f15c8c43 · outbound

This paper cites AIM: Adapting image models for efficient video action recognition.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost AIM: Adapting image models for efficient video action recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.156717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:49.257921Z digest=sha256:9f380174d132ba45ff635247634b2954ef8f365487a127d01080cf07da1b9d0c

Observation 907785b0-b6af-41a1-98b3-a28ef3db9675 · outbound

This paper cites Collaborative video object segmentation by foreground-background inte- gration.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Collaborative video object segmentation by foreground-background inte- gration

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:51.066669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:49.351850Z digest=sha256:da69230a1f17a2336a58c0d6b76fa35df0c792243a14987ef6f68faead045ecd

Observation 5404a6f2-5757-4575-aca4-3a46ccf51623 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.455141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.455141Z digest=sha256:d8b74a872d3c077d649d12b341f95613a7953d780f952886cae319f52ecf7427

Observation b14219f4-0ea2-41f4-8266-4983cc0e55d0 · outbound

This paper cites MobileSAMv2: Faster Segment Anything to Everything.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost MobileSAMv2: Faster Segment Anything to Everything

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.547609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.547609Z digest=sha256:a24097f0c8a78020f4b0dea4fbab45c6fee561706d49e5f31e1f12131003bd1f

Observation 2ee8bb8f-c9f7-4a37-b36b-eaaa7dc82e06 · outbound

This paper cites Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.642417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.642417Z digest=sha256:4aa47dd728c18902f79409e5665a5aa22711d0f627944b8ab9efb65eade9d793

Observation a027915a-2f96-4f65-8176-e2d4128ae780 · outbound

This paper cites Fast Segment Anything.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Fast Segment Anything

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.758396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.758396Z digest=sha256:caf67cbe1c38f0577cd0c324dca5d57374cb595a3e8933ca28fcd83f337fc64d

Observation 19b37adb-16b0-4d3f-ad7a-f5ceaa97255e · outbound

This paper cites EdgeSAM: Prompt-In-the-Loop Distillation for SAM.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost EdgeSAM: Prompt-In-the-Loop Distillation for SAM

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.890149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.890149Z digest=sha256:90aafea1c8ad2e3b74728123c311fdf808e5a55c724776a5a5071a54cbece3e3

Observation d6722ea7-d27d-4d4b-a968-83bc9046d004 · outbound

This paper cites Medical SAM 2: Segment medical images as video via Segment Anything Model 2.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Medical SAM 2: Segment medical images as video via Segment Anything Model 2

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.909797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.909797Z digest=sha256:54b69748dc03debc4c68a13e236800c144916ce425c7f5fc509591357a939f97

Observation e62b112e-18b2-4ce8-8a43-a0b8f4e6efd4 · outbound

This paper cites -” indicates directly combining existing pre-trained models for inference. “Cost.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost -” indicates directly combining existing pre-trained models for inference. “Cost

Reference 61

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:50.813473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:49.988237Z digest=sha256:95f1d78f2e10a83c809396ab2a83ef6acef8e2db2c9c7d390cda611f7a4c704a

Observation 56a8c35a-5af4-4fc7-8199-6c8eecbe5760 · outbound

This paper cites an unresolved cited work.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:52.280421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:48.320982Z digest=sha256:700527f41574ed5738f3215b708d0fb345627bb22720eb8fc57dac3cd0c4b102

Observation 4a59372e-94cc-41ba-9a79-00605792d2b9 · outbound

This paper cites an unresolved cited work.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:52.837961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:52:47.760620Z digest=sha256:91f9a952e09a818e2770e5f127aace65803af2c6a120fe718dee835b0087a35c

Pith citing papers

No inbound Pith citation observations are available.