Pith. sign in

Paper Citation Record · LEDGER

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

As of 17 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2501.13667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13667 v5

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:47:59.927188Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1c3f6c-bbc6-4ba5-8b13-9314fc777a8e · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Xmem++: Production-level video segmentation from few annotated frames

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:01.036615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.638768Z digest=sha256:ae969af7292a66e8322a5e9b1721650f8a40d38629132ba45d41823ef780fdc8

Observation bbc84f67-177f-4a4d-9f71-db45046f149d · outbound

This paper cites RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.644137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.644137Z digest=sha256:e8b60fb32b26a56debb4ff1e52ec5edead6084ac11a77434635228c602a02f2d

Observation 6bf26444-f89a-4366-b158-ec586a8a771a · outbound

This paper cites End-to-end referring video object segmentation with multi- modal transformers.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation End-to-end referring video object segmentation with multi- modal transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:01.019502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.649698Z digest=sha256:cb74946ef6076f00f7720c84febe8d1be0c4ee258ef6fb0fde7f7f79c270beed

Observation fdc8ee81-920a-4426-9562-f90f532b5172 · outbound

This paper cites End-to-end referring video object segmentation with multi- modal transformers.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation End-to-end referring video object segmentation with multi- modal transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:01.001502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.654829Z digest=sha256:28de9c38dee4723a6051a1c93df66005c6c74564544505b110701bb8d91305fd

Observation 987ce447-7a12-4674-bc10-519771e68145 · outbound

This paper cites End- to-end object detection with transformers.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation End- to-end object detection with transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.984032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.659696Z digest=sha256:2858c4fc0931ae873213d591aec791d3fe13582bb1215c54d046bc0115a895a4

Observation 755fb4b8-5a47-4de1-9e55-e8b1463f784c · outbound

This paper cites Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.966930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.664751Z digest=sha256:7c2356b1b0c3e841cb118d4e4347cec7f9dadd73c2008f7a2e1d3ad49f5c16e6

Observation ff279dd1-d879-4b85-adb0-7c78be7d3270 · outbound

This paper cites Putting the object back into video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Putting the object back into video object segmentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.950457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.670130Z digest=sha256:7df83a7eea7bd870ae3a7de41124fc18e6d60081b9ee7bf48d402902f6d91d0f

Observation 4546e337-6ea6-4a03-8f0c-ec6d3e3a81d0 · outbound

This paper cites Segment and Track Anything.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Segment and Track Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.674573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.674573Z digest=sha256:e1b2181775440eb9d4ece4b771044b837b2422abf97dac85090a72306e4df563

Observation 055a0632-966b-491b-95e1-cc2a2fbb4493 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unsupervised Cross-lingual Representation Learning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.679455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.679455Z digest=sha256:405c8726b54c2bd16f8a26bae35cde2159a492dc0de51ade263806868bd05b41

Observation ff018f14-f355-42b4-b98e-04a998bfec9a · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Vision-language transformer and query generation for refer- ring segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.932309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.684899Z digest=sha256:baa43521d1f4f8e5f04b3c7fcd9b365e0491ee3f6016f14edd518bf0193f17cf

Observation 0fa86d86-2c75-4729-a592-455d0adf2496 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.913517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.690233Z digest=sha256:83d8a33c5d6a77acf5ae23bdee2644cc402dae2cbc13101edf75a57066583bec

Observation 5971daea-c441-47f0-b780-85b500c13e4d · outbound

This paper cites Language-bridged spatial-temporal interaction for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Language-bridged spatial-temporal interaction for referring video object segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.896976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.695176Z digest=sha256:a0ceeb86a39a6e9629d84b2e979deaeaa2b9ee4af457a45d39359351a2836750

Observation f94ac849-ef94-41d6-ab2c-b0ef10281cb4 · outbound

This paper cites Unified embedding alignment for open-vocabulary video instance segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unified embedding alignment for open-vocabulary video instance segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.881133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.700026Z digest=sha256:34a4569986f4d34ba1ba31b93c1100ef1aea4d2b32be90f7099bb032436a9a5a

Observation 64500324-77b2-49a1-b524-4c5921b9300e · outbound

This paper cites Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.865317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.704741Z digest=sha256:5c1e8c2ff861886931c02787b8f7fe6fdd6925ce57175bcf1da9e34dc47145a4

Observation 160ea41c-b019-473c-9647-601d3cdd12c5 · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.848532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.709388Z digest=sha256:90bc3364c1e58bc94170c18a600f60704808d80c70bd42c58f55dd1e289a99e7

Observation 7161f8c8-2472-4e51-8763-388d1fb61b66 · outbound

This paper cites Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.714033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.714033Z digest=sha256:b2cbe770b23ba1a6ba756f95688ed9a1b9816ba26cc02e47d16a9bef13f8763a

Observation 69c6c784-e421-4ba8-806c-7309be4bb745 · outbound

This paper cites Segment anything in high qual- ity.Advances in Neural Information Processing Systems, 36,.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Segment anything in high qual- ity.Advances in Neural Information Processing Systems, 36,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.718988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.718988Z digest=sha256:245cf00b93be572f89afd3ebe180ab9a9d49213d2e48f0f07bdb0d9b0d1f9606

Observation dfe47324-1f8c-4d7a-aaad-9288711f0ff0 · outbound

This paper cites Video object segmentation with language referring expressions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Video object segmentation with language referring expressions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.815303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.723933Z digest=sha256:ed47f3515ea1c788796276237548e49a9386142c214cbdfe71d4424bcc36c748

Observation 4e3d5cfb-376c-4f68-a950-94455ae29c78 · outbound

This paper cites Segment any- thing.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Segment any- thing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.797935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.728435Z digest=sha256:351ef9074dabd2fbce024ab82b87a7aae11a029ecb302bd0dafb07c07cd07ed3

Observation b89bf872-4ce4-4000-89c9-2bd0ba6b7222 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Lisa: Reasoning segmentation via large language model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.780736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.733316Z digest=sha256:12b8ce0737fef8bd59b53d88593fa16ef21760407bc502ab7140677a551ae7c9

Observation ced4a7ee-c08d-48ca-af59-a115b9104a63 · outbound

This paper cites Learning to learn better for video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Learning to learn better for video object segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.764725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.737900Z digest=sha256:836eabe62cd7f8c61c5f71545690533b8b6272f08b01b02b320d5ab9d71c723a

Observation 328add4e-aec4-408f-9b64-772b9545be10 · outbound

This paper cites an unresolved cited work.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:48:00.749283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.742330Z digest=sha256:bdbd17c7a2ff049dd227de2228416610fdd3ea1ff02eb00608cce3acfef89184

Observation 5b791fd6-2d71-4e4c-a4a1-6985029cb2a2 · outbound

This paper cites Bidirectional correlation-driven inter-frame inter- action transformer for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Bidirectional correlation-driven inter-frame inter- action transformer for referring video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.733566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.747025Z digest=sha256:16b37aed81e1e195eea9bb477108669dc40618fcba58ba22f4a6cea93fb94016

Observation 9701dfa4-8e8e-4eef-886e-cd73a0ea8d1c · outbound

This paper cites You only infer once: Cross-modal meta-transfer for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation You only infer once: Cross-modal meta-transfer for referring video object segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.717969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.751536Z digest=sha256:72a6e70e8546e17d80ad1609eb5e97df7a3d95a9f0a1d15705d52d77f9bdd564

Observation e5c6324e-ed1f-44bc-9d23-12fb89929b46 · outbound

This paper cites RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.756198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.756198Z digest=sha256:dbdfe3b2fc9b45d8494ac0010e8db119a8684069968be0bbee1b93ae19d41d27

Observation 3e9f7077-2c38-4e6e-99f5-159afcd2d098 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.761037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.761037Z digest=sha256:9925500a0c35ebab0a1456c253006777e114f58afb293ce711594711688000a5

Observation fd5bc5ec-19e1-40f6-b456-e6f50d648946 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.765578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.765578Z digest=sha256:003835c7a8ac0644d0336266404234830dd69f3bbb7e857a0c552ff95dd5995a

Observation 230a054c-9723-4a4a-93ea-8bf4aac4c325 · outbound

This paper cites Decoupled weight decay regularization.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Decoupled weight decay regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.770551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.770551Z digest=sha256:4d0e6a1683f0cb152ddeed5abe374f9859ccc010332afce5896df9d11fb2b718

Observation d947f84c-2012-4c4a-94d4-3e539f1415bd · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.Advances in Neural Information Processing Systems, 36, 2024.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Soc: Semantic-assisted object cluster for referring video object segmentation.Advances in Neural Information Processing Systems, 36, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.682289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.775051Z digest=sha256:c4030e161cacc15a7c5cba3d1f6105364494b22fb1b0ff7960f35774ea538732

Observation 50a38455-8280-4a0f-9131-321f32e3234f · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Generation and comprehension of unambiguous object descriptions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.779445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.779445Z digest=sha256:b615405253c58a5d1067c8cd5ce47dc1ce381fca1495cec329fce73b63466d82

Observation bd1b8c88-a717-449d-a81c-135e74c8e4ab · outbound

This paper cites Visual-textual capsule routing for text-based video segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Visual-textual capsule routing for text-based video segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.655677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.784047Z digest=sha256:33ed1c29b882c41290d46d791782d1659a5bf614d1b2b83e86130554456992c4

Observation 38e25771-d08d-407a-a24a-19c3b8c25675 · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Spectrum-guided multi-granularity referring video object segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.640795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.788681Z digest=sha256:ee003e7fd91d458ab426276fed45b3b5b59832005e19ce779d1098612632e754

Observation 04df44cf-137c-43d2-9023-03e4e5a2394f · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.624776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.793025Z digest=sha256:69c62cc7b7b64ea901a2722ae11201b3073124d4dafac592a20ad31a62d06c1b

Observation ad7d8cdc-6449-4416-bddc-15badaeee255 · outbound

This paper cites Video object segmentation using space-time memory networks.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Video object segmentation using space-time memory networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.609152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.797817Z digest=sha256:a3346f731a3521798eb01ca502c9128868b8bf755b1960f5917141bda57e9a0a

Observation b476880e-13e2-4bc3-ba23-53d3216bf272 · outbound

This paper cites Semantic and sequential alignment for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Semantic and sequential alignment for referring video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.593023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.802324Z digest=sha256:d755b2ba948cb0bcffd7f689568f359b1986421aa1ee9c0734293c551e3aab83

Observation a2d4b4e7-3df5-46c6-88d9-7eb0fb05c1c2 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation The 2017 DAVIS Challenge on Video Object Segmentation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.807087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.807087Z digest=sha256:a3bd8cc1aacb9b07b8ff6d462da552e6dee4cc550b30edb84676ddbec26fc3de

Observation e384604e-d98d-4289-b000-5cd9177d1b1e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Glamm: Pixel grounding large multimodal model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.811931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.811931Z digest=sha256:5dd200a51498e65417a97018720413ac3184602bcd4edee78f17e25a49cd69d4

Observation 36430e9f-d503-4473-b3e4-ed7f044aac35 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation SAM 2: Segment Anything in Images and Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.817076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.817076Z digest=sha256:2679bc78534c328cae47d4a212e168e6b3809963ea33db6942ac0e8d45503e30

Observation 1dd84681-af88-483f-aac6-521cc1ddbc2e · outbound

This paper cites Cus- tomized sam 2 for referring remote sensing image segmenta- tion.arXiv preprint arXiv:2503.07266, 2025.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Cus- tomized sam 2 for referring remote sensing image segmenta- tion.arXiv preprint arXiv:2503.07266, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.821941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.821941Z digest=sha256:3a8d03726a3ca2532046caee1aafb443ac445e660d370432a46c73a639f92958

Observation 92ed7359-2cfb-4b8d-985e-b3e434a333e0 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.567693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.826708Z digest=sha256:0198b1dbb6ff277213626aee00136f505c9c5a85f90eda79e1a4711a93912389

Observation 94e8d278-d03b-4049-8de5-aca91dbf4e0b · outbound

This paper cites Temporal collection and distribution for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Temporal collection and distribution for referring video object segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.552878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.831198Z digest=sha256:2217709c7d293152554c8407dd76fc53d43559826fe5e2d69b8df40536349926

Observation 3065cf5b-7afd-4526-833f-13bb09643fd2 · outbound

This paper cites Samrs: Scaling-up re- mote sensing segmentation dataset with segment anything model.Advances in Neural Information Processing Systems, 36, 2024.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Samrs: Scaling-up re- mote sensing segmentation dataset with segment anything model.Advances in Neural Information Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.537906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.835790Z digest=sha256:a385983ea96369b0bd405842c20593792964bb16c19d4c4b01a0ef0ea3e883ce

Observation a1d6c23c-9577-466b-b38a-327dcc6a4f30 · outbound

This paper cites Asymmetric cross-guided attention network for actor and ac- tion video segmentation from natural language query.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Asymmetric cross-guided attention network for actor and ac- tion video segmentation from natural language query

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.521970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.840270Z digest=sha256:ec6125905b57997d370fd613dc4f3ffed77187e1b5a2f844539296ebc92830fa

Observation 5959649d-ce2c-46b4-9e05-e8c3ca604f26 · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.504760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.845067Z digest=sha256:789537e4e03b9a22cd0362c639f9b436e9df9159e88118d9979377eb494afbaa

Observation 8f1e15ea-a785-4cc4-ace7-b3b1e43c0986 · outbound

This paper cites HyperSeg: Towards Universal Visual Segmentation with Large Language Model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation HyperSeg: Towards Universal Visual Segmentation with Large Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.849646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.849646Z digest=sha256:5803afa24f1f917d8fdda01af43736e6e5097ed1bde93aa6675a6f8335315e10

Observation 99a2b93d-69c9-4020-8819-a18dc653a03a · outbound

This paper cites Multi-level representation learning with semantic alignment for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Multi-level representation learning with semantic alignment for referring video object segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.488432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.854413Z digest=sha256:f0e1e677d6b585c1da1f3a17b1b0707c5dd892bb9f37114130574d2be53f8c9f

Observation dcaacb23-6094-40d8-a4eb-12d4d8299c9d · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Onlinerefer: A simple online baseline for referring video object segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.471795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.858650Z digest=sha256:4ed98ee6c1efcb9788a43457c5b675534972dc124d8e9828bab01872e16d1548

Observation 27dad8b8-5237-45d9-b191-642f246c9a2d · outbound

This paper cites Language as queries for referring video object segmen- tation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Language as queries for referring video object segmen- tation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.455585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.863235Z digest=sha256:c950dbd8a1d4e10ff21b3b8acac775f5b58a15a8e9760c3f93eee4fe901aa136

Observation af69f5b1-cba9-4ea5-bb47-dfc980093eef · outbound

This paper cites Logiczsl: Exploring logic- induced representation for compositional zero-shot learning.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Logiczsl: Exploring logic- induced representation for compositional zero-shot learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.438701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.868031Z digest=sha256:3a46f2c3057163db0fc3608ccc43b2e4cd26f1590cdd4ad2c879f30a30573184

Observation 3690eaf6-802f-4159-9288-f20a2e6b9847 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.422489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.873004Z digest=sha256:bfc0826c0ffab6969d956bb88f03130ff37df695b911970423ab401fe97d3365

Observation d52ffb01-c86d-4286-9866-779efd545e76 · outbound

This paper cites u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.877607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.877607Z digest=sha256:8867a6e970c7b929b0b7814541028defcc6252553581ad319b8504b648b297ad

Observation de3c20d8-e016-42fd-8746-f16c4ab22481 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Visa: Reasoning video object segmentation via large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.405801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.882837Z digest=sha256:a88fe0287ccdf478038bbd7795a60d580291260e327c54ab18af372202063dc7

Observation f36f6979-76b4-4c7c-9717-e82879d29075 · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.390008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.887488Z digest=sha256:3973f12bd6d943985c32a63c391b5421b03618cdc2edff87557b54855fd27763

Observation 39fab523-0cbf-4cde-a3ea-496bc9ffb894 · outbound

This paper cites Modeling context in referring expres- sions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Modeling context in referring expres- sions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.372685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.892095Z digest=sha256:7584308912278f2a982f02c7d019b14fa834c074f1ac9c2d07c56371cbbe505f

Observation cc420157-1065-4e3b-a348-dada5e697f3f · outbound

This paper cites A Simple Baseline with Single-encoder for Referring Image Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation A Simple Baseline with Single-encoder for Referring Image Segmentation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:48:00.007696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.897025Z digest=sha256:987ddd21b58f069ec8fa6554ae6ba88728dcfb2425d44cc0e0e8f656e91dff14

Observation 959aed02-6d41-47d3-9e66-7763d8836a05 · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.901850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.901850Z digest=sha256:c6a75f325978c5cd1ea26d2bdccac98ecbfaa3cd1be641b18c9eed66ef89c149

Observation 015c654e-2242-41cf-a9d4-161540a8f6cb · outbound

This paper cites Surgicalsam: Efficient class prompt- able surgical instrument segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Surgicalsam: Efficient class prompt- able surgical instrument segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.906540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.906540Z digest=sha256:d0d908553fd7b941fa1dc35d36c26f8a5fdde8cfc0c7c6a122eeb977920f7679

Observation 42fa24c1-7262-43ff-885e-fd50ec4fb259 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.911548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.911548Z digest=sha256:fccefa171754b433f086666f0fc2697a3c0928d8fa8387fd5df7b2b05fbb058a

Observation 747b87e0-fc36-405d-a4e0-52a21ae4e60d · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.916575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.916575Z digest=sha256:170d806cbf910024e743edbff608f080f4d083f51cac063497f9f04db8278dff

Observation a569ecdb-9f60-4755-9229-2bedc8ca6b6c · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Deformable detr: Deformable transformers for end-to-end object detection

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.922407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.922407Z digest=sha256:d1ee719355c586e4d6f55d6af066984d3e73030e30b8b2a6e44aca24e9632f20

Observation 0cd8a6d0-c9f7-4798-b2df-3d797451fb83 · outbound

This paper cites Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.326524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T15:47:59.927188Z digest=sha256:47aaf4d8793ad1f0b30ec671319bd631ac6511850d63d4cdd04c97071bede35f

Pith citing papers

No inbound Pith citation observations are available.