Pith. sign in

Paper Citation Record · LEDGER

Mask2Former for Video Instance Segmentation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2112.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:44.060493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T03:06:43.601051Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 164a91ab-6c2d-4928-8067-bfaaf7be41f5 · inbound

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation cites this paper.

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation Mask2Former for Video Instance Segmentation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:25:16.506146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T01:22:29.901174Z digest=sha256:7de8a98044b487f0697a82355cb25c8e15d917c3c3b09943f83e95498987675d

Observation 10f86999-c4bc-44b4-bc1c-5005dc1bc195 · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation Mask2Former for Video Instance Segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.060493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.060493Z digest=sha256:3205744ed48e5b4b9c22295c49aef830c4f5355c6314d8a2cab730891db56a53

Observation cab951f1-9f45-4213-978b-f68eea78264e · inbound

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation cites this paper.

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation Mask2Former for Video Instance Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:23.286782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:50:23.286782Z digest=sha256:ec98596c4433d990290da6ba96b47297ba8a615ef8ae07745db50738c8c74bf5

Observation 87b8979b-f055-487d-850c-6ce6b719beea · inbound

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing cites this paper.

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing Mask2Former for Video Instance Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:13.164766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:13.164766Z digest=sha256:cf007e2cbb96bbb840951edbcf03157f789c68bbf32dcc37be01ef2686ee51c9

Observation bd7d29b2-c6db-4d8c-b6fc-7cec9ca92856 · inbound

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation cites this paper.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.737828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.737828Z digest=sha256:f874968206a5660ae638b43ad9526d52237ccd657f439c32c4c12f4ebd361ce5

Observation d6d28bd7-fe06-4d42-8334-568ec688b40d · inbound

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences cites this paper.

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences Mask2Former for Video Instance Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:32.089540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:37:32.089540Z digest=sha256:513e89b67bedd757ce3628121f25d10d9d33733414cfe3e2b9611e8c85ac5202

Observation 2489c505-f001-449f-97bc-6756f5153e56 · inbound

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting cites this paper.

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting Mask2Former for Video Instance Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:13.254653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:22:13.254653Z digest=sha256:a3c503c57131fe47b422f59f059982e53c23bb09812a03089c4ef488012cd40d

Observation c6ff474e-4aed-44d8-83bf-b09ba6be3630 · inbound

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation cites this paper.

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:08:43.698403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:08:43.698403Z digest=sha256:b6d7bcb9246122161776f5fe3318df101a31e335f4daaf6819b543a83a3e669b

Observation 6f339dd2-819a-4472-a54d-7544ee0bdf16 · inbound

Temporal Cluster Assignment for Efficient Real-Time Video Segmentation cites this paper.

Temporal Cluster Assignment for Efficient Real-Time Video Segmentation Mask2Former for Video Instance Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:11:37.909813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:11:37.909813Z digest=sha256:ca919c6acb57a6d4b5e4527a7a508b7ae757873ad90c4009ac9ac1d17d7fc934

Observation a89c39a0-f55f-4a9b-8ba0-b528648f14c8 · inbound

CObL: Toward Zero-Shot Ordinal Layering without User Prompting cites this paper.

CObL: Toward Zero-Shot Ordinal Layering without User Prompting Mask2Former for Video Instance Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:35.401431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:35.401431Z digest=sha256:7d106e07a46bc5bae09d4bd9a842a8da850c7ae46081b825d7aeb7c61a84b197

Observation 453aadd4-07b8-4fe4-8e23-54c911b8f2fb · inbound

Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System cites this paper.

Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System Mask2Former for Video Instance Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:02.984463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:30:02.984463Z digest=sha256:e13fd25e02540dcd429f5cdf8769ec34d87307edfe062e372f4edaf8615c0aee

Observation 500da900-2e9b-4573-94d9-6a9d882cd9a5 · inbound

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment cites this paper.

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment Mask2Former for Video Instance Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:33:15.913482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:33:15.913482Z digest=sha256:cf31665f72ce0551d29ea45de097f40945d5d891d8df4ab69f7058a8ecdee2e1

Observation 10c61aa5-734c-4aec-b0a5-643ae63bd371 · inbound

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects cites this paper.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Mask2Former for Video Instance Segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.211242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.211242Z digest=sha256:978f363b8b89bae7a881186e54c4021138ea5ba5daff2b8d5cfbfea5ec6fa700

Observation be6a2d55-45cc-495b-a708-0f4ef7620b64 · inbound

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines cites this paper.

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines Mask2Former for Video Instance Segmentation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:56:01.846052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:45:09.913980Z digest=sha256:4c885780f4598c34e419fb65c47068f57025db7e1a3d778e8f3020b162e131da

Observation adf30736-1412-4539-9207-ce3599cc819d · inbound

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes cites this paper.

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes Mask2Former for Video Instance Segmentation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:01.637082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:20:24.032405Z digest=sha256:cbe7c8027d888ac15e565fc19a3c0146fd666fce633490630663076046f3dddd

Observation fc1c1c04-981c-4149-8ea9-74bbae5d0e56 · inbound

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation cites this paper.

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.925722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:10:36.011508Z digest=sha256:77dfedc940f419070380965ee423745039d85817272f9115717ceb1c3b7d9ff0

Observation 7f9d0ede-c8c2-4449-8697-5369740011c7 · inbound

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation cites this paper.

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.493585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:19:56.997875Z digest=sha256:ea3809f807b99185e48db17802ef384aa7b1b0b848503c243f5efe23b6fdb94c

Observation 46758175-4b06-49e8-a5a7-1e1fe1319b2f · inbound

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation cites this paper.

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:54:37.012591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T10:45:35.568212Z digest=sha256:9dcf7d4a5642fff27d9b10e44f0918d790f8d3498ad5ceed602c877ad65c260e

Observation aacbb5a3-64a2-4e90-bd87-ffdedc95b691 · inbound

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation cites this paper.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Mask2Former for Video Instance Segmentation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.602664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:b0177c0ed1fa46bd86e594204295b35c71a5d178674ac9bbc88939f873ca0760

Observation c631030b-3c73-4f78-a06e-796a12b477db · inbound

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation cites this paper.

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:20:46.977826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:20:46.977826Z digest=sha256:00bb2de4c5ff200fe97de308c51e7fe27c918247d9b66f5a2f6eddb457bbfd4a