Pith. sign in

Paper Citation Record · LEDGER

Mask2Former for Video Instance Segmentation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2112.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:44.060493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T03:06:43.601051Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 164a91ab-6c2d-4928-8067-bfaaf7be41f5 · inbound

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation cites this paper.

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation Mask2Former for Video Instance Segmentation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:25:16.506146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T01:22:29.901174Z digest=sha256:ab0081fe45773a5bf1320bb6aadc64bad0b0459e9179684e0bbcb899a9386e16

Observation 10f86999-c4bc-44b4-bc1c-5005dc1bc195 · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation Mask2Former for Video Instance Segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.060493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.060493Z digest=sha256:38a9b3aef2fea610f7d23851a323c90ef7d174de51985e8bf523d6647e75b383

Observation cab951f1-9f45-4213-978b-f68eea78264e · inbound

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation cites this paper.

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation Mask2Former for Video Instance Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:23.286782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:50:23.286782Z digest=sha256:3787e21bc4ee5ebcf9d342e45218b9bdd0423b464a9fa9551e7ec5f2fc44447b

Observation 87b8979b-f055-487d-850c-6ce6b719beea · inbound

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing cites this paper.

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing Mask2Former for Video Instance Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:13.164766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:13.164766Z digest=sha256:53442146d160f022eb4e25e16788cdb5323a558eaf0e1ffbe40d5480992398c9

Observation bd7d29b2-c6db-4d8c-b6fc-7cec9ca92856 · inbound

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation cites this paper.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.737828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.737828Z digest=sha256:41cb55a6fafbb23cd71ad1943782a82879af0f784dee1964e253195223f190c9

Observation d6d28bd7-fe06-4d42-8334-568ec688b40d · inbound

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences cites this paper.

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences Mask2Former for Video Instance Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:32.089540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:37:32.089540Z digest=sha256:e28402d615ef04c6a9cfb89022ba9dec65b6d396be4425a77487ef894a1f28ad

Observation 2489c505-f001-449f-97bc-6756f5153e56 · inbound

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting cites this paper.

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting Mask2Former for Video Instance Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:13.254653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:22:13.254653Z digest=sha256:f4f80e47837fb090bea124f2fb34f7c0d232d81602bccefb955490638e7c1e1e

Observation c6ff474e-4aed-44d8-83bf-b09ba6be3630 · inbound

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation cites this paper.

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:08:43.698403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:08:43.698403Z digest=sha256:db0e6871d0c61aecc0d260d1b228a0e3efa0e51e8bbe690ca6cec71d8fe78222

Observation 6f339dd2-819a-4472-a54d-7544ee0bdf16 · inbound

Temporal Cluster Assignment for Efficient Real-Time Video Segmentation cites this paper.

Temporal Cluster Assignment for Efficient Real-Time Video Segmentation Mask2Former for Video Instance Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:11:37.909813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:11:37.909813Z digest=sha256:0d1298b0e0bbc406aac786fbffdd376b5a3c29ec0012e3352b06b01079718f1b

Observation a89c39a0-f55f-4a9b-8ba0-b528648f14c8 · inbound

CObL: Toward Zero-Shot Ordinal Layering without User Prompting cites this paper.

CObL: Toward Zero-Shot Ordinal Layering without User Prompting Mask2Former for Video Instance Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:35.401431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:35.401431Z digest=sha256:272e86e02ce0b345ae3848dd24e954af4835026677a0c066090431140f6a62d8

Observation 453aadd4-07b8-4fe4-8e23-54c911b8f2fb · inbound

Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System cites this paper.

Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System Mask2Former for Video Instance Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:02.984463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:30:02.984463Z digest=sha256:986261ec6cdec29ff4d61973124a8554f81d90ab69cc5372d1f0b22c0e489972

Observation 500da900-2e9b-4573-94d9-6a9d882cd9a5 · inbound

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment cites this paper.

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment Mask2Former for Video Instance Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:33:15.913482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:33:15.913482Z digest=sha256:2e4fd3937b7cb560a8667352936906cb50052ffd249ff745efcea1580cd62a14

Observation 10c61aa5-734c-4aec-b0a5-643ae63bd371 · inbound

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects cites this paper.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Mask2Former for Video Instance Segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.211242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.211242Z digest=sha256:23336de4511960132dbc8c796501e7c209c78b1abd81d9de2abb0f143978b55f

Observation be6a2d55-45cc-495b-a708-0f4ef7620b64 · inbound

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines cites this paper.

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines Mask2Former for Video Instance Segmentation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:56:01.846052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:45:09.913980Z digest=sha256:854bacfe1ce4d3725ffd899d3cc12dd16528a5c7d59ba92bb5213208fe9536fe

Observation adf30736-1412-4539-9207-ce3599cc819d · inbound

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes cites this paper.

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes Mask2Former for Video Instance Segmentation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:01.637082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:20:24.032405Z digest=sha256:cb72dce5bb8ef659a2f43c0309a8f7fc434f7c87873ba08017c4e4741b884042

Observation fc1c1c04-981c-4149-8ea9-74bbae5d0e56 · inbound

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation cites this paper.

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.925722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:10:36.011508Z digest=sha256:d8ca377d21c94e3494d05b50073a0b77cc679dfebdd5f88b5cb7d5b1f0cb95f2

Observation 7f9d0ede-c8c2-4449-8697-5369740011c7 · inbound

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation cites this paper.

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.493585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T18:19:56.997875Z digest=sha256:ef200b10a785b442650d2aece4edd03ba5746facb26459c066b14b58962a5a44

Observation 46758175-4b06-49e8-a5a7-1e1fe1319b2f · inbound

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation cites this paper.

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:54:37.012591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T10:45:35.568212Z digest=sha256:a22f6059d05be5dfcdb8d6cdc8e90b3803f0687203394777ac64c03913f6d1ac

Observation aacbb5a3-64a2-4e90-bd87-ffdedc95b691 · inbound

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation cites this paper.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Mask2Former for Video Instance Segmentation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.602664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:eb0766e42c270a20c5f070ab2372ded3f856335bc69ab146058dab400cb58491

Observation c631030b-3c73-4f78-a06e-796a12b477db · inbound

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation cites this paper.

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:20:46.977826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:20:46.977826Z digest=sha256:2d25875a01e4b63d994f7a4d76157ac3839855013faae188cdee321cc31032e2