Pith. sign in

Paper Citation Record · LEDGER

Autoregressive Universal Video Segmentation Model

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2508.19242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19242 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:16.689829Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86b416ae-2ac7-42d1-906c-40a7941ad941 · outbound

This paper cites Just read twice: closing the recall gap for recurrent language models.

Autoregressive Universal Video Segmentation Model Just read twice: closing the recall gap for recurrent language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.392102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.392102Z digest=sha256:ab1c29960347fd6139703cb2000e97ddf5b083f97fe40fd4549b2f5d74fce804

Observation ae6d0bf0-1b02-4538-841f-12baaedb2af1 · outbound

This paper cites Tarvis: A unified approach for target-based video segmentation.

Autoregressive Universal Video Segmentation Model Tarvis: A unified approach for target-based video segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.722317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.398111Z digest=sha256:53dfca0904668b94d7cc5d91a0bc5d8d6fda7ff97ddb8f70f34247cba319ba74

Observation ca615088-376e-49bc-9125-482cc92f49a8 · outbound

This paper cites End-to-end object detection with transformers.

Autoregressive Universal Video Segmentation Model End-to-end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.701666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.403264Z digest=sha256:bb88b25d736be202dd586133649070993dcb1353541261746ee37047d151be26

Observation 97f0ece7-c232-4467-a313-477f1bf25b58 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

Autoregressive Universal Video Segmentation Model Per-pixel classification is not all you need for semantic segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.684671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.408380Z digest=sha256:4af18a34eae3b50c306cafe3718a71366f94c9cf97f9c90c52ac8337fff41287

Observation b9172dd7-93d9-4e20-bdb8-7d22ae9c4c11 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Autoregressive Universal Video Segmentation Model Masked-attention mask transformer for universal image segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.666164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.413889Z digest=sha256:b780217d877606de49ebf70ac5a068117329b38546d6ed30e46314ca6a518c37

Observation 7913f846-7918-44f7-a6cc-7bcf07e5d27c · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson- shiffrin memory model.

Autoregressive Universal Video Segmentation Model Xmem: Long-term video object segmentation with an atkinson- shiffrin memory model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.649604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.419859Z digest=sha256:e08607891f55d39b7f7b5c98e0573bfaa3f67ef66379fd4f0821a5c7154dc347

Observation f6c0c31c-4410-4c91-932d-c6b70941c826 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Autoregressive Universal Video Segmentation Model The cityscapes dataset for semantic urban scene understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.632470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.427098Z digest=sha256:7292d54a0b0e1cfd241ac5625c651e8d5e82b7057dd08908c6bc90ddf8eafd08

Observation 4731e06c-f2f2-47e9-bae1-045836d56a9c · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes.

Autoregressive Universal Video Segmentation Model Mose: A new dataset for video object segmentation in complex scenes

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.614476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.431862Z digest=sha256:5ec76077699e6f2cd6507943302f93c0f67d8c9b5bf5412e1b84e163b3e065f5

Observation bb0d6999-317a-475b-9b9a-693a1fadbb33 · outbound

This paper cites The Llama 3 Herd of Models.

Autoregressive Universal Video Segmentation Model The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.436505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.436505Z digest=sha256:8c068805021e27415d3e9a4faaa2f2f31dbf3abd75f09e8a26d2cbfce52dafc8

Observation 2649edb2-2207-4798-9572-e6178d0e3a6c · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Autoregressive Universal Video Segmentation Model Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.441780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.441780Z digest=sha256:eee0bb33673059ab92bc2bc6e8aff0b592fd52666098e41ff59c42f9a2d3a4ef

Observation d91cf4ff-0e1e-4bb8-bc3b-92b85d64b33e · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Autoregressive Universal Video Segmentation Model Efficiently Modeling Long Sequences with Structured State Spaces

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.447019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.447019Z digest=sha256:093706aca65d9debc7f195eafa45e56efefcdb7b31a7b9ab87ff4904a93c25db

Observation 57c9f109-2d46-4f4e-bbcf-e234a9c7b385 · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

Autoregressive Universal Video Segmentation Model On the parameterization and initialization of diagonal state space models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.598398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.453019Z digest=sha256:83bd4f4bb8fe692e4de6d94c12f412dbf6ddc28939e6717e62feca6fd2e424df

Observation 2d6e736f-e66b-42c0-9932-5169ac518a0f · outbound

This paper cites Vita: Video instance segmentation via object token association.

Autoregressive Universal Video Segmentation Model Vita: Video instance segmentation via object token association

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.582327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.457809Z digest=sha256:3a9deebc80649dd7b810e852f7534e1f06c3f105c9e4e3c515e17362318524b5

Observation 4a4bd33f-9d05-4b97-861b-8397af59a42d · outbound

This paper cites A generalized framework for video instance segmentation.

Autoregressive Universal Video Segmentation Model A generalized framework for video instance segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.565401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.462408Z digest=sha256:63371923978e780a4acbea9feda2e1bd60fbd710f0d509c5adae716f4c015b09

Observation d08d82a7-5b73-4a83-ba97-14940e838a40 · outbound

This paper cites Omni-rgpt: Unifyingimageandvideoregion-levelunderstanding via token marks.

Autoregressive Universal Video Segmentation Model Omni-rgpt: Unifyingimageandvideoregion-levelunderstanding via token marks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.549091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.467534Z digest=sha256:75ff97ca764b298fa0147c0397cb901b8a91fd8664275db3081c51c1298c0af2

Observation e5c62569-62d9-46f9-9be6-c0416591949a · outbound

This paper cites Robust and consistent online video instance segmentation via instance mask propagation.

Autoregressive Universal Video Segmentation Model Robust and consistent online video instance segmentation via instance mask propagation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.532803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.472268Z digest=sha256:cef33b227e9d2bdcdf0bb738c67d4d1bb99f5dbc731500339ee3a74a0df083a6

Observation cd06ca60-2ec1-4a76-b8b5-03085f6dc174 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Autoregressive Universal Video Segmentation Model RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.476792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.476792Z digest=sha256:11c20f5ec43f558d5be812a4f50fac1ad41c3f215e4787def2031c4c06fc0fda

Observation c0b308c0-6515-4f70-9e8f-71e5b060850d · outbound

This paper cites Minvis: A minimal video instance segmentation framework without video-based training.

Autoregressive Universal Video Segmentation Model Minvis: A minimal video instance segmentation framework without video-based training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.514993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.482625Z digest=sha256:1601d05d22c9ec0418d4de57e0553d855431138659935231224dfd460d50ba08

Observation 7b45a9d9-c36c-44d2-ad63-e64170754081 · outbound

This paper cites Video instance segmentation using inter-frame communication transformers.

Autoregressive Universal Video Segmentation Model Video instance segmentation using inter-frame communication transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.497831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.488015Z digest=sha256:07e72be73ec1f71dfe9c236597d6eb91daeb1af8b41f8367b6db3b8a13bafb9e

Observation 4825247a-88c9-482d-93e6-d78cb83ab75b · outbound

This paper cites Video panoptic segmentation.

Autoregressive Universal Video Segmentation Model Video panoptic segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.480792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.492725Z digest=sha256:143d5a15d040dadeb99579c9c8b3113c7a4d4eb5a57478b954bbb55f7ede6788

Observation b8594234-570c-40c2-aac4-daf8e1bb48b6 · outbound

This paper cites Tubeformer-deeplab: Video mask transformer.

Autoregressive Universal Video Segmentation Model Tubeformer-deeplab: Video mask transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.464204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.497752Z digest=sha256:0afa478f3bbae5babd88140ac1bb803945bc4eb749e5b8a618441cc931c1ec7a

Observation e0771a88-4d8d-4fe6-ac32-a5f5d94b833d · outbound

This paper cites Visage: Video instance segmentation with appearance-guided enhancement.

Autoregressive Universal Video Segmentation Model Visage: Video instance segmentation with appearance-guided enhancement

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.447234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.502950Z digest=sha256:f6418b85dde298d0c8440aad168a50ae9549044fadf99493548e1a6ed0e0b5eb

Observation 1461a5e3-f4d3-4260-8759-8a302446bdf3 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

Autoregressive Universal Video Segmentation Model Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.425917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.508058Z digest=sha256:59d742ba4478a1ea5fccc5eb5d15d583c5a630b77754aad971fb105e16a1e238

Observation afc8ab21-2689-439d-b37f-0d5b105190ea · outbound

This paper cites Univs: Unified and universal video segmentation with prompts as queries.

Autoregressive Universal Video Segmentation Model Univs: Unified and universal video segmentation with prompts as queries

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.405124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.513868Z digest=sha256:b0748a109918032a3e4a80da9aa37d9c86f0977eb92903b21cb0046e17099fea

Observation fbd9a8f8-6014-45d1-b787-f737545eb36d · outbound

This paper cites Video k-net: A simple, strong, and unified baseline for video segmentation.

Autoregressive Universal Video Segmentation Model Video k-net: A simple, strong, and unified baseline for video segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.387623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.520167Z digest=sha256:0559e2917959909b1bf89cd7326511cbbb538f25b7d92b29df937bf5a76729b3

Observation 5bd90744-5653-4830-812d-1e2451df1eb7 · outbound

This paper cites Microsoft coco: Common objects in context.

Autoregressive Universal Video Segmentation Model Microsoft coco: Common objects in context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.368698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.526647Z digest=sha256:ac2c99a4ef13b51d38b1b25aae400e6dad4c8fc6194d7c8c5ff085cc10d19abe

Observation 51e38b0e-7f0e-4874-bb39-8c9463bf3a93 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Autoregressive Universal Video Segmentation Model Swin transformer: Hierarchical vision transformer using shifted windows

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.352602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.531289Z digest=sha256:b4e1ef57ae04a4a8e6f828d921eeef1d56074b5ba32b6fb0795303ad485d9d14

Observation 891f8809-64ee-4748-857a-a44bfab5c832 · outbound

This paper cites Decoupled weight decay regularization.

Autoregressive Universal Video Segmentation Model Decoupled weight decay regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.536513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.536513Z digest=sha256:9b07adf17bed853b3870f61e6ce123dfd4cd200a518602fd6b322d9cacedd4f7

Observation e0e281c9-c11e-4333-bc78-8b1c9d2113c8 · outbound

This paper cites Language Models are Few-Shot Learners.

Autoregressive Universal Video Segmentation Model Language Models are Few-Shot Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.541193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.541193Z digest=sha256:4967213cfb6da987625ab48c127ed19d8f89899b76dc39e539a045b9faf2c49b

Observation d723adf5-6981-48f6-a702-07d497284ddd · outbound

This paper cites Trackformer: Multi- object tracking with transformers.

Autoregressive Universal Video Segmentation Model Trackformer: Multi- object tracking with transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.326155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.546111Z digest=sha256:4b8248d809ebe0f1c174362bb5e72b3a55082e75a5b312b0ce9b7fed261ce1a0

Observation 39084838-b895-49db-b5e9-23042f0ea2ea · outbound

This paper cites MOT16: A Benchmark for Multi-Object Tracking.

Autoregressive Universal Video Segmentation Model MOT16: A Benchmark for Multi-Object Tracking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.551066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.551066Z digest=sha256:22c08ebb6a8caee877cd77940d91a14a70893c7068aec26af689b947f487fa4a

Observation 96b0c5b9-0c8b-46d3-a322-5775d6c38d19 · outbound

This paper cites Videoobjectsegmentationusingspace-time memory networks.

Autoregressive Universal Video Segmentation Model Videoobjectsegmentationusingspace-time memory networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.309936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.556170Z digest=sha256:1cc05f0e50b18f643dc7747f180ddbebe6f3003354be5c6e54cbee3cbc64bf59

Observation ffe03473-a0b6-406e-9123-b4b74133be84 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Autoregressive Universal Video Segmentation Model The 2017 DAVIS Challenge on Video Object Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.560810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.560810Z digest=sha256:0e5fdd9291b32b3f61fe7fe2817dbd79aea63a0bd1cfcfacf4dbc8f1ed8f6324

Observation 17033fd4-ea4f-40b0-a618-7550bc36d650 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Autoregressive Universal Video Segmentation Model Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.566222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.566222Z digest=sha256:626d765e5a23d6c5779a005a6d10660f227db259e49bdb0a39401dfa1becb5ca

Observation 97f69a1b-300c-4796-bfc5-f0a32e6de72c · outbound

This paper cites Occluded video instance segmentation: A benchmark.IJCV, 2022.

Autoregressive Universal Video Segmentation Model Occluded video instance segmentation: A benchmark.IJCV, 2022

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.292854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.571184Z digest=sha256:a2917ee94f13020b3f229ebf4814db1f09516a2c142883420914603dd8c4f026

Observation a27ccbf2-f2a7-4e03-9083-c16f18856fc8 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Autoregressive Universal Video Segmentation Model Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.277081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.576816Z digest=sha256:481b7046a9f140e67668df327c42d0ba8f63c87ba03b44ef7a4c52be4e05e9c8

Observation ddc4f406-bde8-4145-93bb-a8f4f4d39be0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Autoregressive Universal Video Segmentation Model Learning transferable visual models from natural language supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.261577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.581668Z digest=sha256:05c802aab3805102e11d38e2303eeb8ef3e8502146116225f8422d3f4497de0e

Observation 9cfa00d0-639a-4296-a769-b3d9dbd209c6 · outbound

This paper cites Sam 2: Segment anything in images and videos.

Autoregressive Universal Video Segmentation Model Sam 2: Segment anything in images and videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.244625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.586678Z digest=sha256:e6f1b0c07d70199e3bbaa7adcb1a327162d9821c77ff86aaf1f1ebdc38de18c5

Observation 5689917a-dba1-45b7-8618-18c96592be45 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Autoregressive Universal Video Segmentation Model Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.227593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.591657Z digest=sha256:a27f6fe0fb153373676915561b12df74d23e401d0dabc3b2aafe15f6e2c3e598

Observation 8b14a890-e01d-43b7-a6a8-ee613d1820c0 · outbound

This paper cites Repetition Improves Language Model Embeddings.

Autoregressive Universal Video Segmentation Model Repetition Improves Language Model Embeddings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.597131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.597131Z digest=sha256:c1cd0ce27cb4a9948a13339d48ba2730de62b22123ff1b7488b7a49a8e7816f8

Observation 57c5c273-9677-46b0-bf01-f88a32608fa7 · outbound

This paper cites Sequence to sequence learning with neural networks.

Autoregressive Universal Video Segmentation Model Sequence to sequence learning with neural networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.209269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.602369Z digest=sha256:606093aae37cbbb17d2fd723322e9e02d7a7277f8ee5a9c925caa08a24b51517

Observation 58d6caf9-96d9-4f37-8235-aa98213c07f8 · outbound

This paper cites Attention is all you need.

Autoregressive Universal Video Segmentation Model Attention is all you need

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.192463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.607296Z digest=sha256:69fe5777816b7788599b404807225b4f1f8eac66c8eb881197ba59704e47023a

Observation 471ab83e-d319-43c1-b713-0b7ad505bc3d · outbound

This paper cites Mots: Multi-object tracking and segmentation.

Autoregressive Universal Video Segmentation Model Mots: Multi-object tracking and segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.174807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.612009Z digest=sha256:5f5180cad82cb8c38d218b6a4243aa1e3cfa5fd87a8b9995d18c5d8f30d0b3fd

Observation fe02fd67-d773-4f2d-8c45-b940e4c0ba9f · outbound

This paper cites Max-deeplab: End-to-end panoptic segmentation with mask transformers.

Autoregressive Universal Video Segmentation Model Max-deeplab: End-to-end panoptic segmentation with mask transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.158195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.617095Z digest=sha256:b2e9c33407b924b642320878b5af5d8d4dd308f8a3e422b7c46ff4d1e89844f6

Observation 1b47c41b-c718-4f3a-9541-ffefbc1331ad · outbound

This paper cites End-to-end video instance segmentation with transformers.

Autoregressive Universal Video Segmentation Model End-to-end video instance segmentation with transformers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.141749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.622055Z digest=sha256:1a4b90e43e167af4696fe9bbe9acb33bd34bad4945aca72e9222876072c95558

Observation 88c1a9fb-1e2e-42e3-93d4-4ee22bdb26ff · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Autoregressive Universal Video Segmentation Model Chain-of-thought prompting elicits reasoning in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.122859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.627092Z digest=sha256:6fb79c35154e68a45148b3a55a7cb6de12bcc88c615a233ecdfd152358d95511

Observation 54cc4cc7-8dac-482e-8275-4e8218799bdf · outbound

This paper cites Segment every reference object in spatial and temporal spaces.

Autoregressive Universal Video Segmentation Model Segment every reference object in spatial and temporal spaces

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.105300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.632189Z digest=sha256:fe01b044429a4e5b6d3e14eb2ff5fc1116e99cf0ab21cd8902e046e105389774

Observation 12536e1d-f8d0-46ec-b7a0-6669087dc48b · outbound

This paper cites UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces.

Autoregressive Universal Video Segmentation Model UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.637223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.637223Z digest=sha256:6486987411adc535b8bec201718b0af60983b166f7ff2485a3672d30cf45e666

Observation b8c805a8-9f0d-4803-b494-4103fd5a4c78 · outbound

This paper cites In defense of online models for video instance segmentation.

Autoregressive Universal Video Segmentation Model In defense of online models for video instance segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.086960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.642391Z digest=sha256:3b7498b05421156ffcc0fa837fa1fd030efe817a311dd8fb1828e1c231c60295

Observation 03fabcbb-02d6-453d-814e-9316fd8407af · outbound

This paper cites Online object tracking: A benchmark.

Autoregressive Universal Video Segmentation Model Online object tracking: A benchmark

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.067808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.648003Z digest=sha256:06a77e84bcd007de5a63a4c7f2d9600d422d6cd48669c29ba8c9114d0967acc0

Observation 99af8edc-eff1-4366-8650-221d764ec550 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Autoregressive Universal Video Segmentation Model Efficient Streaming Language Models with Attention Sinks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.653696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.653696Z digest=sha256:d96fc3807f994078c87cc2fa3ad0b0826639911f0e60c6668465eb2e21b49906

Observation 69fb682e-a55a-4b1f-acab-f281f0b2c7b8 · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

Autoregressive Universal Video Segmentation Model YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.659221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.659221Z digest=sha256:64853a18a116610e94123ff49d43c88e7b0d37c17a5f48de0f93f76ce8c791d8

Observation 5eb3c589-7670-4912-8c20-f706e56b01f4 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Autoregressive Universal Video Segmentation Model Universal instance perception as object discovery and retrieval

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.050220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.664517Z digest=sha256:6061635321d1630f4e019d55da59856ce45b6fcc5a1cb7f88284fbf9f53444db

Observation 41f27a09-3114-4a91-aca5-58a7a9d5f584 · outbound

This paper cites Video instance segmentation.

Autoregressive Universal Video Segmentation Model Video instance segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.030729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.669522Z digest=sha256:c310f1701358e092dda2a146ac76d3a38f489bc9cec3e2b747584c8eba7a807c

Observation 7ca2f19f-1bc9-49ea-b5be-af608e24fff5 · outbound

This paper cites Decoupling features in hierarchical propagation for video object segmentation.

Autoregressive Universal Video Segmentation Model Decoupling features in hierarchical propagation for video object segmentation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.011198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.674355Z digest=sha256:8be0f92634a3e000f978588336b57fdaf55bbea58cc256e2063fe1a0bebec488

Observation 2e1593db-fa4d-45ad-a308-0788bd75f222 · outbound

This paper cites Associating objects with transformers for video object segmen- tation.

Autoregressive Universal Video Segmentation Model Associating objects with transformers for video object segmen- tation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:16.992951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.680179Z digest=sha256:1f9917c13990d4bea7aad0be76d4a7f033c9825d415319c3fd9fec92953433c9

Observation d813b1bd-be76-44b6-85a5-45ccd77f5dd2 · outbound

This paper cites Ctvis: Consistent training for online video instance segmentation.

Autoregressive Universal Video Segmentation Model Ctvis: Consistent training for online video instance segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:16.975710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.684828Z digest=sha256:450630b5756597622e543fc16cca7b96c790e76aa7fd7337e3b53a529a04e816

Observation aaf548a7-c725-4cef-932b-04afca00c196 · outbound

This paper cites Dvis: Decoupled video instance segmentation framework.

Autoregressive Universal Video Segmentation Model Dvis: Decoupled video instance segmentation framework

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:16.957792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T15:52:16.689829Z digest=sha256:1235cede07e039d0d55fc482730eb795850ee31c457aeefbc8beb8eeb1e4ad51

Pith citing papers

No inbound Pith citation observations are available.