Pith. sign in

Paper Citation Record · LEDGER

Track Anything: Segment Anything Meets Videos

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2304.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.11968 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:11:06.882718Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

95
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59e406f9-dff0-4aa7-87ed-8e702d7c6fa3 · inbound

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications cites this paper.

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications Track Anything: Segment Anything Meets Videos

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:41:43.492900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:41:43.411128Z digest=sha256:e40f6a355f3b4713d19cc794eeddcf04e030ea5ce71e17be52c1c1959d092c02

Observation 222e03e3-3644-4541-849b-0d4cb32b831b · inbound

On Efficient Variants of Segment Anything Model: A Survey cites this paper.

On Efficient Variants of Segment Anything Model: A Survey Track Anything: Segment Anything Meets Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.551644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:42:24.122342Z digest=sha256:1a9ec2c51261b7b867ca153dc57e2935ea2afbf9a2a53941f4bdba6f47bdcee5

Observation d2c6632d-85fc-494d-a12a-fb32a701b46e · inbound

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos cites this paper.

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos Track Anything: Segment Anything Meets Videos

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T08:02:43.822894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:58:54.015749Z digest=sha256:f14be3587a9f4418a89d3e1945c579e4fd175ec60ad370bddba7ef7c982194ae

Observation 333520f0-fb28-4812-906e-a1d179a6fdfa · inbound

SAM-guided Pseudo Label Enhancement for Multi-modal 3D Semantic Segmentation cites this paper.

SAM-guided Pseudo Label Enhancement for Multi-modal 3D Semantic Segmentation Track Anything: Segment Anything Meets Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:11:06.882718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:11:06.882718Z digest=sha256:a12c7d92ca20c17d83489cb56228cce96cabe3bca29a6896f176b65a3e27463b

Observation 914e81ab-71f6-44d6-a9ed-d965447b7a9e · inbound

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement cites this paper.

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement Track Anything: Segment Anything Meets Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:00.725533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:00.725533Z digest=sha256:c075413097c9a152d4a0307cb813ab3f1b5129b5597398364f5446433935d3ec

Observation 5efc0986-52ca-47c7-b1d9-10c3274f858e · inbound

CU-Multi: A Dataset for Multi-Robot Data Association cites this paper.

CU-Multi: A Dataset for Multi-Robot Data Association Track Anything: Segment Anything Meets Videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:19.959554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:19.959554Z digest=sha256:eaf1aacd6ef4fcd04cfbd43fe0ae33b04762737bd07ad1bacf54c91724aa1826

Observation 59f9617f-a5a4-4939-a621-9dcedf8a1970 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory Track Anything: Segment Anything Meets Videos

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T12:57:17.913143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:44733b6b33a1e4e89b0203cc908708981d83868412bbcbaaea1bbfdb71ab0c01

Observation 9ac3aeba-a833-4a41-8e9e-70ccdfed6daa · inbound

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost cites this paper.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Track Anything: Segment Anything Meets Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.056362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.056362Z digest=sha256:8e439e6a6e43e58430098771dc9e3151182add24392c5724bfec56506c974b45

Observation 340f00e3-ff2c-42a8-af1f-05f4ad41d936 · inbound

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting cites this paper.

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting Track Anything: Segment Anything Meets Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:16.361815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:16.361815Z digest=sha256:be8a5c65c72b76069059646ab7ed4998f4c8d8237d84bce1da5f4cb24002d100

Observation 3da692c5-cc8d-489c-a1d9-cee14765e42b · inbound

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References cites this paper.

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References Track Anything: Segment Anything Meets Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:06.851144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:06.851144Z digest=sha256:3f02f7b779454c0cb95c82d2171b00c06b4d307144bb46fd50caf17e7e22ed1d

Observation a3034808-69f0-46d6-b451-7ef6a9d78fd2 · inbound

R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision cites this paper.

R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision Track Anything: Segment Anything Meets Videos

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:57.411759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:47:57.411759Z digest=sha256:20b0f2393a25d44c7ac29404030f7b4ef37a5a259431257280fcd7359bdeb551

Observation 5962e6ba-592a-4d4b-a0da-765b77fa12db · inbound

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs cites this paper.

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs Track Anything: Segment Anything Meets Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:48.831105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:48.831105Z digest=sha256:5899396fbc39b5dc8e60bdf027f7abd6b3604bac5f48153ad64d01bde5e614e9

Observation 68a3a034-4946-4916-8007-372c479b0fe2 · inbound

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning cites this paper.

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning Track Anything: Segment Anything Meets Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:37.920077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:37.920077Z digest=sha256:93c6a96ac9ccaa9d2d4e62411162db3d5d2d2537c991f364babe024d98093d2b

Observation c66b9629-eb4d-4b1e-b7c0-17d197ee9f7f · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation Track Anything: Segment Anything Meets Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:57.025806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:57.025806Z digest=sha256:d76e9c6a2b70bf112b58769cb925abd08b0e1d7703dc427ce0abd2300422dc80

Observation 9543d116-36d2-4807-8b26-bcfeb0505aa5 · inbound

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation cites this paper.

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation Track Anything: Segment Anything Meets Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:46.955280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:46.955280Z digest=sha256:197e042ff23e767ed8f074b42b6a80455b59c807613c352820530344ce4b66d2

Observation 33b1a0ff-8793-40ba-b7e6-137eb96c29a9 · inbound

Continuous Marine Tracking via Autonomous UAV Handoff cites this paper.

Continuous Marine Tracking via Autonomous UAV Handoff Track Anything: Segment Anything Meets Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:36.662024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:36.662024Z digest=sha256:61d6fe17417563b6fa2818107eeef774679a3062621258378d5e7d2eff4ad8ea

Observation 4d88af1e-0a19-425c-8225-e696cc2273ec · inbound

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition cites this paper.

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition Track Anything: Segment Anything Meets Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:41.282511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:41.282511Z digest=sha256:a8025e6e1c5cf289849d9d4d271160b91d481411777dd8802718e1d0bd437197

Observation be47a831-ff2e-4435-9969-65cef60567e7 · inbound

3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction cites this paper.

3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction Track Anything: Segment Anything Meets Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T22:22:05.096000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:22:05.096000Z digest=sha256:0c2976090747f819b907c9b4522ae7a9ce0147175a27987113509d66cc7bedde

Observation e579c1b4-5e9d-4862-84a0-bcbba4f1976e · inbound

Grouped Speculative Decoding for Autoregressive Image Generation cites this paper.

Grouped Speculative Decoding for Autoregressive Image Generation Track Anything: Segment Anything Meets Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:10.519532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:57:10.519532Z digest=sha256:45b810186e10ef65295e94559785be290c359080c667acd3c4a91d8e1ed093e5

Observation dc07d712-d243-49bb-9167-63d0ab8cde11 · inbound

VoCap: Video Object Captioning and Segmentation from Any Prompt cites this paper.

VoCap: Video Object Captioning and Segmentation from Any Prompt Track Anything: Segment Anything Meets Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:18.011590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:18.011590Z digest=sha256:d5f57d23e2db54c6b1e1e0e134bcdaa3bc6573f5f113ec35e814a8e501be6a0c

Observation 6ebd9c47-9370-431a-b61d-da4d31740695 · inbound

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation cites this paper.

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation Track Anything: Segment Anything Meets Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.690261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T06:53:21.438159Z digest=sha256:fde290aa1369107ed6ab1e47d688ce39492c5039f431976d517e8e2a8a1b6a79

Observation 2cd0e4bc-c030-468d-8127-afb1d449134e · inbound

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors cites this paper.

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors Track Anything: Segment Anything Meets Videos

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:11:34.847331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T06:10:24.636998Z digest=sha256:fda58396a35d1cbfa888f57cbc69ab9c02c195ac096aa0c009b1a990d6aa2ec7

Observation 0a4b6bb3-d735-4366-98cd-29d83974c513 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Track Anything: Segment Anything Meets Videos

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.849678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:43b27192f2b91d3b05f807069c60bf7d55eb146870208d605e2586885b2a1b31

Observation ed35b158-6198-4fd3-b124-f41879769591 · inbound

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping cites this paper.

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping Track Anything: Segment Anything Meets Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.945749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T19:51:58.756605Z digest=sha256:b647d99ec8001ce2a22fb18c4b2ee544cedcf7c75cde00a524e42686c7578f38

Observation 0067fb33-9d7f-4214-83d9-ac1afe7a2e3c · inbound

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis cites this paper.

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis Track Anything: Segment Anything Meets Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T18:24:58.428701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:24:58.428701Z digest=sha256:86026db649e9655b7ec66b697fef20290d4546253416b5bdf1735b605b6446eb

Observation 97ac1bc3-7562-46b6-a5b3-5d9aec6c6867 · inbound

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation cites this paper.

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation Track Anything: Segment Anything Meets Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:51.506643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:20:04.690220Z digest=sha256:5db4c9668a04d2b0140e16304358d132fd2a919bd019d5961c2d953b2e22bea7

Observation 8d1f8330-4715-4b6e-8022-fd92d7f194f9 · inbound

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction cites this paper.

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction Track Anything: Segment Anything Meets Videos

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:00.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:20:54.028846Z digest=sha256:14ab2128bf2905655388f69226e01b82130523b7c12b294658d2c2dc5a6e7d11

Observation 09bf1170-c3c3-41d3-875f-f9003e323fbb · inbound

Do Instance Priors Help Weakly Supervised Semantic Segmentation? cites this paper.

Do Instance Priors Help Weakly Supervised Semantic Segmentation? Track Anything: Segment Anything Meets Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:03.251934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:23:15.172519Z digest=sha256:a150bb46f30812dfc85e620e23c314b119be3039f2f0f5f4344a38db02081ae3

Observation 9f126d71-9aa3-43ad-b798-6358d6c3dcea · inbound

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition cites this paper.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Track Anything: Segment Anything Meets Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.879222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:99a2c4d0a9f2fd58470d1c7f53f2a37851adc85f9a3c9cf82bfb43f6dc9f5c1c

Observation 27a0bc2a-a6db-4545-9e7c-e500ce02400a · inbound

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement cites this paper.

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement Track Anything: Segment Anything Meets Videos

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:11:20.977325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:09:34.131634Z digest=sha256:c43720aa94a590a1e2369063950136839496fee9c835be09fbe1cab86c28a2c0

Observation fe5c1906-bfee-4183-a40a-62fd8c924c85 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Track Anything: Segment Anything Meets Videos

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.687252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:4ff86ed3340164f409cc225d057b0b65972c24525775d485846275f665cba0e9

Observation f6be40f8-8ff4-4b30-8f4d-dee55b3d97e6 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Track Anything: Segment Anything Meets Videos

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.258081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:58553cba1b837a5cb9539d13341ac6d5ac82ecb698e905a21a52ff16df7c273f

Observation f19ecb62-6f13-4fc1-b92c-ab6e4b28bd36 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Track Anything: Segment Anything Meets Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:36.587306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:a2d1135da89dedb6e6ca1812df88f37030413ec3ae747d3d5402428564692dd9