Pith. sign in

Paper Citation Record · LEDGER

Track Anything: Segment Anything Meets Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2304.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.11968 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:29:00.725533Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

95
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59e406f9-dff0-4aa7-87ed-8e702d7c6fa3 · inbound

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications cites this paper.

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications Track Anything: Segment Anything Meets Videos

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:41:43.492900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:41:43.411128Z digest=sha256:318b1935fb31521b161a41a76fbfc5c2e941f683d7d206a54dd09650758228ae

Observation 222e03e3-3644-4541-849b-0d4cb32b831b · inbound

On Efficient Variants of Segment Anything Model: A Survey cites this paper.

On Efficient Variants of Segment Anything Model: A Survey Track Anything: Segment Anything Meets Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.551644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:42:24.122342Z digest=sha256:13e3703b6b21d5e5a1a987c30b2e5d7a5f279ab23f91433a8e61d9a4d7021595

Observation d2c6632d-85fc-494d-a12a-fb32a701b46e · inbound

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos cites this paper.

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos Track Anything: Segment Anything Meets Videos

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T08:02:43.822894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:58:54.015749Z digest=sha256:747c80219629c469a95cc99f1f92dc2414c9b5f0132c4ae162902649aa1d2c83

Observation 914e81ab-71f6-44d6-a9ed-d965447b7a9e · inbound

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement cites this paper.

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement Track Anything: Segment Anything Meets Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:00.725533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:00.725533Z digest=sha256:9135d63ce936d97101686a60a8d5da1d5b60ea7c9f580ee894d6e073e3713e9d

Observation 5efc0986-52ca-47c7-b1d9-10c3274f858e · inbound

CU-Multi: A Dataset for Multi-Robot Data Association cites this paper.

CU-Multi: A Dataset for Multi-Robot Data Association Track Anything: Segment Anything Meets Videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:19.959554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:19.959554Z digest=sha256:7ebc388175eaab232f483da08eaaf6877e7f945f17f578aafbb8df9ec4c6b67c

Observation 59f9617f-a5a4-4939-a621-9dcedf8a1970 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory Track Anything: Segment Anything Meets Videos

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T12:57:17.913143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:d0eb227747128f808c81acaac2a40d16c4cb43a3e68abaedcc9fefa978020043

Observation 9ac3aeba-a833-4a41-8e9e-70ccdfed6daa · inbound

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost cites this paper.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Track Anything: Segment Anything Meets Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.056362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.056362Z digest=sha256:2a244084b315bb792e83bf426a20086e69c29fb4ebb58545e542fbeacdfec81f

Observation 340f00e3-ff2c-42a8-af1f-05f4ad41d936 · inbound

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting cites this paper.

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting Track Anything: Segment Anything Meets Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:16.361815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:16.361815Z digest=sha256:788ccc718332015448adfd3df8fa63efba04a14eefd9a764b83ad62b08f4037f

Observation 3da692c5-cc8d-489c-a1d9-cee14765e42b · inbound

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References cites this paper.

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References Track Anything: Segment Anything Meets Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:06.851144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:06.851144Z digest=sha256:501f7e179f56a4239631059759cb461e6144a505b861d7d8beb86e378bfc6061

Observation a3034808-69f0-46d6-b451-7ef6a9d78fd2 · inbound

R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision cites this paper.

R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision Track Anything: Segment Anything Meets Videos

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:57.411759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:47:57.411759Z digest=sha256:c9f34e2c6dea0eb55d244c25300446ef0b6addd07383f847b4a2dec15f03ea2f

Observation 5962e6ba-592a-4d4b-a0da-765b77fa12db · inbound

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs cites this paper.

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs Track Anything: Segment Anything Meets Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:48.831105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:48.831105Z digest=sha256:5d1fb0f0b65d62c91c2acd82096acccbda9e243f015fa8a75d1d5faa685a6e72

Observation 68a3a034-4946-4916-8007-372c479b0fe2 · inbound

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning cites this paper.

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning Track Anything: Segment Anything Meets Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:37.920077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:37.920077Z digest=sha256:ab483ed5c8d906c20b22507bee8459088fb4d1507b660df70a817974d24ad651

Observation c66b9629-eb4d-4b1e-b7c0-17d197ee9f7f · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation Track Anything: Segment Anything Meets Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:57.025806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:57.025806Z digest=sha256:0556405b8e1a5017b1c5eb58ce8cb0151acd58c56e1cbc922e8ac16a9243a3d9

Observation 9543d116-36d2-4807-8b26-bcfeb0505aa5 · inbound

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation cites this paper.

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation Track Anything: Segment Anything Meets Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:46.955280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:46.955280Z digest=sha256:f0c50857ff69dde3b95a41249744b444d31ec621a6e8137bf264eaa40e571d4d

Observation 33b1a0ff-8793-40ba-b7e6-137eb96c29a9 · inbound

Continuous Marine Tracking via Autonomous UAV Handoff cites this paper.

Continuous Marine Tracking via Autonomous UAV Handoff Track Anything: Segment Anything Meets Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:36.662024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:36.662024Z digest=sha256:1fbe5ff0590e1d7901dd7dded3b31f7b2a162941d1cbedfc26c9e5cf548f8644

Observation 4d88af1e-0a19-425c-8225-e696cc2273ec · inbound

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition cites this paper.

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition Track Anything: Segment Anything Meets Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:41.282511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:41.282511Z digest=sha256:2c4c25304615c0e01c83b806cf3747a868f03a9482475bbabf500e8bf4ccb193

Observation be47a831-ff2e-4435-9969-65cef60567e7 · inbound

3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction cites this paper.

3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction Track Anything: Segment Anything Meets Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T22:22:05.096000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:22:05.096000Z digest=sha256:e712f56044f4ba43d9d33b9bb9a8fc09c1747b5bd68589469f6ed2a35304167b

Observation e579c1b4-5e9d-4862-84a0-bcbba4f1976e · inbound

Grouped Speculative Decoding for Autoregressive Image Generation cites this paper.

Grouped Speculative Decoding for Autoregressive Image Generation Track Anything: Segment Anything Meets Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:10.519532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:57:10.519532Z digest=sha256:a432457e52a48000dce65911f71dec076c89ec9a6803ee8286869c987ad0a446

Observation dc07d712-d243-49bb-9167-63d0ab8cde11 · inbound

VoCap: Video Object Captioning and Segmentation from Any Prompt cites this paper.

VoCap: Video Object Captioning and Segmentation from Any Prompt Track Anything: Segment Anything Meets Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:18.011590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:18.011590Z digest=sha256:c4289a6f174e80a5b0290eefebb8a62a0b7eb4a98e00d8ad5356f8eb103ec583

Observation 6ebd9c47-9370-431a-b61d-da4d31740695 · inbound

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation cites this paper.

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation Track Anything: Segment Anything Meets Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.690261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T06:53:21.438159Z digest=sha256:15095ae49883c6d15c7527c5c8710dc0786e3c307710453cda8efc50771f5371

Observation 2cd0e4bc-c030-468d-8127-afb1d449134e · inbound

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors cites this paper.

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors Track Anything: Segment Anything Meets Videos

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:11:34.847331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T06:10:24.636998Z digest=sha256:2b04ae645374e2a5b805466e73ac70ac49ace21307d4393f06e7833160ae7f37

Observation 0a4b6bb3-d735-4366-98cd-29d83974c513 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Track Anything: Segment Anything Meets Videos

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.849678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:e519b8bedf961f8d27d674e22281bdc73fc5493580bddba532a254e3fb61714d

Observation ed35b158-6198-4fd3-b124-f41879769591 · inbound

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping cites this paper.

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping Track Anything: Segment Anything Meets Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.945749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:51:58.756605Z digest=sha256:a2267aed4add57045ddf49e889795744edd59bf3924c5d52223f6eaface0454f

Observation 0067fb33-9d7f-4214-83d9-ac1afe7a2e3c · inbound

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis cites this paper.

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis Track Anything: Segment Anything Meets Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T18:24:58.428701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:24:58.428701Z digest=sha256:99528d6bac8929868aadc8895a706b9fa8feaa4ccd814247bea23ad60f352ae4

Observation 97ac1bc3-7562-46b6-a5b3-5d9aec6c6867 · inbound

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation cites this paper.

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation Track Anything: Segment Anything Meets Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:51.506643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:20:04.690220Z digest=sha256:a7f7ee2ba943e46aad58a48d048a6b85088c9c183c07a42708f1b1e19815b011

Observation 8d1f8330-4715-4b6e-8022-fd92d7f194f9 · inbound

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction cites this paper.

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction Track Anything: Segment Anything Meets Videos

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:00.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:20:54.028846Z digest=sha256:c465d160f71ba98015ef64983eb1f436f46798b5d048f69c5fb7eaac91e2fe4b

Observation 09bf1170-c3c3-41d3-875f-f9003e323fbb · inbound

Do Instance Priors Help Weakly Supervised Semantic Segmentation? cites this paper.

Do Instance Priors Help Weakly Supervised Semantic Segmentation? Track Anything: Segment Anything Meets Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:03.251934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:23:15.172519Z digest=sha256:cfd8f7ca69376f216244bb1626a67cab92dcf00a86c6ef93019333c65ddddf2f

Observation 9f126d71-9aa3-43ad-b798-6358d6c3dcea · inbound

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition cites this paper.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Track Anything: Segment Anything Meets Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.879222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:e65e0d2f750ee7fe9286c9f51b0b998cc1eb871d4b498532a59de78a25bc786d

Observation 27a0bc2a-a6db-4545-9e7c-e500ce02400a · inbound

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement cites this paper.

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement Track Anything: Segment Anything Meets Videos

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:11:20.977325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:09:34.131634Z digest=sha256:1fe058ca481119e285ee6fe88b6bd69983696aecb0f7afe0c1ca5a03f4580d25

Observation fe5c1906-bfee-4183-a40a-62fd8c924c85 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Track Anything: Segment Anything Meets Videos

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.687252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:2b550f0ad4a415d214e21e31b3ce3c3417ad14b3dfcb7995439914786b8724b6

Observation f6be40f8-8ff4-4b30-8f4d-dee55b3d97e6 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Track Anything: Segment Anything Meets Videos

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.258081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:fe7e97f0defc55bee82c2742f6c328f432fff8d211bd9a4cd485b28504441b49

Observation f19ecb62-6f13-4fc1-b92c-ab6e4b28bd36 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Track Anything: Segment Anything Meets Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:36.587306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:cfdd17cf044a529ddbc3f8585e8645fcf33746efd2e35fa4f83aa774b5afb5ae