Pith. sign in

Paper Citation Record · LEDGER

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

As of 22 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2510.14904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.14904 v4

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:31:56.230128Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T17:27:15.695591Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9d496f5-0135-4d66-9a52-e266713f2900 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:54.671015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:54.671015Z digest=sha256:6cb89f470a714e6c70094bf55f2db5c18f0797c98f0e5b03120d68a192a56aca

Observation f5cb5944-ad01-445f-9ada-49d3245f2506 · outbound

This paper cites Qwen Technical Report.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:54.979216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:54.979216Z digest=sha256:4c362964ccd0f177d4db34a78ca8b1f666e469b9e4e0f1e5d35e968dc2adf694

Observation c1f49450-7e1b-424f-b3e3-caf108aba6aa · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.130506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.130506Z digest=sha256:f60fb8195121c6103be35b9998ca97898eebbe3d70d69ef5bd4d96cd7caeb2ba

Observation 10c61aa5-734c-4aec-b0a5-643ae63bd371 · outbound

This paper cites Mask2Former for Video Instance Segmentation.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Mask2Former for Video Instance Segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.211242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.211242Z digest=sha256:8a5661c8c8cbfec3bd6e6eee084f6a8e1c5e72939e49fd45a876dac592d3cd8d

Observation f0f2e32d-d78b-47c8-962d-16790570e484 · outbound

This paper cites MOT16: A Benchmark for Multi-Object Tracking.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects MOT16: A Benchmark for Multi-Object Tracking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.405737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.405737Z digest=sha256:23d04b5a34bb5e5c6f4e029e4fe81271ec181035ec76fcdd1fa1988c7400cea9

Observation a7f2d587-583a-4d8d-ab65-175b7edd4ba3 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.677230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.677230Z digest=sha256:55687abc9ca1c38ab50584b568b5b96b188961eb73bc9192f84b38bbec97eb18

Observation c086e104-ba0c-4d83-871c-b98f3eecf8f9 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects OPT: Open Pre-trained Transformer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.764860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.764860Z digest=sha256:0261e0a47c2eaea82d108dd8a762da9f1ed8da64ef3b5da6236d70307ac6ebde

Observation 1b28ff36-1485-409a-977a-e2fd82a4212d · outbound

This paper cites Objects as Points.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Objects as Points

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.854600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.854600Z digest=sha256:fdbe7a30837eb508dac2c737a31a3f3327bfe0b13074943db1ec1235d3ed3c2d

Observation df0af9d0-9f62-4bb3-a8cb-6124ff427a92 · outbound

This paper cites We attribute this difference to the numerous objects that disappear for a significant number of frames in the long videos of VidSTG.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects We attribute this difference to the numerous objects that disappear for a significant number of frames in the long videos of VidSTG

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.968091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.968091Z digest=sha256:b9b8ce231c82a895222a1d7c920e2b734f70979cec4e7a934747e0dc9007bfa0

Observation f73a68d5-7146-42a6-9c2a-747a8c1d5d89 · outbound

This paper cites pair of tongs.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects pair of tongs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:56.055146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:56.055146Z digest=sha256:174111346544c0d587e08264686bb87d2cba6fddd21867ba1d7a1d2831f18e8f

Observation e257921a-b5e1-4351-805f-d89416021cfc · outbound

This paper cites bottle":.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects bottle":

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:56.125904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:56.125904Z digest=sha256:58d060a5ae3e13cf95a0c7eeb50d4d12529989b4f31bdb3c735ef070421343a4

Observation c70f8275-a183-47bf-ae31-e5005b7c144c · outbound

This paper cites For VidSTG/VLN/BenSMOT experiments we use video-level tuning for captioning with temporal aggregation, with Tagg = 32/8/8 respectively.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects For VidSTG/VLN/BenSMOT experiments we use video-level tuning for captioning with temporal aggregation, with Tagg = 32/8/8 respectively

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:56.230128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:56.230128Z digest=sha256:1b21c68a5ddcb122bb2aaa0eca729b3c097d982cddffcd868bb25bff6510be5d

Observation 02b19bb7-47ed-45df-af26-832d62963239 · outbound

This paper cites Dreamix: Video Diffusion Models are General Video Editors.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Dreamix: Video Diffusion Models are General Video Editors

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.488204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.488204Z digest=sha256:9a45f03de06cee0b37831d74bb0c839e6af353aba5efb46d12f4fc900411f63e

Observation c31615e6-fe7c-4042-9204-7fcb77738f34 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Gemini: A Family of Highly Capable Multimodal Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.556728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.556728Z digest=sha256:6ae3ce729d70900373bcb2acfa22d2be082a31cdbc0f722560059ec3f3cca093

Observation 484453aa-2b37-403e-91e9-67c6d07ab0b1 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.057137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.057137Z digest=sha256:eb119dc8462b98213dcd2147087193bbdf8d0256f8ca65fabd49f57b4b269f8d

Observation 46532c84-7337-422b-933d-91342ba3ec0f · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:54.874954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:54.874954Z digest=sha256:0a17dc359e61ed9609da62abdda59b1d2b414a4c0a27961ac4649625b36a828f

Observation cc5e43c7-6ca4-42bb-9704-686e15bb2390 · outbound

This paper cites GPT-4 Technical Report.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects GPT-4 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:54.797674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:54.797674Z digest=sha256:8307b09275912881b8b278da3d18871470f458b35e7b15baf5c59af5b22fe4b3

Observation 50f1efc0-2eac-420c-bf07-6137a9194470 · outbound

This paper cites The Llama 3 Herd of Models.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.311502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.311502Z digest=sha256:8f254941873cb0090c3b60203c81970e29aafd126533c42247098eba3413d582

Pith citing papers

Observation f69ea9cf-2ddc-4df9-b101-ba497e1521bf · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

Reference 119

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.697142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:0f6aaf9c0b640b2cdc347bc93f681d5ee204a8cb1fd0d6b401545787d14394a9