Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:31:56.230128Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2510.14904.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:31:56.230128Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T17:27:15.695591Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b9d496f5-0135-4d66-9a52-e266713f2900 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cb5944-ad01-445f-9ada-49d3245f2506 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f49450-7e1b-424f-b3e3-caf108aba6aa · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c61aa5-734c-4aec-b0a5-643ae63bd371 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Mask2Former for Video Instance Segmentation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f2e32d-d78b-47c8-962d-16790570e484 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects MOT16: A Benchmark for Multi-Object Tracking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f2d587-583a-4d8d-ab65-175b7edd4ba3 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c086e104-ba0c-4d83-871c-b98f3eecf8f9 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects OPT: Open Pre-trained Transformer Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b28ff36-1485-409a-977a-e2fd82a4212d · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Objects as Points
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0af9d0-9f62-4bb3-a8cb-6124ff427a92 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects We attribute this difference to the numerous objects that disappear for a significant number of frames in the long videos of VidSTG
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f73a68d5-7146-42a6-9c2a-747a8c1d5d89 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects pair of tongs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e257921a-b5e1-4351-805f-d89416021cfc · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects bottle":
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70f8275-a183-47bf-ae31-e5005b7c144c · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects For VidSTG/VLN/BenSMOT experiments we use video-level tuning for captioning with temporal aggregation, with Tagg = 32/8/8 respectively
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b19bb7-47ed-45df-af26-832d62963239 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Dreamix: Video Diffusion Models are General Video Editors
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c31615e6-fe7c-4042-9204-7fcb77738f34 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Gemini: A Family of Highly Capable Multimodal Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 484453aa-2b37-403e-91e9-67c6d07ab0b1 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46532c84-7337-422b-933d-91342ba3ec0f · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5e43c7-6ca4-42bb-9704-686e15bb2390 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects GPT-4 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f1efc0-2eac-420c-bf07-6137a9194470 · outbound
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects The Llama 3 Herd of Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69ea9cf-2ddc-4df9-b101-ba497e1521bf · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.