Pith. sign in

Paper Citation Record · LEDGER

You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:1911.06644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1911.06644 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:34.042842Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:25:41.714158Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 13de2aab-4ba6-495f-9b9c-b9dac35bac06 · inbound

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection cites this paper.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.042842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.042842Z digest=sha256:3fd6011683bf140c51e7b03d3e3611be922958e26f5dde687262dd0a94f9659e

Observation 69e195b8-e658-44f5-b79e-cc47ee928385 · inbound

Stable Mean Teacher for Semi-supervised Video Action Detection cites this paper.

Stable Mean Teacher for Semi-supervised Video Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:16:59.984438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:16:59.984438Z digest=sha256:34f02577ed228aed83a322fd19f99cf1ce50aba85b612d03d6c18b7f0e33c25a

Observation b1617dc8-b33e-4195-878b-a8bd6d4a1038 · inbound

JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts cites this paper.

JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:55:40.531944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:55:40.531944Z digest=sha256:79242a789c9c6e68aed4d8a5f248d94242b10d653aa6274d005af848aea0e92b

Observation f05c6c33-4616-4aad-b5a9-a3e2e6bc82cb · inbound

Dual Guidance Semi-Supervised Action Detection cites this paper.

Dual Guidance Semi-Supervised Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:01:53.490326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:01:53.490326Z digest=sha256:3880c06b4770cef94a9fce042c6b4367e47c3b28ee16fb831b4b2b572a39155d

Observation 3417e466-8583-4423-9037-0ad64055c02c · inbound

MOVE: Motion-Guided Few-Shot Video Object Segmentation cites this paper.

MOVE: Motion-Guided Few-Shot Video Object Segmentation You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:16.073610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:16.073610Z digest=sha256:2d7ff920bb577385a70a3eed37ffb05177c8f716284d5e2f2d498e5e507b69db

Observation caa151e8-93d8-488a-81eb-a029f452c115 · inbound

VC-FeS: Viewpoint-Conditioned Feature Selection for Vehicle Re-identification in Thermal Vision cites this paper.

VC-FeS: Viewpoint-Conditioned Feature Selection for Vehicle Re-identification in Thermal Vision You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:07.183701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T17:15:27.248533Z digest=sha256:79e4743d0ae2f0dd49f8ff60aee6f0720c070f2f2023ad7146c0d55600ab401c

Observation 82b014a3-0097-423c-8734-ab12b91ca0a5 · inbound

Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors cites this paper.

Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.715580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T05:31:23.544257Z digest=sha256:29fda8410a96ec9a277fbcde858991a05cba3c1176dfa7f52fd07fc7aac89ec7

Observation 9770819f-a305-4b28-a902-e128a8de72c3 · inbound

TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection cites this paper.

TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T15:14:47.300105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:14:47.300105Z digest=sha256:370338db02bd297dcc2c91f3961b4e94f31bf3821fe0e085b1c8399bbecd7c07