Pith. sign in

Paper Citation Record · LEDGER

Audiovisual SlowFast Networks for Video Recognition

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2001.08740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2001.08740 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.185270Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:57:47.660611Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 89b87024-bfb2-4ced-b430-40c7b94df63b · inbound

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models cites this paper.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models Audiovisual SlowFast Networks for Video Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.185270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.185270Z digest=sha256:3b13e162f1ea34ee13ba6bf2802c8b62c84606793445357d660a5ce79b98c5a5

Observation 1e498db1-61bb-4662-9211-c4c727c6e716 · inbound

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision cites this paper.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Audiovisual SlowFast Networks for Video Recognition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.733808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.733808Z digest=sha256:35630e71f6345724106c4f761ea1cf4baf7a259b393bb2bc450aacd6a411cd4f

Observation 0bd3bc9d-c4ef-470e-ac41-5a75624ca8e4 · inbound

DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation cites this paper.

DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation Audiovisual SlowFast Networks for Video Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:12.311803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:12.311803Z digest=sha256:ba5369991cae7a98643b0c3d3c50d430ff164e019986fb5ec825d0d59579b36b

Observation eeeb6297-1c94-4403-b82a-165461958cee · inbound

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding cites this paper.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audiovisual SlowFast Networks for Video Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.288755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.288755Z digest=sha256:07bbd5cdf9fcc1fe9f1d7cde4a66b773ed3fd60f6acfd567c034cdc18800b303

Observation 56600a93-ee73-4a74-9978-01870ce512b2 · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual SlowFast Networks for Video Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.945121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.945121Z digest=sha256:b5ed1930977d1419703b81863c9f36f7c6eafa01b182deac340baf1733b56793

Observation d4f48735-2087-4877-8af9-04073b46d0bd · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audiovisual SlowFast Networks for Video Recognition

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.240129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.240129Z digest=sha256:092148404bc9b007e33e7e651c980cc2aa1ffc57e27e4149b599957dbd2965ad

Observation 54681c95-aeda-4cf9-aca8-04ec0f6c8b98 · inbound

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment cites this paper.

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment Audiovisual SlowFast Networks for Video Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:23.939472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:23.939472Z digest=sha256:748dabf66c19618355b81d2f090534ab0afee965f5c87934261ccb119eb8f95f

Observation 67a76d82-262b-4c59-a43b-2d046a5d4749 · inbound

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection cites this paper.

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection Audiovisual SlowFast Networks for Video Recognition

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:13:16.640679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:10:01.391718Z digest=sha256:070ba5005809efac9f67f27f673b2d1c19ab8e8756f0b89cfb2e886e8f020e1f

Observation 7cbf2acf-08ef-4009-9fd3-f3db5b8f0b6c · inbound

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization cites this paper.

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization Audiovisual SlowFast Networks for Video Recognition

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.544513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:00:44.575524Z digest=sha256:627ebeca7d491f2549242a1f9199e26047b7f9d7a55ab1cc7e5859bb98f70c40

Observation 4809551d-1056-40b8-b724-475f4aee7105 · inbound

USV: Towards Understanding the User-generated Short-form Videos cites this paper.

USV: Towards Understanding the User-generated Short-form Videos Audiovisual SlowFast Networks for Video Recognition

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.827901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T04:52:12.045880Z digest=sha256:83a82c118bf0b4be1b6f9a5de29ed071a0e3136e0ecc2f2e386057dc57004c81

Observation d10bf92b-4f89-4da0-bbc0-aa8103feef67 · inbound

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning cites this paper.

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning Audiovisual SlowFast Networks for Video Recognition

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:06.125333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T21:55:53.630732Z digest=sha256:4e6374208f6053c91f598abc6541c2f8a6182c8dd6c24b0aee5785ffaba8653a

Observation 8a8885ba-b003-423e-a10c-cc72a3821ff8 · inbound

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning cites this paper.

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning Audiovisual SlowFast Networks for Video Recognition

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.661964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:39:37.920953Z digest=sha256:25a80fa1f436a32e4022c6685f7b036c30274e419a82d808b15e709a1052fb0a