Pith. sign in

Paper Citation Record · LEDGER

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

As of 20 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.01237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01237 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:32.991891Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab5ba190-b74b-4850-853f-0225be2800ec · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised learning of audio-visual objects from video

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.775790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.775790Z digest=sha256:2e93b7440190ca721094b81679ff5f04205b351d2a6ac124cd6e718d363f208a

Observation 73b9163a-48bc-4536-91e9-ba49dd15249f · outbound

This paper cites Look, listen and learn.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen and learn

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.820348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.780800Z digest=sha256:b7bd432bfabe7a3be0caf3596199d3eb4e9850f9ecead4f3a975f1dbba457b53

Observation 8721d22e-4783-44c8-99e1-4be0d12cd5fd · outbound

This paper cites Objects that sound.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Objects that sound

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.805349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.785314Z digest=sha256:e8e7159beae5e68a3fde758ab52147f9f145ceb620ea8856a4fa70fa54b77822

Observation abb8d35f-e874-40ce-b211-b62f120a2b8e · outbound

This paper cites Sound- net: Learning sound representations from unlabeled video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Sound- net: Learning sound representations from unlabeled video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.790489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.789973Z digest=sha256:a74333ed9a41ccaf475109a9b1c80e2d76d0423ee452e59ebb7c0d6044309b42

Observation 59c34e70-af9a-468a-94aa-9aac795b5112 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.794959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.794959Z digest=sha256:91ecb8decc3dc95dcc41ae142c962fc5450ee6b7c7296feb65354a6891b46027

Observation 9420bd3c-3674-43c4-8681-a06a2d76b89d · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vggsound: A large-scale audio-visual dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.766400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.799505Z digest=sha256:5e5dc28a3dc6083d51820c588b795a5476675ad6cb5e6e2e61462fa010558850

Observation f19c7b18-3466-41f9-b614-78e182dda387 · outbound

This paper cites Localiz- ing visual sounds the hard way.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Localiz- ing visual sounds the hard way

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.751761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.804317Z digest=sha256:bebdfea4470226fa67382adab6aa8b5fc7cc4644e5734d3cf9e18c1a441ac6c7

Observation aba398b6-c9ce-4d75-a004-151350fab2e0 · outbound

This paper cites Distilling audio-visual knowledge by com- positional contrastive learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Distilling audio-visual knowledge by com- positional contrastive learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.738412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.809033Z digest=sha256:44747e868cac5757383a26cce99314c0ccf177e0e71e30b7f57e09a54b80ab81

Observation 389a656d-d290-4307-9cbf-a00ce46c72ff · outbound

This paper cites Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.724307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.813514Z digest=sha256:9b64198d8f90560cf6623a0511aab70270a2faa8270aa50dd67090620f0d3ad8

Observation cc1b77d1-1f77-4738-9521-0773c9fdb621 · outbound

This paper cites Vision transformers need registers.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers need registers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.711043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.817897Z digest=sha256:bab746fcc6919646d0043207e11c9c9b3964eda7a715e52df1bd405715114a1b

Observation 29d9ea60-eb9f-47fc-998d-7f82934b04ed · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio set: An ontology and human- labeled dataset for audio events

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.697065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.822446Z digest=sha256:e40aea5a572fd109fcf8f5820f747f75749a2c93591a9ec1864dec29153d29b7

Observation 3cd7ed33-9754-4237-a3c0-202be15b27a7 · outbound

This paper cites Audiovisual masked autoencoders.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audiovisual masked autoencoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.683003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.826992Z digest=sha256:f63d068521a94501e623fc61122ddcf509d8df39d03f8c8dd2f3a83bf5ddef57

Observation fa9b2a8a-99e5-4572-96a7-4ba46a39fa6e · outbound

This paper cites Imagebind: One embedding space to bind them all.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Imagebind: One embedding space to bind them all

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.669127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.831640Z digest=sha256:8c4bcfa2935d9a91c73222fe99d537727a82555555e4c4d3737671cf69a1145e

Observation 6c0f4c33-db18-466c-aa5d-712e26f1a4f8 · outbound

This paper cites Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.654970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.835914Z digest=sha256:6394c7d91ba9de0736df4ffa6d76b129e2131fd17c620e650a8e9bd09eb4296b

Observation 7003157a-d225-473d-adb4-ac16868a791a · outbound

This paper cites Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.639798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.840157Z digest=sha256:e8afec550d6681f751235b5154ad63067c341358955e5c9d7f4d62ffd9ae0489

Observation 7485129c-84de-4caa-ae10-4b5e257db1a9 · outbound

This paper cites Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.626057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.844720Z digest=sha256:94a575137a786d5edb3186b69853147d8bfebdea1b522868c17513b27ce05b58

Observation 1e006760-16ec-420d-bc6b-b7abd2a55357 · outbound

This paper cites Jointly dis- covering visual objects and spoken words from raw sensory input.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Jointly dis- covering visual objects and spoken words from raw sensory input

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.611946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.848959Z digest=sha256:d4afee92e09fba70153b2ba657ff697a784fa3ddac4491e43577084b19eff3bf

Observation 78ea52a8-0f76-417b-8399-76e4059b0a34 · outbound

This paper cites Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.596598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.853357Z digest=sha256:d8e1f30e9769d76f1d9cc7c86ad35c950145a27725957ebfc9a812e1eb9e0c0b

Observation 30f027fb-4e56-4c69-b832-00867cb4bbb8 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.581567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.857551Z digest=sha256:8f4fa1b6371845e55b3302023a2c8ba20305a968e8cb7519efc417d5d43f0e49

Observation a20842a8-ea8c-4cea-b304-763c02564db4 · outbound

This paper cites Mavil: Masked audio-video learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Mavil: Masked audio-video learners

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.457498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.862296Z digest=sha256:7912d776a85fa8e40b68848f11bca6537350769bd8e3e5172a9e8e4b59d9cd73

Observation 52136e88-5727-4d12-b8b6-e0e58f7aa1f8 · outbound

This paper cites EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.442031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.866833Z digest=sha256:4f429d6956f7a7e5231514ffda99433cd27ad012482857c045f555c84f4bfee3

Observation a6e05e86-f8c8-4a25-98b3-0ccfdbc1768a · outbound

This paper cites Coopera- tive learning of audio and video models from self-supervised synchronization.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Coopera- tive learning of audio and video models from self-supervised synchronization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.426938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.871032Z digest=sha256:e00840fa4fd434767483e9cadb3332e4d82e4f32f93c92d36eaabedb837bf8a3

Observation 4a87a967-a4e1-4ec5-8d90-9d8ce76aca57 · outbound

This paper cites Cross-attentional audio-visual fusion for weakly- supervised action localization.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross-attentional audio-visual fusion for weakly- supervised action localization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.875357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.875357Z digest=sha256:9db7438397470ff555aa2b82087ddff262fc0bec29d19216bc33d95612721876

Observation 97cfe34f-ac02-4013-af24-f9410fbd6eeb · outbound

This paper cites Siamese vision transform- ers are scalable audio-visual learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Siamese vision transform- ers are scalable audio-visual learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.403313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.879753Z digest=sha256:fa1bf4947580dc1a58807e8e7cdd1ee4c1c3305622761458cb8e8dbf7602064b

Observation 19cb27be-d1c7-44cd-bc3d-7b3d14033aa8 · outbound

This paper cites Vision transformers are parameter-efficient audio- visual learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers are parameter-efficient audio- visual learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.389148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.883993Z digest=sha256:b58b399f99681027203dd5cd999050ba56612a9069e04888117530243aececa0

Observation 4aa27521-bd89-44d0-8dde-73aff2b7ec49 · outbound

This paper cites Active contrastive learning of audio-visual video representa- tions.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.374522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.888136Z digest=sha256:a60d53cfdb7acdee2bcbdec72831eb36e8ccc0b8fd2b8b55680cb2e67d7a7aa6

Observation c48d9fe1-9393-437e-a92b-2beeafada6ea · outbound

This paper cites Active contrastive learning of audio-visual video representa- tions.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.359075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.892420Z digest=sha256:6dacb71fcbcf521da766b24893344eb438e1acad3ac98fbc632dfdf2476b6e87

Observation 261f2fbd-a995-4dd1-8c0e-d4548bedf6c8 · outbound

This paper cites Robust audio-visual instance discrimination.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Robust audio-visual instance discrimination

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.344760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.896858Z digest=sha256:81e417b59617da7a35d41badd4dd46acf2bdf4118602b12aa88dd7c8bacefb60

Observation c4844cf3-ab58-4ec3-88bc-72a32a372b69 · outbound

This paper cites Audio- visual instance discrimination with cross-modal agreement.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio- visual instance discrimination with cross-modal agreement

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.331778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.901142Z digest=sha256:b4c62472e53d8cd0d2e8918d455aaf2dfa238db08f062538676bbca9b0a8510a

Observation ded765ae-f5e6-4a40-aecb-402998a79a28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio-visual scene analysis with self-supervised multisensory features

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.317636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.905131Z digest=sha256:72bef96158051f4cef30a8e8332b8d3916fece08a223d043b22ee765460c8b80

Observation 7e8392aa-f18e-4acf-95a7-682aaddfa66f · outbound

This paper cites Ambient sound provides supervision for visual learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Ambient sound provides supervision for visual learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.303731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.909268Z digest=sha256:13729782210dd3b28cb91d0c594ed5167448cddf1e267f0afe737041592604a0

Observation 1c192f66-1479-475c-a088-ee4d811c364f · outbound

This paper cites On compositions of transformations in contrastive self-supervised learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment On compositions of transformations in contrastive self-supervised learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.289207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.913150Z digest=sha256:dfb53556970db04b7ca442fc225011d45c7af5fba3294f3c5693b074146272e9

Observation 741a9754-e2dc-4f0f-8cf8-db0f6c36e4ed · outbound

This paper cites Broaden your views for self-supervised video learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Broaden your views for self-supervised video learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.274828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.916743Z digest=sha256:0e653941407fdcce64d8d94f6a8e9debf2511dca773099cacc64dc3bd7fda792

Observation 6a6e6f64-3bc0-4896-80db-b4cb6659f2fc · outbound

This paper cites Avlnet: Learning audio-visual language representa- tions from instructional videos.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Avlnet: Learning audio-visual language representa- tions from instructional videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.260253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.920815Z digest=sha256:4927ebb8955d83b3d33a35a3634cd2f290519c5e2bd050881ed6f86d1a964d7b

Observation edc60b02-4d74-4097-b79c-e28efaf0a2bf · outbound

This paper cites Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.245489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.925091Z digest=sha256:03b40240d8d7d19465bca78865902d0782fd66cda4c44601bada1002134b9e5b

Observation 6ff6cfa9-1b25-4b6d-bf08-532115f77f59 · outbound

This paper cites Event-specific audio-visual fusion layers: A simple and new perspective on video understanding.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Event-specific audio-visual fusion layers: A simple and new perspective on video understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.230436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.929279Z digest=sha256:009c6d97679053fcd78535a21286d6099a9437969d5f9228fda17fced55c0d8e

Observation e4892cd6-6be3-40fc-ad05-52d2911f5a50 · outbound

This paper cites From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.214522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.933357Z digest=sha256:d6feee8a5393ea2928f5b903c08648bfc33c0f81b1c76fc29eb9d24f3151df01

Observation 047fe13b-cd2f-4c29-82c8-3751fc36d5ed · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Learning audio-visual source localization via false negative aware contrastive learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.199227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.937457Z digest=sha256:235f802fea6832efdf33a32de91416c4161d2859ee377b3e26aec52458b2e5e8

Observation fc5418b9-2d54-4061-b30e-9488c3298d4a · outbound

This paper cites Multimodal Self-Supervised Learning of General Audio Representations.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multimodal Self-Supervised Learning of General Audio Representations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.941470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.941470Z digest=sha256:3e3006df82b0d1ed2de240dca8a49ab3e6a1a4b169563eacc41c57a320cd26db

Observation 89288667-2be6-4606-8168-8889985cf1d2 · outbound

This paper cites Temporal cue guided video highlight detection with low-rank audio-visual fusion.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Temporal cue guided video highlight detection with low-rank audio-visual fusion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.184876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.946526Z digest=sha256:5f3870129e9690fc2e01605699bd9ea598e0d158f10403fceb2f9bb480efe8ee

Observation 8620266b-ec0c-4d5e-9ab1-4257529d94c2 · outbound

This paper cites Con- trastive learning of global and local video representations.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Con- trastive learning of global and local video representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.171128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.950818Z digest=sha256:cc2cf257acbbef4aed64417272edbf7cbf137a7581e3c2ee4da54a3bd071a136

Observation a0311b32-5c5d-4fc8-abad-76a0f5d08865 · outbound

This paper cites The sound of pixels.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment The sound of pixels

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.156494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.955193Z digest=sha256:d2602fbdca0ee338b980b457a608b2a051c3f73cf64ac269f3c47bcce820e4f6

Observation aa27f171-004a-4ceb-90f0-7c343d55ca27 · outbound

This paper cites Scene parsing through ade20k dataset.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Scene parsing through ade20k dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.959437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.959437Z digest=sha256:50ee4e45c09fa5d91ffb8567d71701ae5fa4a4d6be74a01291c60673e23f0935

Observation 56ae1a59-bbaf-4163-b650-9a12be9fcba5 · outbound

This paper cites Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.133027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.964120Z digest=sha256:521efd23eba349f4a9026b406be6d61c77ba7f4dec225f0615016d872c8ec572

Observation 12498e67-28ca-4db0-b812-1f954bcbcfff · outbound

This paper cites an unresolved cited work.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:26:33.119294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.968662Z digest=sha256:f1c09322aaeba6b4ba3d7db1f07612fa501a15d71f333376d7554f5d5870cc8f

Observation 9bb38671-8523-400e-a3c8-de9b7e29c955 · outbound

This paper cites an unresolved cited work.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:26:33.104530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.973334Z digest=sha256:ec17c60a4d28dd6b6f6cfde40051cdb047650679aa3b3d217f24b4c667fd24fd

Observation da8d6112-834b-4bc6-9fe2-b732aade6310 · outbound

This paper cites di- agonal mean.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment di- agonal mean

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.089873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.977885Z digest=sha256:353b3f103d8333df8d1668438096765391434f9c3e6be5cab17fbf070347086c

Observation dbef4d5f-5e60-45ce-8820-907cc749079f · outbound

This paper cites Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.075447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.982956Z digest=sha256:951d4d5bec0d96ff1d1d492e9462ede56cc653e868c581b65b07ad1aac606402

Observation e08250c1-bba3-4508-8115-66e6321d4225 · outbound

This paper cites writing on blackboard with chalk.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment writing on blackboard with chalk

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.060343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.987325Z digest=sha256:b64f680be4b9998215f44bb0b1d16231ce60c040f4c12c7cdaf04ab6b25b68c3

Observation 521aa4e2-2093-47a7-aa0d-aa4e7c0f997c · outbound

This paper cites For this experiment, we manually annotate the occurrence of the classes throughout the video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment For this experiment, we manually annotate the occurrence of the classes throughout the video

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.045181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:26:32.991891Z digest=sha256:c483f50e3c7a3da3dfcf1d5e49d5b9412d64ddac6e6d0a518dca1d3468d6beb7

Pith citing papers

No inbound Pith citation observations are available.