Pith. sign in

Paper Citation Record · LEDGER

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2507.11967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11967 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:47.086421Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:44.662433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:02:47.349677Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b82f3c9-ed53-401e-b29b-d85ffdd59aa9 · outbound

This paper cites Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:02:47.407214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:44.662433Z digest=sha256:836964688faef96d3cc8f03a6a3113740d171430dff153e015223fab8d18ef06

Observation e914d1db-ca25-4909-abb0-28cfae130dff · outbound

This paper cites We also propose to automatically gener- ate audio-visual-text triplets from unlabeled videos, which are subsequently used for training LG-CA V-MAE.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos We also propose to automatically gener- ate audio-visual-text triplets from unlabeled videos, which are subsequently used for training LG-CA V-MAE

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.780523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:44.810201Z digest=sha256:e9b976fc270c6634a6afbc40e4ffcfdf2788b4b80d82b3f4363eb0daea600ec4

Observation 48449906-1dcf-4cea-a384-8751903f7c49 · outbound

This paper cites an unresolved cited work.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Unresolved cited work

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:02:49.645247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:44.891841Z digest=sha256:3cbb648ea4395392b947842e2fda9fbe16f7a65de9ab14f8a79d40f5db47d40c

Observation 90ee5293-c267-4109-910d-530c82239868 · outbound

This paper cites an unresolved cited work.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:02:49.526984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:44.961943Z digest=sha256:60ae072224478607fcc977cb3418619f877f790e135be7ae17568737a74e7311

Observation 90dd38b2-8ab7-4064-bc64-7caa43145675 · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Contrastive Audio-Visual Masked Autoencoder

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.037290Z digest=sha256:79c427183e81ef470f748e844c6b9dee664a9d234a65a5411e4b095d49218a61

Observation 92d8bf79-fc66-4056-a9b1-5e2c4ae02a20 · outbound

This paper cites Mavil: Masked audio-video learners,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Mavil: Masked audio-video learners,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.361832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:45.118282Z digest=sha256:f39720065582263055091a12e2d0cfeb8de1edbbb2fec5316ad6c918db69b1f2

Observation 0509ec62-d6a4-442b-8677-a82e7db0a7fd · outbound

This paper cites Audiovisual masked autoencoders,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual masked autoencoders,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.178565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:45.248428Z digest=sha256:58d7dde92649e603f2032ed81af303a0fbcea994cf894a136633b6ea71c09c14

Observation fa779b7d-4f7c-489a-a940-a87fd09879ca · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audio set: An ontology and human-labeled dataset for audio events,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.369591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.369591Z digest=sha256:c45d154b275689aecf80942d771aee3bcf05d0e2e38e18c346f61a6c29fbd0a9

Observation 4d1168e6-07ab-4819-86bd-c915906aaf85 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos The Kinetics Human Action Video Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.445056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.445056Z digest=sha256:30fa75e37a0732bb88735d3ad9dbd6045997f3ec9db3936dd8878106957fb2e2

Observation 437a2c0e-0f7f-474a-b31a-b038349e3fd5 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Vggsound: A large-scale audio-visual dataset,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.032112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:45.534477Z digest=sha256:8f99112953045ae9d52207900f697049b4ec184b427a4de650d38e0bfa1a8f3c

Observation a805f781-fdf6-4ca9-833e-05d4e3e771e1 · outbound

This paper cites Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.834728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:45.622514Z digest=sha256:408d8a5f5ca9748bf4d4f4ec5bf2c98d433c2796c0cdf388d427d2e0d267ab4c

Observation 80c6ccf8-756a-4288-a4d3-17eed4ed3939 · outbound

This paper cites DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:02:47.273577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:45.719897Z digest=sha256:00ce0642b60331b063fc23a0780dba93ec0136d46d9fba746fd52521ca94e80f

Observation f32d6421-83f9-4433-a70f-2b1f887c27f3 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Clap learning audio concepts from natural language supervision,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.866977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.866977Z digest=sha256:d48a467f76da3a3c878e498531925948b36f748c50026a395ed8484b30b85ca1

Observation 308d9055-154c-49cf-98d4-6832043d1a37 · outbound

This paper cites Dense-captioning events in videos,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Dense-captioning events in videos,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.623107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:45.967626Z digest=sha256:8388e901aaa00d1a4033b4b93d4a9ae567bfc82d4c8d0775d1dcd2427923040d

Observation 67eb9617-c53e-48fe-8b5f-e4b75ad074fa · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.467621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.050837Z digest=sha256:4545cd2de08b8d8cfbbeb0b3b273bbaaf8008cee27f4990ff6be3fd3214501e8

Observation 51708959-5173-4272-b622-144a7580e5dd · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video- and-language research,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Vatex: A large-scale, high-quality multilingual dataset for video- and-language research,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.312845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.141712Z digest=sha256:e653a98b393f09e00a03d9874519f6a192dc0a0391ca734082173ced213d5608

Observation 1ad27a96-7e47-4f48-a2ac-5410af91390a · outbound

This paper cites Avset-10m: An open large-scale audio-visual dataset with high correspondence,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Avset-10m: An open large-scale audio-visual dataset with high correspondence,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.184192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.219041Z digest=sha256:2b67365e93a77154fba81f1449c930585a84d7b799fdd2ebfba3ca64a547c33a

Observation b7a7a21c-6951-4462-ad65-f7a9b733d473 · outbound

This paper cites Sound event envelope estimation in polyphonic mixtures,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound event envelope estimation in polyphonic mixtures,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.052660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.290229Z digest=sha256:2f2c144c65b8bcd5ad4517ce92342dfa4b97deae1dc6fa61b5c1ed42ffdac18e

Observation 6e7b2256-3a88-40ee-a052-4f261cc9ebde · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.918457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.366677Z digest=sha256:1c6f0d842db5cdf702335f9bb643eba35223f6ce42377e51bca1ee94c17709fa

Observation a5be9aea-0026-43be-80c2-0534256b5b45 · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.431246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.431246Z digest=sha256:3e591b386834f9d64006bd2d029866a811a831e2aa85673b62aac3cc7c5e4a67

Observation 400238fd-6039-465d-a1a4-3597ab50cd9d · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.490467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.490467Z digest=sha256:ebc06b8dbf8f42ed5c54b9a122783773e016f61d7d3e85b64d18ff992cc605c2

Observation 1a5b6cb5-3de1-4f72-bfa6-25d959fd7fb7 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.554864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.554864Z digest=sha256:28d0872220bc4bac828686f976ad1378d3ad55fd556a04a294388299012dc307

Observation 5bb31ecb-e7fc-4aca-95e5-5150ce7a52c4 · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos A simple framework for contrastive learning of visual representations,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.631137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.631137Z digest=sha256:93160f6b7495e479bf68e10314b592f225a5ff49535e2ce3ca6dee1796944a2d

Observation b0fef977-f530-4dcc-aef9-3cbe601b25c4 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Representation Learning with Contrastive Predictive Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.721538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.721538Z digest=sha256:dc925462e5e895e9c6e46af40cd39bd2991fc551cce1d35401a61a135e2fb860

Observation c2c7a069-6bc6-4f75-95bd-94993b383002 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.788088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.808652Z digest=sha256:a170a85c6985caa150d02a8bfe65223bca2fa6b13bbb931e014ad0c89101ff12

Observation 66b550c5-b6d1-4fed-8b0f-f17f9697ef9c · outbound

This paper cites Improved baselines with visual instruction tuning,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Improved baselines with visual instruction tuning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.651647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.889581Z digest=sha256:626a346afc8368d7df6867d2fe78079d3a4a5e07d4c3449d0a577c779480009a

Observation 56600a93-ee73-4a74-9978-01870ce512b2 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual SlowFast Networks for Video Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.945121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.945121Z digest=sha256:b7743d0db164948e71edae4e41a025c1d29d5fb6a7a8f3bfd7a6aaee144c321f

Observation 099ed48e-97e1-4aa7-ad83-aa66bface629 · outbound

This paper cites Masked autoencoders that lis- ten,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Masked autoencoders that lis- ten,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.513305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:46.993705Z digest=sha256:1d00eb2b1c04232ac2b37948771abefae1a8f3c97b3ba875562a371421b55806

Observation 24cb7e99-41c5-435f-9db0-83539c10c1bc · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Adam: A Method for Stochastic Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:47.086421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:47.086421Z digest=sha256:4ebf5e9a263b7b489172eb95ecbe1313b93f363f55245e3e87fab2e06b3e2be9

Pith citing papers

Observation 3b82f3c9-ed53-401e-b29b-d85ffdd59aa9 · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:02:47.407214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:02:44.662433Z digest=sha256:836964688faef96d3cc8f03a6a3113740d171430dff153e015223fab8d18ef06