Pith. sign in

Paper Citation Record · LEDGER

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2507.11967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11967 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:47.086421Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:02:44.662433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:02:47.349677Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b82f3c9-ed53-401e-b29b-d85ffdd59aa9 · outbound

This paper cites Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:02:47.407214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:44.662433Z digest=sha256:f75dfa8e85a904e57a3542a73b432e43b0499a75ac08bae4a3b05a66db516882

Observation e914d1db-ca25-4909-abb0-28cfae130dff · outbound

This paper cites We also propose to automatically gener- ate audio-visual-text triplets from unlabeled videos, which are subsequently used for training LG-CA V-MAE.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos We also propose to automatically gener- ate audio-visual-text triplets from unlabeled videos, which are subsequently used for training LG-CA V-MAE

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.780523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:44.810201Z digest=sha256:8cecbb28f0d145fa4f7cae58afba71e9e91d4f80d7ae78873dd7f613c945035d

Observation 48449906-1dcf-4cea-a384-8751903f7c49 · outbound

This paper cites an unresolved cited work.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Unresolved cited work

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:02:49.645247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:44.891841Z digest=sha256:9e4bcfa9f95666040a68e746bd6e194879c0c6f3e7c83d6c64ef13b1ed87d46f

Observation 90ee5293-c267-4109-910d-530c82239868 · outbound

This paper cites an unresolved cited work.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:02:49.526984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:44.961943Z digest=sha256:580a752225ca1369f83b7073000277c28a4b62a5ef2ff6fe586b458f78005590

Observation 90dd38b2-8ab7-4064-bc64-7caa43145675 · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Contrastive Audio-Visual Masked Autoencoder

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.037290Z digest=sha256:1faadd2a0f07ffbdce8f73139a06b489a5b1da16608cf3a23af2492a51a41091

Observation 92d8bf79-fc66-4056-a9b1-5e2c4ae02a20 · outbound

This paper cites Mavil: Masked audio-video learners,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Mavil: Masked audio-video learners,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.361832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:45.118282Z digest=sha256:e926d362814bdb541695366928715f7640ffea37596823abb6898d24caab8293

Observation 0509ec62-d6a4-442b-8677-a82e7db0a7fd · outbound

This paper cites Audiovisual masked autoencoders,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual masked autoencoders,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.178565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:45.248428Z digest=sha256:39dd955035a3d9c130cca1b81f59c0839b147e3ea80ff88209694fa46a21ee85

Observation fa779b7d-4f7c-489a-a940-a87fd09879ca · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audio set: An ontology and human-labeled dataset for audio events,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.369591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.369591Z digest=sha256:4884aab98315e900a9270570bfe1b2f913488a32635c161e2ef00ccdda3de432

Observation 4d1168e6-07ab-4819-86bd-c915906aaf85 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos The Kinetics Human Action Video Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.445056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.445056Z digest=sha256:ffcd512c50679f1e469de044391ef8d76e9be01ea21a8a5206e4b78646e1eca9

Observation 437a2c0e-0f7f-474a-b31a-b038349e3fd5 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Vggsound: A large-scale audio-visual dataset,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:49.032112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:45.534477Z digest=sha256:31126cf81d08cd78b7688c30a71a7b439114cc711c3d0bf66614c76de40525e6

Observation a805f781-fdf6-4ca9-833e-05d4e3e771e1 · outbound

This paper cites Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.834728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:45.622514Z digest=sha256:9b4fe394625e2d3b705a889b2b5252d069d4cc941db13a5fad24e1b23386592d

Observation 80c6ccf8-756a-4288-a4d3-17eed4ed3939 · outbound

This paper cites DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:02:47.273577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:45.719897Z digest=sha256:6a565de8fe8656fae166a99fda3b9e1c1646caa9f23e6401a230880054b74787

Observation f32d6421-83f9-4433-a70f-2b1f887c27f3 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Clap learning audio concepts from natural language supervision,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.866977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.866977Z digest=sha256:9afb00d321545e370705e4bdbb6ca1892723cc77a1212890873e584c45c830c9

Observation 308d9055-154c-49cf-98d4-6832043d1a37 · outbound

This paper cites Dense-captioning events in videos,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Dense-captioning events in videos,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.623107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:45.967626Z digest=sha256:4b48ea5ac71b65e2626d36806775b2fd45b8757461cc23cc19561dd6fa363a23

Observation 67eb9617-c53e-48fe-8b5f-e4b75ad074fa · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.467621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.050837Z digest=sha256:98fc4da469402cf1eb1e53394c195fce51f6f2aafe8c7727b6f89defe5ef5d82

Observation 51708959-5173-4272-b622-144a7580e5dd · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video- and-language research,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Vatex: A large-scale, high-quality multilingual dataset for video- and-language research,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.312845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.141712Z digest=sha256:849adfac9337a48e3e7e48212baca9dacb7daac28acad2aecd51cd4589fee381

Observation 1ad27a96-7e47-4f48-a2ac-5410af91390a · outbound

This paper cites Avset-10m: An open large-scale audio-visual dataset with high correspondence,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Avset-10m: An open large-scale audio-visual dataset with high correspondence,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.184192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.219041Z digest=sha256:799156610b33ea44672de481d91bdb6432e94b34a970055a30da6e7c8ba5956b

Observation b7a7a21c-6951-4462-ad65-f7a9b733d473 · outbound

This paper cites Sound event envelope estimation in polyphonic mixtures,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound event envelope estimation in polyphonic mixtures,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:48.052660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.290229Z digest=sha256:2d2ca0ae7e9627978f724384223a2c58bf3bed33e4812d977f863b2998fbecc2

Observation 6e7b2256-3a88-40ee-a052-4f261cc9ebde · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.918457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.366677Z digest=sha256:29875c325040094f52d2fa31c42fa75da77c2d4c72a477efef51e842527b898a

Observation a5be9aea-0026-43be-80c2-0534256b5b45 · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.431246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.431246Z digest=sha256:84142a32a10a732409b60ceb2b0c2ef791679b345f778d442f06da9426496433

Observation 400238fd-6039-465d-a1a4-3597ab50cd9d · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.490467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.490467Z digest=sha256:5c33064131843419989359719fb0a0d5408dd9087dd8d6172a6f40b630a04342

Observation 1a5b6cb5-3de1-4f72-bfa6-25d959fd7fb7 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.554864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.554864Z digest=sha256:c8d44b7363dc992fb3368952117f2747dff34946091988d624784639078b6a12

Observation 5bb31ecb-e7fc-4aca-95e5-5150ce7a52c4 · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos A simple framework for contrastive learning of visual representations,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.631137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.631137Z digest=sha256:16b7a37fd35bf50a50084ab5491254621b883e2af77c482d213d7fcd7f6b123f

Observation b0fef977-f530-4dcc-aef9-3cbe601b25c4 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Representation Learning with Contrastive Predictive Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.721538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.721538Z digest=sha256:9011aa9a787531c40bc6e963a78ea0f7cf8e1161579f1efc5e8d26a75496be31

Observation c2c7a069-6bc6-4f75-95bd-94993b383002 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.788088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.808652Z digest=sha256:78cee4f056d96704eeed959eef94a6c3dc9744178136b96b0de57bd2664e72f2

Observation 66b550c5-b6d1-4fed-8b0f-f17f9697ef9c · outbound

This paper cites Improved baselines with visual instruction tuning,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Improved baselines with visual instruction tuning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.651647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.889581Z digest=sha256:ce55ee36c77304dd60c993956177513920d26bc1fc4cf406f8d198c66664546f

Observation 56600a93-ee73-4a74-9978-01870ce512b2 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Audiovisual SlowFast Networks for Video Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.945121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.945121Z digest=sha256:dd95d9427e0a726327dd131bd3005a123ff7ae818544b8e7f37dae26a090c91a

Observation 099ed48e-97e1-4aa7-ad83-aa66bface629 · outbound

This paper cites Masked autoencoders that lis- ten,.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Masked autoencoders that lis- ten,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:02:47.513305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:46.993705Z digest=sha256:68236acd8abfb0acbd531be0e17683808af446c7ba26641b42526fe7ecbb92ab

Observation 24cb7e99-41c5-435f-9db0-83539c10c1bc · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Adam: A Method for Stochastic Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:47.086421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:47.086421Z digest=sha256:4917228e2bf5d798f226047d8ab9df66b1acf4f19ac8f467cfe069724ed6de84

Pith citing papers

Observation 3b82f3c9-ed53-401e-b29b-d85ffdd59aa9 · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:02:47.407214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:02:44.662433Z digest=sha256:f75dfa8e85a904e57a3542a73b432e43b0499a75ac08bae4a3b05a66db516882