Pith. sign in

Paper Citation Record · LEDGER

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

As of 20 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 4 inbound Pith citation observations for arXiv:2501.09291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09291 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:24.131196Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.048189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:16:56.891207Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 974f417b-d9b5-4d9d-917e-128056c9b1a1 · outbound

This paper cites Per- sonalized dialogue generation with persona-adaptive attention,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Per- sonalized dialogue generation with persona-adaptive attention,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.518296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.014439Z digest=sha256:712a4940731d0c31057d2544078f49b4a453c751e525aef982a1c3ef01f2f45e

Observation 2ff738c1-256a-4033-ac6d-f057331c2e68 · outbound

This paper cites Audio captioning transformer,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Audio captioning transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.506292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.019499Z digest=sha256:4dc45dbb4d18322d5ae745ca9fe019dc1926a2e3b99dc65fa643a36077411f96

Observation da985451-d8cc-4ba1-be86-2aa26fec6383 · outbound

This paper cites Automated audio captioning by fine-tuning bart with audioset tags,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Automated audio captioning by fine-tuning bart with audioset tags,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.493629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.023616Z digest=sha256:e881408be9b89c770a8d36d1593bbd91b21dbc7769af1fb26c8dd0cf096e25e3

Observation 52b502a1-bc15-415e-b84a-91db1ab7bc0e · outbound

This paper cites Prefix tuning for automated audio captioning,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Prefix tuning for automated audio captioning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.480282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.027723Z digest=sha256:e9b82e8a9c267be5e264a467283756257ddbadb95c33686451ef6b4272dc1719

Observation 9621f7f6-81f4-490d-94dd-4b96363286fc · outbound

This paper cites EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.468036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.031857Z digest=sha256:d99de7c2344733c25b5f15b6529f98719949de6726989ec17516f708e45c5f99

Observation ae10686b-e2e8-4229-8ca5-2bab65629e0b · outbound

This paper cites WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.455187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.035951Z digest=sha256:da5bdd90d99c3fe805c12a8f317538dee069a6ac020dc25e05dbc1c3c89c68fd

Observation b01b7139-f288-4b6d-a93b-f9f57812430e · outbound

This paper cites CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.442774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.040874Z digest=sha256:03d0ceed3d90b5dbac7cfeca1a33641379c642e134434bfec7e0601037f72526

Observation e2691a3a-ff89-4aed-b951-b25914a5f984 · outbound

This paper cites Enhancing automated audio captioning via large language models with optimized audio encoding,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Enhancing automated audio captioning via large language models with optimized audio encoding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.429866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.044831Z digest=sha256:c96cbad86f05ddefb764f4cb8bcaf56c046dc3abe2c2f84094c4ba07746cc99d

Observation 6c6ef354-ca3f-4c70-aeca-c4fac11cd4bd · outbound

This paper cites Taming Data and Transformers for Audio Generation.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Taming Data and Transformers for Audio Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.048855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.048855Z digest=sha256:95a90a17beebe3e060af95f428eeaa39b1a616eefdb2066837dbaef30284cbe7

Observation 0d03cd14-c60b-4395-a873-daea2797556e · outbound

This paper cites PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.053038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.053038Z digest=sha256:2c1c3b8f583760381a588882dbb4f69d22623e9d0652677ab6fabd0cf365f9f4

Observation 89d09ab9-55c1-4355-b652-a4f2d93695c6 · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.409623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.057056Z digest=sha256:92c5bcd2f48a79203b23f519fa0fd059df4045dbee7831bf071d1c66b82a7576

Observation e4a00442-0441-47f8-9afc-728d2209803f · outbound

This paper cites BEATs: audio pre-training with acoustic tokenizers,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport BEATs: audio pre-training with acoustic tokenizers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.396071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.061563Z digest=sha256:01f40c8b711a92e8e8a902eb8657ff1e66ea299455dc1edad75eabc1042e1df2

Observation 83f552bd-4d01-430b-8a95-2866f08659d6 · outbound

This paper cites CLAP: Learning audio concepts from natural language supervision,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport CLAP: Learning audio concepts from natural language supervision,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.383716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.066683Z digest=sha256:640102b46fe6fd48bd32d1d698cbd513dcefa866dd1d727245ab32be42871e96

Observation cebe1ac3-e3e3-44c6-af8a-aa7f1899a16f · outbound

This paper cites High fidelity neural audio compression,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport High fidelity neural audio compression,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.070730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.070730Z digest=sha256:8a81e7fb0fd1f7afa2e65680c80f12f3b6db90b725d7ab97b59ae2ff03e767b3

Observation 2d913290-c2ef-4d03-9b0a-ef96f4d4f64f · outbound

This paper cites Visually-aware audio captioning with adaptive audio-visual attention,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Visually-aware audio captioning with adaptive audio-visual attention,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.364291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.075092Z digest=sha256:a695ae5671d740abfd9fed193587b18f0b4db3f96a8ee724a66b738a8fc31557

Observation 20f54658-b9e8-4559-ad74-a81046bc808d · outbound

This paper cites A VCap: Leveraging audio-visual features as text tokens for captioning,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport A VCap: Leveraging audio-visual features as text tokens for captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.351991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.079355Z digest=sha256:41b4a76860a895f2bb69878dcc0b53aab18400327695e566d987647092e89810

Observation a2a76d17-790f-49e5-b596-0c17522e13c7 · outbound

This paper cites Multi- granularity correspondence learning from long-term noisy videos,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Multi- granularity correspondence learning from long-term noisy videos,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.339844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.083380Z digest=sha256:a93eecab5537323da080e77a1d38affa2a9e466d32ea1d56da6e4275bd872057

Observation 38cf3bb0-f72c-4514-9954-08c3b7a92adf · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.326475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.087468Z digest=sha256:33fcacbb77e9f49fdd643bd6dfe519e2a050009d38957aa04b43e24a85612723

Observation 269af917-aa57-4f67-9df1-902d0b927e1a · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Audiocaps: Generating captions for audios in the wild,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.314167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.091698Z digest=sha256:c9790ee780456dae80a07544bd676e30e35f2c2b79d5903c3611b4e833e4bfd8

Observation 71c883f2-b2f6-41b9-ac34-2e8ed4e40a38 · outbound

This paper cites Ced: Consistent ensemble distillation for audio tagging,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Ced: Consistent ensemble distillation for audio tagging,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.300965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.095613Z digest=sha256:532093907ca8b155cdce91b401052bd2d8cd89faba1331847a21cb81b1924fc0

Observation b83eef74-2a62-440d-b115-ef3a3b8508a3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Learning transferable visual models from natural language supervision,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.288877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.099959Z digest=sha256:3170dc02374156bd576af147bef27cb7275cef3ac538c95ad5a45fff5ed058ce

Observation 781e6ef9-06a4-4bb9-9206-06a07a0c6bdc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.103831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.103831Z digest=sha256:13555f1ce5a338ca01e9b008d72bc98110f3b3f2f92a32699c11b1b8bbae50ea

Observation 21a33c6f-8470-497e-96d9-64aeb76e88aa · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport LoRA: Low-rank adaptation of large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.276833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.107764Z digest=sha256:eba75aa3c11cf9100a971d8cf492b218de37170960e0527b1becc7aa009e1c36

Observation 47c7585b-c4c7-48d1-b148-a1a3f3ce8dc4 · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport BLEU: a method for automatic evaluation of machine translation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.262313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.111830Z digest=sha256:659b0c9436e10f77bf3070751c8642b783137e978520dcb1ef7af0cc93b247f2

Observation e517c3d9-f19f-419c-b435-54096f732b1a · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport ROUGE: A package for automatic evaluation of summaries,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.247537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.115542Z digest=sha256:e274c3c913fffb16fc886483cbcca27d00b33f9d405aef22dd08a05db8f0ef94

Observation 61bcdc4c-f0aa-4568-8553-6bbc82145c63 · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.234956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.119327Z digest=sha256:5a8ec27e8a2f3bec6800d34e2e4ede4bc2594cd2a9ebcc8aa804db9ecbffb912

Observation c0986ccc-04ee-4e0d-8961-076bf4d492bc · outbound

This paper cites Cider: Consensus- based image description evaluation,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Cider: Consensus- based image description evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.222066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.123054Z digest=sha256:e68ccb4bdcc7cca4f0c476d06b54de49025bbe4526f043e30a147b3d939c9693

Observation 1799ba48-c63a-4819-93bd-47d93c37fe01 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Spice: Semantic propositional image caption evaluation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.209191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.127000Z digest=sha256:ce920cda5cb27a19add49dc9bae74b9f228f39e5be5e02240b8ec220d92641f1

Observation f919a529-ed1e-472e-85d8-a5fd2322b9ad · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Improved image captioning via policy gradient optimization of spider,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.196050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T20:14:24.131196Z digest=sha256:5de87ee80c79a29c9ebde6d60e0b017d07a8eeb884d84f0c15cf74088201e5b7

Pith citing papers

Observation 92e1dd92-2121-47ad-9bc6-fa069eb0e06c · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.048189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.048189Z digest=sha256:d08901f10500ca6fb65a102b0e21014cc8a821eb614d0d20ea58f224c73a151f

Observation 21c84687-edff-4251-a172-5b1d6519ca57 · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.225698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.225698Z digest=sha256:96463334372fd027128693fbd242ebe9ad9c057ad5c76248b5c83ff327669f78

Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · inbound

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction cites this paper.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:57.039584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:16:56.175634Z digest=sha256:a9465f89ba032df3a3c21a779e09a396c277997a33f74caf80d533e0db64010d

Observation 840821bc-8bd9-41ef-9568-2393b100d7b9 · inbound

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition cites this paper.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:71206136b22d8130840d896b785488729b0dda184f726a0fc2ee4994314720c5