Pith. sign in

Paper Citation Record · LEDGER

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

As of 20 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 4 inbound Pith citation observations for arXiv:2501.09291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09291 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:24.131196Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.048189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:16:56.891207Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 974f417b-d9b5-4d9d-917e-128056c9b1a1 · outbound

This paper cites Per- sonalized dialogue generation with persona-adaptive attention,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Per- sonalized dialogue generation with persona-adaptive attention,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.518296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.014439Z digest=sha256:d4bcbec89b32acb0a8fbf53e2d5eb668931e665054be75a563eac5e54acf3221

Observation 2ff738c1-256a-4033-ac6d-f057331c2e68 · outbound

This paper cites Audio captioning transformer,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Audio captioning transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.506292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.019499Z digest=sha256:98caddbac2b80511a15d707a0bd6f27ad6f67dcb850a4cac30f06f9c1512fb24

Observation da985451-d8cc-4ba1-be86-2aa26fec6383 · outbound

This paper cites Automated audio captioning by fine-tuning bart with audioset tags,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Automated audio captioning by fine-tuning bart with audioset tags,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.493629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.023616Z digest=sha256:d2c81bc17bc4b0387ecbf281e90b8c6a227736a993bf476ddf8f4d726b0d9656

Observation 52b502a1-bc15-415e-b84a-91db1ab7bc0e · outbound

This paper cites Prefix tuning for automated audio captioning,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Prefix tuning for automated audio captioning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.480282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.027723Z digest=sha256:434a11cd4f662d71b005209e47fcb5f8938fd280fa05bde2319acaef5423d225

Observation 9621f7f6-81f4-490d-94dd-4b96363286fc · outbound

This paper cites EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.468036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.031857Z digest=sha256:06f334c319d2dbab20a0dd17df3503a3faf19307416b127290554d7f2e1a86c0

Observation ae10686b-e2e8-4229-8ca5-2bab65629e0b · outbound

This paper cites WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.455187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.035951Z digest=sha256:0defeda97970c227e35d71ff523d687dfd9b5bf65a1f6cc15f1fbc771da86d56

Observation b01b7139-f288-4b6d-a93b-f9f57812430e · outbound

This paper cites CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.442774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.040874Z digest=sha256:d298cadeee4117e41c7435d88ad972d105b1aaeae5afedef87d82d79dad7b1e2

Observation e2691a3a-ff89-4aed-b951-b25914a5f984 · outbound

This paper cites Enhancing automated audio captioning via large language models with optimized audio encoding,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Enhancing automated audio captioning via large language models with optimized audio encoding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.429866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.044831Z digest=sha256:36177944579b98bfda41ae545275f8c7c5a40c4c94a6b49a37ab96ea522249d6

Observation 6c6ef354-ca3f-4c70-aeca-c4fac11cd4bd · outbound

This paper cites Taming Data and Transformers for Audio Generation.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Taming Data and Transformers for Audio Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.048855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.048855Z digest=sha256:95a90a17beebe3e060af95f428eeaa39b1a616eefdb2066837dbaef30284cbe7

Observation 0d03cd14-c60b-4395-a873-daea2797556e · outbound

This paper cites PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.053038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.053038Z digest=sha256:2c1c3b8f583760381a588882dbb4f69d22623e9d0652677ab6fabd0cf365f9f4

Observation 89d09ab9-55c1-4355-b652-a4f2d93695c6 · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.409623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.057056Z digest=sha256:d2b864281c1e49a4c09e15aaf968e6b2dbbd281398c42ee817387fda8666674a

Observation e4a00442-0441-47f8-9afc-728d2209803f · outbound

This paper cites BEATs: audio pre-training with acoustic tokenizers,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport BEATs: audio pre-training with acoustic tokenizers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.396071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.061563Z digest=sha256:f422bf035f4092fd145cd326bab8405d80bc8c06227648e9d1ae601ce3361b9b

Observation 83f552bd-4d01-430b-8a95-2866f08659d6 · outbound

This paper cites CLAP: Learning audio concepts from natural language supervision,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport CLAP: Learning audio concepts from natural language supervision,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.383716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.066683Z digest=sha256:d9576a642e3400003ec573ed22c440b31a587a635d4438f60c8059b6f4ea9532

Observation cebe1ac3-e3e3-44c6-af8a-aa7f1899a16f · outbound

This paper cites High fidelity neural audio compression,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport High fidelity neural audio compression,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.070730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.070730Z digest=sha256:8a81e7fb0fd1f7afa2e65680c80f12f3b6db90b725d7ab97b59ae2ff03e767b3

Observation 2d913290-c2ef-4d03-9b0a-ef96f4d4f64f · outbound

This paper cites Visually-aware audio captioning with adaptive audio-visual attention,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Visually-aware audio captioning with adaptive audio-visual attention,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.364291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.075092Z digest=sha256:f5c7c5f1afc4370ec62b92cc78fb0c17a9c77a50acb7aee00221b645144246f4

Observation 20f54658-b9e8-4559-ad74-a81046bc808d · outbound

This paper cites A VCap: Leveraging audio-visual features as text tokens for captioning,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport A VCap: Leveraging audio-visual features as text tokens for captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.351991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.079355Z digest=sha256:48fc66215263adb33b758a04bf6d0b224f0f598ef448c758365cfe849e6a4409

Observation a2a76d17-790f-49e5-b596-0c17522e13c7 · outbound

This paper cites Multi- granularity correspondence learning from long-term noisy videos,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Multi- granularity correspondence learning from long-term noisy videos,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.339844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.083380Z digest=sha256:e8f49ba7c3552f3b81889bc9cf4979bc5f66ab2a12867bbef06a330a31f6a406

Observation 38cf3bb0-f72c-4514-9954-08c3b7a92adf · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.326475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.087468Z digest=sha256:d3325fbad8b1136836e836d119ee0aaa39d67a4dc8cc4eb5dc6fed49ec6f75fd

Observation 269af917-aa57-4f67-9df1-902d0b927e1a · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Audiocaps: Generating captions for audios in the wild,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.314167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.091698Z digest=sha256:df859f2c9319e9ee59ceab8917643bf0247f2d71b1b734ede7b41e1ebfc7f4c0

Observation 71c883f2-b2f6-41b9-ac34-2e8ed4e40a38 · outbound

This paper cites Ced: Consistent ensemble distillation for audio tagging,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Ced: Consistent ensemble distillation for audio tagging,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.300965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.095613Z digest=sha256:7e3a3dbeaea3576c67ad2e3408e6cadb6e1ae3bd295a7e9250889187eacecae8

Observation b83eef74-2a62-440d-b115-ef3a3b8508a3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Learning transferable visual models from natural language supervision,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.288877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.099959Z digest=sha256:d4d2c87c35ad2e7ea4709c6bff38e70bf10bdf4e884b9cb7e9162f934252a5e5

Observation 781e6ef9-06a4-4bb9-9206-06a07a0c6bdc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.103831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.103831Z digest=sha256:13555f1ce5a338ca01e9b008d72bc98110f3b3f2f92a32699c11b1b8bbae50ea

Observation 21a33c6f-8470-497e-96d9-64aeb76e88aa · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport LoRA: Low-rank adaptation of large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.276833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.107764Z digest=sha256:00dd1b52bd33e1449841ab2720761e24cf647c9127c1c460e486d6fbd2a90ece

Observation 47c7585b-c4c7-48d1-b148-a1a3f3ce8dc4 · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport BLEU: a method for automatic evaluation of machine translation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.262313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.111830Z digest=sha256:edffcb01473a8495e9b45e7fdc4b0aabadc963eaaa3392d24c34cfdca2c0ce26

Observation e517c3d9-f19f-419c-b435-54096f732b1a · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport ROUGE: A package for automatic evaluation of summaries,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.247537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.115542Z digest=sha256:1f6291d18bac57afdde7cbbb94a729a9f82d39b01332e52f894dc29b1c16ca3f

Observation 61bcdc4c-f0aa-4568-8553-6bbc82145c63 · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.234956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.119327Z digest=sha256:e02dd732712ed3b987f00ed280bf3fe4005af4e37bf4e4f855c12ddae3555949

Observation c0986ccc-04ee-4e0d-8961-076bf4d492bc · outbound

This paper cites Cider: Consensus- based image description evaluation,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Cider: Consensus- based image description evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.222066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.123054Z digest=sha256:903838fc47bf2c1bf9108919114d2e52d5dc3afb166f117a5ecd5a29ab45e571

Observation 1799ba48-c63a-4819-93bd-47d93c37fe01 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Spice: Semantic propositional image caption evaluation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.209191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.127000Z digest=sha256:fa16a06280457241889652bbc0e18d5d3942a152cba45b8226fdb0367ebada89

Observation f919a529-ed1e-472e-85d8-a5fd2322b9ad · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Improved image captioning via policy gradient optimization of spider,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:24.196050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:24.131196Z digest=sha256:2f10d552c5dd12bfb96a510394227eae9865b0b8a7fa3b8f5b1c608bb50772c4

Pith citing papers

Observation 92e1dd92-2121-47ad-9bc6-fa069eb0e06c · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.048189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.048189Z digest=sha256:d08901f10500ca6fb65a102b0e21014cc8a821eb614d0d20ea58f224c73a151f

Observation 21c84687-edff-4251-a172-5b1d6519ca57 · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.225698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.225698Z digest=sha256:96463334372fd027128693fbd242ebe9ad9c057ad5c76248b5c83ff327669f78

Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · inbound

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction cites this paper.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:57.039584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:16:56.175634Z digest=sha256:2be4e7f5b8aaa833e35b4ac437948cecb8c8c90dad7577ce02981f56f3097502

Observation 840821bc-8bd9-41ef-9568-2393b100d7b9 · inbound

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition cites this paper.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:71206136b22d8130840d896b785488729b0dda184f726a0fc2ee4994314720c5