Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:24.131196Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 4 inbound Pith citation observations for arXiv:2501.09291.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:24.131196Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.048189Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:16:56.891207Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 974f417b-d9b5-4d9d-917e-128056c9b1a1 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Per- sonalized dialogue generation with persona-adaptive attention,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2ff738c1-256a-4033-ac6d-f057331c2e68 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Audio captioning transformer,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation da985451-d8cc-4ba1-be86-2aa26fec6383 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Automated audio captioning by fine-tuning bart with audioset tags,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 52b502a1-bc15-415e-b84a-91db1ab7bc0e · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Prefix tuning for automated audio captioning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9621f7f6-81f4-490d-94dd-4b96363286fc · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ae10686b-e2e8-4229-8ca5-2bab65629e0b · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b01b7139-f288-4b6d-a93b-f9f57812430e · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e2691a3a-ff89-4aed-b951-b25914a5f984 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Enhancing automated audio captioning via large language models with optimized audio encoding,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6c6ef354-ca3f-4c70-aeca-c4fac11cd4bd · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Taming Data and Transformers for Audio Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d03cd14-c60b-4395-a873-daea2797556e · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d09ab9-55c1-4355-b652-a4f2d93695c6 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e4a00442-0441-47f8-9afc-728d2209803f · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport BEATs: audio pre-training with acoustic tokenizers,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 83f552bd-4d01-430b-8a95-2866f08659d6 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport CLAP: Learning audio concepts from natural language supervision,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cebe1ac3-e3e3-44c6-af8a-aa7f1899a16f · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport High fidelity neural audio compression,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d913290-c2ef-4d03-9b0a-ef96f4d4f64f · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Visually-aware audio captioning with adaptive audio-visual attention,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20f54658-b9e8-4559-ad74-a81046bc808d · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport A VCap: Leveraging audio-visual features as text tokens for captioning,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a2a76d17-790f-49e5-b596-0c17522e13c7 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Multi- granularity correspondence learning from long-term noisy videos,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 38cf3bb0-f72c-4514-9954-08c3b7a92adf · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Sinkhorn distances: Lightspeed computation of optimal transport,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 269af917-aa57-4f67-9df1-902d0b927e1a · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Audiocaps: Generating captions for audios in the wild,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 71c883f2-b2f6-41b9-ac34-2e8ed4e40a38 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Ced: Consistent ensemble distillation for audio tagging,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b83eef74-2a62-440d-b115-ef3a3b8508a3 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Learning transferable visual models from natural language supervision,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 781e6ef9-06a4-4bb9-9206-06a07a0c6bdc · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a33c6f-8470-497e-96d9-64aeb76e88aa · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport LoRA: Low-rank adaptation of large language models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 47c7585b-c4c7-48d1-b148-a1a3f3ce8dc4 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport BLEU: a method for automatic evaluation of machine translation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e517c3d9-f19f-419c-b435-54096f732b1a · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport ROUGE: A package for automatic evaluation of summaries,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 61bcdc4c-f0aa-4568-8553-6bbc82145c63 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c0986ccc-04ee-4e0d-8961-076bf4d492bc · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Cider: Consensus- based image description evaluation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1799ba48-c63a-4819-93bd-47d93c37fe01 · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Spice: Semantic propositional image caption evaluation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f919a529-ed1e-472e-85d8-a5fd2322b9ad · outbound
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Improved image captioning via policy gradient optimization of spider,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 92e1dd92-2121-47ad-9bc6-fa069eb0e06c · inbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c84687-edff-4251-a172-5b1d6519ca57 · inbound
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · inbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 840821bc-8bd9-41ef-9568-2393b100d7b9 · inbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.