Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T01:06:45.040464Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2607.09001.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T01:06:45.040464Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f167546c-c4c0-4744-8f5b-c3335d7a8a7f · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition MIR-GAN: Refining Frame- Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee31fcdd-2cb7-416a-a5d6-16690ac91733 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Learning Video Temporal Dynamics With Cross-Modal Attention For Robust Audio-Visual Speech Recognition,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea3be42-e637-48cb-a791-32a2aa241084 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe724f70-47e9-427d-92b7-4e43ce0213ee · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ae49249-600a-4e0d-8ea2-3e0b35093492 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Large language models are strong audiovisual speech recognition learners,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23212bc8-22d4-45e4-80f3-af8260d6fefb · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291ffc70-7eab-4919-b2c7-de4b54986607 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673897b2-02c6-412b-b42f-f4c665320f70 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f81a4102-2ed0-4100-9ae5-2354867d036e · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0226e612-bd6c-4df4-903e-ef63eaf8ea25 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Omni- avsr: Towards unified multimodal speech recognition with large language models,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ab230c-6502-4938-bf79-13b7ddaffdc0 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition V ALLR: Visual ASR Language Model for Lip Reading,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668c63e2-7d92-4074-b494-02f17cbfc5f8 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c009b2-7cc0-4e81-b4a5-35bae2155dce · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Uncovering the Visual Contribution in Audio-Visual Speech Recognition,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759b8958-2b69-44fa-9a63-7f3147880e0d · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Align before Fuse: Vision and Language Representation Learning with Mo- mentum Distillation,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd1ee1b-61a5-4cb7-b8ee-3eaa177c761b · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Optimal Transport for Domain Adaptation,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bfb088-1db6-4d63-9a6c-01e7a1e78bc3 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition From Word Embeddings To Document Distances,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b7f7c9-49ab-4608-aaa4-163cc74b5f90 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Word rotator’s distance,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b16424-c484-4fa8-af63-9c1fce9f8658 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Cross-modal Alignment with Optimal Transport for CTC-based ASR,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd54962-a21f-4501-9935-ec3d23e66a78 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 840821bc-8bd9-41ef-9568-2393b100d7b9 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02ea7ab5-8b32-4ecd-87cd-e36a2972ea81 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Computational Optimal Transport: With Applications to Data Science,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb8549c9-6419-4b36-83d1-811e9f558410 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae23e11-dee8-4f97-8007-8e035b9a828d · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18c4004-6f20-4d08-8545-99bfc3bcb489 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ba792a-fcd2-43b9-a129-43962eb8f5ea · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721de74e-a0fd-40ab-b581-4d492f6e6f44 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d297d3-40a8-43b8-89f8-9a7267209ea1 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3d52b0-2307-442b-9c7e-7e91157c637f · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LoRA: Low-Rank Adaptation of Large Language Models,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6235595-e86e-49db-85aa-5954e299820c · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Sinkhorn Distances: Lightspeed Computation of Optimal Transport,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629c3cf9-fa5f-42be-98c5-2796cc3f325c · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Learning to Align Sequential Actions in the Wild,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 225752e4-744c-478d-91d2-12f97720001a · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition SuperGlue: Learning Feature Matching With Graph Neural Networks,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5253427-f5e2-4868-9c27-cfbf49652fb2 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Unsupervised Learning of Visual Features by Contrasting Cluster Assignments,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74db1391-9cae-403b-a9a2-634eee850db8 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfd9cb0-553a-4bef-8919-9b2a8056e3bd · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition SALMONN: Towards Generic Hearing Abilities for Large Language Models,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b556f8d9-397e-47ee-8ffc-7169f4dd0125 · outbound
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.