Pith. sign in

Paper Citation Record · LEDGER

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2607.09001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09001 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T01:06:45.040464Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f167546c-c4c0-4744-8f5b-c3335d7a8a7f · outbound

This paper cites MIR-GAN: Refining Frame- Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition MIR-GAN: Refining Frame- Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:937253d32af4049a485f604035adc24a4e740e1b6d6dd3101c9abfde6c6661cd

Observation ee31fcdd-2cb7-416a-a5d6-16690ac91733 · outbound

This paper cites Learning Video Temporal Dynamics With Cross-Modal Attention For Robust Audio-Visual Speech Recognition,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Learning Video Temporal Dynamics With Cross-Modal Attention For Robust Audio-Visual Speech Recognition,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:16b128bcbbb61dde7ba9fe450c21e8884ba6bd2608edf504ae7ede738daafb3f

Observation 9ea3be42-e637-48cb-a791-32a2aa241084 · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:fbf7fddcefc018997629274e038e9193cda94a70f8b711395151e81ebba19aeb

Observation fe724f70-47e9-427d-92b7-4e43ce0213ee · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:b7fc9dcdda09a296bb410f62d1bb708029389efc5c3639a872c3e642e221c9a9

Observation 0ae49249-600a-4e0d-8ea2-3e0b35093492 · outbound

This paper cites Large language models are strong audiovisual speech recognition learners,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Large language models are strong audiovisual speech recognition learners,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:0dfe8b328c0d24deb56d2ef3ff93d9ffff60712efcd26fbb3c94ea0945c01c98

Observation 23212bc8-22d4-45e4-80f3-af8260d6fefb · outbound

This paper cites Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:565bce101de97ea183debb3afe64638b7b233030573b33e8e7f527e46708f48c

Observation 291ffc70-7eab-4919-b2c7-de4b54986607 · outbound

This paper cites Hy- brid ctc/attention architecture for end-to-end speech recognition,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:f29315ccdc914b076e3fae2aff5daf4b6771aa4d5b0d76316f203f6931bcd027

Observation 673897b2-02c6-412b-b42f-f4c665320f70 · outbound

This paper cites Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:8778d891ce105ac71db9c60249a19fd128b128340568c54ffacb90e8f9eb4811

Observation f81a4102-2ed0-4100-9ae5-2354867d036e · outbound

This paper cites MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:a16814a7287df505bcc6bfba7aeae9ee81d5d0455757780c9ba5dce2e85a3778

Observation 0226e612-bd6c-4df4-903e-ef63eaf8ea25 · outbound

This paper cites Omni- avsr: Towards unified multimodal speech recognition with large language models,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Omni- avsr: Towards unified multimodal speech recognition with large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:a7aeafcb355044ec30b6934578dd2465115d6b5e9dc1de29648cab04a2ef073f

Observation c8ab230c-6502-4938-bf79-13b7ddaffdc0 · outbound

This paper cites V ALLR: Visual ASR Language Model for Lip Reading,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition V ALLR: Visual ASR Language Model for Lip Reading,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:493e9f77b79837392dd5b512ba955696753559bb8e4488ed8e479bb9eff11141

Observation 668c63e2-7d92-4074-b494-02f17cbfc5f8 · outbound

This paper cites Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:e5df2309574c89202c9ee64f1f5bb1e987ab9dec933f8e7b1de7888a14b40e6e

Observation 49c009b2-7cc0-4e81-b4a5-35bae2155dce · outbound

This paper cites Uncovering the Visual Contribution in Audio-Visual Speech Recognition,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Uncovering the Visual Contribution in Audio-Visual Speech Recognition,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:6d8f173d3b08bb444f9695c14c02c148d05690a3276b7021d44c3c3325b4c287

Observation 759b8958-2b69-44fa-9a63-7f3147880e0d · outbound

This paper cites Align before Fuse: Vision and Language Representation Learning with Mo- mentum Distillation,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Align before Fuse: Vision and Language Representation Learning with Mo- mentum Distillation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:0aa5aca99f5e127bd22cab728f4192e5e5aa2b7835eeb5850f8bacd3f130d594

Observation 8dd1ee1b-61a5-4cb7-b8ee-3eaa177c761b · outbound

This paper cites Optimal Transport for Domain Adaptation,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Optimal Transport for Domain Adaptation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:26cdb2f69afb9f8b10939ccf2ce83dfe57a1934d7f37b16af322ab8ae3aeeaec

Observation 18bfb088-1db6-4d63-9a6c-01e7a1e78bc3 · outbound

This paper cites From Word Embeddings To Document Distances,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition From Word Embeddings To Document Distances,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:5d0d88c0bf70beb77a52d05519b0c1ab7925fbf2b21a2a2ed2028ecdfa0eb983

Observation 17b7f7c9-49ab-4608-aaa4-163cc74b5f90 · outbound

This paper cites Word rotator’s distance,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Word rotator’s distance,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:2a5d530204d8e299770b8be44c4c0de82b396fe91a467fc0e25cb8fa98fe2d85

Observation 87b16424-c484-4fa8-af63-9c1fce9f8658 · outbound

This paper cites Cross-modal Alignment with Optimal Transport for CTC-based ASR,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Cross-modal Alignment with Optimal Transport for CTC-based ASR,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:00e97c44ae2ce2c07f73d76dae1d8e8c2cdcb5ed95fb804bf81462a3aa00e2d7

Observation 8dd54962-a21f-4501-9935-ec3d23e66a78 · outbound

This paper cites Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:42598b5d68ec0377c885db72e41ebcf071e60a0fe34f3551a5b93be564460b29

Observation 840821bc-8bd9-41ef-9568-2393b100d7b9 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:6a58e2e048481ac09c6ff3b4fb468ead1d0fe85ca211fcaab2f5dfa07110fc3d

Observation 02ea7ab5-8b32-4ecd-87cd-e36a2972ea81 · outbound

This paper cites Computational Optimal Transport: With Applications to Data Science,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Computational Optimal Transport: With Applications to Data Science,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:dfe267029ae5098c1abb20f94210d3eb43da4d3e60ef2d289130038a9663e934

Observation fb8549c9-6419-4b36-83d1-811e9f558410 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:6a193bf9a1f26775046bd9d9196fe3d401437e3827d55807932b36ae85018f0d

Observation 7ae23e11-dee8-4f97-8007-8e035b9a828d · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:49f743d76237fb9a08e07be9e158b52597adcf840ab469ce36a3d8bf89a9c31e

Observation c18c4004-6f20-4d08-8545-99bfc3bcb489 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:4afe116e360c6efc15bcc9bd5cab7d581b84abae36ce90eaa6ea8bb1b4e3e38f

Observation a4ba792a-fcd2-43b9-a129-43962eb8f5ea · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:53dc247f44ada06bdc6d2645ef1e9015fe0c9bf9fecf3c3bd62f1c04e2824943

Observation 721de74e-a0fd-40ab-b581-4d492f6e6f44 · outbound

This paper cites an unresolved cited work.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:deec00ce3d227aa550abe406a530a0528ada21efaa6e30c75cd5e257adb0cac3

Observation a4d297d3-40a8-43b8-89f8-9a7267209ea1 · outbound

This paper cites an unresolved cited work.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:58dc34f6bdbff3cac72837dd1d269d4b424951b9a67c754b0e3bc266174e1f46

Observation 6e3d52b0-2307-442b-9c7e-7e91157c637f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LoRA: Low-Rank Adaptation of Large Language Models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:f92e362d29807e06c72e3f2804c22fd30d1b941bf9a87f173a0cf7eb60efc834

Observation f6235595-e86e-49db-85aa-5954e299820c · outbound

This paper cites Sinkhorn Distances: Lightspeed Computation of Optimal Transport,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Sinkhorn Distances: Lightspeed Computation of Optimal Transport,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:2b911f638e08bedaffc7dabfda9a909aa30a1a85a31c36c13ee301f49b0d2fb3

Observation 629c3cf9-fa5f-42be-98c5-2796cc3f325c · outbound

This paper cites Learning to Align Sequential Actions in the Wild,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Learning to Align Sequential Actions in the Wild,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:8bcd382dc16117cbe913191971c503c8edda90d4eeaf8d973202adffc647409d

Observation 225752e4-744c-478d-91d2-12f97720001a · outbound

This paper cites SuperGlue: Learning Feature Matching With Graph Neural Networks,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition SuperGlue: Learning Feature Matching With Graph Neural Networks,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:3414858a16499da2469886e983658d6bd31f4020f99ffea89b1789b2d8baf731

Observation b5253427-f5e2-4868-9c27-cfbf49652fb2 · outbound

This paper cites Unsupervised Learning of Visual Features by Contrasting Cluster Assignments,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Unsupervised Learning of Visual Features by Contrasting Cluster Assignments,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:14db3ce76daf2547a0b5631d5153355e06efa0be04af2fe819ff25a1f5c6d681

Observation 74db1391-9cae-403b-a9a2-634eee850db8 · outbound

This paper cites Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:03c1aafc1e2d7ad031f02636bab17fd7d6eb0a5abdb8b82b9d8ac64c28a8416f

Observation 4cfd9cb0-553a-4bef-8919-9b2a8056e3bd · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models,.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition SALMONN: Towards Generic Hearing Abilities for Large Language Models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:915774384fdadfaaa62c377b170cb8a9a06d5587eab0174762521062e5790b4a

Observation b556f8d9-397e-47ee-8ffc-7169f4dd0125 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T01:06:45.040464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:06:45.040464Z digest=sha256:501906fd71e1159da16c198c54d4bfa1dfcb065cdd8f6427f3a15a3d9dc122c9

Pith citing papers

No inbound Pith citation observations are available.