Pith. sign in

Paper Citation Record · LEDGER

Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.10082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10082 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:54.187906Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.164748Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 282af02d-6efd-4ae8-acd3-9cbcfad85f63 · inbound

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses cites this paper.

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:54.187906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:54.187906Z digest=sha256:f9746ff6b1da37bb5815f6e3c11a7accff2e89af7a6d0b4382be46ac33f29097

Observation a334828b-cd9f-4cb2-a8e6-a0c0914eb90d · inbound

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition cites this paper.

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:12.049729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:12.049729Z digest=sha256:7ff29aa71f673f0ee768a58a5875ad8e8d25e077a49b11e9cbe6218e74d75e3e

Observation ac1134e4-820d-41c1-8371-fd135c4d2cd0 · inbound

HumanOmni-Speaker: Identifying Who said What and When cites this paper.

HumanOmni-Speaker: Identifying Who said What and When Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.527913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T00:57:12.239390Z digest=sha256:51781f863402fa42789b955b89d27552a2e5d1831b49b55339681a386f05c40d

Observation f39153ad-0790-4acc-a295-d007ff5aaa2a · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.166295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:65b5ec31ac31c7f3e77892092e98294ee9e5c55386fda3285023557bc33c0e1e

Observation 7331b4a1-1972-4a67-badd-7c70241832a2 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.598277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.598277Z digest=sha256:83bb12f7793f5c08ee6e4f06c6dbfc8b32579f35ce7ba8e0bb0f1ab09c014509