Pith. sign in

Paper Citation Record · LEDGER

Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.10082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10082 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:54.187906Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.164748Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 282af02d-6efd-4ae8-acd3-9cbcfad85f63 · inbound

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses cites this paper.

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:54.187906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:54.187906Z digest=sha256:607b77507ca303dc403e2e7e8eb0b5f68d4f06cbd681e8d2aff238959d99f28d

Observation a334828b-cd9f-4cb2-a8e6-a0c0914eb90d · inbound

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition cites this paper.

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:12.049729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:12.049729Z digest=sha256:7ff29aa71f673f0ee768a58a5875ad8e8d25e077a49b11e9cbe6218e74d75e3e

Observation ac1134e4-820d-41c1-8371-fd135c4d2cd0 · inbound

HumanOmni-Speaker: Identifying Who said What and When cites this paper.

HumanOmni-Speaker: Identifying Who said What and When Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.527913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:57:12.239390Z digest=sha256:d86e052ebd97e72eec8caed47944b26e1960b126e1c475e9da08f134f076c4aa

Observation f39153ad-0790-4acc-a295-d007ff5aaa2a · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.166295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:4c19f5218b11799da9ceb64362bed8eaea59bc5054fe73cefb6bea66f9330f65

Observation 7331b4a1-1972-4a67-badd-7c70241832a2 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.598277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.598277Z digest=sha256:83bb12f7793f5c08ee6e4f06c6dbfc8b32579f35ce7ba8e0bb0f1ab09c014509