Pith. sign in

Paper Citation Record · LEDGER

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2506.20361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20361 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:51.803073Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:48.450553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:53:52.505268Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cae30958-580b-4a12-898d-bb35d334a24c · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:56.440542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:48.237549Z digest=sha256:437de23a1971482f9d96f0c8b9fe9b7ccea8494b0b81267a234bf5503f58e5b7

Observation 24f8d8b4-f788-404d-b227-2091412a59d3 · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:56.252196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:48.651613Z digest=sha256:272943d0b300737677d5792af6e7f3ee5122b0ca29e4469b419b660c43235662

Observation c0134223-de03-4e42-b025-6ce6f5285cbb · outbound

This paper cites Dataset We analyzed the phonetic decodability window on two datasets separately.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Dataset We analyzed the phonetic decodability window on two datasets separately

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:56.095836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:48.823441Z digest=sha256:dd1602172a0e4bb6f288ddd9078e9276069c80222e3afa726b6b196ff3870d78

Observation 61bea0c4-4e63-4d3f-b9ac-cf73c541fca9 · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:55.892928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:49.013333Z digest=sha256:2c29d47e8970290e91c54952d0fa001e1a46e27ce9cbcebcd1a13b61910bc64c

Observation 7eae74aa-131c-4dc9-8a0f-c363e80322ef · outbound

This paper cites We found that A V-HuBERT’s encoding of speech temporal dynam- ics is dominated by its audio input.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models We found that A V-HuBERT’s encoding of speech temporal dynam- ics is dominated by its audio input

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.743362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:49.188316Z digest=sha256:1c1b471fce383e5e38dcd345e96dafdf9e484d82ae238d7e0c8b55aaf6a6b6aa

Observation 29e31935-0a1a-4314-8027-bc8c1e01f636 · outbound

This paper cites The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.613791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:48.450553Z digest=sha256:671485cf952007741cfb78838e2aa005d7abda1af3000497bc52592a02159c18

Observation dc7c5188-171b-4f40-886a-d53c5020da7c · outbound

This paper cites We would like to thank Biao Zeng from University of South Wales and Hao Tang, Sharon Goldwater from ILCC, Univer- sity of Edinburgh for useful discussion.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models We would like to thank Biao Zeng from University of South Wales and Hao Tang, Sharon Goldwater from ILCC, Univer- sity of Edinburgh for useful discussion

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.556123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:49.377704Z digest=sha256:70c521e5835970f78a707931845a111e2f13828a0dcbafa14f003903fa242afa

Observation 64abd8f1-82eb-4fbe-abb1-9639bad530a5 · outbound

This paper cites Using artificial neural networks to ask ‘why’ questions of minds and brains,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Using artificial neural networks to ask ‘why’ questions of minds and brains,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.394127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:49.498398Z digest=sha256:aff18d94c1ece0af0a162c77fc71cd36d5df501d34bb2b42476b456f58f6416e

Observation b48a8210-51ec-4cc5-9f5f-21253ce644b2 · outbound

This paper cites Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recur- rent automatic speech recognition systems,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recur- rent automatic speech recognition systems,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.241297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:49.674523Z digest=sha256:2997227a479d9549fc1b97bd7ee6ddd4321f4c41c0fa4dd6621cfebcb969eaf3

Observation 8ec8c1a4-95b3-4bb8-a244-04ffd366196b · outbound

This paper cites Hearing lips and seeing voices,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Hearing lips and seeing voices,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:49.888090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:49.888090Z digest=sha256:24c84efda1b41cf0f5626573e4cdd3fc546c8563c6ab3273a954e831ece9d62b

Observation 8cc38be9-dce2-4cfe-aa45-8ac763842dd8 · outbound

This paper cites Multisensory integration: current issues from the perspective of the single neuron,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Multisensory integration: current issues from the perspective of the single neuron,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.026324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.029945Z digest=sha256:72924eed8f223276b77fdfe725cb05b1cb304710fc7a0c0ed3d13b7c9be25f08

Observation 0059e289-68da-490d-b5e5-6050a2e55bba · outbound

This paper cites Multimodal deep learning.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Multimodal deep learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.856937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.185624Z digest=sha256:6200b212ed31b3adc787fff560ace82f15581cda0076f54485dda9cfbc809159

Observation b6c4587a-83bb-48dc-a28a-7e80954ebd17 · outbound

This paper cites On the role of noise in audiovisual integration: Evidence from artificial neural networks that exhibit the McGurk effect,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models On the role of noise in audiovisual integration: Evidence from artificial neural networks that exhibit the McGurk effect,

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:53:52.389153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.297411Z digest=sha256:51fb452a73083b7c44dcd0e5a08309f94df5eff151342b3c53abc70d00922d1e

Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:50.391701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:50.391701Z digest=sha256:6e8dc13a043be035ef980bfe256a336fb5eebe3e700a508d10f7b782f365777f

Observation 28fdc799-899e-4acf-b4c6-6a96d44263e6 · outbound

This paper cites The natural statistics of audiovisual speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The natural statistics of audiovisual speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.662843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.483547Z digest=sha256:a04ea0bb796ee55b4dbb7e38fa1900e82d248caf8796e2bd70f8f539d9b3eab3

Observation 71710634-2011-40e9-a40e-26b6ceaf68bf · outbound

This paper cites Bimodal speech: early suppressive visual effects in human auditory cor- tex,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Bimodal speech: early suppressive visual effects in human auditory cor- tex,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.377216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.593870Z digest=sha256:a8cf6db4ca5d12c1584759dfc55c44754947889289ddd292e7af1fe12e40488d

Observation 7fd2bde6-cde9-4a31-bc63-52dfa385c0da · outbound

This paper cites Visual speech speeds up the neural processing of auditory speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Visual speech speeds up the neural processing of auditory speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.175482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.668168Z digest=sha256:4d03e47a04e9402433586a07077a03002c4e590cb0754c94ae3b008b5d2a8457

Observation d470de32-c3e2-4f2a-9d10-d283a1e640c9 · outbound

This paper cites Asynchronicity between visual and auditory information in audiovisual speech: Evidence from four types of consonant- words/b/,/t/,/k/and/g,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Asynchronicity between visual and auditory information in audiovisual speech: Evidence from four types of consonant- words/b/,/t/,/k/and/g,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.994140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.739263Z digest=sha256:33821b15ff47f79cc8485551c96b21a33a7f5946eedd0696c9ad3e97b621ab07

Observation e307496b-1ff3-4266-ad1e-6a2c3f165244 · outbound

This paper cites A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.840056Z digest=sha256:37d90f8bcfce050a8d58dce572de6ab62f99625c2a08d8dbfb446955a3626e83

Observation 38ebd5c8-3751-4a3f-8bd7-7411414dde71 · outbound

This paper cites Neural dynamics of phoneme sequences reveal position-invariant code for content and order,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Neural dynamics of phoneme sequences reveal position-invariant code for content and order,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.724658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:50.949106Z digest=sha256:fe11beff4b4fd8e8961a97b8e97039669277183764bcd5f3c23f4d04b4b852c2

Observation 123861fc-64e1-40ad-9645-3e6dc382826e · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.030814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.030814Z digest=sha256:fc9d72d70bd1605ccd24f48714c0e7ed1c19187856f5646a619d11af125ef5b3

Observation d15b4bb8-4347-45bc-9b1a-8b7614b2d6f7 · outbound

This paper cites Deep residual learning for image recognition,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Deep residual learning for image recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.105709Z digest=sha256:c824b197ab9b76c54bf2984f3580c63e55cffe0fc48ecf0a8cc19100927a0eca

Observation 22583bb2-f733-4bba-b601-d0c22e9abf22 · outbound

This paper cites Dynamic encoding of acoustic features in neural responses to continuous speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Dynamic encoding of acoustic features in neural responses to continuous speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.448903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:51.187054Z digest=sha256:f611ea81e3d86b7c56aea0d510d21ee5a5052fb51e17da0bfe02d7739a9e480f

Observation e5c540e0-211b-4e72-91e0-98f4ac694089 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Representation Learning with Contrastive Predictive Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.262919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.262919Z digest=sha256:142a8c1b296fc471f5a6565d9f37d62957bb92c76757f891f81dd582d552ec4a

Observation f4d1bdad-7f83-49d2-89a1-2787d06aa1af · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using Kaldi.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Montreal forced aligner: Trainable text-speech align- ment using Kaldi

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.205875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:51.369421Z digest=sha256:24c77349301bfefa61b09422174e041f30ab455434c6cab099a4219abde6caf2

Observation 1fb8ac9e-8336-43ee-b842-f7e5a04fb035 · outbound

This paper cites Prosodylab-aligner: A tool for forced alignment of laboratory speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Prosodylab-aligner: A tool for forced alignment of laboratory speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:52.994338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:51.485127Z digest=sha256:5d3d41e5b3d255b5025111391ce2f0c357c5dac3db3431efb24fa8c1aaaa7735

Observation 216414d2-5e51-46c8-94e4-9dcd976e5bcb · outbound

This paper cites An audio- visual corpus for speech perception and automatic speech recog- nition,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models An audio- visual corpus for speech perception and automatic speech recog- nition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:52.804935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:51.637293Z digest=sha256:32d700e6f5e51c437a316eba27d3ea9e907d4b95898c9bd7f23f418f467d0429

Observation a70d74ae-21d5-4051-92ac-96d592074ef8 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models LRS3-TED: a large-scale dataset for visual speech recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.803073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.803073Z digest=sha256:279050059b1cabc08053095b74b2df84bfc2d0c4191af0db4b82e80a942cc1d3

Pith citing papers

Observation 29e31935-0a1a-4314-8027-bc8c1e01f636 · inbound

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models cites this paper.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.613791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:53:48.450553Z digest=sha256:671485cf952007741cfb78838e2aa005d7abda1af3000497bc52592a02159c18