Pith. sign in

Paper Citation Record · LEDGER

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2506.20361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20361 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:51.803073Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:48.450553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:53:52.505268Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cae30958-580b-4a12-898d-bb35d334a24c · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:56.440542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:48.237549Z digest=sha256:0d17a681dff640a1313fcafcc09551f577f4e02a7747461336c7df72b1247e39

Observation 24f8d8b4-f788-404d-b227-2091412a59d3 · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:56.252196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:48.651613Z digest=sha256:b993d210052023817c5a89a437541c2811d20984eae628ce76537fa09a05f558

Observation c0134223-de03-4e42-b025-6ce6f5285cbb · outbound

This paper cites Dataset We analyzed the phonetic decodability window on two datasets separately.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Dataset We analyzed the phonetic decodability window on two datasets separately

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:56.095836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:48.823441Z digest=sha256:b6e524431e84ef29d38f81d57ddccf457434cb688bf22a5a9946439e4a98950e

Observation 61bea0c4-4e63-4d3f-b9ac-cf73c541fca9 · outbound

This paper cites an unresolved cited work.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:53:55.892928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:49.013333Z digest=sha256:7feb39fd3f62dad3c18bbb2b98719a0b7d8d70c23f5e54c9bd6ae691d4476cb6

Observation 7eae74aa-131c-4dc9-8a0f-c363e80322ef · outbound

This paper cites We found that A V-HuBERT’s encoding of speech temporal dynam- ics is dominated by its audio input.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models We found that A V-HuBERT’s encoding of speech temporal dynam- ics is dominated by its audio input

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.743362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:49.188316Z digest=sha256:8a9773e6e7a155c86758d55d4f33632d5cf0466aa8c537b9e9d1e91ad7de4ad5

Observation 29e31935-0a1a-4314-8027-bc8c1e01f636 · outbound

This paper cites The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.613791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:48.450553Z digest=sha256:604e5ac80aa338b5f8977b9ab202a52b35d3544bd2da01925d1f0542e6598982

Observation dc7c5188-171b-4f40-886a-d53c5020da7c · outbound

This paper cites We would like to thank Biao Zeng from University of South Wales and Hao Tang, Sharon Goldwater from ILCC, Univer- sity of Edinburgh for useful discussion.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models We would like to thank Biao Zeng from University of South Wales and Hao Tang, Sharon Goldwater from ILCC, Univer- sity of Edinburgh for useful discussion

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.556123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:49.377704Z digest=sha256:db10ae3da4a12590bf0e0b93179bb0cdd7a4664cd38a0f591a56e653394ab96e

Observation 64abd8f1-82eb-4fbe-abb1-9639bad530a5 · outbound

This paper cites Using artificial neural networks to ask ‘why’ questions of minds and brains,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Using artificial neural networks to ask ‘why’ questions of minds and brains,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.394127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:49.498398Z digest=sha256:1d087ce55857ba56f211bf1032b20f7f7f56e806c2cf25c22d24e56fcbadcf0a

Observation b48a8210-51ec-4cc5-9f5f-21253ce644b2 · outbound

This paper cites Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recur- rent automatic speech recognition systems,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recur- rent automatic speech recognition systems,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.241297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:49.674523Z digest=sha256:30c355a6f0c840830979968bb01afbc07f2fedc49d871c8c83dfe90d7e7eb4c4

Observation 8ec8c1a4-95b3-4bb8-a244-04ffd366196b · outbound

This paper cites Hearing lips and seeing voices,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Hearing lips and seeing voices,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:49.888090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:49.888090Z digest=sha256:1cddb7e0ce99c6c36ba3d7ada5d1143102d8d5382831919b61254930455edbf0

Observation 8cc38be9-dce2-4cfe-aa45-8ac763842dd8 · outbound

This paper cites Multisensory integration: current issues from the perspective of the single neuron,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Multisensory integration: current issues from the perspective of the single neuron,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:55.026324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.029945Z digest=sha256:ce4696d4de5a2f374fe3619a2a7c08e15294825d3de2787aff45088b590cfd58

Observation 0059e289-68da-490d-b5e5-6050a2e55bba · outbound

This paper cites Multimodal deep learning.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Multimodal deep learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.856937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.185624Z digest=sha256:3aae17d6adb725dcb13b6b19d1e5166aa8dc5be50a8e46c9bb3640815a7a027a

Observation b6c4587a-83bb-48dc-a28a-7e80954ebd17 · outbound

This paper cites On the role of noise in audiovisual integration: Evidence from artificial neural networks that exhibit the McGurk effect,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models On the role of noise in audiovisual integration: Evidence from artificial neural networks that exhibit the McGurk effect,

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:53:52.389153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.297411Z digest=sha256:c21d833eb9e19a70d30cdfe2a136610dfcc3b7a7d9db2b34cc745b6011f5fa69

Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:50.391701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:50.391701Z digest=sha256:e015d141b67043b380a16fded786e88e9149495e6a295d97c9585803145a3627

Observation 28fdc799-899e-4acf-b4c6-6a96d44263e6 · outbound

This paper cites The natural statistics of audiovisual speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The natural statistics of audiovisual speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.662843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.483547Z digest=sha256:29d7ce7eb20a80d7382d3636f27aa06f3f42661a163fb3d1838c07dda714bc55

Observation 71710634-2011-40e9-a40e-26b6ceaf68bf · outbound

This paper cites Bimodal speech: early suppressive visual effects in human auditory cor- tex,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Bimodal speech: early suppressive visual effects in human auditory cor- tex,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.377216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.593870Z digest=sha256:a21b873adcde96acc970805119339edeac01fc3918a01a47b741432eaa91ced2

Observation 7fd2bde6-cde9-4a31-bc63-52dfa385c0da · outbound

This paper cites Visual speech speeds up the neural processing of auditory speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Visual speech speeds up the neural processing of auditory speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:54.175482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.668168Z digest=sha256:7926d105fc1c2f479ea3099ff8fa3f5f0660514269f393d167f0cbeb740e4e90

Observation d470de32-c3e2-4f2a-9d10-d283a1e640c9 · outbound

This paper cites Asynchronicity between visual and auditory information in audiovisual speech: Evidence from four types of consonant- words/b/,/t/,/k/and/g,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Asynchronicity between visual and auditory information in audiovisual speech: Evidence from four types of consonant- words/b/,/t/,/k/and/g,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.994140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.739263Z digest=sha256:635f08d2d4f49bf285aafc59f83e368c197a6b685f582041670eb0d0e6add130

Observation e307496b-1ff3-4266-ad1e-6a2c3f165244 · outbound

This paper cites A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.840056Z digest=sha256:469c81eea3fddb8c3a1bceec6fa8bc28e2be73608b35439d481313a3e17ebf6c

Observation 38ebd5c8-3751-4a3f-8bd7-7411414dde71 · outbound

This paper cites Neural dynamics of phoneme sequences reveal position-invariant code for content and order,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Neural dynamics of phoneme sequences reveal position-invariant code for content and order,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.724658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:50.949106Z digest=sha256:a87df2d6436211bec4bcbe5778362f2516369d425c1b4bb93fc08d341cd8f3bf

Observation 123861fc-64e1-40ad-9645-3e6dc382826e · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.030814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.030814Z digest=sha256:08a05e988213376d51968f12c39dd75c995ccb183498efeac4134286387e955a

Observation d15b4bb8-4347-45bc-9b1a-8b7614b2d6f7 · outbound

This paper cites Deep residual learning for image recognition,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Deep residual learning for image recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.105709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.105709Z digest=sha256:5ea20efb4d131b1f5ecb5514afc9b1ca0063eb7406ee5aadbcb0fd44673c32db

Observation 22583bb2-f733-4bba-b601-d0c22e9abf22 · outbound

This paper cites Dynamic encoding of acoustic features in neural responses to continuous speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Dynamic encoding of acoustic features in neural responses to continuous speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.448903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:51.187054Z digest=sha256:4cef20c4d46045ef41e0400f9d6b3272a875a9fd66b2e6f157839f3237f4aba1

Observation e5c540e0-211b-4e72-91e0-98f4ac694089 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Representation Learning with Contrastive Predictive Coding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.262919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.262919Z digest=sha256:32efcfc1890fa32f11390e287e74bd335f0ec49d09988ebaf798d1b182a6a0bb

Observation f4d1bdad-7f83-49d2-89a1-2787d06aa1af · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using Kaldi.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Montreal forced aligner: Trainable text-speech align- ment using Kaldi

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:53.205875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:51.369421Z digest=sha256:5d543453a7cc98ff463d0d942c4443caa4e9b06b89d069334580462e22af2467

Observation 1fb8ac9e-8336-43ee-b842-f7e5a04fb035 · outbound

This paper cites Prosodylab-aligner: A tool for forced alignment of laboratory speech,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Prosodylab-aligner: A tool for forced alignment of laboratory speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:52.994338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:51.485127Z digest=sha256:5e88ef71c59f5e3073679fadfe17b21df7faa9f6d4635c803f74784344355137

Observation 216414d2-5e51-46c8-94e4-9dcd976e5bcb · outbound

This paper cites An audio- visual corpus for speech perception and automatic speech recog- nition,.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models An audio- visual corpus for speech perception and automatic speech recog- nition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:53:52.804935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:51.637293Z digest=sha256:60f08c907a979b764d56217b6e1619e38208df32662e96a1903ae6ffafdf7ba3

Observation a70d74ae-21d5-4051-92ac-96d592074ef8 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models LRS3-TED: a large-scale dataset for visual speech recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:51.803073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:51.803073Z digest=sha256:55cc81d2feea87ebef9da718aa134c47b5f41107ee620318d28eeaa58c6abedc

Pith citing papers

Observation 29e31935-0a1a-4314-8027-bc8c1e01f636 · inbound

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models cites this paper.

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:53:52.613791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:53:48.450553Z digest=sha256:604e5ac80aa338b5f8977b9ab202a52b35d3544bd2da01925d1f0542e6598982