Pith. sign in

Paper Citation Record · LEDGER

Multi-Grained Spatio-temporal Modeling for Lip-reading

As of 21 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:1908.11618.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.11618 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:13:16.605795Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1fb123a-b994-4db9-a1cf-6d279826eff5 · outbound

This paper cites Improved speaker independent lipreading using speaker adaptive training and deep neural networks.

Multi-Grained Spatio-temporal Modeling for Lip-reading Improved speaker independent lipreading using speaker adaptive training and deep neural networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.342047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.330178Z digest=sha256:1e91d64d265a3dea8c80648dc4e4dcc971346ea0a7cdc4200ada75695668434b

Observation d10a5ed7-5bf3-4b64-be67-3f884ee3769a · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

Multi-Grained Spatio-temporal Modeling for Lip-reading LipNet: End-to-End Sentence-level Lipreading

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T10:13:16.335704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:13:16.335704Z digest=sha256:970c4f54243904e730ef234c0575bc622ac395fcbfd4bf3a55f0eb183d4aafcc

Observation 2c8c86ec-206a-47fb-bc35-314d54824ab0 · outbound

This paper cites The natural statistics of audiovisual speech.

Multi-Grained Spatio-temporal Modeling for Lip-reading The natural statistics of audiovisual speech

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.322733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.344748Z digest=sha256:be18bcdd9e3ffff03535b7913fdfe36ebbb126309f243c032bb704e01234db4d

Observation de53de35-197b-4c63-8a7d-5751cf81757a · outbound

This paper cites Lipreading from color video.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lipreading from color video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.296347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.350281Z digest=sha256:462e8ca6fea35204c4f2a3b5fd78e3d0f1dfb867a1c94a271f4f62211b596552

Observation eb7f2d26-0ca4-4fd3-8b24-4795f649d1c4 · outbound

This paper cites Lip reading in the wild.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lip reading in the wild

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.274433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.356484Z digest=sha256:248bda2a7ffd3c579eea90470a216dc79d946b0df219d0cb9788f61fad3be969

Observation 9000b87c-7555-479f-968c-6449428a50ca · outbound

This paper cites Learning to lip read words by watching videos.

Multi-Grained Spatio-temporal Modeling for Lip-reading Learning to lip read words by watching videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.257495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.364587Z digest=sha256:9b880b259d2ec1a8c1027a53373dca373a00ab29372f6f59c396fb68af9bbd90

Observation 0c9c0936-ef40-4e83-bb45-22f82e3bfaa0 · outbound

This paper cites Lip reading sentences in the wild.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lip reading sentences in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.235879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.379468Z digest=sha256:ae6673a75ac67547a8475e403be9ba17dd993155333d907be71f3be602c3f579

Observation 2ddffe38-9e90-4bc1-a26d-c7ac4cb618c1 · outbound

This paper cites Toward movement-invariant automatic lip-reading and speech recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading Toward movement-invariant automatic lip-reading and speech recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.214393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.389161Z digest=sha256:da3f2ffbc5f5892185dbcccd098c60a8a3d18af94b4a34a9a65bdbdf81230511

Observation 8b9c4c94-bd7f-4e58-b953-fa399ad559db · outbound

This paper cites Long short-term memory.

Multi-Grained Spatio-temporal Modeling for Lip-reading Long short-term memory

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.195559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.400542Z digest=sha256:c23b305ddf9bdd0947bf26c6fa8d7c806d07c435fe9d907d0e3e9eb1319a4af4

Observation 9b155bce-615d-4481-b1ab-5d229fe6e8d1 · outbound

This paper cites Videolstm convolves, attends and flows for action recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading Videolstm convolves, attends and flows for action recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.175556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.410492Z digest=sha256:1af5dd86394330c04b6597e88e45bd997b40cfc3685064854d30c1c281d75716

Observation d19605dc-49ba-43c6-b0b8-7fbccaffdc84 · outbound

This paper cites Audio-visual speech recognition using deep learning.

Multi-Grained Spatio-temporal Modeling for Lip-reading Audio-visual speech recognition using deep learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.152173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.428574Z digest=sha256:19741d4cf3d154124e9eaf4ba1b8c7379054270f812314941df3fd9b11c02060

Observation 430d9fb0-6be7-45d2-b500-6afb1d0bcfdf · outbound

This paper cites End-to-end audiovisual speech recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading End-to-end audiovisual speech recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.126325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.442214Z digest=sha256:0ce104de2e8e7149812a79ee3f0bf1b5536d49cca1fbc399025e23d5707c9560

Observation 3cca9545-f18b-461f-950d-d45d147dd8f0 · outbound

This paper cites An image transform ap- proach for hmm based automatic lipreading.

Multi-Grained Spatio-temporal Modeling for Lip-reading An image transform ap- proach for hmm based automatic lipreading

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.100644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.454167Z digest=sha256:f53fc93ceda98b4db8996157d67f2aa86f08415290396907decab201e784f910

Observation 55ce54b8-4f7a-4136-9cd7-ace573e3b85f · outbound

This paper cites Recent advances in the automatic recognition of audiovisual speech.

Multi-Grained Spatio-temporal Modeling for Lip-reading Recent advances in the automatic recognition of audiovisual speech

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.076930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.470890Z digest=sha256:0259aa1b6b4e7cfcdc2511262eebe2bd53f323cce235b6a1c7419c9f959330dc

Observation f1760879-f233-41ba-aa87-02b2c93af2cb · outbound

This paper cites Lip reading using optical flow and support vector machines.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lip reading using optical flow and support vector machines

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.044818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.479365Z digest=sha256:9285bfc61a5793e2e0d1a5fff5416bc55ffd7351d805be98d6a431ae4c2ea7ec

Observation f94d5605-5501-4d93-882f-a11482410d74 · outbound

This paper cites Convolutional lstm network: A machine learning approach for pre- cipitation nowcasting.

Multi-Grained Spatio-temporal Modeling for Lip-reading Convolutional lstm network: A machine learning approach for pre- cipitation nowcasting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:17.016990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.490705Z digest=sha256:7288b52c35a762d86cd3c92d239f641dee74c7aa338c57034c2078c41c01fde2

Observation 49763019-c3b6-4b02-892b-9d8d8c2749d1 · outbound

This paper cites Two-stream convolutional networks for ac- tion recognition in videos.

Multi-Grained Spatio-temporal Modeling for Lip-reading Two-stream convolutional networks for ac- tion recognition in videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T10:13:16.505615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:13:16.505615Z digest=sha256:19160e8048ae8fdb25a0af977061ac59eddf2e8a9944c5bbd55f6a9585799bd6

Observation 693dd69f-c76b-4e51-832c-a69fa585f26f · outbound

This paper cites Combining Residual Networks with LSTMs for Lipreading.

Multi-Grained Spatio-temporal Modeling for Lip-reading Combining Residual Networks with LSTMs for Lipreading

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T10:13:16.519177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:13:16.519177Z digest=sha256:5c0bc717068e0c22929bd486e0bb175d403d66c45660172de77ff4f16978c273

Observation bd126973-8fec-48e2-b45e-918140376801 · outbound

This paper cites Convolutional long short-term memory networks for recognizing first person interactions.

Multi-Grained Spatio-temporal Modeling for Lip-reading Convolutional long short-term memory networks for recognizing first person interactions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.961031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.526459Z digest=sha256:0483bd96cb25e122944ac097d818f172f98486970334cd02b696598aa12b09f3

Observation 994ea864-7e38-4e17-9fcb-6afa9de28a89 · outbound

This paper cites Improving lip-reading performance for robust audiovisual speech recognition using dnns.

Multi-Grained Spatio-temporal Modeling for Lip-reading Improving lip-reading performance for robust audiovisual speech recognition using dnns

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.937044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.536979Z digest=sha256:e2efeadcdbb79cdea85d2b0008b8e6e456c63e0c58e20b370a766ab1775318d5

Observation 604defd1-f3de-4f81-a1c6-cf891aee9a19 · outbound

This paper cites Hu- man action recognition by learning spatio-temporal features with deep neural networks.

Multi-Grained Spatio-temporal Modeling for Lip-reading Hu- man action recognition by learning spatio-temporal features with deep neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.905106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.542441Z digest=sha256:f0be30d1b7f12d40f3aed49a44eaa0df1ede0fd6666e020492dd01a773ae9d9b

Observation 20992fce-8a4c-495f-bc42-f9a5cf4453e1 · outbound

This paper cites Pre- drnn: Recurrent neural networks for predictive learning using spatiotemporal lstms.

Multi-Grained Spatio-temporal Modeling for Lip-reading Pre- drnn: Recurrent neural networks for predictive learning using spatiotemporal lstms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.876930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.552535Z digest=sha256:2c33a751479afd11e64ba71615424608dcca21da8a48088e5cc2b71974ad31e6

Observation 44855688-7a4f-4c2b-9651-06b751e2e4f1 · outbound

This paper cites Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lrw-1000: A naturally-distributed large-scale benchmark for lip reading in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.856778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.564273Z digest=sha256:5df1086caec02ddb1c3a1b188c41ef7c40f1d27764e2ce254c1dc00ce43bd0aa

Observation a62bb90f-2ddb-45aa-a245-6e0022ff9150 · outbound

This paper cites Learning spatiotemporal features using 3dcnn and convolutional lstm for gesture recognition.

Multi-Grained Spatio-temporal Modeling for Lip-reading Learning spatiotemporal features using 3dcnn and convolutional lstm for gesture recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.839284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.576389Z digest=sha256:24038a6cd6cbd4ebf46974139444322a677b3f31a92f8c73ebf2d97e211b3305

Observation c3836fe0-acb8-4f34-9fd1-e27c4d14bfdd · outbound

This paper cites Adding attentiveness to the neurons in recurrent neural networks.

Multi-Grained Spatio-temporal Modeling for Lip-reading Adding attentiveness to the neurons in recurrent neural networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.812779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.584649Z digest=sha256:cc19122100f3bdce72951ee42881e4a602ebf7fb9c6297602ff681bcc0c68107

Observation 7f9febc2-6774-4a98-8fde-8869dee38b18 · outbound

This paper cites Lipreading with local spatiotem- poral descriptors.

Multi-Grained Spatio-temporal Modeling for Lip-reading Lipreading with local spatiotem- poral descriptors

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.781772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.593574Z digest=sha256:c5c9a2e2916ab3a8bf5d0f2f0d964e07371359033be395f7ba69338002b2f413

Observation 3b50c32d-886b-4b2b-bd68-df5501fe7d0f · outbound

This paper cites Multimodal gesture recog- nition using 3-d convolution and convolutional lstm.IEEE Access, 5:4517–4524, 2017.

Multi-Grained Spatio-temporal Modeling for Lip-reading Multimodal gesture recog- nition using 3-d convolution and convolutional lstm.IEEE Access, 5:4517–4524, 2017

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:13:16.753235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T10:13:16.605795Z digest=sha256:1ee1dc9ff2b03e09e83deae07c9675858434a7d39c0c5a6183f717adb593d7f1

Pith citing papers

No inbound Pith citation observations are available.