Pith. sign in

Paper Citation Record · LEDGER

Emotional Face-to-Speech

As of 14 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 3 inbound Pith citation observations for arXiv:2502.01046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01046 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:50:43.962905Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:27:30.884916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.782299Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fc86a2a-63b2-486e-896e-48f12d1f1a68 · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.604844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.604844Z digest=sha256:44c39ca32c96c8fc9438bf5705862cd98c4419e0503fea784a6f1c1428d7bade

Observation 4a82a1ca-ec06-4862-b264-10040089832f · outbound

This paper cites write newline.

Emotional Face-to-Speech write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.611319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.611319Z digest=sha256:38d34e37eb1f1c7ae05089eefec7b2568ba29d9046974194a9d1e216842a7b51

Observation 8c784a1f-2db3-45d0-982b-b09334faca72 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Emotional Face-to-Speech LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.616656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.616656Z digest=sha256:2df3cb13f2440c2929920a49c830779d06475c561350fa3e896b8ffbf48ad387

Observation 228fd9bc-ab5a-4b86-8cbf-37803986be1f · outbound

This paper cites SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing.

Emotional Face-to-Speech SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.141971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.625383Z digest=sha256:2efe6e5bfa9b7730ebe03406e0765a23a9a461e7bfada389c4f963fe85ee7702

Observation 1ac5b807-2b4e-45ab-ad6b-f84a0dbfcb44 · outbound

This paper cites D., Ho, J., Tarlow, D., and van den Berg, R.

Emotional Face-to-Speech D., Ho, J., Tarlow, D., and van den Berg, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.125362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.630497Z digest=sha256:d278573c2866ab00021cc7dffb7b52a0ce45b7931b73b312ccc87d942880bfb4

Observation bba212cb-3264-4084-8f3b-1641ebdffaea · outbound

This paper cites W., Fidler, S., and Kreis, K.

Emotional Face-to-Speech W., Fidler, S., and Kreis, K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.109829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.635301Z digest=sha256:e9ee8467af8a159dcd9d284e778a3a7b1297c2316319155c11f098457934e72f

Observation 4001e787-5070-46db-ab23-5ef9159ade57 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:45.095205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.640656Z digest=sha256:43cf327b9feaba089daa925031424f6c97dedf06210c2c19baac812fa723ca58

Observation 0ef881d3-5e1b-429f-9015-27998dd27e63 · outbound

This paper cites and Zisserman, A.

Emotional Face-to-Speech and Zisserman, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.079738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.645714Z digest=sha256:c3fd14eb8eb87ea311c8b59f140a19ff3e05710162f7502a9daebce678a67248

Observation 7f473141-0a36-416a-a612-0b4510cb8fe6 · outbound

This paper cites V2C: Visual voice cloning.

Emotional Face-to-Speech V2C: Visual voice cloning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.062125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.651033Z digest=sha256:1aa3bc30cb991771c9743bec40aac40a5d3f39f4a9b55de81427e5e289f050e8

Observation 83a53979-b4a1-44f2-8a66-6c3259c198c0 · outbound

This paper cites S., Nagrani, A., and Zisserman, A.

Emotional Face-to-Speech S., Nagrani, A., and Zisserman, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.045828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.655804Z digest=sha256:0dbb129c44a7b6c573c4a4315b103ddcea09ea14beea73d7414c9e000052288c

Observation 88efc1a8-5fdc-4fe4-89f3-006ecdceea6c · outbound

This paper cites Learning to dub movies via hierarchical prosody models.

Emotional Face-to-Speech Learning to dub movies via hierarchical prosody models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.030003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.661008Z digest=sha256:7ff481854d114a6ad2adcfb6b99e6c7779f13c942869489d9d9dc6401e7ea3ca

Observation 9d01dd86-d44e-436e-b922-49c544605764 · outbound

This paper cites StyleDubber : Towards multi-scale style learning for movie dubbing.

Emotional Face-to-Speech StyleDubber : Towards multi-scale style learning for movie dubbing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:45.013211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.665906Z digest=sha256:24bbc2c436b7ffb5a8de6a104cf996ed71df49e5b731c6d859a12e2cfbc3a61b

Observation 9b997226-46ca-46b9-b267-0d56038f4fa9 · outbound

This paper cites High fidelity neural audio compression.

Emotional Face-to-Speech High fidelity neural audio compression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.994838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.670452Z digest=sha256:7af91e362b70ff40c376ecae97c0fd4e67310817d2022325f737bc9621bf4767

Observation 90c26510-fbdd-45a1-ac15-96b6d63c93ec · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Emotional Face-to-Speech Arcface: Additive angular margin loss for deep face recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.978383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.675081Z digest=sha256:4c0ae53a98753ceb5d6acff1bf4a5b7ab04f1bbc6fce81a08850560bfd177660

Observation df752898-ec6a-4c72-8170-b51272e7ac6d · outbound

This paper cites and Shutov, V.

Emotional Face-to-Speech and Shutov, V

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.963515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.679918Z digest=sha256:6a9d6604388b386ce47ba783e60e434f57423f989bc35b7f29ce38fa4fdd4da1

Observation 201828de-4436-4332-9442-1cd2f16623d7 · outbound

This paper cites Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature.

Emotional Face-to-Speech Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.948604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.684435Z digest=sha256:b542eae35c9db06c582cdfe163314ec2f933ec12d5974a9d10f24145f5e3be5c

Observation 43a1cc26-d1ee-4d6a-aba2-90c7f0720dbe · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation.

Emotional Face-to-Speech Efficient emotional adaptation for audio-driven talking-head generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.933128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.689001Z digest=sha256:0e9fd92a8d5d256f5e306c3af8467571a78fbc9fe391b615df3bc5dc484ec8a0

Observation 6ce5e821-1e50-408b-9497-39e726890cc7 · outbound

This paper cites Improving adversarial energy-based model via diffusion process.

Emotional Face-to-Speech Improving adversarial energy-based model via diffusion process

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.918063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.693418Z digest=sha256:dc315248a4dcc69f5709c71833cf6d9f432967b4a7222e0e51062174b3f5869c

Observation a6996e8a-99a5-4913-8545-3280c0c30699 · outbound

This paper cites Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image.

Emotional Face-to-Speech Face2Speech : Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.901896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.697928Z digest=sha256:7ba606ea2ae271dedce598af991fadd15433056afeaf6d6e84c1efed43919c70

Observation 0dcbacbf-7564-47ef-85b0-061428d49176 · outbound

This paper cites EGC: Image generation and classification via a diffusion energy-based model.

Emotional Face-to-Speech EGC: Image generation and classification via a diffusion energy-based model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.886498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.702510Z digest=sha256:790ed2b0ab6e69a1e8964d5b6570b9337ee902109a1d18054f92aab1a95f1387

Observation 7fc0aca9-365b-4dad-96ae-1f9d0521f517 · outbound

This paper cites Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance.

Emotional Face-to-Speech Emodiff : Intensity controllable emotional text-to-speech with soft-label guidance

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.870210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.707083Z digest=sha256:657bf5a1e22464b92c185006a3ae3c6bb1a45412a9dc9620b341b18ccdf62069

Observation 7870b00c-eed2-4471-bd49-c34b75d4b41c · outbound

This paper cites An investigation of multi-speaker training for wavenet vocoder.

Emotional Face-to-Speech An investigation of multi-speaker training for wavenet vocoder

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.852149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.711763Z digest=sha256:f15a9d1840af9378e230ce561b5185cd74bec7244d592a6f8e5a96bafd47ed3c

Observation bc283b0d-32bc-419e-82b7-98c5ae23e077 · outbound

This paper cites and Salimans, T.

Emotional Face-to-Speech and Salimans, T

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.837142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.716419Z digest=sha256:4dd9aaf12051402f2e164b54687727989e26dd0e52443d2af523395722241bc1

Observation fe0ceb16-2d99-4d55-bc37-d4d864504a47 · outbound

This paper cites Denoising diffusion probabilistic models.

Emotional Face-to-Speech Denoising diffusion probabilistic models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.821971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.721209Z digest=sha256:d29e7606113753a053ea663132edd566bf6a9ae91bb40b7238c070b4644a51a7

Observation b596290c-7592-4d5b-ba2c-cdf081b6ec5e · outbound

This paper cites and Johnson, L.

Emotional Face-to-Speech and Johnson, L

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.806622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.725912Z digest=sha256:8ca1d203b945db2fbff4195ed3f06f24a19943be8488163e98fab7a612f7df8a

Observation 0eaf4e25-4cfb-41a3-9591-b6a03d4ab89a · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.791581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.730525Z digest=sha256:f2e4add07478a57b234349462719a8659cea75b7563567ad89c1b2b3eb7e99e9

Observation e6f77063-e505-4adc-bede-bbd5ddb52ac3 · outbound

This paper cites Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping.

Emotional Face-to-Speech Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.735052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.735052Z digest=sha256:fe15aa3ad82b51292390a543d889e2a61385f6f068cc8940dda8019779196aa2

Observation c1567119-44ca-4a97-bd0a-c4ba830e4cfe · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.740224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.740224Z digest=sha256:251cbbc0f6b3229aa2a2634f9efb749f97b5b7d5c0c432f7bded019d03037e49

Observation 4b7a60f2-fa75-42f0-83c8-bb0b68734653 · outbound

This paper cites Speak, read and prompt: High -fidelity text-to-speech with minimal supervision.

Emotional Face-to-Speech Speak, read and prompt: High -fidelity text-to-speech with minimal supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.765114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.744826Z digest=sha256:f3837a263a0b069ae7759d849ed679bb3e3fa64f8f1fcf0f44adefe81c77b781

Observation b425ebf1-8a0f-488a-a629-1bac92c2f7ba · outbound

This paper cites Deep Directed Generative Models with Energy-Based Probability Estimation.

Emotional Face-to-Speech Deep Directed Generative Models with Energy-Based Probability Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.749850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.749850Z digest=sha256:68da0b7bcb95d59a6d55355d505aa09d0dd0f116db2d911eb6443299f063dd31

Observation 67f83871-f69a-45dd-93dd-0b227ece1eae · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.749956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.755018Z digest=sha256:c26aaa80b5adcd3abed76bc84ffad4cc2718552e0e97a416237e05601137a8bd

Observation 02329e41-fe90-41a0-b8e9-deeb3d81768c · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.734438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.759986Z digest=sha256:d9e0f77cd5507d9e7c6ba072c2ed75fd35508976e327490841d3ab985b2bcdfc

Observation ef8844f7-c50f-409d-abfc-fe157fb86835 · outbound

This paper cites S., and Chung, S.

Emotional Face-to-Speech S., and Chung, S

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.718803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.764727Z digest=sha256:c099327f87e4b9569d542fd478fe591fb17d154d9f7f94e508b150f4e9969bd6

Observation 7badaced-9882-408d-af97-7de45d9d8ef8 · outbound

This paper cites Hear Your Face: Face-based voice conversion with F0 estimation.

Emotional Face-to-Speech Hear Your Face: Face-based voice conversion with F0 estimation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.138143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.769527Z digest=sha256:af8b1d3b5f844a0424109b252e51f4097c46664540a3598f9dd2f5989bac78be

Observation e147df56-4a2d-43a1-a161-0e72be10814a · outbound

This paper cites UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts.

Emotional Face-to-Speech UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.774408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.774408Z digest=sha256:73189b6ac69e0c69e5b293a54d3e352da016577a19ddae4495d549ae8123d22a

Observation 694e3347-2324-4734-a331-2f27283a99a4 · outbound

This paper cites A., Han, C., Raghavan, V.

Emotional Face-to-Speech A., Han, C., Raghavan, V

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.702628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.779447Z digest=sha256:41f8458301ed780fbaac031d889cc8f5c36e056322f6a5320c92b16e83263085

Observation 75876484-b02c-44ce-a574-ebd534e4271b · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.686565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.783888Z digest=sha256:45e8b978ebf4721ad320208be46268bf4b9171f2daf8a26d1c368f32969e9afc

Observation 4eb7e847-dc45-48ec-a169-667fce75cfed · outbound

This paper cites Towards a simultaneous and granular identity-expression control in personalized face generation.

Emotional Face-to-Speech Towards a simultaneous and granular identity-expression control in personalized face generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.668157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.788335Z digest=sha256:6b996356aa8e7338bcdc50d3af0c091c39bdbade610b1075cd1ab90831a140fb

Observation b071e055-cb81-4e53-b369-b905175722f6 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.650702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.792876Z digest=sha256:95a5cead7252431e8fbf4ad3e6e01aec9acea9aa4e81306e4720a088d7c73a0b

Observation 05cb4b77-1554-4915-a91f-04da4147c73a · outbound

This paper cites and Hutter, F.

Emotional Face-to-Speech and Hutter, F

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.635934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.797480Z digest=sha256:9d39ae1b09be08526195ee681c1603803c7fe00adf3184c7698702d2b5391741

Observation 8884668d-d95d-4799-9f20-f312e427d6c6 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

Emotional Face-to-Speech Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.620993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.801954Z digest=sha256:35015e7bb0be33298458658fb7e4558c8530adb5c0e509a53c40de3049196d7b

Observation 487895ef-ea72-49bc-aca9-17699ef1d044 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Emotional Face-to-Speech emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.605341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.806366Z digest=sha256:592dd1e618f23015b83a06f7916bf3a544c5fd6175d185d46e73dd93b7025fde

Observation 95de6b4f-48c9-4fcd-821b-bae05433aa0c · outbound

This paper cites POSTER++: A simpler and stronger facial expression recognition network.

Emotional Face-to-Speech POSTER++: A simpler and stronger facial expression recognition network

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.810875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.810875Z digest=sha256:95997da44864c59a55b1d2935216794c4aebb063d29ff7138c3ca04a37b3968f

Observation 7282e835-8cd3-4b1b-a943-84abdcd98820 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.589517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.816300Z digest=sha256:bd2fad6021252ceea6e1b3ceaf2a883e6a5da33e5f6bb98f12317a11bcffa8ed

Observation 83977cf2-68c9-4aee-8aa0-e5e0e741473a · outbound

This paper cites Concrete score matching: Generalized score matching for discrete data.

Emotional Face-to-Speech Concrete score matching: Generalized score matching for discrete data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.574543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.820798Z digest=sha256:8999e958265433e70c844cd653d0ad1c44f4525c23d97e905cb83c1175a3be7b

Observation 7e9dd3b9-8940-4494-a023-8fe45fcc6085 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

Emotional Face-to-Speech HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.825414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.825414Z digest=sha256:da1d9f378a96000695cdb281adb23f4b12cc3c3839f1fe5131bb31b208fe7993

Observation afdb7d01-8e67-4cb2-a693-ad523e366df3 · outbound

This paper cites Unlocking Guidance for Discrete State-Space Diffusion and Flow Models.

Emotional Face-to-Speech Unlocking Guidance for Discrete State-Space Diffusion and Flow Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.830118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.830118Z digest=sha256:a905201430d6d9a89a59acd065b299d63ed51c9b6067ed46c560745085f11f42

Observation 21379b9e-a8fb-44e7-953f-b299025db48a · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

Emotional Face-to-Speech Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.834886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.834886Z digest=sha256:6451e6bb7c46f1d84cf0242f3e2b80493060d9a02beb51a7c09969cca4c4a528

Observation 4fec4e79-0764-4560-826d-869f4fbe7d28 · outbound

This paper cites Visual form predictions facilitate auditory processing at the n1.

Emotional Face-to-Speech Visual form predictions facilitate auditory processing at the n1

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.559350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.840930Z digest=sha256:7f607ca916f58af2d41d1e9301d643be6283105801b391c99e47fbb1af027a7c

Observation d8b23048-440a-428e-8eeb-b845b9442490 · outbound

This paper cites and Xie, S.

Emotional Face-to-Speech and Xie, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.544780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.845641Z digest=sha256:2852cef260538dcf2e59644ab53e83dff4ebd16fd58847b1585e11131504678e

Observation 5ee9044b-52e0-4d3c-b73f-0b813e45db40 · outbound

This paper cites Hearing faces: Target speaker text-to-speech synthesis from a face.

Emotional Face-to-Speech Hearing faces: Target speaker text-to-speech synthesis from a face

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.530430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.850193Z digest=sha256:d59fc91130c191efa032986b92af0e01a3e8b6725ed74762b0768cd30e6fd34b

Observation e6f466b0-cfd9-4b98-8ff9-fbdf8debc2f3 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

Emotional Face-to-Speech W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.515015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.854922Z digest=sha256:50377a7b769dfaa5ec07fc1d53b1c34d482f28787eb54d13a94541932ec91efd

Observation f319cb0e-356d-4538-b860-072a1b7ee870 · outbound

This paper cites A., Bengio, Y., and Courville, A.

Emotional Face-to-Speech A., Bengio, Y., and Courville, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.500660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.859270Z digest=sha256:ac62f5acc33d6fef064b644b3a48853ff1266bf724465e429cd374dd8a2e36a6

Observation d808a201-85eb-4347-9039-17a490fddc9e · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech.

Emotional Face-to-Speech FastSpeech 2: Fast and high-quality end-to-end text to speech

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.486244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.863822Z digest=sha256:2af31c0765759e038d21327dd2256877c02298b038111a03d75347980d767f05

Observation b467fb4c-8a20-45dc-97aa-341d8420bd84 · outbound

This paper cites J., Jin, Q., and Guo, B.

Emotional Face-to-Speech J., Jin, Q., and Guo, B

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.471369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.868134Z digest=sha256:d6eaa43693814593fcee7a32b59beacf728df80f5bce96217a87360717507990

Observation 1e9a56be-cefb-4282-99bd-0850c636f88e · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Emotional Face-to-Speech Facenet: A unified embedding for face recognition and clustering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.456550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.872724Z digest=sha256:dd9bb4a5f22a86082e015025ee889c80159927577bcf7ef7d4315fd1daa58926

Observation a14e57f9-f64d-4bb3-924b-27c0db7805b6 · outbound

This paper cites NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

Emotional Face-to-Speech NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.441381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.877177Z digest=sha256:bd9830b0153b83591c52c28e1df52129199d936f3523012b120ef5d599789d77

Observation 1e4269af-0220-4bbe-936b-d2ea22cbce0b · outbound

This paper cites Denoising diffusion implicit models.

Emotional Face-to-Speech Denoising diffusion implicit models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.426446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.881486Z digest=sha256:a52c1c5f5a63a6fa33b6a7eb3671c74b6890aad3a8625f267539916d390ae1b7

Observation a4f43710-7e79-4477-b3fc-93cce2bdb6cb · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.411261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.886082Z digest=sha256:14431908f60af8ab897dc3a749756a85f1ede30921aae76ba5d23ad97ebfcb74

Observation 46ef9fa6-b41a-465a-9401-7e7a96969847 · outbound

This paper cites Attention is all you need in speech separation.

Emotional Face-to-Speech Attention is all you need in speech separation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.396768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.890897Z digest=sha256:b51df1d161bd93ab4c33247206c3138863a300338ff90cf0ff8d25497ddc2a77

Observation 40747989-37d4-44d9-9ee2-83348a6edb91 · outbound

This paper cites Score-based continuous-time discrete diffusion models.

Emotional Face-to-Speech Score-based continuous-time discrete diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.382407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.895228Z digest=sha256:59a5905a0d4f66c2ce8a9a1fb2fa45ceeb1117e5bb6b4eb1b48a08918dbbaef2

Observation 67357277-30e9-4a4e-867e-96a6c819c725 · outbound

This paper cites and Fostick, L.

Emotional Face-to-Speech and Fostick, L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.367580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.899637Z digest=sha256:27bb4cfeee6f01c4d861be1083d03d3ef46cf416f0c5841b1b1b126535268b08

Observation 47f5461b-95aa-4381-a1bf-dba0f1d2c3ea · outbound

This paper cites and Hinton, G.

Emotional Face-to-Speech and Hinton, G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.352130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.904043Z digest=sha256:30ba61eefde2894a1d1ee0522334eb92fd16ebe343aba95e46a5c2e9d865908d

Observation 51fafd75-1f5b-4a6c-bde3-d68fa5123445 · outbound

This paper cites and Vanathi, P.

Emotional Face-to-Speech and Vanathi, P

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.336646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.908535Z digest=sha256:cd5de6e4de737ddf29cf8f82f02be152d5ce98f8922868914ae42dc166b5de65

Observation 47d97212-12bb-42ba-ba0d-5eb6f0572c2b · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

Emotional Face-to-Speech N., Kaiser, L., and Polosukhin, I

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.319986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.912912Z digest=sha256:37d7ed4a1924c63e61eb7dcb607e1c67dd08449368bca6b3d61118ccae7201e6

Observation 505ab147-c5c9-486f-80ae-edd969806099 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

Emotional Face-to-Speech Generalized end-to-end loss for speaker verification

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.304950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.917352Z digest=sha256:ba6f2d07d33352abf264adf250a34b877b761962281df9efd20832e15ed80535

Observation 7215bd37-6571-45cc-9288-b4aa56e3d42f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Emotional Face-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.921870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.921870Z digest=sha256:3107a9be692dbf360c37185fc3234645c47a1bc46834cf9ce1f681a83f46b96f

Observation f604beb7-c73f-410f-b68d-c28c69d65f01 · outbound

This paper cites an unresolved cited work.

Emotional Face-to-Speech Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-09T16:50:44.290346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.926776Z digest=sha256:41dbd9f887f89288bd17fe27dea17bb48f9aa5eb730a00e971743c96a1c191cc

Observation 7424b5fb-18ff-4f5e-b46e-3ec3804fc604 · outbound

This paper cites J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R.

Emotional Face-to-Speech J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.274761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.931175Z digest=sha256:f0fc8ddba727877ac70173712dce2dbe1492a9655c1035e432bc40969d960f1d

Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.935681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.935681Z digest=sha256:39acd01bf539f22bd1ed1f574a0a6d37234a566a01908429035ee920f8273355

Observation 0137943d-8457-47dd-bd86-8f5564f3e4b2 · outbound

This paper cites DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation.

Emotional Face-to-Speech DCTTS: discrete diffusion model with contrastive learning for text-to-speech generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.258792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.940309Z digest=sha256:09c1a16b23caba36b48a2fb81f17e6726b2efb50b19d3fdfdca439b18f319717

Observation b1310808-0727-4144-83c7-132e9c8ab306 · outbound

This paper cites FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model.

Emotional Face-to-Speech FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:50:44.007474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.944733Z digest=sha256:0593bd70809ad5fda0f0357f68fb65c9f40d3809bb970d17955b0c80e51dce57

Observation a275ec1a-d076-4ea0-ba38-02192b2dc34d · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

Emotional Face-to-Speech Diffsound: Discrete diffusion model for text-to-sound generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.949596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.949596Z digest=sha256:2a494aeee701f5962bc08316d925a45a35f66d62f169c0de06c7d7bbd89a631e

Observation 7af2d74f-5042-49d0-aaad-507e99131309 · outbound

This paper cites SoundStream : An end-to-end neural audio codec.

Emotional Face-to-Speech SoundStream : An end-to-end neural audio codec

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.232022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.953942Z digest=sha256:5e262992e60bd535ab5a794299cb569972e2b53bd6bcbfcec248e5cb88f4141c

Observation 1f177877-1668-4d84-9463-adb26d58ab72 · outbound

This paper cites SpeechTokenizer : Unified speech tokenizer for speech language models.

Emotional Face-to-Speech SpeechTokenizer : Unified speech tokenizer for speech language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.216139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.958341Z digest=sha256:35cf8138cdb929cec9fb641b34339d56ee33c98eb958b7c89e75b00e793c1434

Observation 1302bf96-acf4-4fef-90e8-34fb688272a8 · outbound

This paper cites Srcodec: Split -residual vector quantization for neural speech codec.

Emotional Face-to-Speech Srcodec: Split -residual vector quantization for neural speech codec

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:50:44.200703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T16:50:43.962905Z digest=sha256:f3e6e2a8d16a15acdec3e19f4512f9c27e5cca25614e35aa1a550c94e9347a73

Pith citing papers

Observation 3c38875d-3e8f-4ddd-a3e0-d4bbc00889c8 · inbound

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing cites this paper.

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing Emotional Face-to-Speech

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:27:30.884916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:27:30.884916Z digest=sha256:7cbc98a5c083d1404a849169c0c90e0c37a1ac24ea7468e485629387b8e8a009

Observation 107ff3ea-8160-4e01-853b-32bfeccd21b6 · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation Emotional Face-to-Speech

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.784053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:560658f30f590485603cf682de1275b6f6e5301eda7579ac5dfe25ebfccff974

Observation cc278dfb-24cf-454f-8d8b-6e0842a3be83 · inbound

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model cites this paper.

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model Emotional Face-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T22:26:14.531002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:26:14.531002Z digest=sha256:8b4c969ccdc0e992a25f5d284c2639ec39c7a61e62967f94c97fcc7e49669c59