Pith. sign in

Paper Citation Record · LEDGER

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2506.16020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16020 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:44.257683Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:39.652813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:53:44.519215Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52b214e2-5f4f-4332-a795-61f44f5b65f1 · outbound

This paper cites As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:53.294893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:39.570507Z digest=sha256:a701f64bab1f5659e9d541408451ecc5dd059d193d67cf185c8fe764ed73f1e8

Observation 1e4c4461-0d32-4f73-83e0-d8b00da30345 · outbound

This paper cites an unresolved cited work.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:53:53.052980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:39.747869Z digest=sha256:5730a1686a5fb7e383de4ba3a7683291bb42f4f13739c4903803cd59b9a9f5a7

Observation 506c1bba-3090-49ed-9ec5-2322f40c60c0 · outbound

This paper cites test- seen.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge test- seen

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:52.830799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:39.855427Z digest=sha256:7ae0c7687950a2f5a596b6a5cea028f626af39978b32d57790435fbda0a0d64d

Observation 277be0cc-4485-447a-a6c9-487e60bc5453 · outbound

This paper cites VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.984203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.204764Z digest=sha256:f7b45dbcd0f37654ba00edb0482c820a64a03b2477ad718748400a38970ec322

Observation ab92c3dd-d228-4d7b-8c13-724f6c3a1018 · outbound

This paper cites an unresolved cited work.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:53:52.547797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:39.987942Z digest=sha256:83f6fe471cf01d67ae760b7c216a62d714c99cea0faf9fd227077244e5b60485

Observation d4bddf8b-9472-435b-b305-b47b5ffe890b · outbound

This paper cites 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:52.268194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.128440Z digest=sha256:ab3fc2390abc6c95a8ff745e21377f92a38c87cb5b96f8f7a42dd13b5332fb46

Observation 52f1e380-921e-4605-b131-3406651a4f90 · outbound

This paper cites Video diffusion models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Video diffusion models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:40.946806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:40.946806Z digest=sha256:9b5335201e16dbe939ba3a30a81212d092f24bd92ebf22c46c58a6b240afff18

Observation d0f9af91-738f-4989-9de4-300298b27e02 · outbound

This paper cites Diffsinger: Singing voice synthesis via shallow diffusion mechanism,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffsinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.637949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.290437Z digest=sha256:7b5ed1bd6833ec6adeb7777d8d4fa23c58b851d39d9715c930bc64f9779a425c

Observation d98e1918-bcd3-461b-ba9e-42dd8f65eb67 · outbound

This paper cites Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.216181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.400396Z digest=sha256:d321d18785347eead1f3b0a7f6dd60472059ebf29bef6e2fb660534d22027ec0

Observation 6020805f-aa77-4bcd-80b9-29d618d53d66 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.874815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.478105Z digest=sha256:4dd79a739237ab5fe7643fec0b1a85e89446da6b544b9644dd183055650ffd64

Observation 09c9318b-5542-4054-b79f-2687f059b699 · outbound

This paper cites An End-to-End Approach for Chord-Conditioned Song Generation.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An End-to-End Approach for Chord-Conditioned Song Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.430171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.561586Z digest=sha256:c2830103f91aac350869dbd39dcdf64635588bedafd8c6433c520f9a3abf67f3

Observation c088f08b-dec3-49b5-b3b5-ef3c8cc19328 · outbound

This paper cites Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.765455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:40.663144Z digest=sha256:af5a1de775571558b7065af653cfcf552748acb81efd8b28e739401910b027a1

Observation 773e026a-481e-4ded-9e08-7b799704bf1b · outbound

This paper cites Elucidating the design space of diffusion-based generative models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Elucidating the design space of diffusion-based generative models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:40.784747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:40.784747Z digest=sha256:cd42b1f63a498036e7375bd47ccb1c0135542c1f8b80793152144782829de44d

Observation f3e481f1-25ae-4764-bc6f-92bfbadd4674 · outbound

This paper cites Score-based generative modeling through stochas- tic differential equations,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Score-based generative modeling through stochas- tic differential equations,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.046653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.697669Z digest=sha256:b2b1b2f0ea5f6ce14f6913897cf7a83e38ed6dc6eea1b468a7e2b625b406991f

Observation c524afbb-bf39-4a08-a685-cf7d4cd61a08 · outbound

This paper cites Learning the beauty in songs: Neural singing voice beautifier,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Learning the beauty in songs: Neural singing voice beautifier,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.499591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.075668Z digest=sha256:5da0a5a2f1f5e214dce461ab32c15190650626b1bd484eba0765daa9cce1419f

Observation 1d54c7ef-8809-4841-9688-b83bfd164986 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Align your latents: High-resolution video synthesis with latent diffusion models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.317937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.185792Z digest=sha256:bfec4567535238d14e87762af322309fe2981481b450e9551cd63a6b3923b530

Observation 99d8ab57-2439-4cf3-a06a-dcfcecd73e90 · outbound

This paper cites An image is worth one word: Personalizing text-to-image generation using textual inversion,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An image is worth one word: Personalizing text-to-image generation using textual inversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.033407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.315721Z digest=sha256:c39d9b765173989bab917a31d50a4dd788d483fa68f4b4f6482c2ef2fdbf7dea

Observation 199a1995-ba54-4999-876c-43f6e67dbf4d · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.783454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.409879Z digest=sha256:32ee0607a977b083d18ad4a8b0af883c274c19c6df74d251227a482fb994e4d6

Observation 1f057377-e4ef-48be-9bb6-1bff537e9148 · outbound

This paper cites Scalable diffusion models with transform- ers,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Scalable diffusion models with transform- ers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.548243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.541143Z digest=sha256:cf955f2a549262a7775d330369c4a311c67cee15a7dec38799a30c4378004183

Observation c9d916d6-6dcf-47b5-962b-70de28f610e6 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Grad-tts: A diffusion probabilistic model for text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.292215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.601848Z digest=sha256:f74a5e8f2e81cbb80678a47ecc5a17876135c7b029986bccc72ff30dca73989b

Observation 953b44ab-1be8-41db-a904-2aa718c0b47f · outbound

This paper cites Novel-view acoustic synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Novel-view acoustic synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.582153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.602814Z digest=sha256:4722e20037be6989534c0fd947315f335f6771230e010abff237319f284b960d

Observation 6b7b15d7-71e0-4583-aecd-21ce5fd23faa · outbound

This paper cites Como- speech: One-step speech and singing voice synthesis via consis- tency model,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Como- speech: One-step speech and singing voice synthesis via consis- tency model,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.915571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.831441Z digest=sha256:49c74a75aa85729e0238d8ca33ff9904989c9d1b48f65334aaed07a039449bb8

Observation 35035fff-63ed-4215-a90d-49f5ca2bf935 · outbound

This paper cites Consistency models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Consistency models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.652424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:41.994604Z digest=sha256:8ac2aeab50cb5254131b66bb5cab699c9e7dd21905a79667e1d237cce145f908

Observation ea639187-d110-4431-83f3-880b40a28e15 · outbound

This paper cites VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.614311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:39.652813Z digest=sha256:31fae3f7cc7342beefb4c2266f5b8ed6b7ab71d085e6e9e288ae03abedaa199f

Observation 762799b2-c25c-48d3-be60-9f847c2a8a4d · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:42.127548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:42.127548Z digest=sha256:60839e992dfc1d95c5ac01a2e58f421ecaa7e9ac71c30c847391e00ea9a9f386

Observation 66992bb6-d77c-48a6-9a8a-0f57d0865c82 · outbound

This paper cites 2.5 d visual sound,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 2.5 d visual sound,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.339568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.249129Z digest=sha256:63ca18aa61096523454bac924229d331f428af4431046b4f2824db985111b05c

Observation aafe6f73-5b3a-4638-86d2-35aef7133339 · outbound

This paper cites Enhancing spatial audio generation with source separation and channel panning loss,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Enhancing spatial audio generation with source separation and channel panning loss,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.071677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.339805Z digest=sha256:6ec84ccdffa0b699dfad36f6b43839146e2943bd8b55293a5dbe0e1f65d491be

Observation 6f054485-c628-454a-8b32-7a6c2efb7461 · outbound

This paper cites Multi-source spatial knowledge understanding for immersive visual text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-source spatial knowledge understanding for immersive visual text-to-speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.807211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.456963Z digest=sha256:5cea809f3f0688b415877e5a5f3bf7771fe1cab6d722488bee9f7c0d486398d6

Observation be610d70-f0ab-47ef-8361-5ed037b37824 · outbound

This paper cites Visually guided binaural audio generation with cross-modal consistency,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visually guided binaural audio generation with cross-modal consistency,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.365066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.753463Z digest=sha256:94bcb60d23a4ae6e464e6b95a140d7f358e582511ac98dc6064a62c526577915

Observation af230ae0-97c4-4739-8772-9a4b993b04f0 · outbound

This paper cites Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.055369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.846974Z digest=sha256:0071bc1e556041bc8d8556a317ec12b84409ff169c9b03e2f142fb97f2ff85ba

Observation c799f48e-d42e-4be5-9312-64db73ad783a · outbound

This paper cites Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.792475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:42.944905Z digest=sha256:49f8a973661a94ea5a7645f04f0861efebcf1d5e15685ce39bde185f8b8a8f86

Observation 6a410bcf-9fb6-404e-828c-75bea8fa0ee0 · outbound

This paper cites Deep residual learning for image recognition,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Deep residual learning for image recognition,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:43.052802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:43.052802Z digest=sha256:1d4cc07a87d6cbd54866c91c4cb60d8da14408fead906257f92706aed500cdd5

Observation e5c29d9c-fdd8-4f27-b17a-584b1dd191b9 · outbound

This paper cites Diffusion schr¨odinger bridge matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffusion schr¨odinger bridge matching,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.625428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:43.170695Z digest=sha256:448e4ecf26b6dab07c6152c77ff37b18474c996b20c3e5cb60028e69ccf3ecfd

Observation dc9a9d56-171f-4a64-914a-df7ef556c927 · outbound

This paper cites Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.384510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:43.271303Z digest=sha256:f2c27111143df74eb547639c99e6649cae26b3ecb66cdcd616db7ece0ec57503

Observation 1489dcf6-c11b-40ad-b9d3-f44e0dcf7b08 · outbound

This paper cites Simplified Diffusion Schr\"odinger Bridge.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Simplified Diffusion Schr\"odinger Bridge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:43.454000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:43.454000Z digest=sha256:1b37b031af3a2e97690f5dea52be4c7075ec6d63b3c3d6b05739b3fb2ff79f73

Observation 4713961d-e9ee-432c-b691-bece85ce1578 · outbound

This paper cites I 2sb: Image-to-image schr ¨odinger bridge,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge I 2sb: Image-to-image schr ¨odinger bridge,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.162553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:43.577638Z digest=sha256:65fe6e204ee5a4a928eed3f0620a02962c0faeefd9d9992b208df1372b185bdf

Observation 6c5ec59b-df2f-4f53-b61d-e06d87539de2 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual met- ric,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge The unreasonable effectiveness of deep features as a perceptual met- ric,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.994718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:43.667055Z digest=sha256:9d4c845eb6a8ee2f822117a7043ea75ee0172e994cae3cd22676a175466f5f94

Observation 69fd5a5a-6235-4337-99af-87d15d8e698a · outbound

This paper cites Visual acoustic matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visual acoustic matching,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.754757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:43.819082Z digest=sha256:8e57b6df7f4173fb95e2de124678e6871b0afb774d97775ba2c0e875b560fd89

Observation 82807bfe-5b9f-44c1-a1c7-768225211a68 · outbound

This paper cites Ima- genet: A large-scale hierarchical image database,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Ima- genet: A large-scale hierarchical image database,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.514262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:43.931528Z digest=sha256:40be317303cd10e85fa42ca8e5b5746718b7ba4ba29fc9c37a55188a5858b887

Observation 1ab969e2-2260-4507-92e2-0b1980acaa83 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.249542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:44.050995Z digest=sha256:6cc38e9b1d71ff76b75787366e1a1719d9b0f3bc5e2754afa6be76f5ac42008b

Observation 2e876622-dc92-42a8-bd94-09c3e13b7a74 · outbound

This paper cites Self-supervised vi- sual acoustic matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Self-supervised vi- sual acoustic matching,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.071008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:44.171093Z digest=sha256:563060e70c55470261df52d80a68d620c740b51fec32dab0cbe6e6a900f8aae7

Observation feda8574-5c2b-480e-bd39-342582047854 · outbound

This paper cites Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:44.843365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:44.257683Z digest=sha256:cca73d3562bab54b9010a5898f5df711fc6ceaef03adc7600fc33aa820794980

Pith citing papers

Observation ea639187-d110-4431-83f3-880b40a28e15 · inbound

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge cites this paper.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.614311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:53:39.652813Z digest=sha256:31fae3f7cc7342beefb4c2266f5b8ed6b7ab71d085e6e9e288ae03abedaa199f