Pith. sign in

Paper Citation Record · LEDGER

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2506.16020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16020 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:44.257683Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:39.652813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:53:44.519215Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52b214e2-5f4f-4332-a795-61f44f5b65f1 · outbound

This paper cites As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:53.294893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:39.570507Z digest=sha256:d186df6b66f430558bd5bbd3cd5cf3b8d78711fa481d73855c3250220e210b47

Observation 1e4c4461-0d32-4f73-83e0-d8b00da30345 · outbound

This paper cites an unresolved cited work.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:53:53.052980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:39.747869Z digest=sha256:3a4611a78a0fc231d9355f8617dc380c09c9a4860748cb2da42708e729b3fd88

Observation 506c1bba-3090-49ed-9ec5-2322f40c60c0 · outbound

This paper cites test- seen.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge test- seen

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:52.830799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:39.855427Z digest=sha256:bc0df44f33dfd90f886aea9de86895b6564036bfdc70774f668444967b322fa7

Observation 277be0cc-4485-447a-a6c9-487e60bc5453 · outbound

This paper cites VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.984203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.204764Z digest=sha256:2639f3b995ac83ab028387529270e59b92dc1ff5a6d77049675b1eae42fb8e8b

Observation ab92c3dd-d228-4d7b-8c13-724f6c3a1018 · outbound

This paper cites an unresolved cited work.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:53:52.547797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:39.987942Z digest=sha256:2f8d5a1a6f04fc976f007e2b56cf70c32f94628e5c7ea58b83ff382835b6512f

Observation d4bddf8b-9472-435b-b305-b47b5ffe890b · outbound

This paper cites 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:52.268194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.128440Z digest=sha256:89ac75680cf34484d959f63656d6902f8cbbd8025192b0017a5158a8446cb14b

Observation 52f1e380-921e-4605-b131-3406651a4f90 · outbound

This paper cites Video diffusion models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Video diffusion models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:40.946806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:40.946806Z digest=sha256:9b5335201e16dbe939ba3a30a81212d092f24bd92ebf22c46c58a6b240afff18

Observation d0f9af91-738f-4989-9de4-300298b27e02 · outbound

This paper cites Diffsinger: Singing voice synthesis via shallow diffusion mechanism,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffsinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.637949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.290437Z digest=sha256:6d878620b34ad40f7ce700ac3bb44d883fa61e859ea060957155482c7b74a80c

Observation d98e1918-bcd3-461b-ba9e-42dd8f65eb67 · outbound

This paper cites Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.216181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.400396Z digest=sha256:0acb8809eb109786369c79148ab3b0c99e41fa98f542e7318c987c779b05fff9

Observation 6020805f-aa77-4bcd-80b9-29d618d53d66 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.874815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.478105Z digest=sha256:afd38b661fe3ad6731c8eab3053e42ef170e1799e0784d7afce2ef0bab8dbdb0

Observation 09c9318b-5542-4054-b79f-2687f059b699 · outbound

This paper cites An End-to-End Approach for Chord-Conditioned Song Generation.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An End-to-End Approach for Chord-Conditioned Song Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.430171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.561586Z digest=sha256:06f4e25096ab5200f9b3c375cad9617db677b91a7f244827527303f6dc132302

Observation c088f08b-dec3-49b5-b3b5-ef3c8cc19328 · outbound

This paper cites Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.765455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:40.663144Z digest=sha256:501847a3935c3e3b715f9d48f3e05daf2302d5569e1f98347d4b7e46b6d8a27a

Observation 773e026a-481e-4ded-9e08-7b799704bf1b · outbound

This paper cites Elucidating the design space of diffusion-based generative models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Elucidating the design space of diffusion-based generative models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:40.784747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:40.784747Z digest=sha256:cd42b1f63a498036e7375bd47ccb1c0135542c1f8b80793152144782829de44d

Observation f3e481f1-25ae-4764-bc6f-92bfbadd4674 · outbound

This paper cites Score-based generative modeling through stochas- tic differential equations,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Score-based generative modeling through stochas- tic differential equations,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.046653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.697669Z digest=sha256:e4218e6883f59b04e62917aeee5e659d37b04aea7c53c9576bfb095c3276e6b0

Observation c524afbb-bf39-4a08-a685-cf7d4cd61a08 · outbound

This paper cites Learning the beauty in songs: Neural singing voice beautifier,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Learning the beauty in songs: Neural singing voice beautifier,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.499591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.075668Z digest=sha256:51fb7a174f8165424d13bbefd5734253108ba78a1104ae3a1e3d9097566f8232

Observation 1d54c7ef-8809-4841-9688-b83bfd164986 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Align your latents: High-resolution video synthesis with latent diffusion models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.317937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.185792Z digest=sha256:8a15ece38eed50bd3f7fa5f32252ac030ce05e6290af82b89697b84ccc3cf84b

Observation 99d8ab57-2439-4cf3-a06a-dcfcecd73e90 · outbound

This paper cites An image is worth one word: Personalizing text-to-image generation using textual inversion,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An image is worth one word: Personalizing text-to-image generation using textual inversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.033407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.315721Z digest=sha256:214378d74d2544319283e4e504ec60f31cd968f7424be4f612bd8dd5393f9a28

Observation 199a1995-ba54-4999-876c-43f6e67dbf4d · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.783454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.409879Z digest=sha256:6bb5baf06dfd4b10d7bd0c13e4102dd610ef4b046cec76bee912d2710f5ea7a4

Observation 1f057377-e4ef-48be-9bb6-1bff537e9148 · outbound

This paper cites Scalable diffusion models with transform- ers,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Scalable diffusion models with transform- ers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.548243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.541143Z digest=sha256:6e0370cad1afdaec313803b20d11cfedb7b48e8f0deac7a31a9fd6c6740b0b3a

Observation c9d916d6-6dcf-47b5-962b-70de28f610e6 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Grad-tts: A diffusion probabilistic model for text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.292215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.601848Z digest=sha256:d880cbc1b434ad942999e0ac3eca768a7d3ea418abdef8bb8033c70da870c27b

Observation 953b44ab-1be8-41db-a904-2aa718c0b47f · outbound

This paper cites Novel-view acoustic synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Novel-view acoustic synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.582153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.602814Z digest=sha256:b084ec9f77578029f978fe7b6efb5249080beceb97b257da36e125ad4fea1a23

Observation 6b7b15d7-71e0-4583-aecd-21ce5fd23faa · outbound

This paper cites Como- speech: One-step speech and singing voice synthesis via consis- tency model,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Como- speech: One-step speech and singing voice synthesis via consis- tency model,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.915571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.831441Z digest=sha256:5b2bb0dbbdb8f2997b16c003762b895d6223e2e0899813f711136df14a1e6a5f

Observation 35035fff-63ed-4215-a90d-49f5ca2bf935 · outbound

This paper cites Consistency models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Consistency models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.652424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:41.994604Z digest=sha256:b51e4ec4a52356491da9e3df4ac6d82ae147bb3a7247a6195f2ba5901e7eacdc

Observation ea639187-d110-4431-83f3-880b40a28e15 · outbound

This paper cites VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.614311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:39.652813Z digest=sha256:50b0e0af73f0a23874d863669c512c804f6efd4be38acff4f416befa8208a708

Observation 762799b2-c25c-48d3-be60-9f847c2a8a4d · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:42.127548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:42.127548Z digest=sha256:60839e992dfc1d95c5ac01a2e58f421ecaa7e9ac71c30c847391e00ea9a9f386

Observation 66992bb6-d77c-48a6-9a8a-0f57d0865c82 · outbound

This paper cites 2.5 d visual sound,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 2.5 d visual sound,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.339568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.249129Z digest=sha256:c014065e0e56bf7cea7468fb253fbbc07aebad1c0ec8cbc73f1e128276a68ee7

Observation aafe6f73-5b3a-4638-86d2-35aef7133339 · outbound

This paper cites Enhancing spatial audio generation with source separation and channel panning loss,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Enhancing spatial audio generation with source separation and channel panning loss,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.071677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.339805Z digest=sha256:f7874e40dc543dbca3f46f264f0d1c4ab3307230e4c7f87c36ffdaea40c631b0

Observation 6f054485-c628-454a-8b32-7a6c2efb7461 · outbound

This paper cites Multi-source spatial knowledge understanding for immersive visual text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-source spatial knowledge understanding for immersive visual text-to-speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.807211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.456963Z digest=sha256:bfb051e7897614cd38a1542588601ee15f1b5ed71d628d02f6a3a2027821d6fd

Observation be610d70-f0ab-47ef-8361-5ed037b37824 · outbound

This paper cites Visually guided binaural audio generation with cross-modal consistency,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visually guided binaural audio generation with cross-modal consistency,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.365066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.753463Z digest=sha256:b9e6185df6a3d4c61e7a94a12a9493a5a9ad94f5a0ad483bc8225f7b19e4ebf8

Observation af230ae0-97c4-4739-8772-9a4b993b04f0 · outbound

This paper cites Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.055369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.846974Z digest=sha256:ea8e928780db35f314efd406a951a8015786fbba97bf1de201717df9bc09833b

Observation c799f48e-d42e-4be5-9312-64db73ad783a · outbound

This paper cites Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.792475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:42.944905Z digest=sha256:de9f4841d197f61332f137bd0ae0106a32dd20b78723c84ba291db0c7e2f57b0

Observation 6a410bcf-9fb6-404e-828c-75bea8fa0ee0 · outbound

This paper cites Deep residual learning for image recognition,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Deep residual learning for image recognition,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:43.052802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:43.052802Z digest=sha256:1d4cc07a87d6cbd54866c91c4cb60d8da14408fead906257f92706aed500cdd5

Observation e5c29d9c-fdd8-4f27-b17a-584b1dd191b9 · outbound

This paper cites Diffusion schr¨odinger bridge matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffusion schr¨odinger bridge matching,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.625428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:43.170695Z digest=sha256:3bd2a7bc6a67e50c58cb26c23d25e5a2e456b6dfba422af2ce2590b30f8729da

Observation dc9a9d56-171f-4a64-914a-df7ef556c927 · outbound

This paper cites Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.384510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:43.271303Z digest=sha256:f9c05c72ff1292efefb6ccd60ddd9571f489e9fec45201eec60ec28c31150bbf

Observation 1489dcf6-c11b-40ad-b9d3-f44e0dcf7b08 · outbound

This paper cites Simplified Diffusion Schr\"odinger Bridge.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Simplified Diffusion Schr\"odinger Bridge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:43.454000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:43.454000Z digest=sha256:1b37b031af3a2e97690f5dea52be4c7075ec6d63b3c3d6b05739b3fb2ff79f73

Observation 4713961d-e9ee-432c-b691-bece85ce1578 · outbound

This paper cites I 2sb: Image-to-image schr ¨odinger bridge,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge I 2sb: Image-to-image schr ¨odinger bridge,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.162553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:43.577638Z digest=sha256:272a6624c0d9f130c70ea2c835a00e92a94c57df90550237ff11197fb7eff6a6

Observation 6c5ec59b-df2f-4f53-b61d-e06d87539de2 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual met- ric,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge The unreasonable effectiveness of deep features as a perceptual met- ric,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.994718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:43.667055Z digest=sha256:737ef5b8b4bfa3e230cfd95ea4bbc05c50a788bbf2a727eb6b39e7339bb312c8

Observation 69fd5a5a-6235-4337-99af-87d15d8e698a · outbound

This paper cites Visual acoustic matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visual acoustic matching,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.754757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:43.819082Z digest=sha256:4687f977229d891be6c2c4fd7b8828e5cdc2d8d2c32dd44eb6653290075fb8a1

Observation 82807bfe-5b9f-44c1-a1c7-768225211a68 · outbound

This paper cites Ima- genet: A large-scale hierarchical image database,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Ima- genet: A large-scale hierarchical image database,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.514262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:43.931528Z digest=sha256:3b4579ec6d46fed52b3c4dc26b67ac5b0aa8bcf2829f67d15c9d9e9532e11442

Observation 1ab969e2-2260-4507-92e2-0b1980acaa83 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.249542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:44.050995Z digest=sha256:74d0d1a804ee52b874f86fe55e30ed2fab43477f5f396ecfd171f4f495e292c7

Observation 2e876622-dc92-42a8-bd94-09c3e13b7a74 · outbound

This paper cites Self-supervised vi- sual acoustic matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Self-supervised vi- sual acoustic matching,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.071008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:44.171093Z digest=sha256:543d366178556b6c3cf997c102cd9d565d8333762698f2b6a2dd11c0818e7ef5

Observation feda8574-5c2b-480e-bd39-342582047854 · outbound

This paper cites Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:44.843365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:44.257683Z digest=sha256:bba9af1fd14345b45f129a21088a4fb92e2abb38559a5e1a76f5d0d0b8398b01

Pith citing papers

Observation ea639187-d110-4431-83f3-880b40a28e15 · inbound

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge cites this paper.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.614311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:39.652813Z digest=sha256:50b0e0af73f0a23874d863669c512c804f6efd4be38acff4f416befa8208a708