Pith. sign in

Paper Citation Record · LEDGER

KVAE: Family of Tokenizers for Multimodal Generative Models

As of 9 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 0 inbound Pith citation observations for arXiv:2608.05798.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05798 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:24:49.346598Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 122 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 738a38b0-545f-4185-a25e-5273ff6c1ee1 · outbound

This paper cites Bitdance: Scaling autoregressive generative models with binary tokens, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Bitdance: Scaling autoregressive generative models with binary tokens, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.884725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.884725Z digest=sha256:93e142c28c87dc32d907701bfd790114c4330141d06c40d93ab3bebdd4b3cfe8

Observation 2cf5ec29-295d-4960-8f0d-451f2c896b44 · outbound

This paper cites OmniDoc-TokenBench.

KVAE: Family of Tokenizers for Multimodal Generative Models OmniDoc-TokenBench

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.890347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.890347Z digest=sha256:03d75edb6a53da5eacc63fb3d2a21a83d04f9cdfb549f1c61b8ffee0d393e2e2

Observation 58e03af8-44a4-4176-9f42-fb71009e891d · outbound

This paper cites AOM Common Test Conditions v5.0.

KVAE: Family of Tokenizers for Multimodal Generative Models AOM Common Test Conditions v5.0

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.895514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.895514Z digest=sha256:50a86e025cb38e7be21f22b16b8e7f6a16a14e0a102eebfc0b23a2d51dc483ef

Observation e0c328e6-e129-4c3b-94e2-3bc18372faf5 · outbound

This paper cites Kandinsky 5.0: A family of foundation models for image and video generation,.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky 5.0: A family of foundation models for image and video generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.900423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.900423Z digest=sha256:363b722f586abccd4473228f861aef4006ea163cfbb397cd119038eb09f1a14a

Observation 1c9e3033-ed5a-4ebb-b5c2-9c9bc343525b · outbound

This paper cites Qwen2.5-vl technical report, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen2.5-vl technical report, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.905419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.905419Z digest=sha256:7d482230eb9fa08399ff7b1ac53d470aba84c83ab42f9fbd3e96992570ed86b8

Observation e7aec70c-cff9-4657-95a8-14e5ec270626 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets,.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable video diffusion: Scaling latent video diffusion models to large datasets,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.915306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.915306Z digest=sha256:692815244be05754094bc10ea6f6f33fd9f1e71fc58594a76936b4ba877a8747

Observation f4e01222-a462-47f1-8769-f251b1d97829 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Align your latents: High-resolution video synthesis with latent diffusion models, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.919954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.919954Z digest=sha256:3070fc2d98745c8ba8b49413688d701b921f7d4a4e81574130407204aa87aa7e

Observation 08a77cb1-62c7-4212-be39-0d505c00db6c · outbound

This paper cites Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A.

KVAE: Family of Tokenizers for Multimodal Generative Models Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.924532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.924532Z digest=sha256:3504a5f07690e4fda21b5308778fe44751fa4f8f43c0f2e3114a81827f2bc969

Observation 2d1a6791-9ac9-4d11-a4a2-c6189f7f9025 · outbound

This paper cites Bradley and Milton E.

KVAE: Family of Tokenizers for Multimodal Generative Models Bradley and Milton E

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.929922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.929922Z digest=sha256:97426a920836edd74fa4808a8e8ccada42c6b4dc0c2273178ad4061788080c86

Observation 1dafd9e2-d524-463c-8765-c667503a3beb · outbound

This paper cites Efros, and Tero Karras.

KVAE: Family of Tokenizers for Multimodal Generative Models Efros, and Tero Karras

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.934742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.934742Z digest=sha256:28f62ffcad20eed1e826c7d79884a06c15c74887c63a6bc0ac84789603b3c206

Observation 8f99be9f-2c46-4c89-a8e8-634720a099f3 · outbound

This paper cites Video generation models as world simulators.

KVAE: Family of Tokenizers for Multimodal Generative Models Video generation models as world simulators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.939667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.939667Z digest=sha256:e4e89a9778f632ea6e5f562f893610fa9296a824d05858f5bf69a145cfdc9e52

Observation ac82de2f-508f-42f8-94f2-b3442585c127 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Deep compression autoencoder for efficient high-resolution diffusion models, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.944198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.944198Z digest=sha256:1a2ef27e5e0bfe4ceabc9fef79fe986db47f8b2262b5b1e8435ef53da327a095

Observation a01b3ac9-6c80-4101-a447-d258f2d05261 · outbound

This paper cites Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.948798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.948798Z digest=sha256:4cb577e1ec974b9d29e27396a7ad45ae5bef155c262d7fd4294efc43a17ee77c

Observation 66f14d0a-fbab-440d-9a23-60e4be4f9113 · outbound

This paper cites Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.953175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.953175Z digest=sha256:b85a858234f8ca00a9645fa61e93c8804e8057ab2d1cda3686c569c5d2ab8563

Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.957682Z digest=sha256:a1c01e9ff15e1242a49ce1d541cf6cd1887f03ce0bb1abd87c7a72945e2260c4

Observation 39300b56-8bb8-49ca-a936-b46c8e8fdf26 · outbound

This paper cites Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana.

KVAE: Family of Tokenizers for Multimodal Generative Models Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.962483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.962483Z digest=sha256:15b3eb65928522dd488401816173436950d3bb811dc476dd22a9ae38a95ac0e6

Observation 3208e4f5-b9a4-4f4e-9113-e86893969d43 · outbound

This paper cites Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V.

KVAE: Family of Tokenizers for Multimodal Generative Models Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.966861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.966861Z digest=sha256:a31d62a38597dbc3bd00c9b55bb52d993a4d013553ae790c23e56ba5c3873b66

Observation 68383886-6f37-4b0b-8661-df2658afe5cf · outbound

This paper cites Adversarial video generation on complex datasets, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models Adversarial video generation on complex datasets, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.971343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.971343Z digest=sha256:17d04f4c08cbfdb71eb65071d497a9e95a873376b6e27e3586c60e9136933c6f

Observation 81dcc4f3-3506-4684-ad39-9ab84632b47e · outbound

This paper cites High Fidelity Neural Audio Compression.

KVAE: Family of Tokenizers for Multimodal Generative Models High Fidelity Neural Audio Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.975675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.975675Z digest=sha256:f9ec1821f53fd2579be86dc2773c935ab87cd489d2ed7201bc67f3cbba329eba

Observation 79456876-3f1f-4d7b-b238-e098f9074619 · outbound

This paper cites Irc-gan: Introspective recurrent convolutional gan for text-to-video generation.

KVAE: Family of Tokenizers for Multimodal Generative Models Irc-gan: Introspective recurrent convolutional gan for text-to-video generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.980726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.980726Z digest=sha256:55811a886e68961a70c03d2969cc8dbb52389d9afd1110c8a1faa40884cbf89c

Observation 90a9e832-bdce-496d-afdd-d01c2ade3351 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

KVAE: Family of Tokenizers for Multimodal Generative Models Taming transformers for high-resolution image synthesis, 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.985253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.985253Z digest=sha256:3ab543c32bc9f63cf87a35035bf7a3395ad8fe43dfba289345ff7c9edff6a5e5

Observation 4502afe0-0465-4f0f-9d01-890af4e52d42 · outbound

This paper cites Stable Audio Open.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable Audio Open

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.989694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.989694Z digest=sha256:3fbddd501b5eb48a172fbdec781ce4041ca96d1512054fbb5bed82381b1dc63d

Observation 5b1c7234-92ae-487b-b756-b6f47c997c35 · outbound

This paper cites Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.994502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.994502Z digest=sha256:b039cbc0ece450c1fbe787ca4422209c830fbdd077a292dfb59c33ceed819b5b

Observation 42f007fe-896c-4768-a598-b7200889aeab · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.998841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.998841Z digest=sha256:c7cdffefb4baa937157ff875165a46236ca08aaa0fc3ce195b349f89f475e496

Observation 9afbd4dd-4573-449e-b4c3-10ea5566898c · outbound

This paper cites Gemmeke, Daniel P.

KVAE: Family of Tokenizers for Multimodal Generative Models Gemmeke, Daniel P

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.003333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.003333Z digest=sha256:271d74fe218b43473d52d3606a6f899b38a840c3b473be367f7cd252515d692a

Observation 6fe1651f-184c-4137-a0e1-f14a18e6255a · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

KVAE: Family of Tokenizers for Multimodal Generative Models BigVGAN: A universal neural vocoder with large-scale training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.007973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.007973Z digest=sha256:45056c894c22f6c215a2d0f42c34b85dd07ea16429476d34283f3db230669bfb

Observation cc499c7a-24e3-4967-9a5e-f6aaf1ad7721 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

KVAE: Family of Tokenizers for Multimodal Generative Models Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.012406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.012406Z digest=sha256:435c15dbd41440fd1de8b4b6e490effaeef4d401e08e0e9815574e2a457b1f17

Observation aee2644b-0f0f-44a3-ac52-1f0c0a3d1f37 · outbound

This paper cites Veo 3.1: Our leading video generation model.

KVAE: Family of Tokenizers for Multimodal Generative Models Veo 3.1: Our leading video generation model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.016907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.016907Z digest=sha256:7f16027db416e64c852e4c4c82da0daedad31adb3c730c1ec7f8a43fcd301b12

Observation b3bd5b17-09ff-45ff-9cea-0b4fa494d047 · outbound

This paper cites Ltx-2: Efficient joint audio-visual foundation model, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ltx-2: Efficient joint audio-visual foundation model, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.021526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.021526Z digest=sha256:eb164d7790431ce4d09e5b6ccef4ba7d93ad6ee2dc775ff40fc391827bb454f7

Observation d25e578e-3c43-4d82-b3b6-a4d102f65412 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.026296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.026296Z digest=sha256:a497fb56cb61c697571cf70690951e0f766b604100973b1feff49b7f7e27ec09

Observation 51815ec0-39d4-4233-b438-7d728e498fc2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

KVAE: Family of Tokenizers for Multimodal Generative Models Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.031152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.031152Z digest=sha256:303c8733f4b7f5c16de02cdb3772020a225c428cbec5ff463c97d8df56a72fde

Observation acfa54a4-7194-4056-9435-82f1d3aada79 · outbound

This paper cites Kingma, Ben Poole, Mohammad Norouzi, David J.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma, Ben Poole, Mohammad Norouzi, David J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.035612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.035612Z digest=sha256:2238da90144f61c814997104f01ee08c07fa086490696992b16ee126acb7e0a2

Observation ac1880fa-d3b4-45ba-b19f-e5d33fbb0059 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.040278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.040278Z digest=sha256:018a207389091198b7bcd5a4c1f2b2eaebdf4237eb0fb3736637ac5ed9598127

Observation 74b0e03a-b622-43c5-9b45-0d0faf6572eb · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.045380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.045380Z digest=sha256:045c9d28afd7d6e5183ce63aca4fe62c7098333aab8322f0d0782e4799263816

Observation 388dabe9-9483-4020-a64f-bf40609068c2 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,.

KVAE: Family of Tokenizers for Multimodal Generative Models Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.049876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.049876Z digest=sha256:7ef90f12170d4ccc95e9c71a4f0cb8c36a64455f0e0b476ba8810b81ee3aab14

Observation 62f3ccf0-45f6-46a3-bae0-92d1aad6f2da · outbound

This paper cites Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001.

KVAE: Family of Tokenizers for Multimodal Generative Models Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.055309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.055309Z digest=sha256:e9ca01001964d081446ab9615543f1f9782d89a59058d83e73caa52be2b54f0b

Observation 2f138c9e-758c-405e-b102-950da176a488 · outbound

This paper cites Video pixel networks, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Video pixel networks, 2016

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.059886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.059886Z digest=sha256:266cf0dded7dbdd987447c06535b603db7d371e506221c923e49de37209cb910

Observation 2b8624e0-78ca-499b-bff4-74489982f7a5 · outbound

This paper cites Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,.

KVAE: Family of Tokenizers for Multimodal Generative Models Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.064527Z digest=sha256:a2b9b6c756edae2c83283a612a440a62944d8d2f38a2ee1cca2bf90846d826d5

Observation 232dab4f-1757-46a5-87c7-363c5b97a577 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioCaps: Generating captions for audios in the wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.069235Z digest=sha256:aba13170755348b195c93ecb44f26129bff015cc96879951618b4c813654774d

Observation 9cddac63-115e-4f59-b7dc-a7ff81cc025e · outbound

This paper cites Kingma and Jimmy Ba.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma and Jimmy Ba

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.073665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.073665Z digest=sha256:1b1c083e65d2d69819031a507650eee3b6d5ad284128b0b78a289c6170d2c2b1

Observation ac8ebf5b-0045-4547-8dd7-3bdecc1b4e8b · outbound

This paper cites Auto-encoding variational bayes, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Auto-encoding variational bayes, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.078139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.078139Z digest=sha256:fa740a0801bc5ecf0045f7661b9a7c8092d1d85437f430d1bb5e1a311e1cf773

Observation eed9165a-132b-46a2-9a61-0628ff3086bd · outbound

This paper cites Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out.

KVAE: Family of Tokenizers for Multimodal Generative Models Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.082831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.082831Z digest=sha256:6be139fd4ef5c6713c7e2e9eba23967e9ed674448bdc32ee615edfbdedd35f04

Observation c3bfd2f4-fbfe-402a-bb65-80e0edfc4a3e · outbound

This paper cites Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024.

KVAE: Family of Tokenizers for Multimodal Generative Models Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.087019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.087019Z digest=sha256:444951a3b9f6dae213dd95f070068ed94ae3ca92df557a9be75f4c80a0728b76

Observation 5c12fc1a-6a0c-4df5-9492-f8bdbf88ce97 · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

KVAE: Family of Tokenizers for Multimodal Generative Models Ross, Bryan Seybold, and Lu Jiang

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.091331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.091331Z digest=sha256:ff003bf6488e096be58632fa01938a53c2657c3483d92610a7e731713adbc966

Observation 11284e29-edbb-49d5-a292-19aed98a3d65 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.095870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.095870Z digest=sha256:a05a6bb747fab6258594e65b0b245f1b1273e57e17e983e8073d3fbb172f1852

Observation 6df06d21-2e6f-4c84-9e10-0f5f86539d0c · outbound

This paper cites Plumb- ley.

KVAE: Family of Tokenizers for Multimodal Generative Models Plumb- ley

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.103658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.100167Z digest=sha256:42ac15eac02cfd8120d040d7a3b75a20376205e184edc6a9e88a51c76deea2c9

Observation 0aaa8769-e8f5-470d-a322-8fcd01577bca · outbound

This paper cites Hunyuanvideo: A systematic framework for large video generative models,.

KVAE: Family of Tokenizers for Multimodal Generative Models Hunyuanvideo: A systematic framework for large video generative models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.105028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.105028Z digest=sha256:27df9cbf03be55dc7b512e4902d5c93ae7e6354e16c090e11eb3331a359698c6

Observation 36e12cc6-4dc9-4036-9ac1-ee3691524038 · outbound

This paper cites Kandinsky video tools.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky video tools

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.078140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.109643Z digest=sha256:dac0f89608e8732b691f0b2448d500aeb3cd923f9a97b414ad05b05dfb739f40

Observation 31a7a776-c7af-4f4f-8bc8-1312dd7dcb93 · outbound

This paper cites Efficient training of audio transformers with patchout.

KVAE: Family of Tokenizers for Multimodal Generative Models Efficient training of audio transformers with patchout

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.062027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.114113Z digest=sha256:d26286ce1c8e35dba03416995662393a128084dd44a852d0bfbc07eaa329e886

Observation 127290f6-9b74-49d9-b344-b37a41322fd7 · outbound

This paper cites Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.046580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.118762Z digest=sha256:fa8ac2f9a625838303f6b71b427ababd14175382517aed99a0fb415a387629a6

Observation 998362ab-e4c9-4bf6-a90c-0ca4b683a950 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

KVAE: Family of Tokenizers for Multimodal Generative Models High-fidelity audio compression with improved RVQGAN

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.031097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.123265Z digest=sha256:6c05dbad9000b471ddb8370b7fc7a1a4b0e700bcd3bb2790cbc4fe412178145f

Observation 5848a0cc-c2f6-482d-954d-dfa79c488168 · outbound

This paper cites REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.015069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.127893Z digest=sha256:2cc3a5dbd37e89952a57cc2e802a4b34d44aed8625aeba42dbb39c8c976da223

Observation cc608454-847f-40a0-8459-3c9ccee14f2b · outbound

This paper cites DiffusionBench: On Holistic Evaluation of Diffusion Transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:50.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.132533Z digest=sha256:d71f74b19d57ab93a1ef8185fa33e94805d6e27a02385c179d704978c965f2cb

Observation 83adf514-2156-4f71-a374-ffa46d3c452f · outbound

This paper cites Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,.

KVAE: Family of Tokenizers for Multimodal Generative Models Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.999575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.137542Z digest=sha256:87371d3bdd079f294cb201f43110465f8c70b2b8e166b184d49d1e7936610189

Observation e5363bd6-7328-439c-add9-f9aedcb225ab · outbound

This paper cites Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.984146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.142269Z digest=sha256:7538ccec3bc2573b6df57e6f168bb4116007f71c92623845b5079f877514f08f

Observation 267aeaba-7451-409e-b2f4-67858bde79ba · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.146653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.146653Z digest=sha256:f78478468a686326c47c580e128807eca199947a8d363c3fa2ed824693bc7698

Observation 66ca06a4-e630-4e22-b35d-8cac8283534a · outbound

This paper cites Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability.

KVAE: Family of Tokenizers for Multimodal Generative Models Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.151448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.151448Z digest=sha256:2c516e76a273b0aae1c73979a4c4c3c4712edcdbf8f9c12a076690e80d50c496

Observation acc391df-336e-415e-be70-460d0f658a11 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.969040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.156389Z digest=sha256:4cb545a56426c1291ff9036ed98c6bdd5cfb718c0bfead5a7b6fc5c3e3cd82ac

Observation b4353548-027b-4ca2-a114-f6b0f9e70e21 · outbound

This paper cites The song describer dataset: A corpus of audio captions for music-and- language evaluation.

KVAE: Family of Tokenizers for Multimodal Generative Models The song describer dataset: A corpus of audio captions for music-and- language evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.954476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.160855Z digest=sha256:24f97b296a166ebb8d3fdf42f0f5713d76f0865ef947c002db25b2c887f05fb5

Observation c0d76d52-e54f-48d8-a3be-b4b6bbb0005b · outbound

This paper cites Balasubramanian.

KVAE: Family of Tokenizers for Multimodal Generative Models Balasubramanian

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.939257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.165459Z digest=sha256:2f0db299843ba7841f76591662890f91b644d7cfde77bf64be67983714f9b36d

Observation b51194c7-d6ad-4f0a-8fe1-844428a199cb · outbound

This paper cites Transition matching distillation for fast video generation, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Transition matching distillation for fast video generation, 2026

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.924914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.170003Z digest=sha256:9b231677012feade5bd1088af5ddfbce51d8ba81bacc36537c3e993b884a3170

Observation d1a2b48e-5f39-4930-9d4b-2590079c194c · outbound

This paper cites Blaschko, Albert Ali Salah, and Itir Onal Ertugrul.

KVAE: Family of Tokenizers for Multimodal Generative Models Blaschko, Albert Ali Salah, and Itir Onal Ertugrul

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.174272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.174272Z digest=sha256:8964f9ef80b43b23dc1286377b7a83476c05a183737fe8599ee2afacf473363a

Observation 81b25390-cc78-405e-975f-b87ef68d0aea · outbound

This paper cites Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.910624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.178680Z digest=sha256:04865db29282cb01539731fc427e9b6ffd7b5896fc262bc12f6b119c2fd6788f

Observation 6377f506-b873-48fb-9081-51f4cfa0c397 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

KVAE: Family of Tokenizers for Multimodal Generative Models Cosmos World Foundation Model Platform for Physical AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.183042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.183042Z digest=sha256:e174c08d576b2d5c840658196e00d3ef2a3bbe7e17b09556aa420c5cf78709f2

Observation 800b8ce9-b089-4322-971e-87cb2867cc2f · outbound

This paper cites To create what you tell: Generating videos from captions, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models To create what you tell: Generating videos from captions, 2018

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.895129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.187278Z digest=sha256:853ab278ed2d4abe93bcbc52a41809b63ab63647d69267a78717852242abbc3b

Observation bbe6bb1f-7dfe-4235-8b32-0602637df5de · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

KVAE: Family of Tokenizers for Multimodal Generative Models Librispeech: An ASR corpus based on public domain audio books

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.880432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.192154Z digest=sha256:4e6c6cceb303baabd7e32d173b31068849c1d6b3d6f784791fd2d611454b15a0

Observation be565eb9-cfe7-46b6-bb9b-8be96549aa1e · outbound

This paper cites Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.864864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.196758Z digest=sha256:cf8ba898fb50242717a99150441992bcee0ed68d366a8f9ba145bcb42d7672ce

Observation 8b31fb12-eda1-4f4a-bf0f-82d1d1e71768 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du.

KVAE: Family of Tokenizers for Multimodal Generative Models Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.849477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.201249Z digest=sha256:1e6ab25fb266f7166669c82e0c5101f87c0025cb48a1a59fb8c6ef3fc7157a5d

Observation ca168e10-08e7-4eab-87da-b326e642041e · outbound

This paper cites Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson.

KVAE: Family of Tokenizers for Multimodal Generative Models Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.833535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.205992Z digest=sha256:199e9cf0273ba1d728859e6a7f1eec92538e2a8b689c18728856b0de9ba6de7e

Observation 0ff90aab-604b-4284-bacd-ad48092a822b · outbound

This paper cites Qwen-Audio-VAE Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Audio-VAE Technical Report

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:49.823936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.210558Z digest=sha256:4f5028fce8f1dc68e77a80406250b2bf38be21b72ad49be21738091fd832683a

Observation 396ca9dd-f10f-4d1b-9a33-54a7aca03ca7 · outbound

This paper cites Qwen-Image-VAE-2.0 Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Image-VAE-2.0 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.215146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.215146Z digest=sha256:cb9462964f9fd1426f3b6debf0ec6daeaedce13c9c58ac1a7cd33312fe52f5a7

Observation a1a0de29-3159-47e1-bae3-88f9a00de7e8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

KVAE: Family of Tokenizers for Multimodal Generative Models Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.817684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.219882Z digest=sha256:080840d06f4250faa56b459aa3d6d3358dbfc30b13c953dc81ff5a3c17ea9937

Observation 56dab5a3-9387-4627-a2a5-79b447587466 · outbound

This paper cites MUSDB18-HQ — an uncompressed version of MUSDB18, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models MUSDB18-HQ — an uncompressed version of MUSDB18, 2019

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.802189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.224296Z digest=sha256:36bc638db16cb9ad3bef6b28ee3a0abb9c936623eea96ce1644f9c7f54e40242

Observation 403980d2-0bb7-4e58-8a4e-b51a022855e0 · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

KVAE: Family of Tokenizers for Multimodal Generative Models EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.787304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.228947Z digest=sha256:dcd65a965be4ceab6174946fbe3d182910416a9dd128441a26668f183601d49f

Observation 908a6963-34df-4061-9ae6-a918534a033b · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.772663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.233612Z digest=sha256:f49d7a5598f1ba51eca32983d2529c1807fac7e1f0e60b6898667bf2e7388c9b

Observation df019351-9721-4388-8474-f420028c64c2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.757578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.238165Z digest=sha256:2c74932b46d7e7ab4e2fd86a3335fc7d9c0747afb3aa8aa421fe0f9d99226684

Observation d2ee0a73-a87d-4d66-8420-375818288256 · outbound

This paper cites Runway gen-4: Ai video generation with world consistency.

KVAE: Family of Tokenizers for Multimodal Generative Models Runway gen-4: Ai video generation with world consistency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.742254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.242655Z digest=sha256:2d52628e1608e85c9ffe7db7513dc8867f33864b2fcb29d735131b094b34cc52

Observation 1381ca67-e28f-4a49-b640-88ed3f1cd99e · outbound

This paper cites Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.725842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.246801Z digest=sha256:799003d78b42951392f65a59beca1ae634e431885d274159e4c8f641ff4d245b

Observation f28f6329-029f-4897-a6f0-4573d0b56e2d · outbound

This paper cites Make- a-video: Text-to-video generation without text-video data, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Make- a-video: Text-to-video generation without text-video data, 2022

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.710306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.251220Z digest=sha256:98314ca49f60cd35302bde46e2452ec737522ff89a7c4ef3cb839ea127cf59f6

Observation a188e79e-5a70-4a33-b517-84db242a8fa8 · outbound

This paper cites What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.255842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.255842Z digest=sha256:0dd40229dc7c5eaa43be4042f61d5d71294ed738914545829c49626d64bad7bb

Observation 6cc40914-4d56-4056-b59a-1fdd5dd4ac21 · outbound

This paper cites Improving the diffusability of autoencoders, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Improving the diffusability of autoencoders, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.693904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.260320Z digest=sha256:3894880901057f7e76ebbf60b0988ca892de8295b4096a6327e53e0d6c45dd27

Observation 76dfbbe9-c306-4d70-bad0-2e7965de56f9 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.678729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.264895Z digest=sha256:cc4024ebc10e95b7cbb9b2e02c42679c32c3293091d891ed98112f4611f789af

Observation 8570bb5d-e8b1-4fe9-8313-b7bdf698ea11 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

KVAE: Family of Tokenizers for Multimodal Generative Models Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.663478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.269466Z digest=sha256:9b6fa343db9985d18f2a8c9de5a063dce9fe630bc28edef5701d32bbf442f869

Observation 19383de8-b3f2-4913-8671-bc2eb2dd16c3 · outbound

This paper cites Scenediffuser++: City-scale traffic simulation via a generative world model, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Scenediffuser++: City-scale traffic simulation via a generative world model, 2025

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.647889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.273993Z digest=sha256:bed2bd9e14714f79693c050032ba6f76803176b96c3da5f0c0ff228830e1c61e

Observation 95573368-f86d-4172-9e9e-868bfa40d269 · outbound

This paper cites Z-image: An efficient image generation foundation model with single-stream diffusion transformer,.

KVAE: Family of Tokenizers for Multimodal Generative Models Z-image: An efficient image generation foundation model with single-stream diffusion transformer,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.632963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.278379Z digest=sha256:b9c6d8b08277cedf397a4ea6b06dae95756a2a53db669df4fe09b192e56bce60

Observation 5405c1df-101f-4983-925d-d8fbf8ec6146 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.617480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.282806Z digest=sha256:9281d0f64f3e7ab4190200ec291d32ca7793a242d4f37e7ffab45ca6c4cd7062

Observation 49e126f2-5c89-408a-bd6a-c64e1f68d805 · outbound

This paper cites Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.602974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.287263Z digest=sha256:e8be0cfc79ebc8f586e44ba555f0e5c7b219da8031e47ce7a9453d5f6f129463

Observation 54f8ed22-adc4-477d-aa77-894691e0c336 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

KVAE: Family of Tokenizers for Multimodal Generative Models HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.291826Z digest=sha256:32c3f050b816232ec450e3680b4a29c29a3dfe4c6c16db34af2373e182ffc13f

Observation fa0cbf27-e6fc-431d-af41-a2baebeba92c · outbound

This paper cites Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.587939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.296716Z digest=sha256:b868d6fca508f0abfa0c72178fa86b4a38de3c997d8e66530356d461647cb118

Observation efee10b7-6d58-44ed-b07b-e9ab1e95acaf · outbound

This paper cites Metaxas, and Sergey Tulyakov.

KVAE: Family of Tokenizers for Multimodal Generative Models Metaxas, and Sergey Tulyakov

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.573136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.301476Z digest=sha256:52f6c3c4ef24486d47d9eb7758aa89db04f2eec7f9acbddf549cd9923eb5c07e

Observation 23c4e403-bc4d-43c7-8fee-7ada83688ad1 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

KVAE: Family of Tokenizers for Multimodal Generative Models Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.306053Z digest=sha256:34d1eb99db8b98695fe8dcd265cc2f7109d0462b7539d96e7f10bd28cf65cc0b

Observation efeb5b68-bd37-45d3-a383-dc5adebe7538 · outbound

This paper cites Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.558721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.310789Z digest=sha256:f4f56953a5d4ddff32c70206aff6dce3645a0e3a7809cdb64d69f46e56649178

Observation f8046108-5b8b-4da1-b9d0-9f8cde99f1d4 · outbound

This paper cites Conditional image generation with pixelcnn decoders, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Conditional image generation with pixelcnn decoders, 2016

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.544106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.315228Z digest=sha256:c1a8b393a4170db48d132c5c07b7f396b42ca233133a25962d9347cc369ce67c

Observation 235a4e58-ae68-4a14-8a35-8677b551f010 · outbound

This paper cites Neural discrete representation learning, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models Neural discrete representation learning, 2018

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.529444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.319830Z digest=sha256:d587906f0561bca1248df052ab89f0e529da481f1fa42dfd645d34b40d04b3e9

Observation fcb04103-a5e0-40db-82c6-a454a4cc5187 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

KVAE: Family of Tokenizers for Multimodal Generative Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.514926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.324257Z digest=sha256:4a0585dd2a1e23cb8ee14a9c2d166f293d034777fa1ad51d0efdc2515b788a86

Observation 2283e7e1-6c9b-45bc-a3ee-a12990febe3d · outbound

This paper cites Generating videos with scene dynamics, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating videos with scene dynamics, 2016

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.499693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.328663Z digest=sha256:9b71e7528e374b6113446150a50d7309359578f1bf612b83e4f93d74a81a3b39

Observation 3d01af25-d7ad-4d26-9e55-82118e8e8c08 · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan: Open and advanced large-scale video generative models, 2025

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.484897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.333015Z digest=sha256:e64be520ba5ed0f349484bac4c3401ca18a9607e1a426da1160fb8bb6ee86f9b

Observation 69657ab6-1c35-43ea-8ef2-6a4476db2986 · outbound

This paper cites Wan-2.2 anouncement.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan-2.2 anouncement

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.470707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.337918Z digest=sha256:857d51209c234fb59840f9f45822fffd9c5889b654dc3e22724a56ab3d1e67cc

Observation 76571650-0562-4483-9884-118e27a89523 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.456314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.342282Z digest=sha256:dd311e86f6cf6ced6fe0ce20c81a01271b42e7200c9f05a85e3818244ca7b7a5

Observation 56c741e4-ba26-43b8-a506-dd18820a4bb4 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

KVAE: Family of Tokenizers for Multimodal Generative Models Videomae v2: Scaling video masked autoencoders with dual masking

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.442381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.346598Z digest=sha256:b48ac697a6751a2ccd160ffbf332c77117fae34f76d6f348b613ab87450a0a36

Pith citing papers

No inbound Pith citation observations are available.