Pith. sign in

Paper Citation Record · LEDGER

KVAE: Family of Tokenizers for Multimodal Generative Models

As of 9 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 0 inbound Pith citation observations for arXiv:2608.05798.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05798 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:24:49.346598Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 122 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 738a38b0-545f-4185-a25e-5273ff6c1ee1 · outbound

This paper cites Bitdance: Scaling autoregressive generative models with binary tokens, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Bitdance: Scaling autoregressive generative models with binary tokens, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.884725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.884725Z digest=sha256:98109bac6fb6696036d7b87bae96b12ba614715f7efbcaea9dffe4416975f8fc

Observation 2cf5ec29-295d-4960-8f0d-451f2c896b44 · outbound

This paper cites OmniDoc-TokenBench.

KVAE: Family of Tokenizers for Multimodal Generative Models OmniDoc-TokenBench

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.890347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.890347Z digest=sha256:bc174088ccbb509722bafc620cbf3cdccc980a45496797018324c4a0de4a2478

Observation 58e03af8-44a4-4176-9f42-fb71009e891d · outbound

This paper cites AOM Common Test Conditions v5.0.

KVAE: Family of Tokenizers for Multimodal Generative Models AOM Common Test Conditions v5.0

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.895514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.895514Z digest=sha256:900214b98089ecfbe758a59edf49e2ab975e16ac1b0c93001801ab217876676d

Observation e0c328e6-e129-4c3b-94e2-3bc18372faf5 · outbound

This paper cites Kandinsky 5.0: A family of foundation models for image and video generation,.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky 5.0: A family of foundation models for image and video generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.900423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.900423Z digest=sha256:32fd8699aaadd25aba662cd88931bee293a319759e752af3898c58e481e58bdf

Observation 1c9e3033-ed5a-4ebb-b5c2-9c9bc343525b · outbound

This paper cites Qwen2.5-vl technical report, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen2.5-vl technical report, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.905419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.905419Z digest=sha256:5e66d1bd855997a39e4357578c6e741d1feee7c4c29071129b1ed2e54d07f32a

Observation e7aec70c-cff9-4657-95a8-14e5ec270626 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets,.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable video diffusion: Scaling latent video diffusion models to large datasets,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.915306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.915306Z digest=sha256:53721205e8d929d3ae75094dc4b1fe40f7f8e9c2f87d4130e243eaf872a2fc2e

Observation f4e01222-a462-47f1-8769-f251b1d97829 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Align your latents: High-resolution video synthesis with latent diffusion models, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.919954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.919954Z digest=sha256:e62f5c3d9373ab3e965ffbf7a62d5b180d3650812ad5380b517def277b41bf01

Observation 08a77cb1-62c7-4212-be39-0d505c00db6c · outbound

This paper cites Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A.

KVAE: Family of Tokenizers for Multimodal Generative Models Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.924532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.924532Z digest=sha256:ea2eb2f956154779fc2cd3036c39871c9d84d3f1b408ebe147224e26a63f85ad

Observation 2d1a6791-9ac9-4d11-a4a2-c6189f7f9025 · outbound

This paper cites Bradley and Milton E.

KVAE: Family of Tokenizers for Multimodal Generative Models Bradley and Milton E

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.929922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.929922Z digest=sha256:cc04709231aa48d0007e02ab6facdab5e3d214faa2c127919ad8627cc7f29a34

Observation 1dafd9e2-d524-463c-8765-c667503a3beb · outbound

This paper cites Efros, and Tero Karras.

KVAE: Family of Tokenizers for Multimodal Generative Models Efros, and Tero Karras

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.934742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.934742Z digest=sha256:766869b24259bf880c4cc795db8838bcb0363aa1fe8e5c80ab46c3d25ba26d7a

Observation 8f99be9f-2c46-4c89-a8e8-634720a099f3 · outbound

This paper cites Video generation models as world simulators.

KVAE: Family of Tokenizers for Multimodal Generative Models Video generation models as world simulators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.939667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.939667Z digest=sha256:132a8826804ad80ae585b83672451c385b45ca0e69bc1fee61cdcda7e170609d

Observation ac82de2f-508f-42f8-94f2-b3442585c127 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Deep compression autoencoder for efficient high-resolution diffusion models, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.944198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.944198Z digest=sha256:65dc3553ef4460dd60d99cee3742547e162b730ad0aca34a932c3fdc88666682

Observation a01b3ac9-6c80-4101-a447-d258f2d05261 · outbound

This paper cites Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.948798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.948798Z digest=sha256:792ce74f3afcdb50005592759cd9bf47087885febcc112a4070f466ee3dc2c16

Observation 66f14d0a-fbab-440d-9a23-60e4be4f9113 · outbound

This paper cites Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.953175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.953175Z digest=sha256:9edbd003d5e91bab71f8e632e2ee7736f02923a64b28082a76a53077dfa8b24d

Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.957682Z digest=sha256:84fdb69ba384a4c3fd80fc66d31ac2d904c237b43bc6e47821ed818cf429e2ce

Observation 39300b56-8bb8-49ca-a936-b46c8e8fdf26 · outbound

This paper cites Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana.

KVAE: Family of Tokenizers for Multimodal Generative Models Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.962483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.962483Z digest=sha256:f3bd5dc627367e3477273d72f282137084b6722ee98488c532c242043a4fad78

Observation 3208e4f5-b9a4-4f4e-9113-e86893969d43 · outbound

This paper cites Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V.

KVAE: Family of Tokenizers for Multimodal Generative Models Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.966861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.966861Z digest=sha256:11462a6a79750dfa3603ccccb651aef2d8d643f0cb51e1e93330e45cd944c130

Observation 68383886-6f37-4b0b-8661-df2658afe5cf · outbound

This paper cites Adversarial video generation on complex datasets, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models Adversarial video generation on complex datasets, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.971343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.971343Z digest=sha256:803cba6dee57322463115fe9a5b887128405ff117836430b7a91ca31cff99d5f

Observation 81dcc4f3-3506-4684-ad39-9ab84632b47e · outbound

This paper cites High Fidelity Neural Audio Compression.

KVAE: Family of Tokenizers for Multimodal Generative Models High Fidelity Neural Audio Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.975675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.975675Z digest=sha256:97c355f09c3998d33f50dc7dae7499ced1d01c6c0967235149d944b9f8563692

Observation 79456876-3f1f-4d7b-b238-e098f9074619 · outbound

This paper cites Irc-gan: Introspective recurrent convolutional gan for text-to-video generation.

KVAE: Family of Tokenizers for Multimodal Generative Models Irc-gan: Introspective recurrent convolutional gan for text-to-video generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.980726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.980726Z digest=sha256:30d76479bce9f9becb8da3d8cfe98ce7c813064879b8b0fcaff355b4b4c51a7e

Observation 90a9e832-bdce-496d-afdd-d01c2ade3351 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

KVAE: Family of Tokenizers for Multimodal Generative Models Taming transformers for high-resolution image synthesis, 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.985253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.985253Z digest=sha256:63ab6449ca89b7cdb267683ceca248f60d2cf3a1f1e53cf7a4a85d11b924bafe

Observation 4502afe0-0465-4f0f-9d01-890af4e52d42 · outbound

This paper cites Stable Audio Open.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable Audio Open

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.989694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.989694Z digest=sha256:601c1115bd54effe3867d67a8fb8949ca814083c495086d85fa0cb8a8c4b09f9

Observation 5b1c7234-92ae-487b-b756-b6f47c997c35 · outbound

This paper cites Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.994502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.994502Z digest=sha256:a8b68697b6c8b3a71827b0e8233ca29deba016da40a2eccde81b521f109c71a7

Observation 42f007fe-896c-4768-a598-b7200889aeab · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.998841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.998841Z digest=sha256:dbb18979a01ce3d4ecaedac0fac371213e851411c1f3cc687e8733ab81e083f8

Observation 9afbd4dd-4573-449e-b4c3-10ea5566898c · outbound

This paper cites Gemmeke, Daniel P.

KVAE: Family of Tokenizers for Multimodal Generative Models Gemmeke, Daniel P

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.003333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.003333Z digest=sha256:5f9fa9ce7cc72d0de5c486b44c74083a450642eaee94bbd45db488f7f7c696a9

Observation 6fe1651f-184c-4137-a0e1-f14a18e6255a · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

KVAE: Family of Tokenizers for Multimodal Generative Models BigVGAN: A universal neural vocoder with large-scale training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.007973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.007973Z digest=sha256:0fbcedf75121d0da3374c1155bde23f28b38e091def732bf3dba37d84d2fdc29

Observation cc499c7a-24e3-4967-9a5e-f6aaf1ad7721 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

KVAE: Family of Tokenizers for Multimodal Generative Models Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.012406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.012406Z digest=sha256:2ac3eee268b1c088eb56492285add8bb27921b7fa012f75230c2d30590c19908

Observation aee2644b-0f0f-44a3-ac52-1f0c0a3d1f37 · outbound

This paper cites Veo 3.1: Our leading video generation model.

KVAE: Family of Tokenizers for Multimodal Generative Models Veo 3.1: Our leading video generation model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.016907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.016907Z digest=sha256:850e1c14805ca1080ebcb419672e5b2496ce8f5cd62276b9384df9a5bd802d8e

Observation b3bd5b17-09ff-45ff-9cea-0b4fa494d047 · outbound

This paper cites Ltx-2: Efficient joint audio-visual foundation model, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ltx-2: Efficient joint audio-visual foundation model, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.021526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.021526Z digest=sha256:0a958caf4df0658f99fc59cddc2078019d1ba10abcf226fea04de7c156fe3927

Observation d25e578e-3c43-4d82-b3b6-a4d102f65412 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.026296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.026296Z digest=sha256:1b4386ed5aad649fa3c65a79b31ae2003fdf25281e5d1b3be8b140588a8aa5d4

Observation 51815ec0-39d4-4233-b438-7d728e498fc2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

KVAE: Family of Tokenizers for Multimodal Generative Models Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.031152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.031152Z digest=sha256:4bc4b98c4b2c87ee4c180f6f362f085fbd83fc244aa8c5799e5d38c08f25ed62

Observation acfa54a4-7194-4056-9435-82f1d3aada79 · outbound

This paper cites Kingma, Ben Poole, Mohammad Norouzi, David J.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma, Ben Poole, Mohammad Norouzi, David J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.035612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.035612Z digest=sha256:095755823359fced3a663bcced39e0900a7649be8d7aeb2202f1a8651fbb6bd0

Observation ac1880fa-d3b4-45ba-b19f-e5d33fbb0059 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.040278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.040278Z digest=sha256:671b92c43c8bfe54bc00f08964d003ef1e1e1d4b5505f908f2721c9c0ce6015a

Observation 74b0e03a-b622-43c5-9b45-0d0faf6572eb · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.045380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.045380Z digest=sha256:37de988cffd123040512871c20222858ed404da1ce2433285cd1c619889a686f

Observation 388dabe9-9483-4020-a64f-bf40609068c2 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,.

KVAE: Family of Tokenizers for Multimodal Generative Models Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.049876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.049876Z digest=sha256:b1a92aec2aae6d5805ccbb09fdc986312041354f8e05dce48a100cbc2a0b1fa9

Observation 62f3ccf0-45f6-46a3-bae0-92d1aad6f2da · outbound

This paper cites Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001.

KVAE: Family of Tokenizers for Multimodal Generative Models Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.055309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.055309Z digest=sha256:b5355154039f57c1e92ce6e0b025782de746cda8610d618a3bc7c7cf9d9fd0f9

Observation 2f138c9e-758c-405e-b102-950da176a488 · outbound

This paper cites Video pixel networks, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Video pixel networks, 2016

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.059886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.059886Z digest=sha256:cb6d0dd534d7446e6ae73ed08e39667fbc1507e23068df3a2d29ef302ebe4eda

Observation 2b8624e0-78ca-499b-bff4-74489982f7a5 · outbound

This paper cites Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,.

KVAE: Family of Tokenizers for Multimodal Generative Models Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.064527Z digest=sha256:6173948305965821795d9b392ba6436c843df00326bc6dbae7b66e3c132d5b7c

Observation 232dab4f-1757-46a5-87c7-363c5b97a577 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioCaps: Generating captions for audios in the wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.069235Z digest=sha256:018b5183a2364e4e1264775142bbef1c3ef8642610d0d3f47f6505fb671042c1

Observation 9cddac63-115e-4f59-b7dc-a7ff81cc025e · outbound

This paper cites Kingma and Jimmy Ba.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma and Jimmy Ba

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.073665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.073665Z digest=sha256:1b071df6ff90268dd225bbf724fcdcb7f881e8896ebc7c73e453e5b109d399f1

Observation ac8ebf5b-0045-4547-8dd7-3bdecc1b4e8b · outbound

This paper cites Auto-encoding variational bayes, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Auto-encoding variational bayes, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.078139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.078139Z digest=sha256:ce904054e89c359b5f576e3fdaf7ca3255c30cdb4768139c9afb15f8884b9066

Observation eed9165a-132b-46a2-9a61-0628ff3086bd · outbound

This paper cites Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out.

KVAE: Family of Tokenizers for Multimodal Generative Models Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.082831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.082831Z digest=sha256:fea9d8398a84d8c4807e6c83f3ea5b67dd37ae66ec1fd8a2c12824bce69fde7e

Observation c3bfd2f4-fbfe-402a-bb65-80e0edfc4a3e · outbound

This paper cites Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024.

KVAE: Family of Tokenizers for Multimodal Generative Models Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.087019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.087019Z digest=sha256:8d015b417ae6fcee824e01969570b0a200c14e4fc39869765e4c39333d4d936d

Observation 5c12fc1a-6a0c-4df5-9492-f8bdbf88ce97 · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

KVAE: Family of Tokenizers for Multimodal Generative Models Ross, Bryan Seybold, and Lu Jiang

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.091331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.091331Z digest=sha256:f40f0e0039b9ef634755c898fd36bde5c478eb1b7396c4cf09aecfaa48e3fb38

Observation 11284e29-edbb-49d5-a292-19aed98a3d65 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.095870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.095870Z digest=sha256:77e091d19dbafc469432b1ac7408f29ef4cafd699938ebe042499d2808c06ad5

Observation 6df06d21-2e6f-4c84-9e10-0f5f86539d0c · outbound

This paper cites Plumb- ley.

KVAE: Family of Tokenizers for Multimodal Generative Models Plumb- ley

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.103658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.100167Z digest=sha256:c49b0690b54b692e0d94acb65fce8979c4648fcc47b97615fd281cf1c6bc43e0

Observation 0aaa8769-e8f5-470d-a322-8fcd01577bca · outbound

This paper cites Hunyuanvideo: A systematic framework for large video generative models,.

KVAE: Family of Tokenizers for Multimodal Generative Models Hunyuanvideo: A systematic framework for large video generative models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.105028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.105028Z digest=sha256:e194b51dbe5a5e485b09be54873097754cbdcccc122d29df5e5884e196eb760b

Observation 36e12cc6-4dc9-4036-9ac1-ee3691524038 · outbound

This paper cites Kandinsky video tools.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky video tools

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.078140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.109643Z digest=sha256:b2341b33f0ecfd7f5c212550cd65a45f21464f1f395e2cddde9b04e9886cb6fc

Observation 31a7a776-c7af-4f4f-8bc8-1312dd7dcb93 · outbound

This paper cites Efficient training of audio transformers with patchout.

KVAE: Family of Tokenizers for Multimodal Generative Models Efficient training of audio transformers with patchout

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.062027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.114113Z digest=sha256:b28bb0f5ed3d6ef4394c953134601e8c827211d2b412552b7124b2a8422eeb6d

Observation 127290f6-9b74-49d9-b344-b37a41322fd7 · outbound

This paper cites Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.046580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.118762Z digest=sha256:5374fad5ce38075ed02e9bc62d8d6003dd5e2e352417465313bc9c833ef38dcd

Observation 998362ab-e4c9-4bf6-a90c-0ca4b683a950 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

KVAE: Family of Tokenizers for Multimodal Generative Models High-fidelity audio compression with improved RVQGAN

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.031097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.123265Z digest=sha256:ccf3497a551e4c912b3a9f679c6453a1e8e1391b291d0a6fcc3fbe2a45bfa08b

Observation 5848a0cc-c2f6-482d-954d-dfa79c488168 · outbound

This paper cites REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.015069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.127893Z digest=sha256:01862ee4eab3e26d7af5a645987abae8f9e580bdf2d413d113d43cd902f28cc9

Observation cc608454-847f-40a0-8459-3c9ccee14f2b · outbound

This paper cites DiffusionBench: On Holistic Evaluation of Diffusion Transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:50.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.132533Z digest=sha256:2a9ed2622a69bf35d9d0223e7e97e0ab5c00566565cad01b2985f901af5d2bdd

Observation 83adf514-2156-4f71-a374-ffa46d3c452f · outbound

This paper cites Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,.

KVAE: Family of Tokenizers for Multimodal Generative Models Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.999575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.137542Z digest=sha256:8284305d7c0f9db111bbd09c40c1838fc0f56e3ba3312d8ba63cebb5869c3f6a

Observation e5363bd6-7328-439c-add9-f9aedcb225ab · outbound

This paper cites Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.984146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.142269Z digest=sha256:7e49923cc1285c167a4f9314a9881ad0556fe31a10d13b1e78e0aa10daa1aca9

Observation 267aeaba-7451-409e-b2f4-67858bde79ba · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.146653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.146653Z digest=sha256:53339cd880fdac931c331260a64c88af04ca9320984cf5a97ad8eb0d4f94b4f5

Observation 66ca06a4-e630-4e22-b35d-8cac8283534a · outbound

This paper cites Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability.

KVAE: Family of Tokenizers for Multimodal Generative Models Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.151448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.151448Z digest=sha256:92f8fc96d3d0118cb8cce95d6a75895e085d3f3cb8b0385e241e780a79b3cf07

Observation acc391df-336e-415e-be70-460d0f658a11 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.969040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.156389Z digest=sha256:348dbcf077f5ab56de47d0d73809bbaa5167f2d34745929d37eca9fdd74e8040

Observation b4353548-027b-4ca2-a114-f6b0f9e70e21 · outbound

This paper cites The song describer dataset: A corpus of audio captions for music-and- language evaluation.

KVAE: Family of Tokenizers for Multimodal Generative Models The song describer dataset: A corpus of audio captions for music-and- language evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.954476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.160855Z digest=sha256:68dd51d86307c07b235f501a95cfb5168a1e0da49426e63e1b3b7e028e00e241

Observation c0d76d52-e54f-48d8-a3be-b4b6bbb0005b · outbound

This paper cites Balasubramanian.

KVAE: Family of Tokenizers for Multimodal Generative Models Balasubramanian

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.939257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.165459Z digest=sha256:6a7b7d53c3d587ebd46c3ad4cc618b2a6c76ed22a2a80f9486db2b6d12c9cadc

Observation b51194c7-d6ad-4f0a-8fe1-844428a199cb · outbound

This paper cites Transition matching distillation for fast video generation, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Transition matching distillation for fast video generation, 2026

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.924914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.170003Z digest=sha256:02c4f730e8e0e52a0be44e595c43c579d28c1b79897254e99a8a3ff343ee1e7c

Observation d1a2b48e-5f39-4930-9d4b-2590079c194c · outbound

This paper cites Blaschko, Albert Ali Salah, and Itir Onal Ertugrul.

KVAE: Family of Tokenizers for Multimodal Generative Models Blaschko, Albert Ali Salah, and Itir Onal Ertugrul

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.174272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.174272Z digest=sha256:dfa5701d4f6ffe50fa29dcad051762eec14c927454548f90edb716f2657c07cb

Observation 81b25390-cc78-405e-975f-b87ef68d0aea · outbound

This paper cites Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.910624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.178680Z digest=sha256:22a851d3c0de0528a94f2511f486a02cceb0465b0271e5ae83ddd3b524d310cb

Observation 6377f506-b873-48fb-9081-51f4cfa0c397 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

KVAE: Family of Tokenizers for Multimodal Generative Models Cosmos World Foundation Model Platform for Physical AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.183042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.183042Z digest=sha256:62146e1b21072a9288cb5ce40a069e6f3f6ee9726e66c22c7c99475e26a6aa65

Observation 800b8ce9-b089-4322-971e-87cb2867cc2f · outbound

This paper cites To create what you tell: Generating videos from captions, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models To create what you tell: Generating videos from captions, 2018

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.895129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.187278Z digest=sha256:5ce2c63415447af35e880a1aab3113f07d16f9890a6e7dc1ca7edec506e470fe

Observation bbe6bb1f-7dfe-4235-8b32-0602637df5de · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

KVAE: Family of Tokenizers for Multimodal Generative Models Librispeech: An ASR corpus based on public domain audio books

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.880432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.192154Z digest=sha256:b4a33677b9cc794d5b38c9f876fecbc4de213c509c2917a5cc9bd42cad112d3a

Observation be565eb9-cfe7-46b6-bb9b-8be96549aa1e · outbound

This paper cites Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.864864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.196758Z digest=sha256:79b69806f62d6e530c7f1e1658e3d3cf311f01ce8c42b45b4676700ba0c43cea

Observation 8b31fb12-eda1-4f4a-bf0f-82d1d1e71768 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du.

KVAE: Family of Tokenizers for Multimodal Generative Models Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.849477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.201249Z digest=sha256:58b6289802c0c03cc818e7a06c3ee60eb255451a09dc4d11ddc957dbdb8d044c

Observation ca168e10-08e7-4eab-87da-b326e642041e · outbound

This paper cites Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson.

KVAE: Family of Tokenizers for Multimodal Generative Models Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.833535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.205992Z digest=sha256:1409713435c4e6111edb2e16aaa4749dea07594ff5a41deca46a409a1b1c3ea3

Observation 0ff90aab-604b-4284-bacd-ad48092a822b · outbound

This paper cites Qwen-Audio-VAE Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Audio-VAE Technical Report

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:49.823936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.210558Z digest=sha256:827a4435b6056ff0f4b44d7d0cc566209ad7313c66564eab4f04955da30dfdce

Observation 396ca9dd-f10f-4d1b-9a33-54a7aca03ca7 · outbound

This paper cites Qwen-Image-VAE-2.0 Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Image-VAE-2.0 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.215146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.215146Z digest=sha256:31f0edac71977ce05205eb5d494e53dc197c507d94b13766c3dc3a5cee9e889a

Observation a1a0de29-3159-47e1-bae3-88f9a00de7e8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

KVAE: Family of Tokenizers for Multimodal Generative Models Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.817684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.219882Z digest=sha256:503c37e90d8939e2e54d60bd6b45bb6ce3d1b331b72cd310fd845db362333c20

Observation 56dab5a3-9387-4627-a2a5-79b447587466 · outbound

This paper cites MUSDB18-HQ — an uncompressed version of MUSDB18, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models MUSDB18-HQ — an uncompressed version of MUSDB18, 2019

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.802189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.224296Z digest=sha256:cc32ee813a653f4d3d02037d81c206343ae05bc0e4381c441e7001c514c46046

Observation 403980d2-0bb7-4e58-8a4e-b51a022855e0 · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

KVAE: Family of Tokenizers for Multimodal Generative Models EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.787304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.228947Z digest=sha256:5ff39cc4bd829004b790b4b7a04f9e13969da24d3d1e47948730276594b22e8f

Observation 908a6963-34df-4061-9ae6-a918534a033b · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.772663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.233612Z digest=sha256:260e3cb7ab7140413b5f158f227fcb6f19780b2b90d9e882c44e6b52129ceb47

Observation df019351-9721-4388-8474-f420028c64c2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.757578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.238165Z digest=sha256:d4bffbfe8a9e8579873054e1d5f81a82a81c4fb6aa634005c4ae0c11315cfaaf

Observation d2ee0a73-a87d-4d66-8420-375818288256 · outbound

This paper cites Runway gen-4: Ai video generation with world consistency.

KVAE: Family of Tokenizers for Multimodal Generative Models Runway gen-4: Ai video generation with world consistency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.742254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.242655Z digest=sha256:391c0b6fdbf180986c08373aa0c4b22309d358467597ba95d16128fcba7bf32d

Observation 1381ca67-e28f-4a49-b640-88ed3f1cd99e · outbound

This paper cites Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.725842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.246801Z digest=sha256:ff6bf89be65de8193927a60cac4456b3ea5516a260233eff0cc572f3928b9895

Observation f28f6329-029f-4897-a6f0-4573d0b56e2d · outbound

This paper cites Make- a-video: Text-to-video generation without text-video data, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Make- a-video: Text-to-video generation without text-video data, 2022

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.710306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.251220Z digest=sha256:65202e7cecb062e649fe3c79865ff814de13ff96b8ec00cd50b68e9c52e82d68

Observation a188e79e-5a70-4a33-b517-84db242a8fa8 · outbound

This paper cites What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.255842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.255842Z digest=sha256:7445841ff41db412459404041c9050e6a0ecd6dbbf66085993cc6caa6078bb7a

Observation 6cc40914-4d56-4056-b59a-1fdd5dd4ac21 · outbound

This paper cites Improving the diffusability of autoencoders, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Improving the diffusability of autoencoders, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.693904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.260320Z digest=sha256:5cc8d309efc9438e88f989e0930079163fd79c550db57bbbebd76d8867058a81

Observation 76dfbbe9-c306-4d70-bad0-2e7965de56f9 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.678729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.264895Z digest=sha256:e77286f7645868d20e1dfa40bf52f66c385e590c48931d13db92c77372d93ee7

Observation 8570bb5d-e8b1-4fe9-8313-b7bdf698ea11 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

KVAE: Family of Tokenizers for Multimodal Generative Models Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.663478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.269466Z digest=sha256:64aa03e380e1d3500e6911c384f38f39aec65b931f11bc7724e494ecbd486c74

Observation 19383de8-b3f2-4913-8671-bc2eb2dd16c3 · outbound

This paper cites Scenediffuser++: City-scale traffic simulation via a generative world model, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Scenediffuser++: City-scale traffic simulation via a generative world model, 2025

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.647889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.273993Z digest=sha256:dd44cba5d4b5bf893d15b73472f56fded8218b7abe930e186efc25a57f455bfa

Observation 95573368-f86d-4172-9e9e-868bfa40d269 · outbound

This paper cites Z-image: An efficient image generation foundation model with single-stream diffusion transformer,.

KVAE: Family of Tokenizers for Multimodal Generative Models Z-image: An efficient image generation foundation model with single-stream diffusion transformer,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.632963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.278379Z digest=sha256:849971526dcc69ca0855ad54d3e696ab3a6b2382383e74bef2d25261ccd115b6

Observation 5405c1df-101f-4983-925d-d8fbf8ec6146 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.617480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.282806Z digest=sha256:81c39df0ac65cb66e2ba98c4a779b590679c8e8c38e09bfe77dba1fcd54eedf1

Observation 49e126f2-5c89-408a-bd6a-c64e1f68d805 · outbound

This paper cites Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.602974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.287263Z digest=sha256:00f4a2e1bbc61f640f1c5402db625d43db3251449ea733569a83c026b49c63a4

Observation 54f8ed22-adc4-477d-aa77-894691e0c336 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

KVAE: Family of Tokenizers for Multimodal Generative Models HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.291826Z digest=sha256:85ba675ecc9acf71dfb5710623ccd12a8fcd23eb8a8c2a8d0d7e962f8fc50741

Observation fa0cbf27-e6fc-431d-af41-a2baebeba92c · outbound

This paper cites Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.587939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.296716Z digest=sha256:a0009468e40caa09265de5b9d68b043a61874d259b5bc74705d9d923e656617e

Observation efee10b7-6d58-44ed-b07b-e9ab1e95acaf · outbound

This paper cites Metaxas, and Sergey Tulyakov.

KVAE: Family of Tokenizers for Multimodal Generative Models Metaxas, and Sergey Tulyakov

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.573136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.301476Z digest=sha256:0b6ef92f8356373c62bfe340440a8e2cbbbb690c1008df6adee5d6f69ec9d8bf

Observation 23c4e403-bc4d-43c7-8fee-7ada83688ad1 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

KVAE: Family of Tokenizers for Multimodal Generative Models Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.306053Z digest=sha256:7633d6862f36fbc676bff58d8dde7b5ae7cb998bf3c4ddb3a2635cfcd056d88f

Observation efeb5b68-bd37-45d3-a383-dc5adebe7538 · outbound

This paper cites Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.558721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.310789Z digest=sha256:5f2343a6fafad0cc772396a9f59c5816560585fe2c89cca92d5654898230a0ca

Observation f8046108-5b8b-4da1-b9d0-9f8cde99f1d4 · outbound

This paper cites Conditional image generation with pixelcnn decoders, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Conditional image generation with pixelcnn decoders, 2016

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.544106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.315228Z digest=sha256:50a7f93c866e8676904c9677befbe2eca5d21941444ecad3e1b1168af0792dff

Observation 235a4e58-ae68-4a14-8a35-8677b551f010 · outbound

This paper cites Neural discrete representation learning, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models Neural discrete representation learning, 2018

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.529444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.319830Z digest=sha256:9eb9fcaa0ef5e8e8f4b671510eff8f60eb53129f32f763eb7aa46beb5d399857

Observation fcb04103-a5e0-40db-82c6-a454a4cc5187 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

KVAE: Family of Tokenizers for Multimodal Generative Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.514926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.324257Z digest=sha256:9d49a3608855a9591d57e783fa607365aca1e21835dbca75390ece58e86f8646

Observation 2283e7e1-6c9b-45bc-a3ee-a12990febe3d · outbound

This paper cites Generating videos with scene dynamics, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating videos with scene dynamics, 2016

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.499693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.328663Z digest=sha256:b4c67c50bd4a94b4a279e1a5238965fc43df6858a33768f9dcecbbd9e86cb60c

Observation 3d01af25-d7ad-4d26-9e55-82118e8e8c08 · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan: Open and advanced large-scale video generative models, 2025

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.484897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.333015Z digest=sha256:c8f58c0dfe38d345dcf927675503fb1dbb1b4641e2f45880c0d47fcabaef838b

Observation 69657ab6-1c35-43ea-8ef2-6a4476db2986 · outbound

This paper cites Wan-2.2 anouncement.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan-2.2 anouncement

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.470707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.337918Z digest=sha256:0d6442cd0123fde80350835511e769e53bd51bd00a36e5649e7b8d8e02f0465c

Observation 76571650-0562-4483-9884-118e27a89523 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.456314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.342282Z digest=sha256:51271e473fd76efc93ebb471f938d153d09b0d3650ea1b66bb56bf22d2ef6e23

Observation 56c741e4-ba26-43b8-a506-dd18820a4bb4 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

KVAE: Family of Tokenizers for Multimodal Generative Models Videomae v2: Scaling video masked autoencoders with dual masking

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.442381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.346598Z digest=sha256:721e79a2e512f713e786c07fc3f8c7ba34246e2508b1c3fae6759f296718a770

Pith citing papers

No inbound Pith citation observations are available.