Pith. sign in

Paper Citation Record · LEDGER

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 7 inbound Pith citation observations for arXiv:2412.01169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01169 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:41:58.965498Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:59:35.563998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T05:51:25.529624Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80172247-fb91-44be-a825-d9e64dc83ff6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.695892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.695892Z digest=sha256:805c423f8a5e7aa1489323b88d59ba3962562bb9d5887199dc9c48c2e90c33b0

Observation 1ff5d55a-d177-4a51-b74d-a7f14a6feaba · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Soundnet: Learning sound representations from unlabeled video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.869427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.701251Z digest=sha256:c4adc478c490e5b708c0f0976e087a5bc1c0a49e1ffe892fbf5186e1aeb4bac1

Observation b7df825a-f920-4572-ba07-5776ae5e7062 · outbound

This paper cites Audiosetcaps: Enriched audio captioning dataset generation using large audio language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audiosetcaps: Enriched audio captioning dataset generation using large audio language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.852892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.706294Z digest=sha256:8a30cd781e1056cb00328e90568e5bc5d2546e1639c09fce272a4254d3039a99

Observation 8f9c18c3-0993-478b-8dfc-def875b09431 · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows One transformer fits all distributions in multi-modal diffusion at scale

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.838101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.710899Z digest=sha256:f341137cdaa8b4f6c3069ffd6e1f7c2e6c9d4bc7a1481673000efc3e2a1bd63d

Observation d3f5ccb4-c2b6-4989-bf17-da0a20f88049 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Coyo-700m: Image-text pair dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.823540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.715454Z digest=sha256:bb3ae7a70f5f63a92a933cce7104ad122018b6d15994afac6485bc13c4ecd384

Observation d7b5ff61-7cc2-4689-a578-0fe393d9e436 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Conceptual 12M: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.809348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.720109Z digest=sha256:4edb3cc66775c499529decc50edbdde9a883330c9c24fe2d12daa42e78c310d2

Observation e26b7d62-9c3c-4cbb-a0ce-e9f6e42f795f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Vggsound: A large-scale audio-visual dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.794621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.724969Z digest=sha256:cee6b6e57a1fbcc4a8625ef27e7127c36408e2a67e2cff25bf5c175c08911af9

Observation de1471bd-58dc-41b1-a440-e190558e99dd · outbound

This paper cites SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:41:59.233860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.729766Z digest=sha256:34e082be884dec22b354f4af6d1410747d653052940f1e18f84ae9e600439800

Observation 2807c2a2-f8ae-432e-87b1-d68cd01983e2 · outbound

This paper cites Scaling instruction- finetuned language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Scaling instruction- finetuned language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.779644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.734983Z digest=sha256:b1fd4167f113a1c9ed3cefe58af0cf15e35ad862ac67d178140acf28f7003b4f

Observation 8b481c2c-1a4a-4843-b2be-526c7a2da070 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Clap learning audio concepts from natural language supervision

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.763042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.739098Z digest=sha256:6b10fb5844c7b7915306cfe955263aea53baf5a49c607e0d448e6f627701a0af

Observation b79bd482-a7fe-456b-aeed-3a34b3f078b6 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.747315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.743809Z digest=sha256:169ec90bfec2b3a53572989f3f307e26fb9a07ba2eb9e65a882c63333d8d5cbb

Observation f700cbe1-38e2-4afa-a9de-ad53700f03f5 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audio set: An ontology and human-labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.732688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.748161Z digest=sha256:38a06a78ffb0f405252f8fd26be83e6ce96669df2c5691472a5ed189ea555348

Observation 2bf2599f-7c1c-4ea2-bc74-109c2f134fc4 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to- image alignment.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Geneval: An object-focused framework for evaluating text-to- image alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.717237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.752616Z digest=sha256:898a2a9d621e1203583e9a0a1cd3c000a798711580c4d73371ce2eff670f67b4

Observation 334aee7e-57ab-48c5-aeb0-47d334d87071 · outbound

This paper cites Text-to-image-2m dataset, 2024.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Text-to-image-2m dataset, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.702327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.756828Z digest=sha256:fb88950077cff61a8ef37d6aa54e76262a62873d98a4d9f5cf10da26062fbd5d

Observation 320147df-2394-42f3-8bc6-94bde91ee8ad · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.761532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.761532Z digest=sha256:951774990a587af31770eae8e2d736570f7f272f7b5fa4d1fd2978559ebcba08

Observation 0c0fc8d5-a86a-4b4b-8f5a-16973d13a3e7 · outbound

This paper cites Classifier-Free Diffusion Guidance.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Classifier-Free Diffusion Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.765962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.765962Z digest=sha256:ac88244741e9b6854fcee4725a8fe5d56d0e75482524b64975229ed11e7a0bb4

Observation c5f08eda-37a5-43bc-9631-bfcc9ae493da · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Imagen Video: High Definition Video Generation with Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.770710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.770710Z digest=sha256:4f200b0f040d6264ba0981337dd3550d0785b26c3b867310092d02a3466ce86d

Observation b93a4ac1-4156-4ca1-b18e-ea5e3eff11dc · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.775653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.775653Z digest=sha256:b847054baba509903f07df5ac01a9bc06f07d4101d2ca152cc39656696c880d8

Observation 6bb9dc43-5aab-4b3e-924c-1c174ac6f688 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.677967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.780711Z digest=sha256:679da88195aceca174a3f88831909a7811eefb4826ccb40c67bfb7558d07ec4d

Observation 007b8214-6836-49b2-841d-1a3950d4afbd · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.785132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.785132Z digest=sha256:252abffca022410d44413dddd0c2e5c849d2b79a7544ef0546853fbda493194b

Observation a98b045d-7610-48c3-81e9-b5ca875f93a8 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audiocaps: Generating captions for audios in the wild

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.662875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.789874Z digest=sha256:960b7267d5d92b241bf3e30acc7b7b27ffcede780f9737864ad2d5d49d2f927e

Observation f0204d47-2aff-49a4-ac9a-2ec73b286ddc · outbound

This paper cites Understanding diffusion ob- jectives as the elbo with simple data augmentation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Understanding diffusion ob- jectives as the elbo with simple data augmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.647193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.794240Z digest=sha256:9a7b833c3a4112e232f31c6d6604293a061184579493d6a6bdef86890324033f

Observation 248e4d2d-532a-4b7e-9089-9afa20e068cf · outbound

This paper cites Equivariant flow matching.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Equivariant flow matching

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.630418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.798876Z digest=sha256:e074652f2744e32ec38eefc9da34fd1e5bd889c00a70fd44ae74c2251a9cf278

Observation 04286985-c320-4d5c-b573-024e5a7f7941 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows AudioGen: Textually Guided Audio Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.803381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.803381Z digest=sha256:8adaddbb37ffb9478a606e5fe4cd3ddb9002e3ad35728f6a5fedf147b7ce3cdc

Observation 29552902-3cb3-4d10-8f0d-7de3493cd21c · outbound

This paper cites Aesthetics for open source, 2023.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Aesthetics for open source, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.615043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.807905Z digest=sha256:98414a71839f5992985742270be2d15109f355326489b82b30bdfd7cf3ed8a01

Observation 7f48e964-7d82-4ece-a0e3-d45e75b654ae · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b- en, 2023.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Laion coco: 600m synthetic captions from laion2b- en, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.600347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.812102Z digest=sha256:1eed976773b5d0cdd8b7f58e0b8ea26ddcb00ec49bbccd6d4f1ebaca1e84a7ca

Observation 2687f756-f306-4d73-94af-7032742ed0eb · outbound

This paper cites Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.816407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.816407Z digest=sha256:a4ac619bc2562b9e7690efa5e2805c8c95b63ef8ec8f81bd1065da75323a9958

Observation b96a9428-f234-4389-a5d8-93b66b7c0b1e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.820882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.820882Z digest=sha256:81f74a72ee8ec75aac8a3178356b1595b82101805381cd8b7f2789906b2c8e50

Observation 88af7c7b-04fe-4c40-b94e-b7f9cab21d57 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.825565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.825565Z digest=sha256:0d6536607ae0502d04c8a1de869fdff7a902a3fc81df30ddf06c9aabb052e33e

Observation 4b3732aa-6928-41e8-a767-e8845c5c7947 · outbound

This paper cites Microsoft coco: Common objects in context.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.830424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.830424Z digest=sha256:6661a32146b89b447550c33a0ea26aa3e2002d8a3fd1b9c178ab491cd859d2d1

Observation d3ac34e2-b2c3-400e-8a9b-cc29a54fdd43 · outbound

This paper cites Flow Matching for Generative Modeling.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.834733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.834733Z digest=sha256:3872efa8e9155b8df56dbcb1428779df2a83ace53e2d2a0eaadfff4448d36556

Observation c630f42a-ffe3-4939-9baf-4c49d857662a · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.839020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.839020Z digest=sha256:ed3340d8e7e6390b1a1240c8da26c4e55b3d8040b61b76667a19a2c0b23910cd

Observation 7d3eb5f8-5726-4468-9ec8-b59e54959409 · outbound

This paper cites Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.555396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.843452Z digest=sha256:ff545c931f9a622a52500d4e8bce5d26a9eaab7ae5dad5704ea8cfa10430fc26

Observation 3cd2133d-490a-425a-8dac-84f5b8158761 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.847351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.847351Z digest=sha256:97abc96c762f3dee0df68dd59c51d731a4823d12d3f459f1d0483fb1da45f445

Observation 59960bc1-3f46-448b-9ed2-7bb3ee4e58a9 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distri- bution.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Discrete diffusion modeling by estimating the ratios of the data distri- bution

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.541623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.851977Z digest=sha256:0b76003b1579847cee3a10c1301df6322bf8dcd6d20c4d6476a30b7a6b568afb

Observation 92dd91c5-768a-4d92-ad19-03d8a2a8f659 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.527405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.856067Z digest=sha256:45b2cd4b37f683fd321d39d2e619e83386c10b75e4d62db7676a0a690e55b599

Observation 941ddad8-4831-4a07-bcd3-a71eff0aa9f2 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal re- search.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal re- search

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.511707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.860248Z digest=sha256:b9a13442b4f43c43173761670dad8f4ac71ceabd72c75f2aa7b74aae48d23df4

Observation f11ce163-9154-4ef4-8e5a-316b398398cf · outbound

This paper cites Image generated using midjourney ai, 2024.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Image generated using midjourney ai, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.496602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.864869Z digest=sha256:b3438a6bdd952853dd22850389088b7c82c8a05ae39b25729a168cbc5202f82d

Observation 40ef35a9-3d82-445f-b56b-c92d13e71590 · outbound

This paper cites Improved denoising diffusion probabilistic models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Improved denoising diffusion probabilistic models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.869147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.869147Z digest=sha256:de566e2f2d6485b8cb186a3d041ff7c5362c6b024f86711a556eb946e9e3b336

Observation 11c577e3-2467-4e9f-a9e8-90ceeaa0b75b · outbound

This paper cites Dall-e 3, 2023.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Dall-e 3, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.472582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.873640Z digest=sha256:2e2da53214c5a52ec02b287e469837acc0ba275b1b8082ffe2a90ed297b95a6c

Observation a63e716a-71d1-4e97-b0ea-c41f5ba8925f · outbound

This paper cites Scalable diffusion models with transformers.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.877953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.877953Z digest=sha256:cb8c3b97800092a90a71bee36cf7706f1f1165786d91dbbcb3d29df37e875ed9

Observation cfd6255d-caec-473c-bb36-78b765dacad4 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.882504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.882504Z digest=sha256:29945eddc09d6e92647f6a1658015490cbd0beeaa02022b01e0bf0670ff3658e

Observation aec40632-a6ea-4845-b3bf-8712a40323e6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Learning transferable visual models from natural language supervi- sion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.887454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.887454Z digest=sha256:ff1f33640bf6b7d5ccb0a1c87237d589dac100069e0eaa20e0318dbd458908b4

Observation 4bed7f85-98df-4e04-b16c-6d7fdda8e690 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows High-resolution image synthesis with latent diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.437189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.892118Z digest=sha256:19216a29d53e860bc4e94d5c18f19f515d21e3a5341d8248f4026289802d92bb

Observation 92c621be-e835-4bdf-b4c0-88af9e8c2088 · outbound

This paper cites Simple and effective masked diffusion language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Simple and effective masked diffusion language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.422384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.896482Z digest=sha256:aee10b1fbca60e18a256bccecf485027a20a022a7b93476bf4dfc6b99acb5fa0

Observation cf004235-c4a9-4120-8a87-d3f24a81a579 · outbound

This paper cites Any-to-any generation via composable diffu- sion.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Any-to-any generation via composable diffu- sion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.407711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.901017Z digest=sha256:f528c6aebc2aad8196c4ab37b812015cabde05d5392992404cb96b768483c789

Observation 80a4ec9c-b2c6-471a-8788-f41b94149161 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.905354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.905354Z digest=sha256:6501db4675032112d7c01591bc007bb50244d8a26ee87cb2b34a1815dce62cf6

Observation f9f188a4-f49a-4e7a-9162-f0e03a9fa852 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.910066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.910066Z digest=sha256:bd5f45c5a0da7f8920bdf366a890752a69759914c412496ff3f1019ecf2e9fba

Observation 86d9bee7-03d3-4c2d-a90d-114cddad3b8a · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Cider: Consensus-based image description evalu- ation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.393341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.914495Z digest=sha256:d5b040d45b51d0b6f012bb264b0a7daf70acba91d7c96bb9c1c3f9360a3857fa

Observation 1ac6539f-4d1f-40a5-bd7a-5ea408400e3a · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Vector-quantized Image Modeling with Improved VQGAN

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.918819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.918819Z digest=sha256:4697b8066f4e715ffbbd8edf4d3e4402d66dcf9591511da1746de9dcd179bdf5

Observation 0b5426c5-4765-4e1b-b98a-be32add155b4 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows TinyLlama: An Open-Source Small Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.923969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.923969Z digest=sha256:fe117541cd5ae01f3ce9247d786d5f7b49fab61b167acb8bb022e2b4cd4eec99

Observation 7af8c7f2-fa1c-4076-8679-52e23dbc0aaa · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.928534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.928534Z digest=sha256:8b99bae0b5b5abfab5561c3b09def28a2e0e14035c82eb54334613b85b64580d

Observation fa0f9525-67a7-420f-9ef7-1b4819366448 · outbound

This paper cites noise level.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows noise level

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.376610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.933160Z digest=sha256:3a9d64665777a2fae1678c351de438e5da8ee1409ae899a563921f91dfb692ef

Observation a3430458-5820-4f57-8419-6ae832680308 · outbound

This paper cites We train Model 2 for 100k steps and Model 3 for 150k steps.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows We train Model 2 for 100k steps and Model 3 for 150k steps

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.360078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.937880Z digest=sha256:1ea15ee71b9a9101c433529e2a17f8cd236165473ded66616dc43181914538d7

Observation 4aadeaf1-2a5a-4820-9d50-4c249a7a6c8a · outbound

This paper cites The learning rate under- goes a linear warmup in the first 1000 steps and a cosine decay throughout the rest of the training.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows The learning rate under- goes a linear warmup in the first 1000 steps and a cosine decay throughout the rest of the training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.342871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.942469Z digest=sha256:1825986b061356b5044173ced9037d57748f0af741742d257a55e61c9a4bcf3a

Observation 101711de-1266-4bde-ac93-14468eade4ea · outbound

This paper cites an unresolved cited work.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:41:59.326555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.947305Z digest=sha256:67ba0bdef2db4d6da7428ba97e5d69e6c46575f1a2263199ae77a7777f8ff6a3

Observation 9c1ca971-3aae-47e7-9d6a-f6202be73e2e · outbound

This paper cites Instead, we sample p(x0 1, x1 3|x0.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Instead, we sample p(x0 1, x1 3|x0

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.311749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.951780Z digest=sha256:dae8583276c9415e75317e3476daf2c568d402355bbbf42a0aa1e7e5aae7caab

Observation 3013141e-f663-4032-a4db-591ae65de4bc · outbound

This paper cites an unresolved cited work.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:41:59.295320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.956113Z digest=sha256:e5a726a1f5c9916f1488c4239b02a583b529892691e27b007e17a23b1d5d26c4

Observation 09941fc5-c0b0-487e-ae04-fccce7b390b7 · outbound

This paper cites an unresolved cited work.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:41:59.279742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.960838Z digest=sha256:85fd92f2474dfc6596517d69ef99aed22c498d8a202786ae3aa527629757fb3c

Observation eac4a76c-4909-4193-a3ca-2a2102097928 · outbound

This paper cites car”, “bird.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows car”, “bird

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.264520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T04:41:58.965498Z digest=sha256:12e586a17c334d35716600c66c9aad9f9f07867076537ecf436eebfee269d70e

Pith citing papers

Observation 52608195-9992-48b1-ab74-4c2502fe686c · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:35.563998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:35.563998Z digest=sha256:438d6f2a243d981f82428dd3bb8361c331d1a48ad9bc33d3661475703f55e122

Observation 43be3b75-d5fa-411d-a840-991ebc006b4b · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.481480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.481480Z digest=sha256:32de2d3d717f2b541e9cbcac5a9b8d9c67e8c6aff4735b8d3d337d3baaf5f5af

Observation 08756ba8-eb7c-4020-943c-5cc4f05a57f7 · inbound

Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces cites this paper.

Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:51.345354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:34:51.345354Z digest=sha256:f83f39e1b946b0c5ed18deea3ca448af5aca056b2821af5c9313e9cb14dd4a0b

Observation 9ebe2114-334c-42b0-ac95-61f35730573a · inbound

Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling cites this paper.

Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:23.676792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:23.676792Z digest=sha256:9e6108b3000ceb280a8e2b37c827c70318b3fd2e8027f9472d5b703c826b772d

Observation e5cec857-e94b-4962-9b6d-4236679b470c · inbound

Flow Straight and Fast in Hilbert Space: Functional Rectified Flow cites this paper.

Flow Straight and Fast in Hilbert Space: Functional Rectified Flow OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T18:00:37.408957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:00:37.408957Z digest=sha256:bdb3085922fc8a4de1b9adea5d6a0f34aa3e632d18bfcc184e41c2404d805b12

Observation 4ced0174-0da9-4e56-bd04-66c6fd7506c5 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:37.010618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:37.010618Z digest=sha256:159c6c3634c84aad88cc4d01d1c8c7ce5651dc75f121cdde022d47fabe457650

Observation 2bcaa1ba-a0f9-462b-8831-b78185c9ddb6 · inbound

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study cites this paper.

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:25.532529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T04:53:06.362430Z digest=sha256:56a57d8b9cd84f29f12892369ab18ab9f7047c86af4ac93977e677be163d33c4