Pith. sign in

Paper Citation Record · LEDGER

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

As of 15 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 7 inbound Pith citation observations for arXiv:2412.01169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01169 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:41:58.965498Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:59:35.563998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T05:51:25.529624Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80172247-fb91-44be-a825-d9e64dc83ff6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.695892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.695892Z digest=sha256:29ccd4e9800a88bc92838ae670c497fa9528e2f13f6fcb1306c2c541e3ed0078

Observation 1ff5d55a-d177-4a51-b74d-a7f14a6feaba · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Soundnet: Learning sound representations from unlabeled video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.869427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.701251Z digest=sha256:aa0a033e0b1c53263f09be3bbcd125e77902d37547b9f365424ec3e5c1200863

Observation b7df825a-f920-4572-ba07-5776ae5e7062 · outbound

This paper cites Audiosetcaps: Enriched audio captioning dataset generation using large audio language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audiosetcaps: Enriched audio captioning dataset generation using large audio language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.852892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.706294Z digest=sha256:b54fe53e7958f67b8e9f24caf77d2b33e8e9d0faeeefbc264f21770791983937

Observation 8f9c18c3-0993-478b-8dfc-def875b09431 · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows One transformer fits all distributions in multi-modal diffusion at scale

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.838101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.710899Z digest=sha256:3f1d31c624c5850d3305ef953d498ef86f72ca250fa856b8210584176bb7ab54

Observation d3f5ccb4-c2b6-4989-bf17-da0a20f88049 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Coyo-700m: Image-text pair dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.823540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.715454Z digest=sha256:6b39da6d8b6d2a81a18948b1195fbb81554c42890e9602bdbfe6b535c99a9b3a

Observation d7b5ff61-7cc2-4689-a578-0fe393d9e436 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Conceptual 12M: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.809348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.720109Z digest=sha256:324be66086c332bf8a6602afcc1dc434dfe227e47c08057c6657bf188b05f738

Observation e26b7d62-9c3c-4cbb-a0ce-e9f6e42f795f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Vggsound: A large-scale audio-visual dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.794621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.724969Z digest=sha256:f5ddbc37f7a93145b518882996ce00d9c4ccfac079739039452cfed9406192ee

Observation de1471bd-58dc-41b1-a440-e190558e99dd · outbound

This paper cites SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:41:59.233860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.729766Z digest=sha256:e652302a4449d67db675030822df9db0f7769cfd7cc366a385f0df9d0813ec98

Observation 2807c2a2-f8ae-432e-87b1-d68cd01983e2 · outbound

This paper cites Scaling instruction- finetuned language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Scaling instruction- finetuned language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.779644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.734983Z digest=sha256:506dc195337e014ca2a8bb6799d6b4b7cdb36c4a4f8f9ed617325f8093093f5e

Observation 8b481c2c-1a4a-4843-b2be-526c7a2da070 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Clap learning audio concepts from natural language supervision

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.763042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.739098Z digest=sha256:257a75d58015facfabade50f7c48f6327dedbeefbc147d3baaf054aa2e132969

Observation b79bd482-a7fe-456b-aeed-3a34b3f078b6 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.747315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.743809Z digest=sha256:ef509e8e0a8a7599e08a389e3743c3742591c0cb1ce4c1d589aede408826459b

Observation f700cbe1-38e2-4afa-a9de-ad53700f03f5 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audio set: An ontology and human-labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.732688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.748161Z digest=sha256:433ce56b594ed5cc646b83d54b459dc7bcda638b8a54bd9b266874dbb38b91f5

Observation 2bf2599f-7c1c-4ea2-bc74-109c2f134fc4 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to- image alignment.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Geneval: An object-focused framework for evaluating text-to- image alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.717237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.752616Z digest=sha256:5628fba7794619e8bc4810e313e3a16821ade2f010334516452c2c2d724f7620

Observation 334aee7e-57ab-48c5-aeb0-47d334d87071 · outbound

This paper cites Text-to-image-2m dataset, 2024.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Text-to-image-2m dataset, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.702327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.756828Z digest=sha256:d5aa68e7dcf565564db1d657d99ce8c57d05c0db7ea0c1acb70c0ccd40a6627b

Observation 320147df-2394-42f3-8bc6-94bde91ee8ad · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.761532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.761532Z digest=sha256:951774990a587af31770eae8e2d736570f7f272f7b5fa4d1fd2978559ebcba08

Observation 0c0fc8d5-a86a-4b4b-8f5a-16973d13a3e7 · outbound

This paper cites Classifier-Free Diffusion Guidance.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Classifier-Free Diffusion Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.765962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.765962Z digest=sha256:ac88244741e9b6854fcee4725a8fe5d56d0e75482524b64975229ed11e7a0bb4

Observation c5f08eda-37a5-43bc-9631-bfcc9ae493da · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Imagen Video: High Definition Video Generation with Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.770710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.770710Z digest=sha256:4f200b0f040d6264ba0981337dd3550d0785b26c3b867310092d02a3466ce86d

Observation b93a4ac1-4156-4ca1-b18e-ea5e3eff11dc · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.775653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.775653Z digest=sha256:b847054baba509903f07df5ac01a9bc06f07d4101d2ca152cc39656696c880d8

Observation 6bb9dc43-5aab-4b3e-924c-1c174ac6f688 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.677967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.780711Z digest=sha256:e2c1a604f3fdb6c1b1d5a5e801c649cc9fcf1d50202d879397f55c221dbe0919

Observation 007b8214-6836-49b2-841d-1a3950d4afbd · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.785132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.785132Z digest=sha256:252abffca022410d44413dddd0c2e5c849d2b79a7544ef0546853fbda493194b

Observation a98b045d-7610-48c3-81e9-b5ca875f93a8 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audiocaps: Generating captions for audios in the wild

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.662875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.789874Z digest=sha256:dae80f245c20a6fa13734485836a5f024d370c49939d64275d2490ccbdf09379

Observation f0204d47-2aff-49a4-ac9a-2ec73b286ddc · outbound

This paper cites Understanding diffusion ob- jectives as the elbo with simple data augmentation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Understanding diffusion ob- jectives as the elbo with simple data augmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.647193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.794240Z digest=sha256:3d7def0bc13028d680f340d0c6d2def0773c2c9bdb116b08a42fb4a75502ff86

Observation 248e4d2d-532a-4b7e-9089-9afa20e068cf · outbound

This paper cites Equivariant flow matching.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Equivariant flow matching

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.630418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.798876Z digest=sha256:25f4ad7903005ad489c293f3322b5d6de5d014e825bd9c99629fe32a1aa83e12

Observation 04286985-c320-4d5c-b573-024e5a7f7941 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows AudioGen: Textually Guided Audio Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.803381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.803381Z digest=sha256:aca258d8d6b297175e5ec8e8e0dbf33457b90778f35b570bcd964002dd18b67f

Observation 29552902-3cb3-4d10-8f0d-7de3493cd21c · outbound

This paper cites Aesthetics for open source, 2023.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Aesthetics for open source, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.615043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.807905Z digest=sha256:a8a4569c91a311829d588cbfc462218e9c644d05b581a58c5db97582abf8c02d

Observation 7f48e964-7d82-4ece-a0e3-d45e75b654ae · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b- en, 2023.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Laion coco: 600m synthetic captions from laion2b- en, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.600347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.812102Z digest=sha256:6463e4bc47288a1200341676d6eb1fedf22ada11c9a5f6baff8f6303dba1c6b6

Observation 2687f756-f306-4d73-94af-7032742ed0eb · outbound

This paper cites Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.816407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.816407Z digest=sha256:a4ac619bc2562b9e7690efa5e2805c8c95b63ef8ec8f81bd1065da75323a9958

Observation b96a9428-f234-4389-a5d8-93b66b7c0b1e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.820882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.820882Z digest=sha256:81f74a72ee8ec75aac8a3178356b1595b82101805381cd8b7f2789906b2c8e50

Observation 88af7c7b-04fe-4c40-b94e-b7f9cab21d57 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.825565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.825565Z digest=sha256:0d6536607ae0502d04c8a1de869fdff7a902a3fc81df30ddf06c9aabb052e33e

Observation 4b3732aa-6928-41e8-a767-e8845c5c7947 · outbound

This paper cites Microsoft coco: Common objects in context.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.830424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.830424Z digest=sha256:6661a32146b89b447550c33a0ea26aa3e2002d8a3fd1b9c178ab491cd859d2d1

Observation d3ac34e2-b2c3-400e-8a9b-cc29a54fdd43 · outbound

This paper cites Flow Matching for Generative Modeling.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.834733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.834733Z digest=sha256:5e4d8151109edf8ac2a556b5b93ee9b5037064861404450fabe087a552a169b8

Observation c630f42a-ffe3-4939-9baf-4c49d857662a · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.839020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.839020Z digest=sha256:ed3340d8e7e6390b1a1240c8da26c4e55b3d8040b61b76667a19a2c0b23910cd

Observation 7d3eb5f8-5726-4468-9ec8-b59e54959409 · outbound

This paper cites Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.555396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.843452Z digest=sha256:d36cb5948e3629adba0e7a40d45363b04a0b0bafe12d59842a26c0908d0142af

Observation 3cd2133d-490a-425a-8dac-84f5b8158761 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.847351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.847351Z digest=sha256:97abc96c762f3dee0df68dd59c51d731a4823d12d3f459f1d0483fb1da45f445

Observation 59960bc1-3f46-448b-9ed2-7bb3ee4e58a9 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distri- bution.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Discrete diffusion modeling by estimating the ratios of the data distri- bution

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.541623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.851977Z digest=sha256:e9f583d8d3fc464cdf8b2fee9ee619102885c1af11267aa2171963905cacf3c1

Observation 92dd91c5-768a-4d92-ad19-03d8a2a8f659 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.527405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.856067Z digest=sha256:86a78005c63edd2b3a5c2c3eacee1585482e4608ae86a1d0d7ffd198dbf6b164

Observation 941ddad8-4831-4a07-bcd3-a71eff0aa9f2 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal re- search.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal re- search

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.511707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.860248Z digest=sha256:d49dcc1b76f33672c3390812d4b348bceec56d6ed56ea6ffd42f70075fd56c77

Observation f11ce163-9154-4ef4-8e5a-316b398398cf · outbound

This paper cites Image generated using midjourney ai, 2024.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Image generated using midjourney ai, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.496602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.864869Z digest=sha256:1dfd45834fde302450b3bf527797aaf7457afcb2af4a45900f8b3500c9b93abf

Observation 40ef35a9-3d82-445f-b56b-c92d13e71590 · outbound

This paper cites Improved denoising diffusion probabilistic models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Improved denoising diffusion probabilistic models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.869147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.869147Z digest=sha256:de566e2f2d6485b8cb186a3d041ff7c5362c6b024f86711a556eb946e9e3b336

Observation 11c577e3-2467-4e9f-a9e8-90ceeaa0b75b · outbound

This paper cites Dall-e 3, 2023.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Dall-e 3, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.472582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.873640Z digest=sha256:7e580155a1d80081a277bc7323ce11dca40ca4accc6dab9f91c4835aa93a1af9

Observation a63e716a-71d1-4e97-b0ea-c41f5ba8925f · outbound

This paper cites Scalable diffusion models with transformers.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.877953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.877953Z digest=sha256:cb8c3b97800092a90a71bee36cf7706f1f1165786d91dbbcb3d29df37e875ed9

Observation cfd6255d-caec-473c-bb36-78b765dacad4 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.882504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.882504Z digest=sha256:29945eddc09d6e92647f6a1658015490cbd0beeaa02022b01e0bf0670ff3658e

Observation aec40632-a6ea-4845-b3bf-8712a40323e6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Learning transferable visual models from natural language supervi- sion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.887454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.887454Z digest=sha256:ff1f33640bf6b7d5ccb0a1c87237d589dac100069e0eaa20e0318dbd458908b4

Observation 4bed7f85-98df-4e04-b16c-6d7fdda8e690 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows High-resolution image synthesis with latent diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.437189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.892118Z digest=sha256:72bcdb99f75a6ae1fab178564a6ea9dd8162474c049635b4ab19252d76201aed

Observation 92c621be-e835-4bdf-b4c0-88af9e8c2088 · outbound

This paper cites Simple and effective masked diffusion language models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Simple and effective masked diffusion language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.422384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.896482Z digest=sha256:176554dffb0625aeeefbdf118b9fc4d1ecb87b87fa08e1f76786a88167f8f433

Observation cf004235-c4a9-4120-8a87-d3f24a81a579 · outbound

This paper cites Any-to-any generation via composable diffu- sion.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Any-to-any generation via composable diffu- sion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.407711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.901017Z digest=sha256:4890b67a82d08aa87b66629234cd161ae26ce0c862033e6c585293e0bc7902d3

Observation 80a4ec9c-b2c6-471a-8788-f41b94149161 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.905354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.905354Z digest=sha256:6501db4675032112d7c01591bc007bb50244d8a26ee87cb2b34a1815dce62cf6

Observation f9f188a4-f49a-4e7a-9162-f0e03a9fa852 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.910066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.910066Z digest=sha256:bd5f45c5a0da7f8920bdf366a890752a69759914c412496ff3f1019ecf2e9fba

Observation 86d9bee7-03d3-4c2d-a90d-114cddad3b8a · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Cider: Consensus-based image description evalu- ation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.393341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.914495Z digest=sha256:efc6f33958b792857fe77e08fd7d157d95d6fa6451252c6d4b2054fe6a320428

Observation 1ac6539f-4d1f-40a5-bd7a-5ea408400e3a · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Vector-quantized Image Modeling with Improved VQGAN

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.918819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.918819Z digest=sha256:4697b8066f4e715ffbbd8edf4d3e4402d66dcf9591511da1746de9dcd179bdf5

Observation 0b5426c5-4765-4e1b-b98a-be32add155b4 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows TinyLlama: An Open-Source Small Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.923969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.923969Z digest=sha256:fe117541cd5ae01f3ce9247d786d5f7b49fab61b167acb8bb022e2b4cd4eec99

Observation 7af8c7f2-fa1c-4076-8679-52e23dbc0aaa · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:41:58.928534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:41:58.928534Z digest=sha256:8b99bae0b5b5abfab5561c3b09def28a2e0e14035c82eb54334613b85b64580d

Observation fa0f9525-67a7-420f-9ef7-1b4819366448 · outbound

This paper cites noise level.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows noise level

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.376610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.933160Z digest=sha256:67d2332b07776007e19352bc265bc8f09f55781dd32a8c34cb5fb06ea2538ca3

Observation a3430458-5820-4f57-8419-6ae832680308 · outbound

This paper cites We train Model 2 for 100k steps and Model 3 for 150k steps.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows We train Model 2 for 100k steps and Model 3 for 150k steps

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.360078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.937880Z digest=sha256:ff8e794a8e87c2e0cb1536df1f6674f5cc392d4a170bb896ffb023b346b3ec51

Observation 4aadeaf1-2a5a-4820-9d50-4c249a7a6c8a · outbound

This paper cites The learning rate under- goes a linear warmup in the first 1000 steps and a cosine decay throughout the rest of the training.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows The learning rate under- goes a linear warmup in the first 1000 steps and a cosine decay throughout the rest of the training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.342871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.942469Z digest=sha256:878a9e14f65315c91a27c76d98fd677276f2e0b0dae3a6a7c3a53ec347310bef

Observation 101711de-1266-4bde-ac93-14468eade4ea · outbound

This paper cites an unresolved cited work.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:41:59.326555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.947305Z digest=sha256:a830b961da2025f2e83d111045cb00cb76130924591909ad73f566cad85a7689

Observation 9c1ca971-3aae-47e7-9d6a-f6202be73e2e · outbound

This paper cites Instead, we sample p(x0 1, x1 3|x0.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Instead, we sample p(x0 1, x1 3|x0

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.311749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.951780Z digest=sha256:dc1bfd192c284791a15c261946cb4fe4b65533bf26b4c1f2c03137ac2b9efbb0

Observation 3013141e-f663-4032-a4db-591ae65de4bc · outbound

This paper cites an unresolved cited work.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:41:59.295320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.956113Z digest=sha256:19ce88b70f615cca45931a0bc4af6229728e4dc8aafec2632af7222e348082b5

Observation 09941fc5-c0b0-487e-ae04-fccce7b390b7 · outbound

This paper cites an unresolved cited work.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:41:59.279742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.960838Z digest=sha256:9cb8c08574b5a3a776e5eb24ae571d99727d841548865c76674913f6c3efe096

Observation eac4a76c-4909-4193-a3ca-2a2102097928 · outbound

This paper cites car”, “bird.

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows car”, “bird

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:41:59.264520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T04:41:58.965498Z digest=sha256:bd71b1d556d9fc7d256a12916b16c6f4d037ac3f74027a92bdaff6694b54bd48

Pith citing papers

Observation 52608195-9992-48b1-ab74-4c2502fe686c · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:35.563998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:35.563998Z digest=sha256:438d6f2a243d981f82428dd3bb8361c331d1a48ad9bc33d3661475703f55e122

Observation 43be3b75-d5fa-411d-a840-991ebc006b4b · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.481480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.481480Z digest=sha256:32de2d3d717f2b541e9cbcac5a9b8d9c67e8c6aff4735b8d3d337d3baaf5f5af

Observation 08756ba8-eb7c-4020-943c-5cc4f05a57f7 · inbound

Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces cites this paper.

Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:51.345354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:34:51.345354Z digest=sha256:fa64a0949dd37ecfc22f281df2ff042e9199750b1e38ec41328594354020ce64

Observation 9ebe2114-334c-42b0-ac95-61f35730573a · inbound

Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling cites this paper.

Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:23.676792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:23.676792Z digest=sha256:73946245d056b745caa2d50c958631b2ec22c5486176fa3278777cb99fa69180

Observation e5cec857-e94b-4962-9b6d-4236679b470c · inbound

Flow Straight and Fast in Hilbert Space: Functional Rectified Flow cites this paper.

Flow Straight and Fast in Hilbert Space: Functional Rectified Flow OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T18:00:37.408957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:00:37.408957Z digest=sha256:bdb3085922fc8a4de1b9adea5d6a0f34aa3e632d18bfcc184e41c2404d805b12

Observation 4ced0174-0da9-4e56-bd04-66c6fd7506c5 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:37.010618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:37.010618Z digest=sha256:159c6c3634c84aad88cc4d01d1c8c7ce5651dc75f121cdde022d47fabe457650

Observation 2bcaa1ba-a0f9-462b-8831-b78185c9ddb6 · inbound

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study cites this paper.

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:25.532529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T04:53:06.362430Z digest=sha256:eebc46b5c08efaa8d90da7a951fe085b4631d6f1a741919ede28e2b8bc8c4fff