Pith. sign in

Paper Citation Record · LEDGER

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

As of 9 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 26 inbound Pith citation observations for arXiv:2505.23661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23661 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:18.590800Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:37:16.161813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.237442Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6035d8ec-a76e-4ea3-b87f-954d67ad2cc0 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.668718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.668718Z digest=sha256:9a061f4a44798c69062d281ca4f414bb23ec04d76dbb6a5e088ed2ce68daaaaa

Observation d4a14748-5626-46a2-b2f0-d7828f34232f · outbound

This paper cites Visual instruction tuning, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Visual instruction tuning, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.783250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.783250Z digest=sha256:3f18f61426d059108fcbd8d3a0df4e5ecc4f54a35e23d998af42c261885dd08d

Observation 578b26f6-e097-41ca-9f47-f8e52869ffde · outbound

This paper cites Improved baselines with visual instruction tuning.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.894781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.894781Z digest=sha256:653d88499b4529494db8fa1127489df71c90c5d5463ad13818f10f1365215b1d

Observation 7054ed70-b2e0-4bd7-8a9a-eef25feca87a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.987323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.987323Z digest=sha256:4016243565ceb993e4c252c240835d3f6f44ab7dd703c6f77efea91053edd67b

Observation 0d1ff2ac-df0c-4395-b94e-d6748d81253f · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.043657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.043657Z digest=sha256:5b3415e40be5f6c71a04906d834ed957db045900f265f9faa3addff01e4ceaa5

Observation 3bf26c07-2194-4332-bacf-f1af516967b0 · outbound

This paper cites Qwen2.5 Technical Report.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.101067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.101067Z digest=sha256:4893e4001cb4eefd935155b5a43160f64690e639a92432371c40832d7bc068d8

Observation fd256ac3-fec7-4db9-a8f9-e796a66a9bb7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.184621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.184621Z digest=sha256:627ed97e13e5a99e64f6bc5f41559ebb256f59b45dfb7431f149c502019b6758

Observation 48873b2e-ed5e-451d-ab14-b7be49289677 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.275342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.275342Z digest=sha256:e9000e2be49c94377c5a0f6f5a58ea5474736bfce94b84c1f171303bf6cbbd92

Observation a7932a8d-5df6-4975-9d6e-a08bf667421e · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.368913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.368913Z digest=sha256:0c85523cb8d7ebe66896233b2f785cd31893b600e2f62ea723aff6d04998ff12

Observation 3f3ea260-958c-40a7-ba07-2f6de7deef77 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.561728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.561728Z digest=sha256:00df50ea77ac7ac75e6e8ee521663e5d1e60ca010e55ae0b6c147fcc6c238177

Observation cc33ce15-5a5a-40cb-91cd-f8ffe9b554a6 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.780066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.780066Z digest=sha256:ef75f9fd59afcffc2690e913ba1b5ee149f95bddd83bfcaafd14e2974544a97e

Observation 9ee0b66b-1621-422d-baf0-e6c2443cb82b · outbound

This paper cites Zero-shot text-to-image generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Zero-shot text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.915135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.915135Z digest=sha256:49bc9418f55e2b6d35f6541266d1575d9eeb34f08d6b26b15a43304c5cbdaa77

Observation 4ea54124-5ae1-4e4f-a7df-a3851213fbfe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.029102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.029102Z digest=sha256:5e0d9754e14e41d0c883b8e3975db525c621dfc0da113ad51a7316c298a61c65

Observation 8f5966fa-82c8-4c95-b511-30fee3c82857 · outbound

This paper cites Improving image generation with better captions.Computer Science.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improving image generation with better captions.Computer Science

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.132584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.132584Z digest=sha256:b7078de677fa205c1a3a0f035c7ed6afd406203ade135839cab77ff3a6b05bf0

Observation 2ed071f9-acf5-477a-852a-ccd165990bc6 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.230577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.230577Z digest=sha256:276bbbbba55623bec62d0c304598a343d078d20ec258eae26649b869522b140c

Observation c87388e8-1a6a-44c0-a5e4-6a5d94ef3bc6 · outbound

This paper cites GPT-4o System Card.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.337586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.337586Z digest=sha256:470bb6b35b5a132008e02fb71f174ec3cabe5b93062ebf4f1819bcc13c4c2e04

Observation 34ab3c70-925c-4147-91a7-c49f487e21fa · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.494263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.494263Z digest=sha256:d326cead45c15fb44d622ba35f1acecf6e4d9db553c752c9da5b5937b6991532

Observation fc3aa865-95bc-4299-a5b2-ae5b1f852c80 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.606841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.606841Z digest=sha256:587ae24256a092112aa0e503be68efbb6e5d863baa59bfa224a06f83816a1178

Observation a4ff7499-3701-4273-bee7-aaf6c91ab774 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.742428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.742428Z digest=sha256:344166a1333261bcaa2675477e41300c512261b9d67ee26ef24025ba48254ed9

Observation feee7fbc-8865-4c9e-a27c-61780c8fae68 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.810243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.810243Z digest=sha256:72b6914b931fffc60299c0d97e587a03d76eef0efa4e9fcf604a58a6b9aabdbe

Observation e20cbcff-517f-4cda-a019-6c68bd4f3b08 · outbound

This paper cites Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.891013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.891013Z digest=sha256:602be13647dca4630dd4c48ca3f00ae8575808746f23e0dbcb40e2a2e1139619

Observation e06dda0f-75f6-4bd1-8c8d-770a35ed0620 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.957297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.957297Z digest=sha256:1be271418b35e5db6110a8fe25630b004b3de80719e807161765f23a9770508e

Observation c7590a71-3091-42fb-ac32-7b3bb6302515 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.039181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.039181Z digest=sha256:ae9921a379ba7ec5290f4a8d1c97c93c7808637c0916a5386f49abbf4de3bbe0

Observation e9dea184-515f-4a16-8c6f-2a99ab13ca1e · outbound

This paper cites Generative multimodal models are in-context learners.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Generative multimodal models are in-context learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.110694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.110694Z digest=sha256:829bd56c70e17a41c4c8b3571e36845b811b038bd5cab7b51d1b6dd760db946d

Observation 5c9ce947-b55b-4545-be2e-997b370492f3 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.148096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.148096Z digest=sha256:9d79bea78121bc0108ee556dbc5fec061daa81b20e9685fcff20f505ae5fce54

Observation 189049c2-f882-4d64-b9bb-2f825baa2e5f · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.179360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.179360Z digest=sha256:49684d75162cff202d91eee1247b3288acf94bf9f29d1e88bc1ee1ca6521cd12

Observation 10b76f1f-2460-40a6-a315-cdc7ee7d006d · outbound

This paper cites Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.204122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.204122Z digest=sha256:3e974f01ab8f1c09c266854b38a89f84949353abc10746cc0bed6ab751e04999

Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.245982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.245982Z digest=sha256:6035b673524d75d94bc59b7d1965bdc18286f0089851901e74c3cb670d5c6972

Observation 7fcac129-639a-4db5-b51a-66727b97a7e1 · outbound

This paper cites Transfer between modalities with metaqueries, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Transfer between modalities with metaqueries, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.671731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:15.331127Z digest=sha256:6633064c4b56433da5e1bbe7b396402686f5f99898124b7e97abbfb220f70a39

Observation 1c947697-6a78-4fde-b22c-7779b8e343b1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.407565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.407565Z digest=sha256:68ec65c7fc2229aa231a67dd24d319931e36cf861626f763366aa52993147c50

Observation 41ef8a11-9e40-44c8-b07d-8ea86291d395 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Learning transferable visual models from natural language supervision, 2021

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.481465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.481465Z digest=sha256:1e4acbadfabf8106906c320d12b9afefef649a1f61f9ee2d7b4581fccb886c16

Observation d49b3c90-0a92-413f-a3f3-805b76b068b4 · outbound

This paper cites Sigmoid loss for language image pre- training.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre- training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.558349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.558349Z digest=sha256:cb9d4a01195b83c0ba4ab5501bc5437a6d35b43b3f3ef3f4689859a0d48f8d29

Observation eaefd3ea-0f7b-4612-aeaa-352193c3059c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.611853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.611853Z digest=sha256:0b12572e7784d481a5d3889706ab3080b22dd3142280c5523616a60fe2534526

Observation 5573576b-e62c-4c13-8f53-a4148241707c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.658112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.658112Z digest=sha256:500d7127d44a2e67c3aaa96f9e6f1a66f7d4e2d4652bceadf0836bb87bcedd5a

Observation 76af517e-7ec2-40fd-977e-99a915ce80c3 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.495287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:15.707332Z digest=sha256:7a978a0808117edf16bcd249bff826fc439768e087ba9fa48678dc53056cfc53

Observation 92061989-3c66-4e66-9699-a0e2baa14917 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.737384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.737384Z digest=sha256:923ae3588e6bcdac5adeeb62bbf19046c0ca9d33bd3418801586cf3aa109f10f

Observation 50d89353-e2df-46d1-9085-6e6aed608311 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.790154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.790154Z digest=sha256:8eaf87c1280a210ff068d38a8fda6b4f67cbc6e2386fd145744fb9e9cb9ad335

Observation 39b5e017-0143-4b75-9849-cd0125d4f9fd · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.884310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.884310Z digest=sha256:2f81ab5a471187aa0b8a5bb550948d0a48b356b6ec84b3b2ac679c30ced49780

Observation 07610958-a4b0-4e79-be63-b8bf11b1e917 · outbound

This paper cites Scalable diffusion models with transformers.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.937556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.937556Z digest=sha256:5b29da7275b5f87eeeb2624ea26f55738de3df68e2a7479076cff5b5cf490079

Observation 2077f18a-effd-48eb-be04-092b0ee7208c · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.037523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.037523Z digest=sha256:0da258e67583cce10add9606f9cc691017ed6681f3dcbc14a98821c651f45da8

Observation ec1f466d-522f-403c-9f96-6012750e4d9b · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.103086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.103086Z digest=sha256:754405b0d15d4f400932c75f117ff0683817e89740e31f748b1cddb4eaad1037

Observation 9ab9c62b-e79b-4b6b-97f6-a13c8b0f88f7 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation U-net: Convolutional networks for biomedical image segmentation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.284865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.284865Z digest=sha256:af09fcf7791b7a5202c2617d048c7a4b1b0c227454db0c63419f30db03ece8ff

Observation c60d4210-7d90-44e4-b06c-ea52c35bfb5e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.361544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.361544Z digest=sha256:33ec4885bc10aa299d0d377cfca861cf5abaf77b2ec5656151e2a1fc0978471a

Observation 5ea1f9f1-6990-4fd9-b2fa-c90bcf99f027 · outbound

This paper cites Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.273660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:16.438696Z digest=sha256:95414cbe337b5848bb1905a4aa42331ee075f0ffc4a64207039278348632351e

Observation 4dfaf339-61e1-49d1-8b58-eda1097cfc44 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flux.https://github.com/black-forest-labs/flux, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.556293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.556293Z digest=sha256:c20f07a558ea6d79aa351e3f5c9732f30e23690df074465e4b78a5ba666bd6a5

Observation 19350a6c-032d-4ce6-8ec1-3fb0e20ebb8e · outbound

This paper cites Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.150952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:16.644623Z digest=sha256:469a9ce41dc83d70dff6ea2312f2ba163f16158e00d82b4a17edb99815a8efcd

Observation f46da92d-f9b5-4608-99af-db1e4c6520d2 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.743577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.743577Z digest=sha256:0069931809778204143c66775881c0682789ab35337a8229f6be931870c948f4

Observation 76282589-2e37-47ca-b250-694a41e03f99 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.826082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.826082Z digest=sha256:34dcb57dc95f8e80f9b3b59fbedb66787647fd4a5a5f2c031b61612eed40d361

Observation 38a85164-1543-46b7-b883-2d733592397b · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.907537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.907537Z digest=sha256:07db0e3ca92bd71d00afdbb7e372f254c06d6813da6614c98714c8c07b6327ba

Observation c8d77a9a-3258-4780-8bc4-d4aadae3a339 · outbound

This paper cites Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.012891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.012891Z digest=sha256:95089b922cd29dd4408103a8e3d01652fad26efff577fa191765efe6bb533ab3

Observation e6378ad2-21fc-4c2c-939a-1eaa4ae8a8dd · outbound

This paper cites F-LMM: Grounding Frozen Large Multimodal Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation F-LMM: Grounding Frozen Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.054823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.054823Z digest=sha256:024c69fda2a9fe45b7fdc0c3d7c46c04b8739c332be104fff93b7cb2f7f012a8

Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.119373Z digest=sha256:6fe5c248be1115bd52497e01054de338aa15a20eef1a8e3adfa940a2fc90b328

Observation 3e059d4e-8c73-46df-a29c-1c8a4bb510f2 · outbound

This paper cites Scaling Laws for Native Multimodal Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling Laws for Native Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.154880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.154880Z digest=sha256:924cafaba0de27f35d9b8179c3752f834904764393b10d07b42cba7e6d2beda4

Observation 13c031c1-c6c1-43ea-8e5c-f94f3a17d559 · outbound

This paper cites Qwen2.5-VL Technical Report.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.248177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.248177Z digest=sha256:fc3804318ed01f3c741bfec9d6f0a6e13f39bd9e010bb8c5747216716cee1865

Observation f667d273-2f2e-47ac-892a-b5ff3481590c · outbound

This paper cites Classifier-Free Diffusion Guidance.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Classifier-Free Diffusion Guidance

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.315211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.315211Z digest=sha256:b30063f6ff27da4ac8c3612e39f5d1779e1e88f2fe7bf47c7fd494a33b9bab5f

Observation 20ea9665-0bbe-4b92-b2dc-81d09ba090cf · outbound

This paper cites Decoupled Weight Decay Regularization.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Decoupled Weight Decay Regularization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.387643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.387643Z digest=sha256:95c473014b76ab38d095c00931046ed6a0d044a676489eb63856d41b7b20a29d

Observation bef601de-0ec9-4569-b8aa-0709d22d821b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.429682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.429682Z digest=sha256:17e7ad737ce3bdd9531ae7f071ae09faa24e9a5ab53edc5b5edd16a8739c795d

Observation d648e502-14c5-431f-815f-920dd602f54d · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models, 2022

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.506293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.506293Z digest=sha256:f633c7600de5e938a793b43c19cc5fad9a7a1a9166bb3efb065fffaa1d36378e

Observation 0762f47d-604d-422f-92ef-14c98bdbfb6a · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.567850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.567850Z digest=sha256:556eb3afd0c923d4f11b99c1069eedca91393f0bf3ea5cc818a236a906c6dde7

Observation e6d7ba6e-4e7f-484b-8866-57a44e081f07 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.634220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.634220Z digest=sha256:3364253f2bbfc12a090800a291641a9931661bb00acf24930a9959af18939549

Observation b128eb78-e9c5-4581-a077-ba2022a1ad64 · outbound

This paper cites Flow-grpo: Training flow matching models via online rl, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flow-grpo: Training flow matching models via online rl, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.685541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.685541Z digest=sha256:4e90f81c79cc85083a10503a38bd91d667f2482771c3840e09ad269e15fedba4

Observation cdcc7b72-102b-4a33-a9df-31ce5984a324 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.764755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.764755Z digest=sha256:3e8841c0cbf1f52a167ae50ea0e0d8ad909fd8a65dccb5ca46b401366592a35d

Observation 15e45610-e6f7-481a-aac5-478ccd9f27d4 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.824218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.824218Z digest=sha256:5c9448d5bb0186d47f8248e4c6c202dba31f145844524307618e32b8bfa760d1

Observation 62f640ca-edc3-42c1-af76-465ed47dc62d · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.919986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.919986Z digest=sha256:cc3ae64845e8a1dae1d9262e229e55d41a16459685f8eab59fc0ed1924a13792

Observation 010c2984-c358-4ce9-b2c1-2177b04ea6d5 · outbound

This paper cites text-to-image-2M: A high-quality, diverse text–image training dataset.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation text-to-image-2M: A high-quality, diverse text–image training dataset

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.017543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:17.986317Z digest=sha256:ce45a1d081e997de8dd2010257286173f3d0eca2d1791bc5ee790f8e2216e302

Observation 538e3db9-6637-4525-a4cb-c950e8931524 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.086468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.086468Z digest=sha256:9eeb5c0ea0579dfbeebf83d9c167afd8d2c4f21a64ae42266c02faca1f3e6b4d

Observation cb5006ae-20b2-4e04-bedc-c6914bae2c8f · outbound

This paper cites Megalith-10M: A dataset of 10 million public-domain photographs.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Megalith-10M: A dataset of 10 million public-domain photographs

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:19.908808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:18.163452Z digest=sha256:44c78d73151c8591681d3c39e39645c3b3e3313906e2bb12ab384b9228634eec

Observation c1e9cfc8-45c1-4728-bf5b-8f6682335c2e · outbound

This paper cites RedCaps: Web-curated image–text data created by the people, for the people.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation RedCaps: Web-curated image–text data created by the people, for the people

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:19.834909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:18.256979Z digest=sha256:d77ca79794792c980c924f390bacce0749a5a9b8a40eef2ad39b6e2a8537920f

Observation 44765e00-e05b-477d-8811-bd652baf00a5 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.323666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.323666Z digest=sha256:2476e772ca9f5b57b381cfc75e1f0ee0de94d53bc89bde379770c63cfbd75db3

Observation 0c6db8f2-4aab-47be-90c5-330ac5c78f80 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.427453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.427453Z digest=sha256:28b00645915b9ed094bd7732ade6da3d07840feacd74f22a8c943370fadccfbe

Observation eea42e66-daf8-4d68-9b58-bb75b8c65925 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.464184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.464184Z digest=sha256:bba037040670af54d8ac99df63503ba870ea6be79b060449f8fbd79388c1a3f1

Observation 47103e9b-bdb7-42c9-94ba-44a69aa07525 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.542352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.542352Z digest=sha256:48f7428a3733e09d4c8cde396aa3de2243c9eeaf3a4e517f5456f96e41bcdb68

Observation 29dcb26a-57e5-44e0-8b57-e43c8271223a · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.590800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.590800Z digest=sha256:ac3c0785dc4f36e90030c9c7327ea34b83ee2393b65d495244cfaff943baf280

Pith citing papers

Observation 6b93d0b8-5e55-41ec-8c49-81e8ff292d4f · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.498867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:287a742224fe2337a7258ad4942ea9427bfd241abb883bb5e3d1858a0b022dab

Observation 2440898c-8e79-4572-8e74-09e812db5517 · inbound

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation cites this paper.

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:37:16.161813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:37:16.161813Z digest=sha256:9cc1712567de96e9c0e6bad9775ba2fe8c5d36b872c7ee39fc9afcd9f8256671

Observation e6e71b26-80af-437d-8272-7790425ea856 · inbound

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing cites this paper.

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:06:50.057017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:04:00.773443Z digest=sha256:4a8121da9dd32d03b7c9af1e6151f6fd243c411d185c1569bded74cad64086a4

Observation 2dd93c32-9928-4ffc-a9d7-b37b27e91874 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.243005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.243005Z digest=sha256:2c0a1c9909ea2d509f160f649d3852e5e1a0f2a3fb435df48adef53ad725ca03

Observation 1120e7a6-8acc-4195-9c58-944c4448e8fa · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:14.264753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:14.264753Z digest=sha256:0623b9c4a99c31f05e87c480825c6fa3c34651d48e234b85d09d7e5fe95161b8

Observation 21aa7041-6a31-430c-b826-375928f15e2f · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.393690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:012742105b0d716316ad386b5efed477b42a00c0b67143fe8eefdd914e394d5a

Observation c3b208eb-af74-4035-9ed4-2c9adbb461d9 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:ee1e25c50bd57cb27a2bf89408ea1cc833480a689ed289a1fda489ef43a8bee6

Observation 85e93ea3-9f3f-415a-aa55-0158c10054e5 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:03.172221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:b732a997e91fbc5399ccbf39842e8958b59c8cb4958e6e7845325f6023e05c9d

Observation 5b0f89c4-257c-4a73-a01f-98ebefdd4451 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.268707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:34e9af11c9b1d32f2e94b9b34749b88c366d64e6930c7bdb6b0c475dd996c7e4

Observation c96aa05d-13c8-4ef0-88af-bcbcf5319edf · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.580506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:ee93e6cc43d4754aed5a62dd868e847d6a507081aa8d9f3725be8a0831763866

Observation d8b082f0-af4d-4067-a289-d77dd57f1849 · inbound

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens cites this paper.

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:24.805355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:58:03.974053Z digest=sha256:82c19e50976bf227fa294e13861142ae5f4a88cf624a54d198c61ac7305d2be3

Observation 519af2ee-a3a0-4291-83e1-c444c03131fb · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.050035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:ee1c87b516bf4acb66046d5c21eed7a40bb1eec1ff4eb412643fda21f5c621b8

Observation 921dab5b-a36c-4778-adfb-20bfbf7fbcb0 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.195508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:5a64fd51bb206127203f7b8e2d6ae69935747f6fa19d40f3eb76364115db774f

Observation ac2fb108-0f9b-431b-8e5c-909ddd734640 · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:07.974566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:06de417e36daa971867ccaacd659468b4aa2a23cfd2f12d902ff21dccae30d20

Observation 333d691b-d2e0-4bd4-9cd5-d6ba23abcb42 · inbound

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision cites this paper.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.358702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:29d033506f17774170d62192b6673e8885752a9d05d4594f3a17047be88b753e

Observation 008a8304-022a-4fbb-9b37-34e5fbe1fa6f · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.898621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:b8bab6edfa3f496f9a64f0eab2e48adda6a663a136fbc2e6186aeda6daec8207

Observation 4473075d-f5ca-4df6-8206-6ce5471a2566 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.389113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:cc104ec629d7e6f039c918274a41b80ba5050e7b57d88f0f526c02e66ffbabbf

Observation 6a177bd3-2801-445e-91e8-250c9704f213 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.301808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:c4033d4c96d77935bcebda3eb450c9737080cfffc9acd2cb92fa58f2d16f3938

Observation a318c7ae-ef0e-480b-9093-0fd3fadac828 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.345303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:03e3bb55c3533fbc6c21708a005c1c3f72e5814f6ade44ed4c91663b26080954

Observation 56505dc3-693d-40eb-bc29-692615926173 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.896947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:46af1326adbb7b5cfff50d9bfd70b94d6cf5b258ea19ee4b4f1f6b54285a38a5

Observation 8243c92a-4a9d-47b3-a384-fee137bb55cc · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.697635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:89e4a709addbd4e11ac05581302ec0c03d6d0ec004c94e897ba50e829336083d

Observation 2e827ec3-e9be-4230-b11a-982702ef9b13 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.439470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:7c6af49d0b0ca5c24452a65e106b2279920552c61452d9d4f0d42ca0397c1708

Observation 962dcbc7-89cc-4983-8fef-605c13eafe22 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.238735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:d488b9d629a24e79f4ea7d09c33efe6283a44f907ec5c127577823454a6ce375

Observation 4335c52f-26bc-460c-b559-203232a2b955 · inbound

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation cites this paper.

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T05:11:19.001184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:11:19.001184Z digest=sha256:3738b784533f26ab1be3b7a462ebd298e41cd4ba6cbc2fe6f1ccd19856cbe8e5

Observation 8ae6621d-7e99-4606-8967-b66ff2bf033f · inbound

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation cites this paper.

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:18.617071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:18.617071Z digest=sha256:e35c8f10158c2f0cefe2cb5860e2f56089b22109dc438c35944972ea9e73a6ea

Observation 5d74a6e2-0343-4150-a5f2-213031a8a0d0 · inbound

Test-Time Curriculum for Open-Set AIGC Detection cites this paper.

Test-Time Curriculum for Open-Set AIGC Detection OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:23.340701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:23.340701Z digest=sha256:995e06dece6b5dd3fe8f4e320a77efd636d5bedc6cbebd99cf3d855d1f14cfa1