Pith. sign in

Paper Citation Record · LEDGER

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

As of 15 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2605.12305.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12305 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:48:04.997796Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:27:25.307245Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact32
  • verified fuzzy17
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 184fb529-11ff-4f91-bc77-d47362a88a87 · outbound

This paper cites Qwen2.5-VL Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.763962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:c779b5b6d16ba5c3e0ebc5924bed0cc806f27fdbdf8a9f555d5867c1cac87323

Observation d02cc122-6a9e-45fd-b276-cb6cb638b2e9 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Emerging properties in self-supervised vision transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.861704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:c063fbc2d2d6b3a36f625869bda4093f904f1308fdf62b770bcefd1fa6e97b40

Observation 34142fe9-7a68-445d-94dd-dc4d9f9e8886 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.690835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:ffbb5b5a1c2509321d87877fab1695bfac42a2e222443e856f02a856b1fcb093

Observation ec5f7b26-8613-497c-9838-d9e2a579dbb4 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.743962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e5b778335098011574a701242ecf85532607f5ee9791ab7e52555923835ab432

Observation a77391f2-5a52-42be-a84f-9723756559f7 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.720119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:b6c6191726118e8907bad5ba9beb2321e45447c12206bb92ab06a23adefc7a23

Observation 6859090a-7d47-493b-8db7-75ed1d5ad84a · outbound

This paper cites Gemini 2.5 flash image.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Gemini 2.5 flash image

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.807997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:d7f18994980915189db97e3da15474ead80ef6a8b24487a6353f8faf372b4f3c

Observation bde4c4cb-f641-4bc1-8f6e-8da44201c409 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.754291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e2a96af62e8f1412aa725ac239829d96d79ed2f2387ad4f5a1158a2dce208260

Observation c15bc8fe-bdfd-4bf4-8cdd-171527205dce · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.814309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:dc71dead3588c54b9c85f106bc3690e76c7ae0c41cb7193a29f409f20a1c5f41

Observation a915e8f3-18a6-447e-9509-ea2dccc0f5aa · outbound

This paper cites Seed1.5-VL Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.697607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:1afd3e0e72d3ba1dc0d592ed42ee1b96e0019b081dba8d832083c120100b8d84

Observation 95701d1e-0e24-4826-bd44-f3772ff1fe0c · outbound

This paper cites Chameleon: Hierarchical clustering using dynamic modeling.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Chameleon: Hierarchical clustering using dynamic modeling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.842703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7615c29af72d385ac81bfa0505f18fb9933dc75def3f45dabd317fc1de81950d

Observation f8dd4f0a-f458-4477-b5b8-369d13a354ed · outbound

This paper cites Segment anything.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Segment anything

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.803268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8d0d782b62b73cc734c421eb5e7e36c256e1028e7ac7fdf1518f4ca1a7b74848

Observation a41534cd-f50b-4f6e-9449-24649a562a6c · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Flux.https://github.com/black-forest-labs/flux

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.794238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:52fcd0949df35e647e17906015b40ad23534fed92cd29e0ad28e9c4bd8eaacaa

Observation 691e6ee3-268f-4a8a-85c2-3db6d39c2708 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.757588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:c61e50e64f9fb34d97a5974b39d990f0ef5a02db8f1a6d874b425137eeedd316

Observation ed4986ff-549e-4aee-8c4a-8a21cccce593 · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.782843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:c5297d351415b63e17d1c468ef1535f29c87d021328c585784a87513faeb5e16

Observation 87b58bcf-d74a-4665-b981-7d4557b10129 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Describe Anything: Detailed Localized Image and Video Captioning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.777196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7fca3bd9281a38aa3f2f5ff4203f10171fbc85be9fadfb8a0c11bab8b5043265

Observation 40d95ae1-9ffc-4cf6-8974-1b2b6cbd2541 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:bb2abd369c14df38bb48d95514d8fc4e2caa764d8df20d18288cd0f9e10c8935

Observation d995ee84-a320-4df2-89ed-2763e9389dc6 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.785528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e92f866ddf9420b416dc3dd2ae3b9bd4454fb4aa2d0db12b35036e6497bf0cf6

Observation 6e416529-8161-44fd-8d2f-6e6778ffb2ef · outbound

This paper cites Visual Instruction Tuning.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Visual Instruction Tuning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.790888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:3e9b7b98244994b969a996d36c3964a585ddab30dff9e562a51308da05fe4f90

Observation c6a591f9-b05a-4d3d-8321-b25e1edf1490 · outbound

This paper cites Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.781426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:fac7369f8e0ee469f39def60e1067c1cf4815a4dcff06d37503e930f66c70926

Observation cbf2701d-f968-4984-b387-5143eccc9a18 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.788227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:95bcbb76f4e625ca610260ef28a878568899828a3a3a516fecf94fe44b4edd1b

Observation 79b11037-e058-406e-a1c6-9b940fedf1ec · outbound

This paper cites Dreamo: A unified framework for image customization.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Dreamo: A unified framework for image customization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.773926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:862b532f629dc2168b66ad06a8ec32ecfbb145e9b2bfc4e674cf00203baf1b6c

Observation 0da8606e-f67f-4aae-85d1-6590d4590e00 · outbound

This paper cites GPT-4 Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation GPT-4 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.737992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:9d1f09ac3cccdc95f3695d26fcc0387dfabb4c2ca365fea53827938f357c0bd7

Observation a830e551-45e3-4a76-807b-4c882d120aa2 · outbound

This paper cites Gpt-4v(ision) system card.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Gpt-4v(ision) system card

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.798877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:c399198b9e9cfebd0e797d7a1a16e8a1f71236811249d1745f64ed957202d2e0

Observation 8df37de9-e8ca-4b22-ab98-0957be847c64 · outbound

This paper cites Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.847133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:f083c66310ca44acf35d8f0237c02f1de94a7a521bf389c6766c80f5d5a7c192

Observation ca46de99-96da-4d03-9951-b21a76c762d2 · outbound

This paper cites an unresolved cited work.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Unresolved cited work

Reference 25

Resolution
parse uncertain
raw_fallback, observed 2026-05-13T10:12:37.851695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8b5d1b6e0f1db62729b147571d3418bf260063571cf6ca57ba6d394c765a5ad3

Observation 74bb8080-a6af-4291-9980-545ac4c94525 · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.770498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:a5d1d758861a4a2468195b709c60f67cf98df0206417219d4380dc336cc7bb93

Observation e6de04df-cdf7-4de2-bf8a-5f4e2345d248 · outbound

This paper cites DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.767075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e89808ba324205b49c47c7f89c4bf85e81e4824facabe48dea85e067bdc62af2

Observation b9cc249d-4f56-47f3-9fad-7f0522c1758b · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.760735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:62da83f049229c34a1afedfd64a5c897073c30ba3dea3891c73640a82853406d

Observation 35079abb-46b9-499f-8a98-7330939ce200 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.710492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:94ffc1e379488833f9d92be43fcae2555b7de76b48971b9ad53ef867eba3fdfd

Observation 29f1fd73-3ec2-40f1-831a-49f943a99bd9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Learning transferable visual models from natural language supervision

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.856835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:b805a524cfc31127c1a8f6e37e0a6ff54a67aed01c59cfd53d7166ea8265a594

Observation c8fe1854-57a6-4bd3-be70-5172928e8262 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.869908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:77eda29e21700dc1b01717dd26bf8fd2c643cdc742c916d91a29f1dc9552fbbc

Observation 05a1c1d1-8de2-4676-b1fe-d306af5cb477 · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.701026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:fff081a74ae3d26dab00ecd6fa370199859fd97712c6806a89734a453547f5c7

Observation d4776120-3d4a-4c3d-8792-798b1b8513d2 · outbound

This paper cites Generative multimodal models are in-context learners.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Generative multimodal models are in-context learners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.865865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:1754ca6bfe270a0d28bb976503eef780c261e0e33e9cafa074f86292e0710238

Observation 815dab61-6538-43db-838b-ac5235468c31 · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.726708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:6f2efe846f903baaa6065d955f81b5f77f4572ffef222c7f1775d16b8fea47f5

Observation 1a4ee54e-4b6d-44b7-b286-383a13866abd · outbound

This paper cites Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:22.732596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:173d7f036b588ad08356979c585f516ba4e84763e5c93ddfa9c9602ca5029de0

Observation c181c187-ec0c-430d-af82-7e915d741b23 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:02:41.552159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:951404e5d9eee60f5af84da12b1efa81f6b60ce3c115d933854ab7a5cfd4639f

Observation 34e5f6b8-aa7f-488f-a273-e231abb719ad · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Emu3: Next-Token Prediction is All You Need

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.729608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:633fc89a0a9cdd1999f8063eddc02e9fab0e8e6c12a8900e713c93632dd09e31

Observation d30ce7dc-45e1-4262-abef-f8d354840325 · outbound

This paper cites Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.779971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:f8e167eb2a872a6cb6423b5595d40628c13381acff6e3db5c27eb9f70496f466

Observation e44053e8-a5c4-4de8-b981-70520fe6bcf9 · outbound

This paper cites Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.838638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:dfe69a09abc6eb1deef420f770cb51491e762eeb755a741099cc968e27cfbc42

Observation 460aecfc-05ad-4a34-a506-df0bed331fab · outbound

This paper cites Qwen-Image Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Qwen-Image Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.722761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:3b2e0670b8199a64047157817e25f734fb06202da8aff7097836c7f261166582

Observation 9ce679b5-b696-4cd5-a973-71cae6ab9f6d · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.833070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:0a2d060d2f9daa1808c789076e47e57de74899481a99039b391fd14a0e04d6a5

Observation 4f78e979-c64b-43cb-990a-aa9b56877b56 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.735182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7a6510726a8887bc110c82d0b1880356fd260423a64924428003533e02bb947f

Observation b6af3b54-1388-4e76-b3c0-4b9ffa81aff5 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.714030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:3d7815d66dabe4f1a4ec61325bfa5faca48dc7cd418079517adeeabb02536f16

Observation ed38d77d-eb02-489f-859a-713aecd61eca · outbound

This paper cites Dreamomni2: Multimodal instruction-based editing and generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Dreamomni2: Multimodal instruction-based editing and generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.750928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:9fd1b72b783128c34e106df65eb13c1f6735ebb70ff4d6763f93f4c1473066a6

Observation 8aa35615-b44f-4bbe-8c39-2e55c8afceb0 · outbound

This paper cites Omnigen: Unified image generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Omnigen: Unified image generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.825560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:c2b4227425c3aa9301a0068a47cc8dad486d0b5c9e97000b07671b5a02515003

Observation eb73f5d1-3058-4b12-84d0-33a396d6d6f0 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.707578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:a56f7a52e79b5f9a61617b25751afae8bf2214c40eb3a0ec6ceb1e01316bbd1b

Observation 3c2fbb64-7882-4437-97cb-ce9e9ec09fd4 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Show-o2: Improved Native Unified Multimodal Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.747295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:5fca3d9a846d03d5cc0309819edadc9e5cfa1abf095526aa8684243e1afc730e

Observation ccbd9b75-0f48-4282-9465-0574330b9c6d · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.741046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:ea3596310f8013dc753a1c122180a348ea30de1d8ffd0d93e2bad380a5700df3

Observation 9f65ce3f-0c10-407a-ab3a-167a1e28387d · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.717435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:76724e70c84150507b7c56492ceb252c54d79f74d97f4be98f8172e08aa4bc8f

Observation facabc92-ffa3-4be5-9c3b-ec6978ae69b0 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:aac56e6a08213b49d49638dbfb49a83c5c2781ab95606965e0252fa8e11c11df

Observation 81863fae-9bfc-4f09-891e-b9bcbbe42cf2 · outbound

This paper cites Image First.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Image First

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.820894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7343bdfacc502423789187e3662ec20ebbaf7e38efd80618b8e5dbcafbb56cac

Pith citing papers

Observation 46ed4f89-c7f1-4cdd-abe3-920c94449164 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:25.307245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:25.307245Z digest=sha256:e6309351afc71e9e47226e6c02ab6639dbb000a9a03556e4c2611d30a39d9b82