Pith. sign in

Paper Citation Record · LEDGER

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

As of 9 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 3 inbound Pith citation observations for arXiv:2502.05415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05415 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:32:32.091286Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T04:50:01.364000Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T04:50:53.692347Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved71
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0d853e5-6e7e-4ad0-a0cf-cb73588cdd9c · outbound

This paper cites Nocaps: Novel object captioning at scale.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Nocaps: Novel object captioning at scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.758583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.758583Z digest=sha256:9722bdc5c18e92625af3a9cf2c8a84038308972c006fbef9353af47a41ea0214

Observation 81370336-c46c-44d0-b036-0c6841e1c9ea · outbound

This paper cites Meissonic: Revitalizing masked generative transformers for efficient high- resolution text-to-image synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Meissonic: Revitalizing masked generative transformers for efficient high- resolution text-to-image synthesis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:32:33.017870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:32:31.764015Z digest=sha256:995a205947df1f5bceaacb4bad3e5f8a756e2f5792a1a771c6ff24ebfbdc2ffc

Observation 660cb863-51a3-4175-9aff-4bcbcf7a17e1 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.768624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.768624Z digest=sha256:da97638141d853010b199d923b37a1be594f13e702de891fd34009e21425b2a6

Observation 8ba8980a-52ae-4ca7-a561-d16dfe2a572e · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.773644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.773644Z digest=sha256:d2223979eb186d6ae96cb7a04808f4776661079f67064c5527c996207b64319d

Observation 0de2ff5e-cf58-484e-bb1b-2f9b1e62810e · outbound

This paper cites Maskgit: Masked generative image transformer.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Maskgit: Masked generative image transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.778117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.778117Z digest=sha256:40f2fe53c6999ec88facc9b76d7b636dec778d243f6335bb63617a26b55d371e

Observation 24e8bcc5-a1d0-4987-94ad-3c0953077f90 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.782645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.782645Z digest=sha256:f2a1b10eb80b93e21209c64b156c5a15033c25d774c559c978596a537e734f05

Observation 4b93af3c-892d-40ac-8a96-9a1069643208 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.787962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.787962Z digest=sha256:87d4faf2a1ee7ebfba560b341374daaf50d26524d2191e139c366049ae31aee2

Observation 447310c9-9bc5-4483-a619-43fa9df5f104 · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.792915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.792915Z digest=sha256:cc614b78ba3d64159fadfb737beae496d09d825626d2a218f0a28b8bc89e3e52

Observation 2c6f69ac-a1d0-44e7-8f5a-a00b949aa260 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.797843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.797843Z digest=sha256:ec058abb123f04ae4cf9492d088c86d886c209a362189347f0844f886f62b29f

Observation f2fef40c-a55f-42f0-ae25-362f1dcb8acc · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.802454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.802454Z digest=sha256:59517216af57f11e82e7643de3ecff92b60f4262fa4a056f839a186803eda77f

Observation 32b8e202-2fec-44c2-b1d3-315b148233df · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.806550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.806550Z digest=sha256:f9b3891267e8d19ef393b2cb731baf547c336ef29e3149acdccfdee90511fd27

Observation b7f1a086-c2f6-48f2-b148-5bf3f1e9c6a0 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.811276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.811276Z digest=sha256:3ef1740067de2b0c1b500991f0ca633b416b6cb52be470e8f5ef421fa0744705

Observation deafb6bc-1dc2-4018-ad54-f99353b76ffe · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.815634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.815634Z digest=sha256:30f7909b15e08d00369b13af798ea8ecb9888e731117782f9851a4a56b6d1f94

Observation 6fc7ee8d-a6bb-4657-924c-a52eb0d95681 · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.820303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.820303Z digest=sha256:fe8e367aaa10680f03d7e4e8ae12cff05b1275efb514eeba02098826172d8211

Observation 04b9c00b-afcd-4f63-99ff-6a1562670bb1 · outbound

This paper cites Scaling Diffusion Language Models via Adaptation from Autoregressive Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Scaling Diffusion Language Models via Adaptation from Autoregressive Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.824667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.824667Z digest=sha256:13f9289009aed3e5e1b3feed2e29f670699ca2feb6ab01a04c06c4f6ce2d4710

Observation f4b9ecba-5e8e-4185-9508-22dd5d8ecb71 · outbound

This paper cites Distillation of Discrete Diffusion through Dimensional Correlations.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Distillation of Discrete Diffusion through Dimensional Correlations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.829284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.829284Z digest=sha256:b626c4875854be9b705f959df6264f597c95e77c42901b93cdb22e861ea0c3bc

Observation 3970dfbc-d387-4fe7-ae33-5ac3a22f2475 · outbound

This paper cites Multistep Consistency Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Multistep Consistency Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.834011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.834011Z digest=sha256:db41a4ff55ac625ffcf1c92273539a165332cd76a909ba128533add71f1da860

Observation e509b2e8-5a80-41aa-9ab3-49933a812192 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.838794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.838794Z digest=sha256:c6f236660adb3071c040340f2dacf34c2b51356b6546232335df8535dfa5aadd

Observation 6f439845-ce92-4046-b555-40aebbe79065 · outbound

This paper cites Classifier-free diffusion guidance.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Classifier-free diffusion guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.843547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.843547Z digest=sha256:2a29441a0416416c3ffbd0dcbd9710bfa158e6933e1c4116db348a9b52064c56

Observation 2ec9ac62-e955-4cfd-b58b-fe558420c712 · outbound

This paper cites Denoising diffusion probabilistic models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Denoising diffusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.848040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.848040Z digest=sha256:abf94b74e00ea27cba8494a473d25f501f19d31a5eea7e338b6f8d67005f21be

Observation 904bd2c7-d05c-470d-8b50-5c64dbb4f304 · outbound

This paper cites CLLMs: Consistency Large Language Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding CLLMs: Consistency Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.852651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.852651Z digest=sha256:7287b092e22581f8ba6f7ff9218ade861cafee9ce18a4babec08d4c95f93c31d

Observation 526fbf8b-d74f-470b-a624-48c9600b9ca7 · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.857081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.857081Z digest=sha256:7f0ec800e2546da1cbc191b17830ee1fbfd94048fc1d54ac15d7cf100187d5fe

Observation 9d9897c6-5d50-4db0-afda-4117e9e9b6bb · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.861883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.861883Z digest=sha256:4da0c18bdfedeb3662cf4ae3aa0cdf315c047995b639e2fe0521e40689c4c5a3

Observation d51b8ec5-4492-4816-98c9-34dfbd9619af · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Evaluating Object Hallucination in Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.866468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.866468Z digest=sha256:da5f8b9fc2101f17ff498d5d32fc3bb0cb0a8b3d28756cd5e09a7f0bbd03c987

Observation fbd23343-f99f-4090-a82f-8fe3ccfc52e2 · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.871239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.871239Z digest=sha256:befed713d6fa798100f9288bafef616ded64e9b08436f862e33b62a1be615669

Observation 8bc4c100-3750-4c4e-a460-f2075022732a · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.875847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.875847Z digest=sha256:8281ec84daa04b2d0b30bb18b66d929aaf73d5593eb7c62532c306a9118097d5

Observation 7338d1e5-19e6-4933-a70a-8ed36c5d9927 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.880312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.880312Z digest=sha256:8355057ff054658b9304e860dde2a60d3445c017e72709ef3525c69eafa8fc60

Observation b9b81f87-738e-4787-8cfa-862e55dfb9b9 · outbound

This paper cites Microsoft coco: Common objects in context.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Microsoft coco: Common objects in context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.884801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.884801Z digest=sha256:81a37a38f0feebcb8db7ecfb2848193c3e3649f0fb9ac4d7411f45a1d0fa16fb

Observation 0c2adbeb-1897-479c-9aa0-3f3b135984a4 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.888953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.888953Z digest=sha256:15f26fd044e26ad03d889307126bb9a69dd402c9ad0d0753d5aa5e1ea661f19e

Observation 0bb7cda8-5ca1-4b95-b1e2-e0ae48961921 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.893442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.893442Z digest=sha256:62cd41b4147b51301cb69107ff1e1d70b82fb288d4715db60c661aace2ee702c

Observation 8a29e46c-953f-4a2e-a1d9-b7fc85736ce8 · outbound

This paper cites Improved baselines with visual instruction tuning.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Improved baselines with visual instruction tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.897650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.897650Z digest=sha256:5aedad13c6fc8fa892cbd3d798dd84a69ea583c9fee4dc5cf3e91fca0d441b7c

Observation 9559d9d4-9dfe-43b9-9bb2-05b99a055959 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.901831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.901831Z digest=sha256:fbe5136d45dda2fbfe86aec321bd9f4bee3ebddac59b111ed6ece9281e3ae514

Observation 99a115c1-da86-4fcd-898b-fb4d6056903f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.906235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.906235Z digest=sha256:fc874830cbcac8e6a5f92055dff3085ee109506971b8fe25886dc9b344ba9996

Observation 37b9de32-2054-4735-97c5-7001583a8849 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.910383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.910383Z digest=sha256:33de513e54fc852ebcacba7bd3183531f64e54acd4049f0e286561651bc13702

Observation 2b54b27c-9509-481b-a262-86094d3f6613 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.914446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.914446Z digest=sha256:b9f081af33b94c328ed3e532f187bf86a9547a4a93bd039e7ff0c2fccfadda07

Observation 154ef14f-8bb0-49d3-9d41-1c806bfbcb1e · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.919095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.919095Z digest=sha256:805d6a47a70ce574bd8eccf13f4ef280406c43116b35954d108fa240295d5238

Observation a53df81c-9326-421e-9334-7ba9ffc0be04 · outbound

This paper cites Scaling up Masked Diffusion Models on Text.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Scaling up Masked Diffusion Models on Text

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.923602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.923602Z digest=sha256:9037e1a82994c54f69677ec9151fabe87111c42ff3f9d24c42cc24a1004f8bcb

Observation 7ee81add-873a-445f-8cfa-5e4e30c60307 · outbound

This paper cites Large Language Diffusion Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Large Language Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.928082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.928082Z digest=sha256:71b29fe56bd2c68f07fd566db44e95a4950151b0d46f76a2aabf2f4e3a97dbb5

Observation faf83f71-33e7-400e-b8b6-2d18f7aba3c9 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.932524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.932524Z digest=sha256:d560637c5123fdf9050a3c76f98803e5f158cab87fd35249b7df8e2df6328860

Observation efdb6574-064c-41e5-96d3-6eaaf667541a · outbound

This paper cites Plummer, Liwei Wang, Christopher M.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Plummer, Liwei Wang, Christopher M

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.936845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.936845Z digest=sha256:8a281d98e1c3dc0cb37a35cd41e9e96b2fd9a1fe56c29966d3b6b490037a6015

Observation 9439238b-6840-4291-b5e4-af7475faaf31 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:32:32.889853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:32:31.941265Z digest=sha256:7f9f13645ec288da766c2113f8268551a5a17fe9d167d63db4a55828c482eb65

Observation 6f290ff0-26dc-4bc4-98f1-a88a07611163 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.945668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.945668Z digest=sha256:7a0c03c08b5863c31f2b5f496faba6585d03e0df13d44e9476294cd92e8d7326

Observation b15d7f19-78dc-4a73-b46e-59ace3e2603c · outbound

This paper cites Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.950210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.950210Z digest=sha256:a880a63f4c5ad5d7260f10c155a8857c6410fec6d98eb85d4743938d8fb08803

Observation d4414e88-df61-42a9-acea-b3fb5aba8644 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding High- resolution image synthesis with latent diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.954647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.954647Z digest=sha256:033d5149456b6d59f11058d999869c7e1a8f614540bf86ea5f626d487f5dc3fe

Observation ec808a2c-c23b-4aa9-ab1f-86ead260abe7 · outbound

This paper cites Adversarial diffusion distillation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Adversarial diffusion distillation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.958916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.958916Z digest=sha256:a810fc9b07ac082d5e66b31d54dc4834b4a5bfd12adab1f8f3a05c935397377f

Observation bca5b52d-d0b8-4121-bcde-b442ee225a54 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.963024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.963024Z digest=sha256:6fd5038c279aa0875666b259a3592a09dc9e4d8cde38b52ddd204b7d9b193ebc

Observation 25ab360e-4429-41ae-a6ad-5a5080358342 · outbound

This paper cites Improved Techniques for Training Consistency Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Improved Techniques for Training Consistency Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.967460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.967460Z digest=sha256:6d80d6f5ab6140fc3f58d340c312b55a69e21e02036c119e49d630df20c40aab

Observation 91f9e05f-97a1-40fe-b076-dcc217254a34 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Score-Based Generative Modeling through Stochastic Differential Equations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.971891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.971891Z digest=sha256:73e782687756e9552ac850224184f972619f7ba1fba945a797bfe9e78918c467

Observation fde81fc3-1829-4bcf-afa4-9d8fe87a585f · outbound

This paper cites Consistency Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Consistency Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.976432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.976432Z digest=sha256:3cb62edcd4259ea8befd92940b1f3cdd1f73865621928ee169e9d1f2d02691a2

Observation da308618-fe28-4b9e-b2f2-a01872e681ec · outbound

This paper cites Score-based Continuous-time Discrete Diffusion Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Score-based Continuous-time Discrete Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.980841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.980841Z digest=sha256:374541d848cb523c4cabd227871ebea7516239bc8da8921c9eb1f3f63fca61fb

Observation d944ede1-2ced-4366-a6b8-b905feae19c2 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.985055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.985055Z digest=sha256:0f5b58c6249b009a0886b80f6a0517fbca98eac738d2d6bf199a44f6e9c9dd16

Observation f0445cb3-7820-40e6-b720-5805db2ee054 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.994068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.994068Z digest=sha256:0c7428c8cedf60cf9cb8e76079d4442de71bd71b42ccc540c1015b075924597f

Observation e171689a-5604-40d0-858a-16ad27ff22a9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.998283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.998283Z digest=sha256:b4e8be541afdd7ac42adf21a51ba5c6de12f63f5e07ae6d9ef83d9395cef1b13

Observation cc2ea35a-8e48-4050-a500-c80da3423dc7 · outbound

This paper cites Neural discrete representation learning.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Neural discrete representation learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.002935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.002935Z digest=sha256:0417e0d4fd4ad62d950b193d41a7eaeb185f210bd0978ed7a8791635b5152c87

Observation c3eb76dd-d6d1-4240-bc7e-ca7d8b745a8a · outbound

This paper cites Phased Consistency Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Phased Consistency Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.007195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.007195Z digest=sha256:0a87835e2ab4f089fae7bc5d4c641d1d6d326456fa4ed811924aa006398fb6b7

Observation 957c7356-cfe8-41d5-80ff-3173c2fa1573 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Emu3: Next-Token Prediction is All You Need

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.011924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.011924Z digest=sha256:9ebac0c1fd07dd4c29080954d4ca67c0e84cda632d08652d36efb362881ffe71

Observation b9955bfc-4bff-4a03-a99c-a5322c44cf4e · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.016320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.016320Z digest=sha256:005bd29cbed5f9c81ecf02ac427cc0e6d2a78ba6d877236bda1c797eb171af35

Observation eec755af-05c5-4871-b90c-f65d67021776 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.021100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.021100Z digest=sha256:b603fbbfd6084c6c0e57adc939bb94e65fd07f51188d2bbbe2c87ddf3db17e1d

Observation 2637f880-89d3-4bf3-9d63-9a6f15a0c72a · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.025426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.025426Z digest=sha256:849fac7a669c73e49e4fb5d3b5992da6e3f11bc21b2f34f4efe3bc3cbb5d63c8

Observation d09d9d59-0121-4557-a18b-4f60abad04df · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.029916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.029916Z digest=sha256:f258b834d899056039f4531b8cbf7b6d0808e6a9b5508bd793f4a747f4aa069c

Observation 47906f3b-9956-47f1-93f2-5695d6760804 · outbound

This paper cites TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.034210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.034210Z digest=sha256:01967d8f5480073cae08ac02138ce73daf54888af8f54d01fefa56d8884a604d

Observation c8dd303b-553f-4400-907f-4116e1dbe800 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation,.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Imagereward: Learning and evaluating human preferences for text-to-image generation,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.038683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.038683Z digest=sha256:6184aeebc37444640b78737dbd202e6582e21764be6533cbfa55ff34465d70d2

Observation 228f90ae-71d0-4972-859f-eaf95c1c2ce2 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.048292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.048292Z digest=sha256:2a5077845ed0579ed1f3d11221b8137e7bdcd6a3df88e6cc9569a33f526ca6bd

Observation e2fb8539-93e2-4302-a281-4b762bf26f5c · outbound

This paper cites Dream 7b, 2025.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Dream 7b, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.052671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.052671Z digest=sha256:9ff7a5b63fe2ce944a36f9be2b0e675015d1724532cf5b7dc033e0964e87dca4

Observation 517edfec-b44d-4c4b-980b-8c2136a0a728 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:32:32.831695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:32:32.056762Z digest=sha256:e12da7b2cf43512168119fa5b32f88c757796aff73778e879d124a9ae07e6114

Observation 1a6e184c-5fbd-4ef6-8ca3-597789589df7 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.061072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.061072Z digest=sha256:1c9e02a488d2f450d644f451978c8f4a9bf44fe600d7ee4f973eae7777713d72

Observation f00eb8ad-2912-44e4-af1c-794378240ca5 · outbound

This paper cites Magvit: Masked generative video transformer.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Magvit: Masked generative video transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.065132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.065132Z digest=sha256:51f13ffe15c17ae839a22be4f840f2f75040c5caeb9b1bc08e40c961fcacf588

Observation f29aac15-1bc5-49c4-bd81-1bf84505ecd2 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.069403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.069403Z digest=sha256:11b3b2913589f37cf0c1b9ffd049ff7f62cc412bafbdf2f49d6346958d22e46e

Observation cdafda51-e48f-43b0-8445-fe23c13e09e6 · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.073558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.073558Z digest=sha256:c8e63db756398988735daa03b7d9f976bdf61b23f7d0ca3927527b48b038df07

Observation 2107e3c7-cfcc-4997-88dd-2b6d3c25a642 · outbound

This paper cites Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.077872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.077872Z digest=sha256:d8fd0465c46f3e5698b40a6b704764bd5a89a1e32e6574f61389ec09268eed6c

Observation d2777fa4-c676-4236-b850-1174e4de9cdd · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.082440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.082440Z digest=sha256:4a12b591c9951362a9c018e209d8d29b50584b5a75ba7c95d0a165a396528c9a

Observation 5a606780-af7b-4afa-af80-0f12f843693b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.086866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.086866Z digest=sha256:c3cdef21443577dc69b2af154711594b0113f7fcc30c126003fe3820fc597840

Observation 22b36fa5-7f81-4cef-a74a-5cb88c287ac6 · outbound

This paper cites LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.091286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.091286Z digest=sha256:926d289b5fc28fef0769fc02e615f01c0d39988b659ea370a0446b5c5a050518

Observation 5b85435b-31ba-42f6-9292-a343ba9f3741 · outbound

This paper cites ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.042993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.042993Z digest=sha256:708f65a18ae6a95ecb2ce9e0d8451dc3ef87681d6e23941948b0afee75b9eb83

Pith citing papers

Observation 6c1abd88-de48-4b76-9c4e-e768c8a03ddf · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:42:54.637399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:658ea3f58600d8b14f4ddc9128d88327bbc4cc0ee6a110eaeb02d327f927cc14

Observation 27cb4660-3136-4758-bd6b-81c5094cc310 · inbound

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning cites this paper.

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:42:57.108742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T14:42:56.565621Z digest=sha256:cb192f4c8374fedf60b1dee52457ddcec335564a210fe10f67061813ee12bcac

Observation c40071c7-5406-4dd1-aded-3fdf24d131b2 · inbound

Adversarial Concept Distillation for One-Step Diffusion Personalization cites this paper.

Adversarial Concept Distillation for One-Step Diffusion Personalization UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:50:53.697000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:50:01.364000Z digest=sha256:3f55f72732c2e7695dd454440e987ec5b0c89828d23710e4414026a413aa71d2