Pith. sign in

Paper Citation Record · LEDGER

Jodi: Unification of Visual Generation and Understanding via Joint Modeling

As of 8 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 3 inbound Pith citation observations for arXiv:2505.19084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19084 v1

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:22.585226Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:38:39.360472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:44:19.355063Z

Reference resolution

100 of 121 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d01027f-d400-4688-8fad-0eb0ba28e6a8 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.023881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.023881Z digest=sha256:84c1225844d7f01a61983ec5140456d123855f2d84abf18118d1a545042e4813

Observation fd6e124a-d999-445a-a6bc-8e4ceb437b5e · outbound

This paper cites Building normalizing flows with stochastic inter- polants.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Building normalizing flows with stochastic inter- polants

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.085409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.085409Z digest=sha256:81b1e847cfd1e33fcbfad28b34ba9cf2189908353b4c09e55d2082f83056b241

Observation a5191bef-b846-4461-8686-53a8bd4b77f4 · outbound

This paper cites SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:25:23.226183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:20.166543Z digest=sha256:b35ec2b78882c844ae20aeb25149489fca1e22a82147d7af419c95a262213497

Observation 15bd8a39-1ea5-4cfb-934a-7ffc4a46fc79 · outbound

This paper cites Contour detection and hierarchical image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Contour detection and hierarchical image segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.269453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.269453Z digest=sha256:6b4279151e99d41a1ef7ec6f5447d3bb6c8fe3dd8510f72f12a5127fdd3fc098

Observation 0f7ea959-4290-480d-9d5c-248382f79c16 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.356553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.356553Z digest=sha256:be651c6c70096b1973b74fe9ef1203f8842002b5e46529a913e42600e7b0d638

Observation 502e35c8-1da1-43ee-98b4-f909b3f00b77 · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling One transformer fits all distributions in multi-modal diffusion at scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.438427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.438427Z digest=sha256:0316a0e19eee070e710b32598f306ef4b37c55db66cddd1118aa7c0145682ca4

Observation 79e3745c-1ebb-4f90-985a-310525f35c9b · outbound

This paper cites an unresolved cited work.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.520178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.520178Z digest=sha256:d98e200a2d1de19e2e644076ee3f00b6f376be47090edea863e077aaf7ecebd4

Observation 8e387625-7020-4d8a-a362-9e4e3bb53d7e · outbound

This paper cites an unresolved cited work.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.740440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.740440Z digest=sha256:81a87121f19d96ab1c4d604dd831c8a16ff0c879cb43ff03f8b3c04d3accc81d

Observation d1a859c1-ac95-496e-b17f-eeb7c2fcfb13 · outbound

This paper cites Intrinsic image decomposition via ordinal shading.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Intrinsic image decomposition via ordinal shading

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:20.900857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:20.900857Z digest=sha256:1b46f06daa1e5c6accf41a8132bfa803475a57e8802f45d0e85f9ee774a20997

Observation ebba4899-bf1a-4efb-b1db-2d78d61b6e55 · outbound

This paper cites Colorful diffuse intrinsic image decomposition in the wild.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Colorful diffuse intrinsic image decomposition in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.067772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.067772Z digest=sha256:b90b23a863a393c180e8717638cc97ed77341eb8ffea50934c9ba61f384de7be

Observation 1b9a9ad4-5879-41ef-aeee-ce67b278bf62 · outbound

This paper cites End-to-end object detection with transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling End-to-end object detection with transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.233841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.233841Z digest=sha256:54c075515ccd3d3d232831252dbb0c7e96df249422bb11a3feaeedc7c0004358

Observation a1fa9929-cb5f-498f-86c0-c5d7d6dec632 · outbound

This paper cites Artists as experts in visual cognition: An update.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Artists as experts in visual cognition: An update

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.361271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.361271Z digest=sha256:6c3a34166c57c7d7afa116a818e77159279dc579dce2895ad3b3abc883097c9d

Observation d45b59d7-5999-4ec0-8189-892cbe76d218 · outbound

This paper cites Learning to generate line drawings that convey ge- ometry and semantics.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Learning to generate line drawings that convey ge- ometry and semantics

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.398707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.398707Z digest=sha256:2bc7df94dcb0c03d0a5106b32c85ad30bc56f40340354f7ef01f2b465dece7f1

Observation e627f4f8-64db-4e6e-9475-432074c64545 · outbound

This paper cites Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.463272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.463272Z digest=sha256:f1ea98b285257adc4a1ab1e318e0a6a82791aa93a89f0be5858a18743949943e

Observation d218a445-6150-48c6-963b-e6f49d8505a7 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Deep compression autoencoder for efficient high-resolution diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.609686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.609686Z digest=sha256:d6f5b18635d066f82abaf7cb184b2c0a616d82532b4b4b96cfea45132f4d0d80

Observation 456e047a-67e1-4267-b9d9-fd4357d45170 · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Anydoor: Zero-shot object-level image customization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.659977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.659977Z digest=sha256:62deae299a517a10c50926142092da9e7e0ed59fd2ac96073467584dd5bf4529

Observation ab528964-950d-4149-ab96-c82f5e05e0cd · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.664689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.664689Z digest=sha256:0106b59ca3163cc6d7de7ed3d5009ac5c00e2bdf748d262038ae567bfecdec0f

Observation 2d08a53e-606c-4d05-b200-576eb63a0664 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.733104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.733104Z digest=sha256:ea9540444d2e5b7cfe6fcc4a60e4f708ab905b678857197c3632907f0437ec17

Observation 3004283c-7d1e-4f43-ad3c-f69b81981ff7 · outbound

This paper cites Idadapter: Learning mixed features for tuning-free personalization of text-to-image models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Idadapter: Learning mixed features for tuning-free personalization of text-to-image models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.789469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.789469Z digest=sha256:cc3e22148551b88cf5eafb0176e9e7692be42b90f8774f80e5a31a0cd2fde4a4

Observation c334c49e-2a5d-4c1f-ab37-8a2fecaaac47 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.869510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.869510Z digest=sha256:a77aea3a23f3876e0eba11a9c47511de82e048c1c76410974962d2205b127369

Observation 742df795-4d6a-46ce-9df0-3bbc4911245e · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:21.956584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:21.956584Z digest=sha256:90d67508e36e218ea2039853bc6989da8e8b1cb5fb23d54598abb0286abe853e

Observation c9a7aaa5-c1c9-45d0-bccc-974431354446 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Imagenet: A large-scale hierarchical image database

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.030338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.030338Z digest=sha256:e863a967fe96b075742c49119bdcb157d2b8ffe070539f64e8115e161c5669d4

Observation 19e0a22e-b91e-4bc0-8275-974ff4701e59 · outbound

This paper cites NICE: Non-linear Independent Components Estimation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling NICE: Non-linear Independent Components Estimation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.092599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.092599Z digest=sha256:ef27ce05af4d2960aa3f7911dc484ce7a3d6a78a4f053683b38fd9f4cf2b8b8a

Observation e1d25dda-9a3f-4f47-bbb3-9cb702117117 · outbound

This paper cites Evaluative and generative modes of thought during the creative process.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Evaluative and generative modes of thought during the creative process

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.177856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.177856Z digest=sha256:18098bfe10a22de09b9acc33b79fc816e0cb04bd1c154e6d6e0e56bc389bd02c

Observation 451cda10-37cc-4624-ab4b-2565953a710a · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Scaling rectified flow transformers for high-resolution image synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.229147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.229147Z digest=sha256:cf30874999515dfe5b308434f048617b49a60f163002098173c1fca8b8ff636a

Observation c300ea3e-c3eb-4861-946e-6dbe3e3c14ac · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Taming transformers for high-resolution image synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.233468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.233468Z digest=sha256:c53c4424aa23bc217a5b354e4869e2750be8175982fb79a26b2dd9f87926f255

Observation 919ea3c9-db4c-4b48-ab61-07804eda1506 · outbound

This paper cites The surprisingly powerful influence of drawing on memory.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling The surprisingly powerful influence of drawing on memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.238046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.238046Z digest=sha256:01e47b31700215ade0183d3ee7b7869b7c4bffed21f6ebe2b01bc3e6f7bd3368

Observation d2a5110c-c8d9-4efa-bc82-2ef113613f0a · outbound

This paper cites UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:25:23.130205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.241809Z digest=sha256:0a72230f7324dcc8e3ee2f1009661c414e0a11ef84a4243e2f4f426be1c3c8db

Observation f2637597-011e-4ff8-a947-2307739055e9 · outbound

This paper cites Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.247783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.247783Z digest=sha256:099822ee07dca73b08c274ea51e084b5c0c22933671b94e1483c7ced6e95c693

Observation 4cfd5468-25f0-4cf3-9061-b3120997a6ab · outbound

This paper cites pexels-portrait.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling pexels-portrait

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.251756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.251756Z digest=sha256:f1886e8bd8137b0e38da6dcaa19eeb1236d8c9938f296968a059cf5f4466d0fc

Observation b9425971-ce73-4695-b7c9-2666ab61252d · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.256655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.256655Z digest=sha256:1c88ae65ece9c2e09d2458f9327988fc4dfaf967056e6ebb1ccbf2e65fa14d43

Observation 82a39014-5b84-4f2a-b501-b3ef600bbcc1 · outbound

This paper cites Generative adversarial nets.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Generative adversarial nets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.261526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.261526Z digest=sha256:4c5622483f9fd7b358fced9cf71452c7531c975be5fe1b47bd87af5f313525cf

Observation b2cdb417-a08b-4101-a158-e3e32d41f347 · outbound

This paper cites Pulid: Pure and lightning id customization via contrastive alignment.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pulid: Pure and lightning id customization via contrastive alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.266396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.266396Z digest=sha256:67795ad2e1dcd97f6ade2be97c3e9386f00f31693eaba0f8e8809ad5591b38cd

Observation 26c8e1b9-7923-4a9b-a678-10a051576392 · outbound

This paper cites Svdiff: Compact parameter space for diffusion fine-tuning.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Svdiff: Compact parameter space for diffusion fine-tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.270883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.270883Z digest=sha256:d4967d4e3f15445d31833a47adb377c93c8dfc05a008fdf2f4493fbb02e3f5d9

Observation cfdaebcf-74cb-4790-8fa6-2bfa9d0d0e4b · outbound

This paper cites Transformer language models without positional encodings still learn positional information.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Transformer language models without positional encodings still learn positional information

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.276191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.276191Z digest=sha256:97b85802e95e734bfeff6a1f5bf7d962d30043c6ba081f7c54d43c8c5f957d35

Observation cc1ec785-7b50-4e25-80c2-e7ce58fe012f · outbound

This paper cites Lotus: Diffusion-based visual foundation model for high-quality dense prediction.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Lotus: Diffusion-based visual foundation model for high-quality dense prediction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.280980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.280980Z digest=sha256:0cd8690aa67f0998d1886e40d8f1fa53c5f537e9c29264f1ee7038aa44758e4e

Observation d36af7db-ff5e-4872-b15b-6a7a128f33d8 · outbound

This paper cites Mask r-cnn.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Mask r-cnn

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.285397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.285397Z digest=sha256:d510f67433d4d1e66a10de1744ca1e522821a4fe7628dd9d8ed99d3adf1263a1

Observation dc5859f8-72e4-4d0f-92e7-42cf56f6af5f · outbound

This paper cites Deep residual learning for image recognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Deep residual learning for image recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.289875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.289875Z digest=sha256:38937b7de3533f0d644e93e6664a9b7c4d2c224dc31e9f1741718526d31d7c75

Observation 7e05d896-e9fe-4625-9376-779459bde742 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.294337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.294337Z digest=sha256:9d3ff19bf63d60c56b61027c05245ce11aaf1ed664fd1bbabe8b562e4298c2ab

Observation 0a924934-76a8-44cb-a44c-61e7dae76d92 · outbound

This paper cites Denoising diffusion probabilistic models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Denoising diffusion probabilistic models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.299329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.299329Z digest=sha256:d2abe208bf78793343bb26f345d7f89ed8cf116bdab6b1c73a4e6d33c5278614

Observation 588a842f-3af9-44e0-a282-3f7e67eeb00e · outbound

This paper cites Classifier-Free Diffusion Guidance.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Classifier-Free Diffusion Guidance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.303298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.303298Z digest=sha256:f9e1f30dca765c567c30bd0e87c6ebdc711113d7dc9cd33a4d75fcb5cec95f8c

Observation b920f5c5-b7ba-4a81-a2df-94ea4fcba894 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Oneformer: One transformer to rule universal image segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.307733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.307733Z digest=sha256:49fd86f0735cb0d15446e3d956c063c97117ae4d19204a054c4183be43771866

Observation 5c156863-3be1-423c-b09a-805b9a3a4ea6 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling A style-based generator architecture for generative adversarial networks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.312506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.312506Z digest=sha256:92318c9602b70682860bb60fd0f510d414c7853bb73453580dbcc0c31522e23b

Observation 2a57b08c-6051-4221-922b-f2714d0086aa · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.316794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.316794Z digest=sha256:4aa4edf8e56518f94e0ff751088c9140312cd8944b36514dad61625616540787

Observation 00bf20dc-1533-48c0-a6f9-3f3d7b491d0f · outbound

This paper cites The impact of positional encoding on length generalization in transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling The impact of positional encoding on length generalization in transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.321110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.321110Z digest=sha256:d73f2646e02d8438661f789477194ef8484e51c34f3e77e8a725f17dab9a5695

Observation afd4cb01-edd3-4521-a8dc-75d8958e415e · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Repurposing diffusion-based image generators for monocular depth estimation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.325277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.325277Z digest=sha256:8ed6f003d8630f2a121149d154aa001af5e7c9947696bb161f373fd8c9f4ba4e

Observation e274ca9c-b875-43ec-92fd-4143dcfacc17 · outbound

This paper cites Auto-Encoding Variational Bayes.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Auto-Encoding Variational Bayes

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.329422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.329422Z digest=sha256:ae74a740a04499685d91ec1584946ac6e68e515c405c427e15bba970afdf3842

Observation 811f385e-1963-4c3e-bc59-606dd577116e · outbound

This paper cites Segment anything.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Segment anything

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.333436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.333436Z digest=sha256:1b7db5a811f173fd070371590426ead983952632ff35ba46780833800ea84382

Observation b6ad9372-3893-4659-8528-2b773675776b · outbound

This paper cites Evaluation of cnn-based single- image depth estimation methods.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Evaluation of cnn-based single- image depth estimation methods

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.337682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.337682Z digest=sha256:cb62b0d4e6ae396b53a7a1bcb909b5ad64b0b84f72e4f28ab427b0d19848364c

Observation 238b022c-551c-4b10-ac53-bdc8f0b4133f · outbound

This paper cites Intrinsic image diffusion for indoor single- view material estimation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Intrinsic image diffusion for indoor single- view material estimation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.342567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.342567Z digest=sha256:b1b4bbf5fdeee9ebbae91b74bcf8699a2d622e8d78e55df0b14b08045cb2239e

Observation c7917e06-628a-4d5b-b870-84caef2c7d8e · outbound

This paper cites Artists as experts in visual cognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Artists as experts in visual cognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.995509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.347021Z digest=sha256:60e281a3126889dbe5ac5489464876e1145f7ab5fb5b33b182b07c2e4d2de1a2

Observation d6f11f3c-cfe9-4b85-ae82-90d39917510c · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Imagenet classification with deep convolutional neural networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.351696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.351696Z digest=sha256:1ecf452a30207d406c668fbf02e8d7b5fcc2ccf71900eb255afb574d7ff981d1

Observation a3265254-8ec6-4d35-8701-b11f459d5907 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Multi-concept customization of text-to-image diffusion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.970690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.355703Z digest=sha256:6a30a41c8a9774d6de10cfa864d172f39344c23a969898a9da1d6e88ab8b870c

Observation b1704c97-8bcd-42c0-9e37-2153ce3b547e · outbound

This paper cites One Diffusion to Generate Them All.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling One Diffusion to Generate Them All

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.360463Z digest=sha256:8daa328f69faaa216786f646fbff2840afeed4957cd21bb7896c8877a0722707

Observation aafa071c-ad5f-438b-a2c7-8344e83d445e · outbound

This paper cites Gradient-based learning applied to document recognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Gradient-based learning applied to document recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.364853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.364853Z digest=sha256:d587760fc4c06a317fd0dfcafc44820f9f53aaef3b0271bb8bcc0fd90fb0380e

Observation d1d375e9-e3f1-4e95-a7d8-5a16e3ce944e · outbound

This paper cites Exploiting diffusion prior for generalizable dense prediction.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Exploiting diffusion prior for generalizable dense prediction

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.946089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.368903Z digest=sha256:25a16c0134965c6a69dc2087e4453e31305fa83da42bc6b1630fb6069b092fa1

Observation 507d85d1-50d0-4329-8c30-dc4717ec4569 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.373168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.373168Z digest=sha256:249cdd3d4300116d2b3f529dc093965aa02280a0e93f54989a87af0a9cbe399c

Observation 0e19f149-40c9-4c0f-a95f-f08d3513ec3d · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.378612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.378612Z digest=sha256:e507fcaf2492a7330d2e92641fedda6b88c39f741be5aa6aef06b3a1681c0d55

Observation e493cd5d-fe29-4316-8fac-c760f4498989 · outbound

This paper cites Uniformer: Unified transformer for efficient spatial-temporal representation learning.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Uniformer: Unified transformer for efficient spatial-temporal representation learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.919950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.382933Z digest=sha256:aec18182dcb54922abf7508e788ac92c4dcb231d9d8b81b5894bcdf8bf761a93

Observation b66de16d-e7a1-485d-8e67-33dcedf2e603 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Gligen: Open-set grounded text-to-image generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.386866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.386866Z digest=sha256:45b249168446ecae430b2d5a8713c9b3d35d07891a986d6b16ddf230c47a99bb

Observation 46c467f9-78f9-418d-bd47-fdb1a38194df · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Photomaker: Customizing realistic human photos via stacked id embedding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.390909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.390909Z digest=sha256:a05b72005ed4540d364cefc3a255443c4c65f2c720b7aeb83b642c66df47654d

Observation ba11db8f-ae42-4811-9e68-11470bdeffb9 · outbound

This paper cites Dual Diffusion for Unified Image Generation and Understanding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Dual Diffusion for Unified Image Generation and Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.394934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.394934Z digest=sha256:5080ce9be8bd036cd51f67e7500d25b915d7f0cd47575ccbed5903251095c837

Observation 3caba455-76f1-4bac-ae07-9f4787975fa9 · outbound

This paper cites Pixwizard: Versatile image-to-image visual assistant with open-language instructions.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixwizard: Versatile image-to-image visual assistant with open-language instructions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.880370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.399786Z digest=sha256:c7344b2bbe963a8da5cf3bc779c3924509bd70555ac65d1e0e657df00ceb7b34

Observation 27f252d0-8c15-4085-8ccb-812893f5a6b6 · outbound

This paper cites an unresolved cited work.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.406222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.406222Z digest=sha256:5ea78645483c5a83806cca063fa5f6a3cfae0ddc42b1caaadfae2e5512ec799f

Observation 5de40839-52c6-484e-92d3-ea71107275ed · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.410437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.410437Z digest=sha256:a4d675c0078e23bef1852798a04c36674baf89fe0347236eed50577c176fdd4c

Observation 9f9e7c7d-a976-4d9d-b215-9c9a3bdcf9d6 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.415198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.415198Z digest=sha256:36d30d8a81cfe7b285443e791a0299c17c2d64e06fd044a5007243201fd64d03

Observation 91527882-fa98-463f-aa4b-9ab807cdd40c · outbound

This paper cites Came: Confidence- guided adaptive memory efficient optimization.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Came: Confidence- guided adaptive memory efficient optimization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.839289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.419555Z digest=sha256:f98701de46e42334ba846283285efff878619d4696227c91ae6e87fdc6fd601f

Observation 6950d879-870c-4d4c-88fe-fa61443f0703 · outbound

This paper cites T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.823402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.424091Z digest=sha256:5c421462f2017a5da04e164bbfa0ad5db60c2fd01f8d992245c3f1e9ac01d3a1

Observation 1aec3c3b-0be3-4a5a-8f91-3c1392630109 · outbound

This paper cites Novelai improvements on stable diffusion, 2022.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Novelai improvements on stable diffusion, 2022

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.806790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.428891Z digest=sha256:75e534b50380e1ffab16bb1327dfdf146ab95441cbea13042ee4a6ef1327deed

Observation 8253015f-aebc-4a08-ab85-6d855b7596ed · outbound

This paper cites pexels-photos-janpf.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling pexels-photos-janpf

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.790583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.433006Z digest=sha256:129bf84aa144b50f397d9dc2b2acf3814178991649dc094d38f636ab89bc808e

Observation 56a94a99-f0eb-4314-905d-95b0d1394218 · outbound

This paper cites Scalable diffusion models with transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Scalable diffusion models with transformers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.437358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.437358Z digest=sha256:695f9be6900a4a9ceff1e2cbda085ef40f03f00d97666aa8162decf265908184

Observation 6a4b3458-4594-4c27-9982-d82d246bb741 · outbound

This paper cites Ld-znet: A latent diffusion approach for text-based image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Ld-znet: A latent diffusion approach for text-based image segmentation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.764838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.441752Z digest=sha256:b9a01a54222637c79af96e56e9186d45455de5f3194927707ce8575816a0f83d

Observation 536825bf-274c-4163-9677-33421549d7ab · outbound

This paper cites Unicontrol: A unified diffusion model for controllable visual generation in the wild.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Unicontrol: A unified diffusion model for controllable visual generation in the wild

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.749161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.445661Z digest=sha256:ff730119f79213219bce45d70dad1f723582a8f61f4697eb7f309189e5bb986e

Observation d51d9076-fb7b-4258-ab6c-d6be16c01cc8 · outbound

This paper cites Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.450236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.450236Z digest=sha256:0f11bd7d65e60d1a550a24d0f5088314ba053db1738fdd892314d33c829ced23

Observation 705116ba-f314-4f5a-ad56-5f5e5e3dcaae · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling High-resolution image synthesis with latent diffusion models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.725331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.454569Z digest=sha256:b8993a9b0fdafcb125f592d3eb7f70c24fe0e3ca633220fd90f5774cb35729eb

Observation ad4a8edb-860c-40a0-8a2e-2add47940beb · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling U-net: Convolutional networks for biomedical image segmentation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.459677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.459677Z digest=sha256:d679939fec9aad359c51395de4f62d0a1bc6699c88abdade94fb6aff159d2035

Observation 24b5f3a9-af8b-492a-9e43-d8bbef430a82 · outbound

This paper cites Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.464421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.464421Z digest=sha256:74a1f5cc3ba860b37dd705ea30612bc555a7b1a09c2d30ae55dccdfe27a6fee6

Observation a3a4b2a4-2ab8-4b2d-b88c-b7220352a623 · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Photorealistic text-to- image diffusion models with deep language understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.469175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.469175Z digest=sha256:67234bab6f72ef33627db0c4ee1cb5dbbf4a33c5c7642216d6dec2cadf28ba25

Observation 2a99b0ee-69d2-4ef3-a72d-e31ed09614f8 · outbound

This paper cites Ziplora: Any subject in any style by effectively merging loras.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Ziplora: Any subject in any style by effectively merging loras

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.474251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.474251Z digest=sha256:2f10e988a72ca2e9587e35393c89e4924ca2d380c6a802f0a28e7c4e512b0fe9

Observation 1337a926-c7a1-421f-9c38-17ae3e0d4cbb · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Indoor segmentation and support inference from rgbd images

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.478989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.478989Z digest=sha256:aff0aa98ce6f0ed83641bc53e2156aa4a18d80f2e6097bc8c88e2390e2a214de

Observation 33529430-2f5e-499e-9e35-4a05e2ef5d84 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.483167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.483167Z digest=sha256:f3e0b6d88cbc18dfd45c80c343cd9c5e9ca447f2a3ebbaa5db9aa6f26053cea0

Observation 309b4ac6-122a-4401-8f2d-fbbc8b0a387f · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Deep unsupervised learning using nonequilibrium thermodynamics

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.487987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.487987Z digest=sha256:17484e7cd11b9644cb685318961d271fc8f26c628ec5c41b588d10ab862b31d3

Observation 9223c140-660f-4f3c-bcc3-1aefd4a45fff · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Generative modeling by estimating gradients of the data distribution

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.492170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.492170Z digest=sha256:8cb07be90d5388dcdc25cdc708bc0885cd8bf85f80304597b2900ce52b876faf

Observation 4fe99b11-47c5-45fe-b27f-3e25421545bc · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Score-based generative modeling through stochastic differential equations

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.497026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.497026Z digest=sha256:1097bf8b69e6869722e7a8357b8aebfd5d851a8fe7a7073dcd44d71e55dde4fb

Observation d110e3d8-a2b7-49a3-93a2-764e51034878 · outbound

This paper cites Pixel difference networks for efficient edge detection.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixel difference networks for efficient edge detection

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.625294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.501389Z digest=sha256:aed575e7cafb7c92807d1091bdfa2692004fd5e44194b93590e1adeb21b15266

Observation 9ad2a344-228e-47eb-8db8-265851a07630 · outbound

This paper cites Going deeper with convolutions.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Going deeper with convolutions

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.506402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.506402Z digest=sha256:cae161fd0b6dcaa5067ae9c931a71f39f230c632d9546717eb3ccb56fc79c0eb

Observation 4d27d7ac-3902-4056-abd8-684fa6c21255 · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.511560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.511560Z digest=sha256:3e79dce71af395cd1fb03640d999b2ac81472ded1c547e21677f6e9f21a92276

Observation 579ed44c-06ed-4cfa-8bf3-9eb1ed1e39e9 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.516710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.516710Z digest=sha256:e75f534f395c7dc5956c6b7e38bfe4be468525711fd84268fe6663ce86f3bf45

Observation 308e4df1-654c-4046-91a0-dc895d6fc5e7 · outbound

This paper cites Pixel recurrent neural networks.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Pixel recurrent neural networks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.521960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.521960Z digest=sha256:44756753ba44edc92f31b3a2ef8d05d0d6573c0898981b9a852ec9a6c0bc184c

Observation 704e821b-f21d-4751-ac9a-61d05220b4e9 · outbound

This paper cites DIODE: A Dense Indoor and Outdoor DEpth Dataset.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling DIODE: A Dense Indoor and Outdoor DEpth Dataset

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.528715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.528715Z digest=sha256:6317738658fca1be557f6c53d49cf24654f1c9b3eeea18f629d831a959a8ca35

Observation b2f1a476-2042-4334-832d-01582ae434cc · outbound

This paper cites Attention is all you need.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Attention is all you need

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.534394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.534394Z digest=sha256:a358d49fe9f953a0a9bde74839896213ee70bea8ca89f2d91d51112b86a0d319

Observation a9fa7bfd-06dc-4760-afd8-4232533b08e8 · outbound

This paper cites MMGen: Unified Multi-modal Image Generation and Understanding in One Go.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling MMGen: Unified Multi-modal Image Generation and Understanding in One Go

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.539339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.539339Z digest=sha256:d4403eab366cb354e455e33e4294c551042113911c392aac85a76e3b750faad3

Observation 8b67ff78-b67f-4e1f-b116-53a3f2a43fe5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.545579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.545579Z digest=sha256:6bdbc51013f13d63773a03b6e536679d53c2b3fc8e752a2f8743b148711bc026

Observation 6dabafcd-fe87-4a07-88d0-f7ce63c883d1 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.551869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.551869Z digest=sha256:dacbebcedea01a4c4cbaef52fce6d38f4426c309d30a5961a374eac8309f07d2

Observation 28729510-947b-4d84-9e2f-ba67af8485af · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Emu3: Next-Token Prediction is All You Need

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.556883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.556883Z digest=sha256:101a351928f2c19e0530646e1f0f651de9f0cdd741b631408260be88a4663271

Observation eed5e9f7-4f3d-49e0-890f-3445f1aa2a37 · outbound

This paper cites VILA-u: a unified foundation model integrating visual understanding and generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling VILA-u: a unified foundation model integrating visual understanding and generation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.577141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.562233Z digest=sha256:4c442026250ec366ee82f8287137873487c72170de0f74cb4be1c7c1f8732f5a

Observation 2ef3d780-38c4-4624-995f-ee8b35e92789 · outbound

This paper cites Infinite-id: Identity-preserved person- alization via id-semantics decoupling paradigm.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Infinite-id: Identity-preserved person- alization via id-semantics decoupling paradigm

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.559406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.568323Z digest=sha256:030dae451a6a69131db7bebaf4abef2dcfa6dfa383d554f1bc9364282780672d

Observation 69893bf6-8d16-49ce-b7c5-8e94aad087d7 · outbound

This paper cites OmniGen: Unified Image Generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling OmniGen: Unified Image Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.573243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.573243Z digest=sha256:0840968869e1bc985d92f7790a126a7b08a241b17fa1effe14bf9282a9d092d1

Observation 2978706d-23ba-4e0d-b4dd-3c23fee37374 · outbound

This paper cites SANA: Efficient high-resolution text-to-image synthesis with linear diffusion transformers.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling SANA: Efficient high-resolution text-to-image synthesis with linear diffusion transformers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.579652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.579652Z digest=sha256:7e97cd433124d35f0c421208729e1797675788411f433c026622f6c33b0da2e4

Observation d8910e70-353f-4bd1-b693-250a93b310ce · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Show-o: One single transformer to unify multimodal understanding and generation

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:23.533825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:22.585226Z digest=sha256:eff8b049d2c6a118342e69bfafc24f07435e6b98f4bd825725c3805c002c0763

Pith citing papers

Observation bd0bfbbc-735c-4fab-8c87-1ea894e2b05c · inbound

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors cites this paper.

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:07.834476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T20:05:21.723724Z digest=sha256:33188102415314ec504880d9c4b4d1cf869844ee7d98531ab2b6e209da617d86

Observation 31191285-971e-4c0e-862f-aad9e6c2fdd3 · inbound

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study cites this paper.

Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:25.664198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:53:06.362430Z digest=sha256:03a3a5c43d8b7a0f35b93020321ebacd84d6c6099b29eb84676d24d5baab9f38

Observation fbb715a8-beff-41b8-b7c4-03b22b9d9b89 · inbound

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception cites this paper.

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception Jodi: Unification of Visual Generation and Understanding via Joint Modeling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.357185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:38:39.360472Z digest=sha256:3f6242cf868071032f762833f792fa6d2e4d6e2501f27d1a2afd04107fa4ad30