Pith. sign in

Paper Citation Record · LEDGER

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers

As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2505.00482.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00482 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:47:32.086124Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e20639d2-b967-495e-8357-f167766ded85 · outbound

This paper cites Depthformer: Multi- scale vision transformer for monocular depth estimation with global local information fusion.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Depthformer: Multi- scale vision transformer for monocular depth estimation with global local information fusion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.359818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.740489Z digest=sha256:dc4254bc3d0d6d0dad45b701133a7c6b08fefdade2098ddde92141c8c2cc7d43

Observation 638bf7f6-2956-4ca9-bfa5-26245c5c6d99 · outbound

This paper cites Multimae: Multi-modal multi-task masked autoen- coders.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Multimae: Multi-modal multi-task masked autoen- coders

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.343791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.745882Z digest=sha256:b51d0d90e57860ad5022dfc2a63b48db85662bb85c62740610c34c81d8cafc3b

Observation 539feb88-6b35-41e6-8ce0-4251eec86a4b · outbound

This paper cites 4m-21: An any-to-any vision model for tens of tasks and modalities.Advances in Neural Infor- mation Processing Systems, 37:61872–61911, 2024.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers 4m-21: An any-to-any vision model for tens of tasks and modalities.Advances in Neural Infor- mation Processing Systems, 37:61872–61911, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.327777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.751082Z digest=sha256:87e69412bf1c69c56928d97ba3e4cb4a02be72a3888529135fece300c5b4342f

Observation cfb5ebd8-5431-4fb5-b5a0-49ba2da2c7fd · outbound

This paper cites Multidiffusion: Fusing diffusion paths for controlled image generation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Multidiffusion: Fusing diffusion paths for controlled image generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.311923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.756124Z digest=sha256:0551a1be489eb51b728c174479e8cece71aaae9b06c54a8b28a90621bc8e599e

Observation 3bac5b32-4eb9-4221-bf1c-6d950f95d84e · outbound

This paper cites Se- mantickitti: A dataset for semantic scene understanding of lidar sequences.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Se- mantickitti: A dataset for semantic scene understanding of lidar sequences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.761070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.761070Z digest=sha256:13ae645a6d5487da9ed22becb0c38b4f2aec067dc9a5adcdb8849ee41f8a4e52

Observation 84c0527e-61dc-4832-96bb-ae5d9348171d · outbound

This paper cites Loosec- ontrol: Lifting controlnet for generalized depth conditioning.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Loosec- ontrol: Lifting controlnet for generalized depth conditioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.765916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.765916Z digest=sha256:fc086f83f6468fb14248fbd581c9cb48b82c712240d7ba0dd1e0c3bee9129833

Observation 277bc16c-9d08-438e-b729-c3ddcb67f6da · outbound

This paper cites Flux.1.https://huggingface.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Flux.1.https://huggingface

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.271520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.770700Z digest=sha256:10bda8a4afb127eb4563cda82f306154e530be959b2bf1387c2dca15eb1ac4e0

Observation f6408f9e-5835-4446-ac12-295d3e3fe99c · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.Advances in Neural Information Processing Systems, 37:24081–24125, 2025.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.Advances in Neural Information Processing Systems, 37:24081–24125, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.255156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.776031Z digest=sha256:6a29a3a05d411c9c913d252258661cbc92be32f37e201b37f96d96fc25bbc8b7

Observation a12b82ee-2136-43d4-999c-d8b68e4a8ee2 · outbound

This paper cites Text2tex: Text-driven tex- ture synthesis via diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Text2tex: Text-driven tex- ture synthesis via diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.238261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.780461Z digest=sha256:f2e894d2fb8d118de36599c863607c6a612b80cb62501317d2a1b8d603eacc51

Observation e3322d9e-bd64-4fc3-91e3-778378b88a1f · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.785181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.785181Z digest=sha256:5bc7ed0e466a7d5742d4dcc8f326b5f813c0835a1ceb70a04d1e184dbbcd25c7

Observation e424f2f6-1aae-42a9-a2ce-834c84b7f418 · outbound

This paper cites Neural ordinary differential equa- tions.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Neural ordinary differential equa- tions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.221679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.790267Z digest=sha256:b4960e1139d8007dd43eb8c38204cb9ea0cf1b99153c0947957ddd5db894516c

Observation 6ab98a38-4b5a-4ca4-8e11-d3c8ca25a96f · outbound

This paper cites Deep diffusion image prior for efficient ood adaptation in 3d inverse problems.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Deep diffusion image prior for efficient ood adaptation in 3d inverse problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.205553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.795385Z digest=sha256:c08b0845ff61644ff54862efd42ee74326a2e3d59c57fda424f2fd44ac95297d

Observation 5b0afb4f-3199-4fbb-b843-be69d0908e23 · outbound

This paper cites Improving diffusion models for inverse prob- lems using manifold constraints.Advances in Neural Infor- mation Processing Systems, 35:25683–25696, 2022.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Improving diffusion models for inverse prob- lems using manifold constraints.Advances in Neural Infor- mation Processing Systems, 35:25683–25696, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.188360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.800111Z digest=sha256:23fc457536ad26c7f83cde717cfcaa89e81d29c95e4c565dae1913debd46a494

Observation 44d0262f-e37b-4fe7-931f-c3215bc8e681 · outbound

This paper cites Solving 3d inverse problems us- ing pre-trained 2d diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Solving 3d inverse problems us- ing pre-trained 2d diffusion models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.804493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.804493Z digest=sha256:59be3cab43b7ca59f2eaaff37a69feab9e5713b76fbf505e89ad23b48fd78365

Observation 0a47916b-fd92-4230-8d66-d34143d1b414 · outbound

This paper cites Latentpaint: Image inpainting in latent space with diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Latentpaint: Image inpainting in latent space with diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.158873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.809715Z digest=sha256:f4082cdcc273876b74254ab71750edf942421f2ea9924758a7591df013e52ce6

Observation d25b22bf-6406-464c-ae7c-b8a6b27e4b97 · outbound

This paper cites DiffEdit: Diffusion-based semantic image editing with mask guidance.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers DiffEdit: Diffusion-based semantic image editing with mask guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.814438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.814438Z digest=sha256:81cd9d0b5c721ccff7eaf0141710da6bd8839bf4ab6c2664b1c24d8ee99c9e8b

Observation 7508e38b-4b78-4cd9-a750-98697494ab7b · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.142025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.819346Z digest=sha256:f8fc864baad0e2cc2ab04ec2306a5f11b6e19cf821b0f592c5a89d8f3b2ff421

Observation f77ecb42-ba28-4636-8774-acdf2170ad25 · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scaling vision transformers to 22 billion pa- rameters

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.124854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.823696Z digest=sha256:eee532fc894e3935dfe87c69b2f030550e5614ec8649159ff392eb44d0a3e5dc

Observation d6ddda39-23d1-4c91-8f80-b2fcab394fec · outbound

This paper cites Scaling rec- tified flow transformers for high-resolution image synthesis.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scaling rec- tified flow transformers for high-resolution image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.106749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.828210Z digest=sha256:f67037e46c30f284eefa1e4b74567391cd07b5ac2d2bf66b0d081be47c731aad

Observation 88ba2f57-a753-471b-9475-c7cb6b4f5e86 · outbound

This paper cites The pascal visual object classes (voc) challenge.International journal of computer vision, 88:303–338, 2010.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers The pascal visual object classes (voc) challenge.International journal of computer vision, 88:303–338, 2010

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.084975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.832964Z digest=sha256:4c7ad7bcae37dcabaf888aa4216ab0d4b0270a6567efe96c4929119c8a271f4d

Observation 184ad83e-704c-4dc8-8e76-83fd531d14c9 · outbound

This paper cites Geowiz- ard: Unleashing the diffusion priors for 3d geometry esti- mation from a single image.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Geowiz- ard: Unleashing the diffusion priors for 3d geometry esti- mation from a single image

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.069501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.837638Z digest=sha256:536d04144be558b157f7166b1ab501191c0fb5e45c87ecf21b39b0f7c5440a41

Observation 34108945-3372-47fa-9433-be9dffdcba9b · outbound

This paper cites Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Ommer.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Ommer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.053680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.842853Z digest=sha256:b272e3c7f1d0b9fe865347622f79e9a91b126bc31c5cc4943cfde1cc3eb53f10

Observation 70114f14-a7c6-4740-9e1e-5bf642aac3bb · outbound

This paper cites Efficient diffu- sion training via min-snr weighting strategy.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Efficient diffu- sion training via min-snr weighting strategy

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.038231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.847919Z digest=sha256:c4c1e4717d1c87afd6890588310fedda7bc0b550af78ce0eb2d1e21c25db80e3

Observation bab29a7a-832b-4f96-9616-516c93de7b41 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.022264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.852592Z digest=sha256:20936bdaed1ba1e2e6a01d4faac18c8c957eb3cbb6558435d23cc867b5a51f53

Observation b6b1a09d-5fe0-40f2-817d-6c5c76184e18 · outbound

This paper cites Classifier-Free Diffusion Guidance.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Classifier-Free Diffusion Guidance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.857329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.857329Z digest=sha256:fd930586b94e862196abb3584da8dc466b51745d3782a0ac2858e42d3fa284fc

Observation 7a2cdd4b-fcd0-4adf-b77c-fdc4440b7dd4 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.862188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.862188Z digest=sha256:4100ac9bbef702745abcc074ddc140c7412796f7d275e5c3e51e8e5d46aeb9f3

Observation f12d4fa1-d866-45ac-9bdd-a24740c0f68b · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.990748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.866706Z digest=sha256:ec3c86c4e984ff30c8df7686f98ae26fb0800c3a88f66103ceb023ece75b46db

Observation 2b127d2a-7d10-404e-bd23-f06a61b4a54e · outbound

This paper cites Zero-shot depth completion via test-time align- ment with affine-invariant depth prior.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Zero-shot depth completion via test-time align- ment with affine-invariant depth prior

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.972436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.871529Z digest=sha256:cdd585f38ce10f6d559e869d5230cefea4709c8c23dd6000ad02a98178178e6e

Observation 8d32c6f0-a73f-4804-b121-2370953719a9 · outbound

This paper cites Mixture of Diffusers for scene composition and high resolution image generation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Mixture of Diffusers for scene composition and high resolution image generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.875973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.875973Z digest=sha256:5ba85d1ab2daa63e41c70aceb053f2fcaad524abbc254970edd203ccb8deb1df

Observation 46006c99-c6b8-431d-b66d-c74fba81ad8c · outbound

This paper cites Dy- namicstereo: Consistent dynamic depth from stereo videos.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Dy- namicstereo: Consistent dynamic depth from stereo videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.956075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.881223Z digest=sha256:71a1f19d7f06d534dbadfa014d472863276edd39d6cf24b8274afae225b51f8e

Observation a0db4282-6c47-4564-b9b0-aba4f28d0466 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Imagic: Text-based real image editing with diffusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.885577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.885577Z digest=sha256:ddd887e690dfe2d5c83561134d3b240f51b9c6d489a845276fde9185b818a023

Observation 48ec99ae-962e-4da9-843e-006bf6899dc1 · outbound

This paper cites Repurpos- ing diffusion-based image generators for monocular depth estimation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Repurpos- ing diffusion-based image generators for monocular depth estimation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.929179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.890786Z digest=sha256:baf7f8dc93cd53bff1b3470a6128297bf50ad0d279cae86d825f1d9703436048

Observation f05808a1-c3c4-4c9c-b47a-ef711f6c1036 · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class im- age classification.Dataset available from https://github.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Openimages: A public dataset for large-scale multi-label and multi-class im- age classification.Dataset available from https://github

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.913766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.895569Z digest=sha256:1b61ebe6c9d475700551f02a7981ab4f272022aafb102ebe7c44dbfb21808846

Observation 83f18812-e8e0-41ca-9a8e-bcc549e6a6ee · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.900304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.900304Z digest=sha256:987cc417699bde893867c04ab96a74291931fdf815c67762f130d70b21d019d5

Observation cd96c19c-110a-46e9-82c2-97154f571b78 · outbound

This paper cites A sim- ple approach to unifying diffusion-based conditional gener- ation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers A sim- ple approach to unifying diffusion-based conditional gener- ation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.885523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.904970Z digest=sha256:d767b2427bf0421c5ef2834268185158e8d48f5dd652c9ee1be8158736f50b7d

Observation e76b6efc-0845-4b3f-bbe2-6e796cae985d · outbound

This paper cites Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.867109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.910360Z digest=sha256:69e5ee3eb2ba93a4433ddb7415942ff5cdc71ee8ee35f152d2d8530608061621

Observation fc484d44-5e99-4cf7-b3ae-b2a343c091cb · outbound

This paper cites Revisiting stereo depth estimation from a sequence- to-sequence perspective with transformers.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Revisiting stereo depth estimation from a sequence- to-sequence perspective with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.847315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.915037Z digest=sha256:d88e22af1f768e75fe921d086353faff7f16296aae7b684e8b5ee6ab7f07e190

Observation 08077ad5-380a-4e0f-9a05-2a2f72ba0fdd · outbound

This paper cites Microsoft coco: Common objects in context.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Microsoft coco: Common objects in context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.829741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.919396Z digest=sha256:7794af2aa86f9fb29b727a4a5191b9cf3e4583cb6d807331dcd9584ef771ef09

Observation a80c5ed8-f222-492f-b0c2-8904a0560559 · outbound

This paper cites Flow Matching for Generative Modeling.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Flow Matching for Generative Modeling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.923802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.923802Z digest=sha256:5f0c816c05d64d6b5ca835c30090cc06fe344d1cee0e337b8d14b4151c6b6570

Observation 7470cf35-68c6-4d78-b1df-11819e2f20c7 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.813347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.929218Z digest=sha256:61eef87c1233bc382e05eef5cc150affdeedcadd4762a88a55f88872b19e610e

Observation 698df1ed-81f7-4a64-9031-172d40c1a499 · outbound

This paper cites Zero-1-to- 3: Zero-shot one image to 3d object.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Zero-1-to- 3: Zero-shot one image to 3d object

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.794549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.934046Z digest=sha256:4b66a041313fedd00c0c45cfdba476bd6f5336f24101771429d2fba72385035b

Observation 6fbf4387-b655-4f54-884a-02153e21f2d8 · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.938971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.938971Z digest=sha256:51dfe9c66f78172be0018de7b457ff5ca3c3d20caced85cac64af09fdfa763e5

Observation 3ccc817f-b3bc-4391-bfa2-8f67f3c91286 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.778339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.944083Z digest=sha256:933b952006f28019503a3e5ed6da68eeaa5bdcc3334bf8446d4efbef9dc1c477

Observation b6f02291-c849-4ad9-9595-ddabf8f720c2 · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Repaint: Inpainting using denoising diffusion probabilistic models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.948760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.948760Z digest=sha256:594c1b9c81bd46ede36f7aed3208724527d142596841f155828a254300ca5dd8

Observation 9b57c80d-87e4-499f-b8a8-4fd6d514eea3 · outbound

This paper cites Readout guidance: Learning con- trol from diffusion features.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Readout guidance: Learning con- trol from diffusion features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.750508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.953585Z digest=sha256:8252f374f636c0470e30a48379e22357f8ba62bd65d25e7277ed74a83c257413

Observation 9180c32e-ce26-43b9-95d9-8ec625af36ce · outbound

This paper cites Scalable diffusion models with transformers.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scalable diffusion models with transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.958883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.958883Z digest=sha256:fc38d70bbb307316eaeb38695f88c88bcc32f9a55c92ddedc820a43d527cdd3d

Observation 0e50cd36-d621-4350-8f5d-8dc2b3f090b5 · outbound

This paper cites Pexels, royalty-free stock footage website.https: //www.pexels.com.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Pexels, royalty-free stock footage website.https: //www.pexels.com

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.717285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.963503Z digest=sha256:1b02a525cc7ab3e5352da95fdb76cd908d4ffa3a5742bff22a1cbd172c0b74f6

Observation 6bdc975b-33d0-4546-ade7-154e0eed6cea · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.700647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.968322Z digest=sha256:f1f42eb963763e158e38fb66ed99656bcbeb62815f3ed3326270b70edfa88cef

Observation 08656bfd-de6b-4f5b-aabe-23ef6ff4240b · outbound

This paper cites an unresolved cited work.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:32.685202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.972585Z digest=sha256:1a0fd2312892c862f60b5b463b3ac79286eb7660f52dba3c8af312e2694ffc53

Observation 6807bea5-e543-4fd7-9e27-1ac9e589dc5f · outbound

This paper cites Vi- sion transformers for dense prediction.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Vi- sion transformers for dense prediction

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.668627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.977327Z digest=sha256:3b414ede4ee40b7a2993ec30b9745509d121369ad0c111126c5d9364b3ac150c

Observation 27fa0ec4-24af-4d52-b849-db6ee5980a96 · outbound

This paper cites Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.652645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.981665Z digest=sha256:2fd5f771617bd927ffbcfad751fc1853e23ff7158170897fd3c37bc733cab9e4

Observation c03e09bd-8d5f-44ac-ada3-fc50d73bfa3b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers High-resolution image synthesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.635760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.986378Z digest=sha256:ae4c41ad6b85450b080c05fc1e0af805b40d561d9db3fd0a5d976fb6f9c96617

Observation f852eb49-cb3a-4f9b-876d-7de63aafd134 · outbound

This paper cites Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.619105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.990852Z digest=sha256:cb0854bd9707d9a1e06def0ae7430dd84048b79e34ec85b92ffc6082f22dbf5d

Observation fe6d243f-9a3d-4f55-abd5-1c4e86eedf5d · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.600796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.995542Z digest=sha256:3f3ac548ea806474d679c7ad927dd6b3f418c74b8f639b661e6428b306126ef0

Observation 0d498dfd-7bad-4841-8be7-3e54115a63db · outbound

This paper cites A multi-view stereo benchmark with high- resolution images and multi-camera videos.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers A multi-view stereo benchmark with high- resolution images and multi-camera videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.582793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.000071Z digest=sha256:a7b6db94fe01302b16e2109949acdcd1b26699ba22dd8ba416214e823c756a7c

Observation 43486afc-c5c8-41fe-8f25-6cb4cb57ae12 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Indoor segmentation and support inference from rgbd images

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.563874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.004897Z digest=sha256:48d2c3e0944d62409a806ba3a22282ae5108ed9b12ea1006fb7848cbe0ecaef0

Observation dc3e3101-6f71-4093-851b-55de20ba80a1 · outbound

This paper cites Generative modeling by esti- mating gradients of the data distribution.Advances in neural information processing systems, 32, 2019.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Generative modeling by esti- mating gradients of the data distribution.Advances in neural information processing systems, 32, 2019

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.546773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.010202Z digest=sha256:68ac16da0ac05d667b195df919424b4a04d1057b6a2b05e1c74e7d89b8c85741

Observation 1c97dfc6-f6ac-4964-904f-ac71d1c4f08d · outbound

This paper cites Improved techniques for training score-based generative models.Advances in neural information processing systems, 33:12438–12448, 2020.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Improved techniques for training score-based generative models.Advances in neural information processing systems, 33:12438–12448, 2020

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.528974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.014636Z digest=sha256:b086e63ff6eba94fcd2be5375bba7310809178c4e49ca42fd6e7df69ff394175

Observation 164c3f2f-ccec-4470-842d-d00f27b0c700 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Score-based generative modeling through stochastic differential equa- tions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.019532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.019532Z digest=sha256:945610db4e77365dc588c9f31076ad28229816c2406f4e75e2f9d1b098a8571d

Observation 02f508d3-5f8b-45e4-b086-ada5b5d12922 · outbound

This paper cites LDM3D: Latent Diffusion Model for 3D.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers LDM3D: Latent Diffusion Model for 3D

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.024582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.024582Z digest=sha256:eebf0ac63d2dc4261cd69d4d2b88f4eb5aa88c46075e69946d006cd134dde124

Observation 52086cdb-08af-45e3-8b09-c4ecca06e572 · outbound

This paper cites Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.030323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.030323Z digest=sha256:06138cd6dd2b9fcb6996b5ce84f1174fc819055ba30773fe54b1705221c2e09e

Observation 8b9390b1-9fc4-41f2-a5e7-c474b946b25b · outbound

This paper cites Soundbrush: Sound as a brush for visual scene editing.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Soundbrush: Sound as a brush for visual scene editing

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.503412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.035308Z digest=sha256:716a637225b0cdefa3cda09fddbc4df148da3d64ffc84484b9c58fe6a53fab65

Observation e47614c9-1030-4b72-98f6-d4312ed2ba1c · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Plug-and-play diffusion features for text-driven image-to-image translation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.486346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.040085Z digest=sha256:ecb3dcf0fadb0ed1f6512d88025e8e33734ce8203b7affcbffabe0d7d4ce4c4a

Observation 7064dcd2-c2b2-4a6e-a4f7-ed06b613ff2a · outbound

This paper cites DIODE: A Dense Indoor and Outdoor DEpth Dataset.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers DIODE: A Dense Indoor and Outdoor DEpth Dataset

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.044798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.044798Z digest=sha256:e4bf2a78e25fad0afb8cc3bcd069eb8dd74e327709a3ecc6e5135428ba92ceec

Observation d40fd372-e52c-4e32-b2af-708f6e2f0ac4 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.050298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.050298Z digest=sha256:54ec987960149ee409b4b6f2f5df70729f749f9bd5cf56f08380e55bb1b097a0

Observation d941c7ab-f341-46d4-a091-db4be8757017 · outbound

This paper cites Irs: A large naturalistic indoor robotics stereo dataset to train deep models for dis- parity and surface normal estimation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Irs: A large naturalistic indoor robotics stereo dataset to train deep models for dis- parity and surface normal estimation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.460079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.055018Z digest=sha256:9e69b19247c8904e13e79b61c3afd7e359a8c8391173a3c0c69fb3168c388bf1

Observation 70dec2f1-5e85-4e52-a191-5b990e93bf10 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.443142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.059412Z digest=sha256:7a1b838c545cbab0bf0f9e45d7389c646b8e9ec2acf271978f6bc8e4d5bd2408

Observation b913b190-caf6-4d10-ac67-96ebffd7c4a4 · outbound

This paper cites Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2025.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2025

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.424683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.064200Z digest=sha256:07edc3b3351dd86abc9ae2106febec2745b4620448a4c11c4c113680614e9850

Observation 78552e6d-0d9a-4028-bcda-1cd8f69fa44e · outbound

This paper cites Paint- it: Text-to-texture synthesis via deep convolutional texture map optimization and physically-based rendering.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Paint- it: Text-to-texture synthesis via deep convolutional texture map optimization and physically-based rendering

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.404820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.068423Z digest=sha256:63fca6acb6b60c36c635631b70d5ce36a7f6a4d9676ce0b521f7cc2b42ea28c5

Observation 061a8559-6029-4eb0-b3fd-39aa4cc40872 · outbound

This paper cites Metta: Single-view to 3d textured mesh reconstruction with test-time adaptation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Metta: Single-view to 3d textured mesh reconstruction with test-time adaptation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.386796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.072809Z digest=sha256:1a56d066b70267851e5c31b529bbcc3ab46839701ad34f283d13566690522539

Observation 66dc40ec-204f-4850-be39-b048a42aa84f · outbound

This paper cites Joint- net: Extending text-to-image diffusion for dense distribution modeling.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Joint- net: Extending text-to-image diffusion for dense distribution modeling

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.370386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.077128Z digest=sha256:ee1c0bf1b336106815f0a63dcabc9490b4b8526eba4f592f99c39958ea7adc12

Observation 0ad6a0fd-d2b8-47e9-98b0-17b328268a18 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Adding conditional control to text-to-image diffusion models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.351727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.081667Z digest=sha256:9f048be460723cf7c27c959b987dad8945a6763eb181663282916ce9964fb53b

Observation 4264956c-fc13-4bcd-a31c-4e2bc28cb58d · outbound

This paper cites TNBMM CMBDL LJUUFO CBMBODJOH B MFWJUBUJOH QPUJPO CPUUMF GJMMFE XJUI TIJNNFSJOH CMVF MJRVJEu t1BTUB XJUI NVTISPPNT BOE CBDPOu t.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers TNBMM CMBDL LJUUFO CBMBODJOH B MFWJUBUJOH QPUJPO CPUUMF GJMMFE XJUI TIJNNFSJOH CMVF MJRVJEu t1BTUB XJUI NVTISPPNT BOE CBDPOu t

Reference 73

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T04:47:32.191857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.086124Z digest=sha256:f615ebb906d8370a50356e1bb15d79536c2b53d94295c742fe3f965da11affaf

Pith citing papers

No inbound Pith citation observations are available.