Pith. sign in

Paper Citation Record · LEDGER

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers

As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2505.00482.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00482 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:47:32.086124Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e20639d2-b967-495e-8357-f167766ded85 · outbound

This paper cites Depthformer: Multi- scale vision transformer for monocular depth estimation with global local information fusion.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Depthformer: Multi- scale vision transformer for monocular depth estimation with global local information fusion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.359818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.740489Z digest=sha256:7aefc15b007e91a184bfdab5be3ac43f05f762f75a9286a5bc3a4f00f4a38af0

Observation 638bf7f6-2956-4ca9-bfa5-26245c5c6d99 · outbound

This paper cites Multimae: Multi-modal multi-task masked autoen- coders.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Multimae: Multi-modal multi-task masked autoen- coders

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.343791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.745882Z digest=sha256:483b08dc94f5e3dab863160d225cc8bc4e19176234e58b9d5b51474129496297

Observation 539feb88-6b35-41e6-8ce0-4251eec86a4b · outbound

This paper cites 4m-21: An any-to-any vision model for tens of tasks and modalities.Advances in Neural Infor- mation Processing Systems, 37:61872–61911, 2024.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers 4m-21: An any-to-any vision model for tens of tasks and modalities.Advances in Neural Infor- mation Processing Systems, 37:61872–61911, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.327777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.751082Z digest=sha256:ff6f6bb78741a988000f620088748f8daf92835404f8e4c8ddb7a5eab9afd055

Observation cfb5ebd8-5431-4fb5-b5a0-49ba2da2c7fd · outbound

This paper cites Multidiffusion: Fusing diffusion paths for controlled image generation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Multidiffusion: Fusing diffusion paths for controlled image generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.311923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.756124Z digest=sha256:6308036eca3a130f372de62589543405e24ea6df17e3050468ef5bff72a08c54

Observation 3bac5b32-4eb9-4221-bf1c-6d950f95d84e · outbound

This paper cites Se- mantickitti: A dataset for semantic scene understanding of lidar sequences.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Se- mantickitti: A dataset for semantic scene understanding of lidar sequences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.761070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.761070Z digest=sha256:2d1c10dbc5090c3983e9e642a64c704cf8812b390439e418b2bc8c9a183dde8d

Observation 84c0527e-61dc-4832-96bb-ae5d9348171d · outbound

This paper cites Loosec- ontrol: Lifting controlnet for generalized depth conditioning.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Loosec- ontrol: Lifting controlnet for generalized depth conditioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.765916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.765916Z digest=sha256:c84cd2f91fe6ee7e1d41771a54089f604d175d4ba598c65cd1852dcfbd77bbbe

Observation 277bc16c-9d08-438e-b729-c3ddcb67f6da · outbound

This paper cites Flux.1.https://huggingface.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Flux.1.https://huggingface

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.271520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.770700Z digest=sha256:aba34e5eeb5610b4826ac2061f75e646b1ea615cb99cdedd2f8c27e412ea880e

Observation f6408f9e-5835-4446-ac12-295d3e3fe99c · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.Advances in Neural Information Processing Systems, 37:24081–24125, 2025.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.Advances in Neural Information Processing Systems, 37:24081–24125, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.255156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.776031Z digest=sha256:d01072cc2b7a225b7aaf7a556686f7c811dc5340b1e3c2d962421faf6c9333a0

Observation a12b82ee-2136-43d4-999c-d8b68e4a8ee2 · outbound

This paper cites Text2tex: Text-driven tex- ture synthesis via diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Text2tex: Text-driven tex- ture synthesis via diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.238261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.780461Z digest=sha256:489fb7c4fa4e15ead33002ee08d73796127b524b67ccc73ace456faebab63571

Observation e3322d9e-bd64-4fc3-91e3-778378b88a1f · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.785181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.785181Z digest=sha256:aa38c4133e7f9d521ee4b695a721d1e50eb445e2c23ffdb8f33f3e8fc272ed43

Observation e424f2f6-1aae-42a9-a2ce-834c84b7f418 · outbound

This paper cites Neural ordinary differential equa- tions.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Neural ordinary differential equa- tions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.221679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.790267Z digest=sha256:456c420361049c63b7140d583d4045675c9764879af78a18625122d5142464d7

Observation 6ab98a38-4b5a-4ca4-8e11-d3c8ca25a96f · outbound

This paper cites Deep diffusion image prior for efficient ood adaptation in 3d inverse problems.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Deep diffusion image prior for efficient ood adaptation in 3d inverse problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.205553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.795385Z digest=sha256:0b8d0179eeba2990593ab211eb10faa7b9296894c85cf720f79b8e5fb5d0e5ab

Observation 5b0afb4f-3199-4fbb-b843-be69d0908e23 · outbound

This paper cites Improving diffusion models for inverse prob- lems using manifold constraints.Advances in Neural Infor- mation Processing Systems, 35:25683–25696, 2022.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Improving diffusion models for inverse prob- lems using manifold constraints.Advances in Neural Infor- mation Processing Systems, 35:25683–25696, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.188360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.800111Z digest=sha256:f9c7e6b7b70c1081d1fde1e90fd1e1950e7fd724305a33878a866c83157605ec

Observation 44d0262f-e37b-4fe7-931f-c3215bc8e681 · outbound

This paper cites Solving 3d inverse problems us- ing pre-trained 2d diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Solving 3d inverse problems us- ing pre-trained 2d diffusion models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.804493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.804493Z digest=sha256:6604725791a4e7b99b4d1fb6b0812834d44ee3efe8947b40b17ccbe407332c7a

Observation 0a47916b-fd92-4230-8d66-d34143d1b414 · outbound

This paper cites Latentpaint: Image inpainting in latent space with diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Latentpaint: Image inpainting in latent space with diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.158873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.809715Z digest=sha256:3386f919dc790928d8c0c553b673abdbe8448e7ab40472d116638c349a85a98b

Observation d25b22bf-6406-464c-ae7c-b8a6b27e4b97 · outbound

This paper cites DiffEdit: Diffusion-based semantic image editing with mask guidance.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers DiffEdit: Diffusion-based semantic image editing with mask guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.814438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.814438Z digest=sha256:aad4facd5428b805e3a298e97d81c3d4bda570a2957c086b1640904c3bf273ea

Observation 7508e38b-4b78-4cd9-a750-98697494ab7b · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.142025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.819346Z digest=sha256:ee55cf3e3338772c6681bcbed54cceb5c53232bc927ce509856c56be6280fb10

Observation f77ecb42-ba28-4636-8774-acdf2170ad25 · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scaling vision transformers to 22 billion pa- rameters

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.124854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.823696Z digest=sha256:4ac43605a5f81a5538b3c4ddd6f5540b872bef57403ec20ffa84510649187039

Observation d6ddda39-23d1-4c91-8f80-b2fcab394fec · outbound

This paper cites Scaling rec- tified flow transformers for high-resolution image synthesis.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scaling rec- tified flow transformers for high-resolution image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.106749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.828210Z digest=sha256:b784097797c9afcdf5e9dc05796bdba50af6b93214f1a6ba6bd6518a0f8d0497

Observation 88ba2f57-a753-471b-9475-c7cb6b4f5e86 · outbound

This paper cites The pascal visual object classes (voc) challenge.International journal of computer vision, 88:303–338, 2010.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers The pascal visual object classes (voc) challenge.International journal of computer vision, 88:303–338, 2010

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.084975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.832964Z digest=sha256:d2ecc5ebf27705efdf757b8bae9756ef1e21d482dcfbff578116392effc3725f

Observation 184ad83e-704c-4dc8-8e76-83fd531d14c9 · outbound

This paper cites Geowiz- ard: Unleashing the diffusion priors for 3d geometry esti- mation from a single image.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Geowiz- ard: Unleashing the diffusion priors for 3d geometry esti- mation from a single image

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.069501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.837638Z digest=sha256:6e342a7ebd8929aa464502c89640496af61d122fa14a8f96f2c10c6ddcccdb5e

Observation 34108945-3372-47fa-9433-be9dffdcba9b · outbound

This paper cites Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Ommer.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Ommer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.053680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.842853Z digest=sha256:8456ed832eefee7e410f56985d07336025493bbfc7cddaa1e672cff9aaf18330

Observation 70114f14-a7c6-4740-9e1e-5bf642aac3bb · outbound

This paper cites Efficient diffu- sion training via min-snr weighting strategy.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Efficient diffu- sion training via min-snr weighting strategy

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.038231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.847919Z digest=sha256:67aa23553b7b7bab31b427428c0c33d66bc2256a5e571f8bef4b8906ee905ffb

Observation bab29a7a-832b-4f96-9616-516c93de7b41 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:33.022264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.852592Z digest=sha256:b1e37de18d80a434999023c5144ce8d73ec19a79bc890942098e172887602aa7

Observation b6b1a09d-5fe0-40f2-817d-6c5c76184e18 · outbound

This paper cites Classifier-Free Diffusion Guidance.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Classifier-Free Diffusion Guidance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.857329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.857329Z digest=sha256:e0ada06fccf70876c0a8ce36dcf245356796d1b63bea3af689f276a4656fd4a7

Observation 7a2cdd4b-fcd0-4adf-b77c-fdc4440b7dd4 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.862188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.862188Z digest=sha256:26473bd43377625492657c42825891cfaff4b610536d6217bf444b99b3e9592b

Observation f12d4fa1-d866-45ac-9bdd-a24740c0f68b · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.990748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.866706Z digest=sha256:9bc0edd30a7b89c76d08ef18a82232d23b7c972ee8506807f1163683371fa357

Observation 2b127d2a-7d10-404e-bd23-f06a61b4a54e · outbound

This paper cites Zero-shot depth completion via test-time align- ment with affine-invariant depth prior.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Zero-shot depth completion via test-time align- ment with affine-invariant depth prior

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.972436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.871529Z digest=sha256:06daa38a1a250f694d818995ae688834832502923cc342e3513ded1247c82f67

Observation 8d32c6f0-a73f-4804-b121-2370953719a9 · outbound

This paper cites Mixture of Diffusers for scene composition and high resolution image generation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Mixture of Diffusers for scene composition and high resolution image generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.875973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.875973Z digest=sha256:1d620d4d1545e077bb7230c23a3f79c9b7fe6e1ab8bb2269eb1bc2b2cc3b9020

Observation 46006c99-c6b8-431d-b66d-c74fba81ad8c · outbound

This paper cites Dy- namicstereo: Consistent dynamic depth from stereo videos.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Dy- namicstereo: Consistent dynamic depth from stereo videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.956075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.881223Z digest=sha256:ad56d59d0427f7b277dd0ada851f2ced799d6b22bd7df5689ac3a80aec574d42

Observation a0db4282-6c47-4564-b9b0-aba4f28d0466 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Imagic: Text-based real image editing with diffusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.885577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.885577Z digest=sha256:ac55223825af754fa5c55a15b4bcf5558d013e03742c03a69d2fc03410cca002

Observation 48ec99ae-962e-4da9-843e-006bf6899dc1 · outbound

This paper cites Repurpos- ing diffusion-based image generators for monocular depth estimation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Repurpos- ing diffusion-based image generators for monocular depth estimation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.929179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.890786Z digest=sha256:3ba64cb4152a2a41f9e20198cf7d7f87160be06a28d31a1b120adca1e43c3ccf

Observation f05808a1-c3c4-4c9c-b47a-ef711f6c1036 · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class im- age classification.Dataset available from https://github.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Openimages: A public dataset for large-scale multi-label and multi-class im- age classification.Dataset available from https://github

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.913766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.895569Z digest=sha256:ae54fac5ae803891237f5ad16d001d497112211eaaed84b7ca8000e49b444adf

Observation 83f18812-e8e0-41ca-9a8e-bcc549e6a6ee · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.900304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.900304Z digest=sha256:071c608ca6924bc59cbae3c1131b2da7f3f9199060755de3c40088a179d370b9

Observation cd96c19c-110a-46e9-82c2-97154f571b78 · outbound

This paper cites A sim- ple approach to unifying diffusion-based conditional gener- ation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers A sim- ple approach to unifying diffusion-based conditional gener- ation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.885523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.904970Z digest=sha256:55be925bd96f84fd1d9f3132f637e83b29d1d26aacab8f0380fd20c527b7d68b

Observation e76b6efc-0845-4b3f-bbe2-6e796cae985d · outbound

This paper cites Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.867109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.910360Z digest=sha256:94e13f5338e0ddc84d5434d3b64d033cebc528dea871369f09d707e65b8e0c51

Observation fc484d44-5e99-4cf7-b3ae-b2a343c091cb · outbound

This paper cites Revisiting stereo depth estimation from a sequence- to-sequence perspective with transformers.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Revisiting stereo depth estimation from a sequence- to-sequence perspective with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.847315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.915037Z digest=sha256:3066c0751afe76d8865f0d28543d8510b07170ace6f2eb5cd314712ee08dce8f

Observation 08077ad5-380a-4e0f-9a05-2a2f72ba0fdd · outbound

This paper cites Microsoft coco: Common objects in context.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Microsoft coco: Common objects in context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.829741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.919396Z digest=sha256:815ee359ae40d6c4e332da9900a6adbc4734e0c932c1a0c9a67246266f5b4ce9

Observation a80c5ed8-f222-492f-b0c2-8904a0560559 · outbound

This paper cites Flow Matching for Generative Modeling.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Flow Matching for Generative Modeling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.923802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.923802Z digest=sha256:0f8d43ab4fd85cddf69eec01d075f22ec09cf3822d2d962f13bf8b03db742d2c

Observation 7470cf35-68c6-4d78-b1df-11819e2f20c7 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.813347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.929218Z digest=sha256:6e1813800aaa6262073bb5b31a76ad649cc8ef03df2ada1d15e036900383881d

Observation 698df1ed-81f7-4a64-9031-172d40c1a499 · outbound

This paper cites Zero-1-to- 3: Zero-shot one image to 3d object.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Zero-1-to- 3: Zero-shot one image to 3d object

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.794549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.934046Z digest=sha256:dd7c1133fb184509e288ad2a84604e78388de63f6bd33348943a9daa6c8cb622

Observation 6fbf4387-b655-4f54-884a-02153e21f2d8 · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.938971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.938971Z digest=sha256:2e1024828ff329e694e887b213e524063b1020c522f82b4f14c9ff85cdb06470

Observation 3ccc817f-b3bc-4391-bfa2-8f67f3c91286 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.778339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.944083Z digest=sha256:62a2296061957cbd7eb508cb81f819106050e49a318ea3e0d60c6784e89f1e18

Observation b6f02291-c849-4ad9-9595-ddabf8f720c2 · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Repaint: Inpainting using denoising diffusion probabilistic models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.948760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.948760Z digest=sha256:b498b734a6ae873ca081cae285380a4643628bd4aa740110501695d6a08d07ed

Observation 9b57c80d-87e4-499f-b8a8-4fd6d514eea3 · outbound

This paper cites Readout guidance: Learning con- trol from diffusion features.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Readout guidance: Learning con- trol from diffusion features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.750508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.953585Z digest=sha256:96af08559814981e802478e3f3dac0684fd299e1ef91c7537be7e2a6f6ab6630

Observation 9180c32e-ce26-43b9-95d9-8ec625af36ce · outbound

This paper cites Scalable diffusion models with transformers.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Scalable diffusion models with transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:31.958883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:31.958883Z digest=sha256:41ecc6babb0c5dd280eea825d87dfd9b38658d11db88a003bacf338a4a297c66

Observation 0e50cd36-d621-4350-8f5d-8dc2b3f090b5 · outbound

This paper cites Pexels, royalty-free stock footage website.https: //www.pexels.com.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Pexels, royalty-free stock footage website.https: //www.pexels.com

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.717285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.963503Z digest=sha256:68d411828ed144d930da37c190f409ec0c367025519faaf396de989eb5952cdd

Observation 6bdc975b-33d0-4546-ade7-154e0eed6cea · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.700647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.968322Z digest=sha256:bc1c64983ff90b316375a6b027e412f8a8c641c3b3a9cd8bbd3100098b2e3b75

Observation 08656bfd-de6b-4f5b-aabe-23ef6ff4240b · outbound

This paper cites an unresolved cited work.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:47:32.685202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.972585Z digest=sha256:18d3606d551ca49bb6b4df96e8bfeff5174fbb6c8dfc9439ca6afade2e7d19af

Observation 6807bea5-e543-4fd7-9e27-1ac9e589dc5f · outbound

This paper cites Vi- sion transformers for dense prediction.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Vi- sion transformers for dense prediction

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.668627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.977327Z digest=sha256:de73e171cfdf39d21ab4e1208fd239d139747270e0c58d58757c05b5412ce29a

Observation 27fa0ec4-24af-4d52-b849-db6ee5980a96 · outbound

This paper cites Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.652645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.981665Z digest=sha256:b76626870022d88a5630311da06a58b400a6cc3be12c539cf4970acf7c08111a

Observation c03e09bd-8d5f-44ac-ada3-fc50d73bfa3b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers High-resolution image synthesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.635760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.986378Z digest=sha256:fab945c6f36e35ed98d92ab239bcf7595980209bfd26d8e871f8e8cf1ce0dff6

Observation f852eb49-cb3a-4f9b-876d-7de63aafd134 · outbound

This paper cites Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.619105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.990852Z digest=sha256:d3914f66aa1e8555353b02b723293ce9225ce7dfc842c634f18d2cb3bab6dd14

Observation fe6d243f-9a3d-4f55-abd5-1c4e86eedf5d · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.600796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:31.995542Z digest=sha256:b614f0d33079899a33fd1561c63c48d2f7a27bf95e9568deb6449b7fe957d6e8

Observation 0d498dfd-7bad-4841-8be7-3e54115a63db · outbound

This paper cites A multi-view stereo benchmark with high- resolution images and multi-camera videos.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers A multi-view stereo benchmark with high- resolution images and multi-camera videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.582793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.000071Z digest=sha256:a2ce9582fd995b7ebdb3c18d8bf53287f12652eacbdeedc95f3be9948f564367

Observation 43486afc-c5c8-41fe-8f25-6cb4cb57ae12 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Indoor segmentation and support inference from rgbd images

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.563874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.004897Z digest=sha256:2fcc07fde60e0ec5ec021580bf9991d2f7890d8815674bbc3cd8340f02a313b0

Observation dc3e3101-6f71-4093-851b-55de20ba80a1 · outbound

This paper cites Generative modeling by esti- mating gradients of the data distribution.Advances in neural information processing systems, 32, 2019.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Generative modeling by esti- mating gradients of the data distribution.Advances in neural information processing systems, 32, 2019

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.546773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.010202Z digest=sha256:b7bd09dfd0a9fc7f23d98ab6950f9f7ce9059de2a7f1f5e747b40c22f0c1d4c5

Observation 1c97dfc6-f6ac-4964-904f-ac71d1c4f08d · outbound

This paper cites Improved techniques for training score-based generative models.Advances in neural information processing systems, 33:12438–12448, 2020.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Improved techniques for training score-based generative models.Advances in neural information processing systems, 33:12438–12448, 2020

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.528974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.014636Z digest=sha256:17bed4b50f6740e02130edc3ac3b5f6a1fb768fe6c634be12cdce5168e4ea607

Observation 164c3f2f-ccec-4470-842d-d00f27b0c700 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Score-based generative modeling through stochastic differential equa- tions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.019532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.019532Z digest=sha256:0aed1e038d72973af44facab07de715b7b1c9720ab5106af129a970406503745

Observation 02f508d3-5f8b-45e4-b086-ada5b5d12922 · outbound

This paper cites LDM3D: Latent Diffusion Model for 3D.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers LDM3D: Latent Diffusion Model for 3D

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.024582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.024582Z digest=sha256:93b14ca843472b7312f95373630c17fb2878388ba6e96eeb3ffff62370553dd1

Observation 52086cdb-08af-45e3-8b09-c4ecca06e572 · outbound

This paper cites Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.030323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.030323Z digest=sha256:fb1efa9d85a59d2420423638460fd3a43cc136cadcf4e06152a9611fc190f0da

Observation 8b9390b1-9fc4-41f2-a5e7-c474b946b25b · outbound

This paper cites Soundbrush: Sound as a brush for visual scene editing.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Soundbrush: Sound as a brush for visual scene editing

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.503412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.035308Z digest=sha256:3aae9355eb67c298c4d2ecf64168d5975a16ad6119fd112dad90c4907076d755

Observation e47614c9-1030-4b72-98f6-d4312ed2ba1c · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Plug-and-play diffusion features for text-driven image-to-image translation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.486346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.040085Z digest=sha256:f24ad81696825e46c3e81f01097bee237dc6c5f494516ddaa77fac4446e12fb9

Observation 7064dcd2-c2b2-4a6e-a4f7-ed06b613ff2a · outbound

This paper cites DIODE: A Dense Indoor and Outdoor DEpth Dataset.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers DIODE: A Dense Indoor and Outdoor DEpth Dataset

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.044798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.044798Z digest=sha256:753d1a5236aad367d8326087af5e1f1d92db7bba8a51f09f266b80bd22a833ac

Observation d40fd372-e52c-4e32-b2af-708f6e2f0ac4 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:47:32.050298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:47:32.050298Z digest=sha256:96c5ed731c4d81fce8a0f9fb919d5a8eb69a2b270dc70b777177634b5bc8ab86

Observation d941c7ab-f341-46d4-a091-db4be8757017 · outbound

This paper cites Irs: A large naturalistic indoor robotics stereo dataset to train deep models for dis- parity and surface normal estimation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Irs: A large naturalistic indoor robotics stereo dataset to train deep models for dis- parity and surface normal estimation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.460079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.055018Z digest=sha256:488c4a35a0711503023a847adb78b53e5066e845a6a8c4cca273f1500fd5aeb3

Observation 70dec2f1-5e85-4e52-a191-5b990e93bf10 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.443142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.059412Z digest=sha256:f3e3e19a34e9ca8908a4e850be771b31419efe32d71da5a206300e402ea4c4d7

Observation b913b190-caf6-4d10-ac67-96ebffd7c4a4 · outbound

This paper cites Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2025.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2025

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.424683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.064200Z digest=sha256:7d940f44e6ac68a1821cea8ab4208e42521bd077f1029a54f811fbf0ad1f1a44

Observation 78552e6d-0d9a-4028-bcda-1cd8f69fa44e · outbound

This paper cites Paint- it: Text-to-texture synthesis via deep convolutional texture map optimization and physically-based rendering.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Paint- it: Text-to-texture synthesis via deep convolutional texture map optimization and physically-based rendering

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.404820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.068423Z digest=sha256:d5d10698f0894786267481eadddf194754a123cbcb9a41f699a98c0e97ea6190

Observation 061a8559-6029-4eb0-b3fd-39aa4cc40872 · outbound

This paper cites Metta: Single-view to 3d textured mesh reconstruction with test-time adaptation.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Metta: Single-view to 3d textured mesh reconstruction with test-time adaptation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.386796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.072809Z digest=sha256:ad2217dfb89654dd48e2c7a7d8fbb44a802b355981fc17cb5dfb7bba6040ecd7

Observation 66dc40ec-204f-4850-be39-b048a42aa84f · outbound

This paper cites Joint- net: Extending text-to-image diffusion for dense distribution modeling.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Joint- net: Extending text-to-image diffusion for dense distribution modeling

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.370386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.077128Z digest=sha256:ed2cd2df186906615d7207974b16fba2a1050bac6f24f8f80973f2e283499a54

Observation 0ad6a0fd-d2b8-47e9-98b0-17b328268a18 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers Adding conditional control to text-to-image diffusion models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:47:32.351727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.081667Z digest=sha256:e6c66c254b0985d8dbbd525acedbf00c536a7d2012e5ee1ae29bb11ba44d326c

Observation 4264956c-fc13-4bcd-a31c-4e2bc28cb58d · outbound

This paper cites TNBMM CMBDL LJUUFO CBMBODJOH B MFWJUBUJOH QPUJPO CPUUMF GJMMFE XJUI TIJNNFSJOH CMVF MJRVJEu t1BTUB XJUI NVTISPPNT BOE CBDPOu t.

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers TNBMM CMBDL LJUUFO CBMBODJOH B MFWJUBUJOH QPUJPO CPUUMF GJMMFE XJUI TIJNNFSJOH CMVF MJRVJEu t1BTUB XJUI NVTISPPNT BOE CBDPOu t

Reference 73

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T04:47:32.191857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:47:32.086124Z digest=sha256:712e0121acc6b0fc84f5c25331c9d29aba8f7ef7688526d70f58fdc9d22ed7bf

Pith citing papers

No inbound Pith citation observations are available.