Pith. sign in

Paper Citation Record · LEDGER

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

As of 16 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 4 inbound Pith citation observations for arXiv:2412.03558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03558 v3

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:19:24.924471Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:57.655659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T20:03:57.298307Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e235fb43-7527-4703-8188-ee4147644d52 · outbound

This paper cites Neural rgb-d surface reconstruction.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Neural rgb-d surface reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.470319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.470319Z digest=sha256:56d3066a51c50acb59bd574b67899b5d1ba907b68489ac8fbba3a010e9de0bd5

Observation 6e2bcb3f-4e3d-4f15-9319-b75e40814483 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.477120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.477120Z digest=sha256:0d135fa545dcae1e49e2c0dea52decc683c90270a5160dd902f6ec7be7723d03

Observation a26b3620-e4d2-481e-ae07-4d385608d747 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.483193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.483193Z digest=sha256:895df37f2288adbc13902e3e15b37a495ae8304bbad152f2cb97834102c53a9d

Observation 1e37fa38-5f5e-4003-808c-a92f919788e9 · outbound

This paper cites Single-view 3d scene reconstruc- tion with high-fidelity shape and texture.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Single-view 3d scene reconstruc- tion with high-fidelity shape and texture

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.489199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.489199Z digest=sha256:26320df515ee8332a56915379ccc902367dfe333892e61d71d073a084656bda9

Observation 8f6e5658-3c90-42d0-b9d5-457b21d44953 · outbound

This paper cites ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.494401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.494401Z digest=sha256:6e32b2f4194721ae6c31163b6e1b9ffb905a1852b535cd3cd2c3976671d6ae42

Observation c02f3fd2-ad73-4121-8653-e66339e13dda · outbound

This paper cites Buol: A bottom-up framework with occupancy-aware lifting for panoptic 3d scene reconstruction from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Buol: A bottom-up framework with occupancy-aware lifting for panoptic 3d scene reconstruction from a single image

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.500289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.500289Z digest=sha256:6dc2d810c7f39de33fd79e6b539137dce93f366e810ca79af9dd2fbb40b23611

Observation 239bf3cb-c715-4bd7-9433-650664bdb1d0 · outbound

This paper cites Panoptic 3d scene reconstruction from a single rgb image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Panoptic 3d scene reconstruction from a single rgb image

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.506633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.506633Z digest=sha256:e20af54c3ff655ad536f4674a6112bafae19f959fbbc5a32be1a4978664a3679

Observation 09b903b6-84d8-4024-975f-5276c71ff7d4 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.512113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.512113Z digest=sha256:5d15141072b67a7aa77f9072b4477799151142c0a191c98ad07dc806fec48f06

Observation 94ab5b39-77d7-4abd-aa7f-914bbc3bc830 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Objaverse: A universe of annotated 3d objects

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.517348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.517348Z digest=sha256:f29c4923b68c56f826b8fc54d4ec73ec58e395c8b0c8fa0d6b26ca1f7d084e77

Observation 1905d084-a603-4c44-bcba-fe7c466462dc · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Objaverse-xl: A universe of 10m+ 3d objects

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.257528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.523049Z digest=sha256:4bd380978118fcccf309fd82ee17c696b6e5ecf80cad67a0682e79677e76c905

Observation 1228f12e-7aaa-4928-a7dd-409e7e3bbfcb · outbound

This paper cites Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.527729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.527729Z digest=sha256:371f94c5c548f385d56ad4f4eeb047ef5d108f878a0e64a4b9d9bfe470724e8a

Observation c68931e9-63d4-46ab-ba3e-2eec7022ddd0 · outbound

This paper cites Tela: Text to layer-wise 3d clothed human generation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Tela: Text to layer-wise 3d clothed human generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.532509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.532509Z digest=sha256:484ba406058bce5e935eab226f0eb4726613fa632abda9c52b40b34a9f733b7c

Observation ba352958-b5c6-444b-a471-1507a2a77061 · outbound

This paper cites Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.538986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.538986Z digest=sha256:f6534ef3bacb9adc1cc405f6e26018469b2e32bcf4b61e81744b5eec1c47104b

Observation 595beabd-0d2e-40fb-8395-fd5b260a82c8 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.544028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.544028Z digest=sha256:6272094f088ecddcdcd588b339a6abf0ad934c5b97b50e72a17cfa508ede1916

Observation 034999a5-8e4a-43a0-8fb2-22483d787c19 · outbound

This paper cites 3d-front: 3d furnished rooms with layouts and semantics.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3d-front: 3d furnished rooms with layouts and semantics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.548492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.548492Z digest=sha256:572b5c61c3217da7acf79e676c347fdeebd27bb6d3b5b07ee3a4e6d990a696aa

Observation 5c6d46ed-3de8-4569-956d-101af54ab5ac · outbound

This paper cites 3d-future: 3d fur- niture shape with texture.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3d-future: 3d fur- niture shape with texture

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.192855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.553425Z digest=sha256:efac5c4a92a7db71a8e1c9dc7d15bf0b232025288e20011abb439ebfa4a2d7e3

Observation aba27113-874c-4f8d-9746-34ba865c5f5b · outbound

This paper cites Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.174826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.559147Z digest=sha256:ea593d61335e21292a58ce62c52ac1e1fce64e877b622770746a2eef9957b0ed

Observation a063ab1c-879f-4b75-9296-6f0737c091c9 · outbound

This paper cites Learn- ing 3d object shape and layout without 3d supervision.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Learn- ing 3d object shape and layout without 3d supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.156177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.565160Z digest=sha256:02854ff04ceb7dbe17b633595ce2d284c266dd80df95212ccc46037a9da97780

Observation 3ad51016-5d78-4210-bcce-c97abcc04d6b · outbound

This paper cites Roca: Ro- bust cad model retrieval and alignment from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Roca: Ro- bust cad model retrieval and alignment from a single image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.136435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.570414Z digest=sha256:ccd1f110f7eca767f658033cc2006655db463851e19815d8ab04bdff65b26eba

Observation 7d3f575f-4163-4bcf-adc6-c83283dab3c2 · outbound

This paper cites threestudio: A unified framework for 3d content generation, 2023.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation threestudio: A unified framework for 3d content generation, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.118486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.575723Z digest=sha256:ee3f912cba2dbc0c4cd3a98a5efed9535088c077370ad7bf4ca8875efff5ddfd

Observation 3d281103-59ca-4932-ac68-e7c9aa24e9bd · outbound

This paper cites REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.580426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.580426Z digest=sha256:4a04f5197f70381f2d1baf881e97f3e7a21d28270a3fbd55dce5b446427ae767

Observation 5fa7230e-a9e5-43f4-99e9-d81214e9651e · outbound

This paper cites Classifier-Free Diffusion Guidance.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Classifier-Free Diffusion Guidance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.585916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.585916Z digest=sha256:c5d34608d2d338cd86495ac635c663a62f6ad2082377cf74c3cce6cfb9d3027f

Observation 69933ff9-1972-4ba0-befc-fd8bbbb83f8f · outbound

This paper cites Denoising dif- fusion probabilistic models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.592499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.592499Z digest=sha256:ed04fd879872c25bdabbea82f4c0ce81bd31ea440c55ccbb90f781ef1f5a66b6

Observation 5d2bf12b-69e0-4362-b26e-f0b1ed265da8 · outbound

This paper cites LRM: Large Reconstruction Model for Single Image to 3D.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LRM: Large Reconstruction Model for Single Image to 3D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.598267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.598267Z digest=sha256:af089658e46ba8d23742a9615544b5abda0cf2934d290a79395491f9e60ecb1a

Observation 2018dd2e-4c69-4d98-9878-fcb0e982f5fe · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.603228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.603228Z digest=sha256:16ad8a4e86babb51da3bf4a7111db0f46b775f57b1d5cdea5bece00d7c93333b

Observation 96c46aaf-8813-4829-8cc1-e09b9da2b740 · outbound

This paper cites MV-Adapter: Multi-view Consistent Image Generation Made Easy.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation MV-Adapter: Multi-view Consistent Image Generation Made Easy

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.608784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.608784Z digest=sha256:18a8baa8f359dd8dbae7b924f1c29fd715ed19cf10cbd4b4413025c8e8c0f028

Observation 3097edca-607d-4802-b4d4-0549bfba768e · outbound

This paper cites Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.613872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.613872Z digest=sha256:0ceb932c429a08500e5c2cedf20b902a4dc5d6605fb2ee06aeee85ee4a09bbc9

Observation 36675878-c973-4ca6-8952-84f4c963f89c · outbound

This paper cites an unresolved cited work.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:19:26.078879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.618869Z digest=sha256:e95415c39ebe5df80c5d08beba808a71592a12baea715ac7c29b1f8f32001f44

Observation 7109ec4f-ecb9-494d-8c84-0b40e7bc6c64 · outbound

This paper cites Shap-E: Generating Conditional 3D Implicit Functions.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Shap-E: Generating Conditional 3D Implicit Functions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.624279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.624279Z digest=sha256:87102578e9c3e28e4fe1c5f745caff535c1dd4ff16e539872be25d44467da879

Observation bb96c0ff-5170-430f-9afd-99b179c31daf · outbound

This paper cites Auto-Encoding Variational Bayes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Auto-Encoding Variational Bayes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.629396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.629396Z digest=sha256:dd19dfa2e1f944c84f3768caebbd7d2aa3464bb0a4483f54604d15354d2e6125

Observation 6c3d10de-360d-47b4-8c78-afe325186f7a · outbound

This paper cites Segment Anything.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Segment Anything

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.634421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.634421Z digest=sha256:da2cdabc1a1ef3bf12a0c91e2a9064ba1523858f1a9234e211b4db4dafeec322

Observation 54f7fe89-3a06-41ea-b913-8f226b7e47b6 · outbound

This paper cites Mask2cad: 3d shape prediction by learning to segment and retrieve.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Mask2cad: 3d shape prediction by learning to segment and retrieve

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.062304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.639549Z digest=sha256:7888cddfb29956b2c3fe152d0ad1a81214fa0718e6af8a70bb5b94240eab5b88

Observation 6a93a930-8502-4b47-854d-d93ef3fb6723 · outbound

This paper cites Patch2cad: Patchwise embedding learning for in-the- wild shape retrieval from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Patch2cad: Patchwise embedding learning for in-the- wild shape retrieval from a single image

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.040938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.645115Z digest=sha256:d92faaf5046ecb7d83dac0471411fc208e5643c7d1473bd24113a1cb980b4f46

Observation 4f859191-a6a3-40d4-b8d9-8bccf58e4bd8 · outbound

This paper cites SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.651197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.651197Z digest=sha256:72083224006cb1284657de7cb00d9e45d432aef6500f32bb2b3c6afa6b4488e1

Observation 75e3988e-fddb-49f1-b8ac-d10803d98fcc · outbound

This paper cites CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.657182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.657182Z digest=sha256:7da74f4b36be1a0a173fc009a68fd4f85a9a49e7377676a5db97e71888eade61

Observation 175e174b-e84c-42c9-9015-f4f6c702b2f2 · outbound

This paper cites TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.662780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.662780Z digest=sha256:dffdeb8b4543edf0569f3ab932a9db9ed0d526d797d480f96c93aaefb024d47c

Observation 4e14715a-3bd3-4850-a60d-10d45d5eac70 · outbound

This paper cites Part123: part-aware 3d reconstruction from a single-view image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Part123: part-aware 3d reconstruction from a single-view image

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.024483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.668803Z digest=sha256:b6545da7e3e42cc3b2442025ee27ff1f9617f95484b55c2d8416857735345cf9

Observation 98b75164-459c-48c0-9671-a6fdb90a1060 · outbound

This paper cites Towards high-fidelity single-view holistic reconstruction of indoor scenes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Towards high-fidelity single-view holistic reconstruction of indoor scenes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.006488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.673792Z digest=sha256:d6102be23583d35a713cd4207a9ba54f67228fe8b0581b3fd937cad2c4f6ad26

Observation d0fb7e2a-d33e-4a7b-8772-ac2af0b768ca · outbound

This paper cites One-2-3-45++: Fast single im- age to 3d objects with consistent multi-view generation and 3d diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation One-2-3-45++: Fast single im- age to 3d objects with consistent multi-view generation and 3d diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.990012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.678786Z digest=sha256:a01e3c899b6199915c5b1f26c4e89449c3e0f12f4e1e24a691fa46cefd56754d

Observation cda20c72-bd78-4aab-a7fe-d0a97c94dbd3 · outbound

This paper cites One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.973929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.683680Z digest=sha256:bfc03a078d243afa6b31ce56a392ef1a2808c9c27d3f9710d086ea976c1dae4d

Observation 47648afb-9b6b-4fc5-8461-efd9a8bac810 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.688612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.688612Z digest=sha256:246fea518b3896561171ef21b496a0436d191f06c2219b976b9124cb408213a8

Observation 13d820e1-6bc3-44ec-9861-0c422768086e · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.694510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.694510Z digest=sha256:b7eec6517535f9d43fea0a974d311665664bea9e42bbd0d69a155f6f842cc827

Observation e927f0aa-afb3-4d16-a9d1-5d3f42350950 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.700568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.700568Z digest=sha256:2e2b1811c53137085243ee00b566d895626bf9f7cdc2317d043e2a300cf5ac6b

Observation ddea7528-a219-40ec-ba4c-d59833c031cc · outbound

This paper cites Wonder3d: Sin- gle image to 3d using cross-domain diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Wonder3d: Sin- gle image to 3d using cross-domain diffusion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.706386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.706386Z digest=sha256:1466fc955174acdc0673f6518c7f5a24f3918a91830fe139e12d6f627b85ac8a

Observation 019a47af-dc24-4573-902e-3ea9154f3954 · outbound

This paper cites Marching cubes: A high resolution 3d surface construction algorithm.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Marching cubes: A high resolution 3d surface construction algorithm

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.946424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.711368Z digest=sha256:1ffe9e93390a900a9131ecb162d8f64b666ce7ab1b044c7153761e6ffee7704c

Observation 666a8032-ff54-4e60-b1a5-b0abaec7e79d · outbound

This paper cites LT3SD: Latent Trees for 3D Scene Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LT3SD: Latent Trees for 3D Scene Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.716309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.716309Z digest=sha256:b3a8f1936c038eef4f63341edc5ff741586b0847d9967947864ad7eed6edf6be

Observation e9f1308a-1369-420f-a75d-8593b327dbac · outbound

This paper cites GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.928013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.721777Z digest=sha256:f3bc0e9762ae84d3ed4b3456e6a8efc88f1a2409891696ffc940cd8f6563db43

Observation 54ec66ed-b173-448b-b89e-1c23a5b7e666 · outbound

This paper cites Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.910886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.727319Z digest=sha256:e064213260121a751ef274f50059b6604b49aa19cd113bdb1c2202361e31f89d

Observation 0c0282a3-f776-4280-8189-4e258e7b3018 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.731797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.731797Z digest=sha256:ae3d615dde9a1b3f2c594ee82521aecdcd2283e962d6a06ca12c5d86f88ce33b

Observation e24e5a4d-0044-4f4d-a413-796cca81d393 · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Atiss: Autoregres- sive transformers for indoor scene synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.737191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.737191Z digest=sha256:84cbab23fa2b527d36f15b356eb18f2b4a3b96beb4c7e7f905c87600319ae880

Observation 66b52f76-11e1-42e9-89e2-5415084b1480 · outbound

This paper cites Scalable diffusion models with transformers.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scalable diffusion models with transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.742085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.742085Z digest=sha256:76dc7316ac10953c089214f2d5e2d090608a835fc42fc1e6fb2049b99f5f8870

Observation b8817661-99de-421b-b153-8fd0e05d1cf8 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.746930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.746930Z digest=sha256:f7e102000fd588486f997fe6d55dcfd03f3403ce681ff8fa05fc5ce9de88b324

Observation ea90cd2f-9f21-4682-bce3-cd66911dcd04 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.752142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.752142Z digest=sha256:35b11fad293da4a2f2bf23ffb9dcb33084b04e4affea483a382f4ba388830681

Observation 0aee56db-70e7-46d5-986f-f0326180cebe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.756723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.756723Z digest=sha256:e321fcaff8b1696bd771cee27a7bf3e2595c001258481e0ed2b3d6f417c30ef0

Observation 5910e255-592b-4ac9-9f78-605d0f2f09c8 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.761588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.761588Z digest=sha256:c677ed0022d0919ccb31fbd855ea1702ba65a90efc39379695d2d2f1cb9f4513

Observation b1c76509-8895-41e3-b0d9-73ab3a7b02e6 · outbound

This paper cites L3DG: Latent 3D Gaussian Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation L3DG: Latent 3D Gaussian Diffusion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.766543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.766543Z digest=sha256:dae52d2bdccc779c8d7f17e6afaa1c4f5a83e271b7f192129af4e53f7d34731b

Observation 1b3cd6f3-fc8a-4551-9bd2-193a6ce60af8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation High-resolution image synthesis with latent diffusion models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.771589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.771589Z digest=sha256:cf207460a3fcb35cc93bc94414c1a41390cfbefcfe8fe0bb8f236b5245ebc905

Observation 2efe727b-e8ed-4d56-85b9-9871961d6cf0 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.776858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.776858Z digest=sha256:3c42ae11a2b18d318e278f21d7265c832c5c37e877430183c4a80a2360ca1b4b

Observation c28f9010-5b4c-4dc1-8fbf-5bfff51f1762 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Deep unsupervised learning using nonequilibrium thermodynamics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.781731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.781731Z digest=sha256:718f93f66d8e2a7a1960bbea143bfa8d5754b94552aba31bdd0cd39651e8b3ea

Observation b00e9ecb-db77-40ea-ba34-1c2ce12e8eb3 · outbound

This paper cites Denoising Diffusion Implicit Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Denoising Diffusion Implicit Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.787127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.787127Z digest=sha256:1646d176bd80073ed7447baf6feb9e2bdea884d29a994cd7bcb7ddce82c90004

Observation 2f01a310-0bf5-4fbd-8e7f-fb13c0a66753 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.793007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.793007Z digest=sha256:de82854b98a7ab199e0fd44f4a4bb1839294c29ec66a575bcd6faa53b20e4da7

Observation 85b091e5-86cd-459a-81a0-68a15fcd8638 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.797905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.797905Z digest=sha256:4d12c5650bf39eb50f9a444e91c280eac7b7c9f031be889ee326848b68539157

Observation 021566ec-dc81-4c05-aca8-dcac1652e558 · outbound

This paper cites Diffuscene: Denoising diffu- sion models for generative indoor scene synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Diffuscene: Denoising diffu- sion models for generative indoor scene synthesis

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.803354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.803354Z digest=sha256:b122ae4574e74a69794726e44f788c72e0d3b22cae40c5eba5aa36ce974e01cc

Observation 9305c0fb-9476-48f6-a03b-e66c1acd340a · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.811998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.809546Z digest=sha256:75b786c933021147c0d0259a99e39e065aca2bbdbf6096a2eaee6e36de7a3346

Observation 2387520d-6b36-4f6e-9be4-bfe118c6ff55 · outbound

This paper cites TripoSR: Fast 3D Object Reconstruction from a Single Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation TripoSR: Fast 3D Object Reconstruction from a Single Image

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.814273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.814273Z digest=sha256:5bf67e74db65f2459fef144c1cdb28a5092cbc99567b0fa6ea20cb72cf5d550f

Observation 93544d0d-ed6c-4fc2-b764-d07cf498ddd6 · outbound

This paper cites Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.793990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.819595Z digest=sha256:1d800d45b512ccbea83009af4826b7656be00e0259d56fd94a625cf6bfb1c789

Observation 62907540-adfa-4392-b6dc-bd2287856f7c · outbound

This paper cites Neus: Learning neural im- plicit surfaces by volume rendering for multi-view recon- struction.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Neus: Learning neural im- plicit surfaces by volume rendering for multi-view recon- struction

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.776076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.824339Z digest=sha256:3ab137618ad493954d2d5ab3c555d2b8f03c3c976d48edf6555deb2fa6f40aa8

Observation 57786f34-0475-4391-a4cc-3fc6e4136d45 · outbound

This paper cites CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.829216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.829216Z digest=sha256:eac055103ea67fc95f7eafe8dcdcb8a2c6b072bdde7d2ac5ce7b348308a82a6f

Observation 7bd221cb-40b3-47df-bb91-eaf98485e198 · outbound

This paper cites Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.833706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.833706Z digest=sha256:16c0f8e8600e19f8d735d430d7a9d40f27e1b16ab80758d3fcd7356116c5f1b7

Observation 0f8074ca-6a1c-4761-9079-777815bf6a79 · outbound

This paper cites Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.839041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.839041Z digest=sha256:2b8b003ed3ffecdac66072fcc69a7b77933969d7a1444477c2a1ab97a68fc60d

Observation fa8909cb-667a-4e9d-babe-276b712589e8 · outbound

This paper cites Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.845335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.845335Z digest=sha256:af1aa51f6fae0c9aa45958327cfc7703eac6f01145a0357e75fd0abe76932b9e

Observation fba5c971-0aef-460f-a9ab-59ce9745411b · outbound

This paper cites Blockfusion: Expandable 3d scene gen- eration using latent tri-plane extrapolation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Blockfusion: Expandable 3d scene gen- eration using latent tri-plane extrapolation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.758795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.850436Z digest=sha256:51ab553e94b03fe50d64397dfab4eb16c7520eaacabd9a265bcdc0000b731c15

Observation 9776c8a8-70a4-4eb2-bb2f-4ef7cc0b03b8 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.855123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.855123Z digest=sha256:e6717ddc8c361c0bae319abdbe310694fb23d19bc9c5c4ea64119bfd66b102ee

Observation ba131941-921c-4897-b769-bb2b2e7efed5 · outbound

This paper cites GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.861066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.861066Z digest=sha256:8c1a1f7bbc4ee4b6cd307ba3f1ead251cab3b07678a5cc5d215ab502dc7d54d0

Observation 919b44c2-db8c-4e0a-828a-906f139bf3bf · outbound

This paper cites Hifi-123: Towards high-fidelity one image to 3d content gen- eration.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Hifi-123: Towards high-fidelity one image to 3d content gen- eration

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.866504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.866504Z digest=sha256:49149e8ab67f57112c0a9dc1279ce33b7deb258344c6f5ab095d63b64b019bbe

Observation 7a460b27-9180-4f30-b0c1-b4ee2e977304 · outbound

This paper cites 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models.ACM Transactions On Graphics (TOG), 42(4):1–16, 2023.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models.ACM Transactions On Graphics (TOG), 42(4):1–16, 2023

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.873513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.873513Z digest=sha256:0f795199b046fc929f21e8f1788a4f0e823db7e62bf13c86012439c13b01d36f

Observation 803b17c4-b102-4c33-89ff-fa26962839e1 · outbound

This paper cites Holistic 3d scene un- derstanding from a single image with implicit representation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Holistic 3d scene un- derstanding from a single image with implicit representation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.720197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.878453Z digest=sha256:b9f9dbff208c1279708f3fdd24b2dcc36818c7e60c41c944d86b07ef02dcc114

Observation 5e6ee03e-db88-4d40-9444-2807829d1d27 · outbound

This paper cites Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.704010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.883741Z digest=sha256:522befedf931a2a01d24e7a5e3190f2c92c5c31abf32e455541a47c1ac6bf0da

Observation b54915e5-e69d-4a60-9f04-67cacd5c000f · outbound

This paper cites Uni-3d: A universal model for panoptic 3d scene reconstruc- tion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Uni-3d: A universal model for panoptic 3d scene reconstruc- tion

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.685715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.888271Z digest=sha256:cafced034f05fed85a2a56baf5d325807d45c1a05ab6607a1a2918e452829555

Observation e329b9b4-f34f-4268-9b47-398e1a188dc8 · outbound

This paper cites Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.667464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.893202Z digest=sha256:3171dff898b04afb030e87352c3ba330a35231c294239b77f406ee93cdba9aab

Observation 23183069-f4b1-4e97-bea8-398ee8cea31f · outbound

This paper cites Zero-shot scene reconstruction from single images with deep prior as- sembly.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Zero-shot scene reconstruction from single images with deep prior as- sembly

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.649304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.897649Z digest=sha256:32caa5e4333ad4cb53e25414ca3f93db0ced6ace9416cd846b9147d6e879ea81

Observation e4009483-3653-444a-b7f4-ee62d19a5571 · outbound

This paper cites Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.630479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.903549Z digest=sha256:d1e6cf016cdc2a48171ef40051e56d59f9cacd38a2f1d4bf95f7cda8f8689c0a

Observation d0bc4880-5eca-48fc-adce-dc3268818d0a · outbound

This paper cites Following scalable 3D object generation methods [35, 71, 78, 80], we firstly trains a V AE to com- press 3D geometric representations into a low-dimensional latent space.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Following scalable 3D object generation methods [35, 71, 78, 80], we firstly trains a V AE to com- press 3D geometric representations into a low-dimensional latent space

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.611579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.908649Z digest=sha256:2251e9d5738c75fc2845065c3ee1690571368c373566cb9e0e173cf7c31357d2

Observation 41c32b14-175a-423f-977a-abc73eef6e11 · outbound

This paper cites we trained MIDI to simultaneously generate up to N = 7instances.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation we trained MIDI to simultaneously generate up to N = 7instances

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.594995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.914229Z digest=sha256:27468821a643135d0a6c63c7aae5f2df406125ae5a86618feb7c677c766bbe33

Observation c9d72829-63f6-48ce-ac3e-402ab691a560 · outbound

This paper cites compositional generation methods.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation compositional generation methods

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.577736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.919288Z digest=sha256:8245c398f37656233cd42d8ec733191eea932da1a3378ecaa0d5a63d930cbb74

Observation f827c2f4-932a-458f-9e23-5d056656da2f · outbound

This paper cites an unresolved cited work.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:19:25.560639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T22:19:24.924471Z digest=sha256:e75143cfcb8512021547780bd89f30039393458987613d1e85fe9b1107083df3

Pith citing papers

Observation ea8209b2-d077-4e41-8e51-922df5ca57d8 · inbound

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers cites this paper.

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:57.655659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:57.655659Z digest=sha256:ba34e05ccd16e958769525dc918b3696f697d4e69452562cd1d1af45cca9c39b

Observation 5582339e-cde9-40f7-9916-29b6ffea2854 · inbound

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts cites this paper.

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:11:40.108634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T17:11:26.524195Z digest=sha256:c83ca6ac7bad7639412864dcefa0fa497c8ca7a01787219ba04270f6e5da9d82

Observation 22391441-1ba5-497e-b491-b8c7d723cbec · inbound

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System cites this paper.

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.649538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T19:14:48.069637Z digest=sha256:efbf75bcb462fbd5fbe623356c4cc82187df4800eeb0a6b6d3d93ad7f332c0f2

Observation 4cdd353a-d9dd-482d-86a7-fe77014bcc66 · inbound

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures cites this paper.

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:57.299791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T04:30:56.107507Z digest=sha256:f2937ebeb3c5d4b6e1ab9346f4a56a53897546285835aa5a25ac2b623738d02c