Pith. sign in

Paper Citation Record · LEDGER

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

As of 15 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 4 inbound Pith citation observations for arXiv:2412.03558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03558 v3

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:19:24.924471Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:57.655659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T20:03:57.298307Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e235fb43-7527-4703-8188-ee4147644d52 · outbound

This paper cites Neural rgb-d surface reconstruction.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Neural rgb-d surface reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.470319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.470319Z digest=sha256:56d3066a51c50acb59bd574b67899b5d1ba907b68489ac8fbba3a010e9de0bd5

Observation 6e2bcb3f-4e3d-4f15-9319-b75e40814483 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.477120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.477120Z digest=sha256:f5619fe95e0782a74cc9c7e98b0daf909f2e18a5526a59805fc76ef572bb2a00

Observation a26b3620-e4d2-481e-ae07-4d385608d747 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.483193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.483193Z digest=sha256:895df37f2288adbc13902e3e15b37a495ae8304bbad152f2cb97834102c53a9d

Observation 1e37fa38-5f5e-4003-808c-a92f919788e9 · outbound

This paper cites Single-view 3d scene reconstruc- tion with high-fidelity shape and texture.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Single-view 3d scene reconstruc- tion with high-fidelity shape and texture

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.489199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.489199Z digest=sha256:26320df515ee8332a56915379ccc902367dfe333892e61d71d073a084656bda9

Observation 8f6e5658-3c90-42d0-b9d5-457b21d44953 · outbound

This paper cites ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.494401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.494401Z digest=sha256:6e32b2f4194721ae6c31163b6e1b9ffb905a1852b535cd3cd2c3976671d6ae42

Observation c02f3fd2-ad73-4121-8653-e66339e13dda · outbound

This paper cites Buol: A bottom-up framework with occupancy-aware lifting for panoptic 3d scene reconstruction from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Buol: A bottom-up framework with occupancy-aware lifting for panoptic 3d scene reconstruction from a single image

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.500289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.500289Z digest=sha256:6dc2d810c7f39de33fd79e6b539137dce93f366e810ca79af9dd2fbb40b23611

Observation 239bf3cb-c715-4bd7-9433-650664bdb1d0 · outbound

This paper cites Panoptic 3d scene reconstruction from a single rgb image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Panoptic 3d scene reconstruction from a single rgb image

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.506633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.506633Z digest=sha256:e20af54c3ff655ad536f4674a6112bafae19f959fbbc5a32be1a4978664a3679

Observation 09b903b6-84d8-4024-975f-5276c71ff7d4 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.512113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.512113Z digest=sha256:5d15141072b67a7aa77f9072b4477799151142c0a191c98ad07dc806fec48f06

Observation 94ab5b39-77d7-4abd-aa7f-914bbc3bc830 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Objaverse: A universe of annotated 3d objects

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.517348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.517348Z digest=sha256:f29c4923b68c56f826b8fc54d4ec73ec58e395c8b0c8fa0d6b26ca1f7d084e77

Observation 1905d084-a603-4c44-bcba-fe7c466462dc · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Objaverse-xl: A universe of 10m+ 3d objects

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.257528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.523049Z digest=sha256:ce997e8985ec5fd712a5b9bcd18388faad5bff4d46c79662c6d824bf184b88b4

Observation 1228f12e-7aaa-4928-a7dd-409e7e3bbfcb · outbound

This paper cites Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.527729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.527729Z digest=sha256:371f94c5c548f385d56ad4f4eeb047ef5d108f878a0e64a4b9d9bfe470724e8a

Observation c68931e9-63d4-46ab-ba3e-2eec7022ddd0 · outbound

This paper cites Tela: Text to layer-wise 3d clothed human generation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Tela: Text to layer-wise 3d clothed human generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.532509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.532509Z digest=sha256:484ba406058bce5e935eab226f0eb4726613fa632abda9c52b40b34a9f733b7c

Observation ba352958-b5c6-444b-a471-1507a2a77061 · outbound

This paper cites Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.538986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.538986Z digest=sha256:f6534ef3bacb9adc1cc405f6e26018469b2e32bcf4b61e81744b5eec1c47104b

Observation 595beabd-0d2e-40fb-8395-fd5b260a82c8 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.544028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.544028Z digest=sha256:6272094f088ecddcdcd588b339a6abf0ad934c5b97b50e72a17cfa508ede1916

Observation 034999a5-8e4a-43a0-8fb2-22483d787c19 · outbound

This paper cites 3d-front: 3d furnished rooms with layouts and semantics.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3d-front: 3d furnished rooms with layouts and semantics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.548492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.548492Z digest=sha256:572b5c61c3217da7acf79e676c347fdeebd27bb6d3b5b07ee3a4e6d990a696aa

Observation 5c6d46ed-3de8-4569-956d-101af54ab5ac · outbound

This paper cites 3d-future: 3d fur- niture shape with texture.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3d-future: 3d fur- niture shape with texture

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.192855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.553425Z digest=sha256:e54302131917ed2a19a2834da156ac8ba2fa46e0e0b472ee85eb260af5039573

Observation aba27113-874c-4f8d-9746-34ba865c5f5b · outbound

This paper cites Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.174826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.559147Z digest=sha256:a414e1384aa324f832c9fd6901dd92cefc15ae9fed94df3398728da84036a176

Observation a063ab1c-879f-4b75-9296-6f0737c091c9 · outbound

This paper cites Learn- ing 3d object shape and layout without 3d supervision.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Learn- ing 3d object shape and layout without 3d supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.156177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.565160Z digest=sha256:a7073f896f786df98a1807850599b15b28303ee3edef3351e0b045bab89e412c

Observation 3ad51016-5d78-4210-bcce-c97abcc04d6b · outbound

This paper cites Roca: Ro- bust cad model retrieval and alignment from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Roca: Ro- bust cad model retrieval and alignment from a single image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.136435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.570414Z digest=sha256:6e083e3752450bbc7155c448ddcb8396b8a73f765afefe952b8b4da5364e88dd

Observation 7d3f575f-4163-4bcf-adc6-c83283dab3c2 · outbound

This paper cites threestudio: A unified framework for 3d content generation, 2023.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation threestudio: A unified framework for 3d content generation, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.118486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.575723Z digest=sha256:054571cfcb239a8da201349aa09a0cef646cc0a3acf6bad3c1a9e902a0b8f077

Observation 3d281103-59ca-4932-ac68-e7c9aa24e9bd · outbound

This paper cites REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.580426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.580426Z digest=sha256:4a04f5197f70381f2d1baf881e97f3e7a21d28270a3fbd55dce5b446427ae767

Observation 5fa7230e-a9e5-43f4-99e9-d81214e9651e · outbound

This paper cites Classifier-Free Diffusion Guidance.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Classifier-Free Diffusion Guidance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.585916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.585916Z digest=sha256:c5d34608d2d338cd86495ac635c663a62f6ad2082377cf74c3cce6cfb9d3027f

Observation 69933ff9-1972-4ba0-befc-fd8bbbb83f8f · outbound

This paper cites Denoising dif- fusion probabilistic models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.592499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.592499Z digest=sha256:ed04fd879872c25bdabbea82f4c0ce81bd31ea440c55ccbb90f781ef1f5a66b6

Observation 5d2bf12b-69e0-4362-b26e-f0b1ed265da8 · outbound

This paper cites LRM: Large Reconstruction Model for Single Image to 3D.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LRM: Large Reconstruction Model for Single Image to 3D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.598267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.598267Z digest=sha256:fa62f345c6ec6335baf57296993789b709afcd65cfd22bd7f233155811fd8d6c

Observation 2018dd2e-4c69-4d98-9878-fcb0e982f5fe · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.603228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.603228Z digest=sha256:16ad8a4e86babb51da3bf4a7111db0f46b775f57b1d5cdea5bece00d7c93333b

Observation 96c46aaf-8813-4829-8cc1-e09b9da2b740 · outbound

This paper cites MV-Adapter: Multi-view Consistent Image Generation Made Easy.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation MV-Adapter: Multi-view Consistent Image Generation Made Easy

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.608784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.608784Z digest=sha256:18a8baa8f359dd8dbae7b924f1c29fd715ed19cf10cbd4b4413025c8e8c0f028

Observation 3097edca-607d-4802-b4d4-0549bfba768e · outbound

This paper cites Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.613872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.613872Z digest=sha256:0ceb932c429a08500e5c2cedf20b902a4dc5d6605fb2ee06aeee85ee4a09bbc9

Observation 36675878-c973-4ca6-8952-84f4c963f89c · outbound

This paper cites an unresolved cited work.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:19:26.078879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.618869Z digest=sha256:60b073559bce6d3ef42078feaa2a090d6af139db19f050a76d80a56eb316d2e0

Observation 7109ec4f-ecb9-494d-8c84-0b40e7bc6c64 · outbound

This paper cites Shap-E: Generating Conditional 3D Implicit Functions.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Shap-E: Generating Conditional 3D Implicit Functions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.624279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.624279Z digest=sha256:87102578e9c3e28e4fe1c5f745caff535c1dd4ff16e539872be25d44467da879

Observation bb96c0ff-5170-430f-9afd-99b179c31daf · outbound

This paper cites Auto-Encoding Variational Bayes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Auto-Encoding Variational Bayes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.629396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.629396Z digest=sha256:dd19dfa2e1f944c84f3768caebbd7d2aa3464bb0a4483f54604d15354d2e6125

Observation 6c3d10de-360d-47b4-8c78-afe325186f7a · outbound

This paper cites Segment Anything.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Segment Anything

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.634421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.634421Z digest=sha256:da2cdabc1a1ef3bf12a0c91e2a9064ba1523858f1a9234e211b4db4dafeec322

Observation 54f7fe89-3a06-41ea-b913-8f226b7e47b6 · outbound

This paper cites Mask2cad: 3d shape prediction by learning to segment and retrieve.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Mask2cad: 3d shape prediction by learning to segment and retrieve

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.062304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.639549Z digest=sha256:f2dccf5fe350981c39c9bc1e94d37527098b9348f034ef165b399c78e9d0be44

Observation 6a93a930-8502-4b47-854d-d93ef3fb6723 · outbound

This paper cites Patch2cad: Patchwise embedding learning for in-the- wild shape retrieval from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Patch2cad: Patchwise embedding learning for in-the- wild shape retrieval from a single image

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.040938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.645115Z digest=sha256:81ee3397b64dcf65f21aba24578f123c54c5de2987f48f0b2ed776a26accb6db

Observation 4f859191-a6a3-40d4-b8d9-8bccf58e4bd8 · outbound

This paper cites SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.651197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.651197Z digest=sha256:72083224006cb1284657de7cb00d9e45d432aef6500f32bb2b3c6afa6b4488e1

Observation 75e3988e-fddb-49f1-b8ac-d10803d98fcc · outbound

This paper cites CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.657182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.657182Z digest=sha256:7da74f4b36be1a0a173fc009a68fd4f85a9a49e7377676a5db97e71888eade61

Observation 175e174b-e84c-42c9-9015-f4f6c702b2f2 · outbound

This paper cites TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.662780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.662780Z digest=sha256:dffdeb8b4543edf0569f3ab932a9db9ed0d526d797d480f96c93aaefb024d47c

Observation 4e14715a-3bd3-4850-a60d-10d45d5eac70 · outbound

This paper cites Part123: part-aware 3d reconstruction from a single-view image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Part123: part-aware 3d reconstruction from a single-view image

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.024483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.668803Z digest=sha256:bf40877e5685d800a5e9295cce53016d7f4040d77f574a4bf9c0a7478f8380cb

Observation 98b75164-459c-48c0-9671-a6fdb90a1060 · outbound

This paper cites Towards high-fidelity single-view holistic reconstruction of indoor scenes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Towards high-fidelity single-view holistic reconstruction of indoor scenes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.006488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.673792Z digest=sha256:63c2f959f1db288f1a3b3b5c7721422f42051ed769a67fec6d16f84153e3e694

Observation d0fb7e2a-d33e-4a7b-8772-ac2af0b768ca · outbound

This paper cites One-2-3-45++: Fast single im- age to 3d objects with consistent multi-view generation and 3d diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation One-2-3-45++: Fast single im- age to 3d objects with consistent multi-view generation and 3d diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.990012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.678786Z digest=sha256:81f4e326a4eea4cd45ad6117266e300f80ca6cc9b278c2cb16a89d11642eeac8

Observation cda20c72-bd78-4aab-a7fe-d0a97c94dbd3 · outbound

This paper cites One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.973929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.683680Z digest=sha256:d524709b85a37d68979f3447af770a758c30beebb27d9a801b4c7401c31ecedb

Observation 47648afb-9b6b-4fc5-8461-efd9a8bac810 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.688612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.688612Z digest=sha256:246fea518b3896561171ef21b496a0436d191f06c2219b976b9124cb408213a8

Observation 13d820e1-6bc3-44ec-9861-0c422768086e · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.694510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.694510Z digest=sha256:b7eec6517535f9d43fea0a974d311665664bea9e42bbd0d69a155f6f842cc827

Observation e927f0aa-afb3-4d16-a9d1-5d3f42350950 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.700568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.700568Z digest=sha256:2e2b1811c53137085243ee00b566d895626bf9f7cdc2317d043e2a300cf5ac6b

Observation ddea7528-a219-40ec-ba4c-d59833c031cc · outbound

This paper cites Wonder3d: Sin- gle image to 3d using cross-domain diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Wonder3d: Sin- gle image to 3d using cross-domain diffusion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.706386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.706386Z digest=sha256:1466fc955174acdc0673f6518c7f5a24f3918a91830fe139e12d6f627b85ac8a

Observation 019a47af-dc24-4573-902e-3ea9154f3954 · outbound

This paper cites Marching cubes: A high resolution 3d surface construction algorithm.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Marching cubes: A high resolution 3d surface construction algorithm

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.946424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.711368Z digest=sha256:6b89512a57a5fbd3b9652a93a288789291735d032aa35e6940a1eea724e07c11

Observation 666a8032-ff54-4e60-b1a5-b0abaec7e79d · outbound

This paper cites LT3SD: Latent Trees for 3D Scene Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LT3SD: Latent Trees for 3D Scene Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.716309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.716309Z digest=sha256:b3a8f1936c038eef4f63341edc5ff741586b0847d9967947864ad7eed6edf6be

Observation e9f1308a-1369-420f-a75d-8593b327dbac · outbound

This paper cites GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.928013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.721777Z digest=sha256:9c4ebfc48175122e7453c1e23b574520bf297d4acbac4a7cdc43d4999d2700ca

Observation 54ec66ed-b173-448b-b89e-1c23a5b7e666 · outbound

This paper cites Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.910886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.727319Z digest=sha256:251497a9e0e82cb85b0041799a901b2a51b0b0d153edf10b96fd3c19e2f52f53

Observation 0c0282a3-f776-4280-8189-4e258e7b3018 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.731797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.731797Z digest=sha256:ae3d615dde9a1b3f2c594ee82521aecdcd2283e962d6a06ca12c5d86f88ce33b

Observation e24e5a4d-0044-4f4d-a413-796cca81d393 · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Atiss: Autoregres- sive transformers for indoor scene synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.737191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.737191Z digest=sha256:84cbab23fa2b527d36f15b356eb18f2b4a3b96beb4c7e7f905c87600319ae880

Observation 66b52f76-11e1-42e9-89e2-5415084b1480 · outbound

This paper cites Scalable diffusion models with transformers.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scalable diffusion models with transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.742085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.742085Z digest=sha256:76dc7316ac10953c089214f2d5e2d090608a835fc42fc1e6fb2049b99f5f8870

Observation b8817661-99de-421b-b153-8fd0e05d1cf8 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.746930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.746930Z digest=sha256:f7e102000fd588486f997fe6d55dcfd03f3403ce681ff8fa05fc5ce9de88b324

Observation ea90cd2f-9f21-4682-bce3-cd66911dcd04 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.752142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.752142Z digest=sha256:35b11fad293da4a2f2bf23ffb9dcb33084b04e4affea483a382f4ba388830681

Observation 0aee56db-70e7-46d5-986f-f0326180cebe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.756723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.756723Z digest=sha256:e321fcaff8b1696bd771cee27a7bf3e2595c001258481e0ed2b3d6f417c30ef0

Observation 5910e255-592b-4ac9-9f78-605d0f2f09c8 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.761588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.761588Z digest=sha256:c677ed0022d0919ccb31fbd855ea1702ba65a90efc39379695d2d2f1cb9f4513

Observation b1c76509-8895-41e3-b0d9-73ab3a7b02e6 · outbound

This paper cites L3DG: Latent 3D Gaussian Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation L3DG: Latent 3D Gaussian Diffusion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.766543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.766543Z digest=sha256:dae52d2bdccc779c8d7f17e6afaa1c4f5a83e271b7f192129af4e53f7d34731b

Observation 1b3cd6f3-fc8a-4551-9bd2-193a6ce60af8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation High-resolution image synthesis with latent diffusion models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.771589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.771589Z digest=sha256:cf207460a3fcb35cc93bc94414c1a41390cfbefcfe8fe0bb8f236b5245ebc905

Observation 2efe727b-e8ed-4d56-85b9-9871961d6cf0 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.776858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.776858Z digest=sha256:3c42ae11a2b18d318e278f21d7265c832c5c37e877430183c4a80a2360ca1b4b

Observation c28f9010-5b4c-4dc1-8fbf-5bfff51f1762 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Deep unsupervised learning using nonequilibrium thermodynamics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.781731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.781731Z digest=sha256:718f93f66d8e2a7a1960bbea143bfa8d5754b94552aba31bdd0cd39651e8b3ea

Observation b00e9ecb-db77-40ea-ba34-1c2ce12e8eb3 · outbound

This paper cites Denoising Diffusion Implicit Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Denoising Diffusion Implicit Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.787127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.787127Z digest=sha256:1646d176bd80073ed7447baf6feb9e2bdea884d29a994cd7bcb7ddce82c90004

Observation 2f01a310-0bf5-4fbd-8e7f-fb13c0a66753 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.793007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.793007Z digest=sha256:de82854b98a7ab199e0fd44f4a4bb1839294c29ec66a575bcd6faa53b20e4da7

Observation 85b091e5-86cd-459a-81a0-68a15fcd8638 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.797905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.797905Z digest=sha256:4d12c5650bf39eb50f9a444e91c280eac7b7c9f031be889ee326848b68539157

Observation 021566ec-dc81-4c05-aca8-dcac1652e558 · outbound

This paper cites Diffuscene: Denoising diffu- sion models for generative indoor scene synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Diffuscene: Denoising diffu- sion models for generative indoor scene synthesis

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.803354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.803354Z digest=sha256:b122ae4574e74a69794726e44f788c72e0d3b22cae40c5eba5aa36ce974e01cc

Observation 9305c0fb-9476-48f6-a03b-e66c1acd340a · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.811998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.809546Z digest=sha256:4705116dca059ddafad887ff39773fe31a388f0596ca71cbc7f7f3638441c799

Observation 2387520d-6b36-4f6e-9be4-bfe118c6ff55 · outbound

This paper cites TripoSR: Fast 3D Object Reconstruction from a Single Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation TripoSR: Fast 3D Object Reconstruction from a Single Image

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.814273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.814273Z digest=sha256:5bf67e74db65f2459fef144c1cdb28a5092cbc99567b0fa6ea20cb72cf5d550f

Observation 93544d0d-ed6c-4fc2-b764-d07cf498ddd6 · outbound

This paper cites Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.793990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.819595Z digest=sha256:f1e17dc491c6d5fd664a566dcc623b1171dc09ea6d9f63375ca834a470e3eac4

Observation 62907540-adfa-4392-b6dc-bd2287856f7c · outbound

This paper cites Neus: Learning neural im- plicit surfaces by volume rendering for multi-view recon- struction.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Neus: Learning neural im- plicit surfaces by volume rendering for multi-view recon- struction

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.776076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.824339Z digest=sha256:3c18d0d758e6ac4d06da59c03a2b9776a39ef3f2618824a7897d8fd123f8a464

Observation 57786f34-0475-4391-a4cc-3fc6e4136d45 · outbound

This paper cites CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.829216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.829216Z digest=sha256:eac055103ea67fc95f7eafe8dcdcb8a2c6b072bdde7d2ac5ce7b348308a82a6f

Observation 7bd221cb-40b3-47df-bb91-eaf98485e198 · outbound

This paper cites Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.833706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.833706Z digest=sha256:16c0f8e8600e19f8d735d430d7a9d40f27e1b16ab80758d3fcd7356116c5f1b7

Observation 0f8074ca-6a1c-4761-9079-777815bf6a79 · outbound

This paper cites Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.839041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.839041Z digest=sha256:2b8b003ed3ffecdac66072fcc69a7b77933969d7a1444477c2a1ab97a68fc60d

Observation fa8909cb-667a-4e9d-babe-276b712589e8 · outbound

This paper cites Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.845335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.845335Z digest=sha256:af1aa51f6fae0c9aa45958327cfc7703eac6f01145a0357e75fd0abe76932b9e

Observation fba5c971-0aef-460f-a9ab-59ce9745411b · outbound

This paper cites Blockfusion: Expandable 3d scene gen- eration using latent tri-plane extrapolation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Blockfusion: Expandable 3d scene gen- eration using latent tri-plane extrapolation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.758795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.850436Z digest=sha256:e5aa9a1dac03fdfa10f60cfe1a3786f76368db0c24e0d6a067ede87eb71af0e7

Observation 9776c8a8-70a4-4eb2-bb2f-4ef7cc0b03b8 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.855123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.855123Z digest=sha256:e6717ddc8c361c0bae319abdbe310694fb23d19bc9c5c4ea64119bfd66b102ee

Observation ba131941-921c-4897-b769-bb2b2e7efed5 · outbound

This paper cites GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.861066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.861066Z digest=sha256:8c1a1f7bbc4ee4b6cd307ba3f1ead251cab3b07678a5cc5d215ab502dc7d54d0

Observation 919b44c2-db8c-4e0a-828a-906f139bf3bf · outbound

This paper cites Hifi-123: Towards high-fidelity one image to 3d content gen- eration.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Hifi-123: Towards high-fidelity one image to 3d content gen- eration

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.866504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.866504Z digest=sha256:49149e8ab67f57112c0a9dc1279ce33b7deb258344c6f5ab095d63b64b019bbe

Observation 7a460b27-9180-4f30-b0c1-b4ee2e977304 · outbound

This paper cites 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models.ACM Transactions On Graphics (TOG), 42(4):1–16, 2023.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models.ACM Transactions On Graphics (TOG), 42(4):1–16, 2023

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.873513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.873513Z digest=sha256:0f795199b046fc929f21e8f1788a4f0e823db7e62bf13c86012439c13b01d36f

Observation 803b17c4-b102-4c33-89ff-fa26962839e1 · outbound

This paper cites Holistic 3d scene un- derstanding from a single image with implicit representation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Holistic 3d scene un- derstanding from a single image with implicit representation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.720197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.878453Z digest=sha256:6baf5ca4c646e60252705b06ba90eb798153d68e76ed57f8682a460f82af11e1

Observation 5e6ee03e-db88-4d40-9444-2807829d1d27 · outbound

This paper cites Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.704010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.883741Z digest=sha256:6f38b87ba864a119b6e4c69a73add4dc41df656ba065ed2140b00edeb4a11fa6

Observation b54915e5-e69d-4a60-9f04-67cacd5c000f · outbound

This paper cites Uni-3d: A universal model for panoptic 3d scene reconstruc- tion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Uni-3d: A universal model for panoptic 3d scene reconstruc- tion

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.685715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.888271Z digest=sha256:ebb56a324c6f68fd163eebba7e230fdf6b0df8e60480b34e280950b004aa176b

Observation e329b9b4-f34f-4268-9b47-398e1a188dc8 · outbound

This paper cites Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.667464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.893202Z digest=sha256:2f3cd43845adba60c348bc54f538f6ee8dc5e805f71a1dd26492d42e2ae418f4

Observation 23183069-f4b1-4e97-bea8-398ee8cea31f · outbound

This paper cites Zero-shot scene reconstruction from single images with deep prior as- sembly.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Zero-shot scene reconstruction from single images with deep prior as- sembly

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.649304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.897649Z digest=sha256:e0f5e50e298379c15a24a16977eb26e04ce172f4076ca7ab5667ebbca63ef75e

Observation e4009483-3653-444a-b7f4-ee62d19a5571 · outbound

This paper cites Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.630479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.903549Z digest=sha256:7efd18f65831b5239d6f2adfe40202d251c0ae0388e2168a59b4e7051691c255

Observation d0bc4880-5eca-48fc-adce-dc3268818d0a · outbound

This paper cites Following scalable 3D object generation methods [35, 71, 78, 80], we firstly trains a V AE to com- press 3D geometric representations into a low-dimensional latent space.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Following scalable 3D object generation methods [35, 71, 78, 80], we firstly trains a V AE to com- press 3D geometric representations into a low-dimensional latent space

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.611579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.908649Z digest=sha256:56773c942df431a813e9383ff8e10682de8396a79e8f9a23d20c6dcc14eacfd6

Observation 41c32b14-175a-423f-977a-abc73eef6e11 · outbound

This paper cites we trained MIDI to simultaneously generate up to N = 7instances.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation we trained MIDI to simultaneously generate up to N = 7instances

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.594995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.914229Z digest=sha256:dbba77a60d28464602d79faa6b03a08cfaeea3034d5f287206af60f914c698fe

Observation c9d72829-63f6-48ce-ac3e-402ab691a560 · outbound

This paper cites compositional generation methods.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation compositional generation methods

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.577736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.919288Z digest=sha256:647312816bb521452cc99b4cbd9ef52aecde306d90bd720fd078c9bea46694e6

Observation f827c2f4-932a-458f-9e23-5d056656da2f · outbound

This paper cites an unresolved cited work.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:19:25.560639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:19:24.924471Z digest=sha256:613e7be7795106c53427a8652dd2269d6892a590e4a5c77aa7136f4446197302

Pith citing papers

Observation ea8209b2-d077-4e41-8e51-922df5ca57d8 · inbound

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers cites this paper.

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:57.655659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:57.655659Z digest=sha256:ba34e05ccd16e958769525dc918b3696f697d4e69452562cd1d1af45cca9c39b

Observation 5582339e-cde9-40f7-9916-29b6ffea2854 · inbound

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts cites this paper.

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:11:40.108634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T17:11:26.524195Z digest=sha256:99960223840a7c928de41ea8ffce447b6e627117e0f626f736729f0b8de5acd4

Observation 22391441-1ba5-497e-b491-b8c7d723cbec · inbound

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System cites this paper.

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.649538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T19:14:48.069637Z digest=sha256:a501961d3e75d1312b284ee087aa14b87ace2363405651fb705a5e04813c6fbd

Observation 4cdd353a-d9dd-482d-86a7-fe77014bcc66 · inbound

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures cites this paper.

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:57.299791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T04:30:56.107507Z digest=sha256:73611f10bc4a82f02013b5d089c619a6eb141d3044641b8f972bc0a0c2b0dcb8