REVIEW 3 major objections 5 minor 87 references
MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read MUVOD is a multi-view video dataset whose panoptic masks keep 459 object instances identifiable across time and across camera views.
desk verdict MUVOD fills a real gap in multi-view video segmentation, but the paper must fix inconsistent image counts and show the XMem-based ground truth is not circularly inflating its own baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a spatio-temporal annotation pipeline that turns sparse manual input into dense ground-truth masks: for the initial camera, keyframes are manually boxed and masks are produced by a promptable segmentation foundation model, then refined; those masks are propagated to neighboring cameras with a long-term video object segmentation model, with view-frustum similarity, triangulated 3D point clouds, and projected-point prompts providing geometric guidance; bidirectional temporal propagation fills the remaining frames. A depth-layer ordering for static and environmental objects lets the pipeline treat occlusion deterministically by overlaying masks from farthest to nearest. On the evaluation side, the paper's metric averages the standard 2D video object segmentation score, the mean of region similarity $J$ and contour accuracy $F$, over the $N$ cameras being used: $J\&F^N = \frac{1}{N} \sum_{i=0}^{N-1} (J\&F)_{c_i}$.
What would settle it
Manually re-segment a random sample of frames from multiple annotated views, then compute the mean IoU between those fresh manual masks and the dataset's propagated masks; if that agreement is below roughly 90% IoU, the propagated masks are too unreliable to support the benchmark comparisons.
Extended reading notes
Core claim
The paper's central claim is that MUVOD provides a resource that combines synchronized multi-view video with dense, instance-level, view-consistent annotations in real dynamic scenes. It contains 459 instances across 73 categories, with each object assigned a semantic label, a unique instance ID, and a motion status of dynamic, static, or environmental; dynamic and static objects are treated as things, while environmental objects are treated as stuff. The annotations are produced by a semi-automatic pipeline that manually refines keyframes of one initial camera and propagates masks to other views and other frames, using depth layers to resolve occlusions deterministically. On top of the dataset, the paper defines a multi-view video object segmentation task and a camera-averaged metric, and it reports a baseline reaching 79.4% J&F on three views averaged over all scenes. The paper further presents a 3D segmentation benchmark of 50 objects and evaluates four state-of-the-art methods on it, with the best method averaging 78.8% IoU across the four object difficulty classes.
Load-bearing premise
The propagated masks are accurate enough to count as ground truth even though only keyframes of one camera are manually refined.
Editorial extensions
If this is right
- 4D object segmentation methods can be trained and compared on real dynamic scenes with many objects instead of isolated humans or street views.
- The camera-averaged J&F metric gives the field a common yardstick for spatio-temporal mask propagation across viewpoints.
- The 3D segmentation benchmark exposes where current lifting methods lose accuracy, particularly on occluded, small-scale, and complex-structure objects.
- Downstream applications such as 4D scene editing and VR/AR interaction become testable because object identity is already linked across space and time.
Reading between the lines
- Beyond the paper, a human-agreement study on a held-out set of frames would test the central accuracy assumption, since the pipeline only refines keyframes of one camera.
- Beyond the paper, the depth-layer representation suggests a testable variant: estimate depth ordering automatically and measure how much annotation time and mask accuracy change.
- Beyond the paper, the sparse-view experiment implies that camera-rig geometry is itself a performance variable, so future benchmarks might report results as a function of view density rather than for a single rig configuration.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents MUVOD, a multi-view video dataset of 17 realistic scenes with panoptic-style instance annotations in space and time, a semi-automatic annotation pipeline (manual keyframes, SAM, XMem spatial/temporal propagation, LightGlue geometric cues), an evaluation metric J&F_N for multi-view video object segmentation, and a baseline that applies XMem across views and time. It also introduces a 3D object segmentation benchmark with 50 target objects from 12 scenes and evaluates ISRF, SA3D, SAGA, and Gaussian Grouping. The central claims are that MUVOD is a large, diverse, richly annotated multi-view video dataset and that the proposed benchmark enables meaningful evaluation of 4D/3D object segmentation.
Significance. The dataset addresses a real gap: prior multi-view video datasets are limited to specific domains such as street scenes, furniture, or human-only scenes. The inclusion of 17 scenes from four sources with 9 to 46 views, 459 instances, and 73 categories, together with a defined task, a clearly stated metric, a baseline, and a 3D benchmark with four state-of-the-art methods, is a potentially useful community resource. The paper provides a dataset URL and the benchmark is a concrete contribution. The paper is not a theoretical derivation; its value depends on whether the released ground-truth masks are accurate and consistent. That dependency is currently unmet, so the significance is conditional on the requested validation and on reconciling the reported dataset scale.
major comments (3)
- [Abstract; Section I; Section III.C] The headline dataset scale is internally inconsistent. The abstract states that MUVOD provides 7830 RGB images (30 frames per video), but Section I says only three camera views per scene are annotated and that each view contains 21 frames with manually refined ground-truth masks, and Section III.C describes sampling keyframes every 10 frames from a single initial camera. Under the stated protocol, 17 scenes times 3 views times 21 frames gives 1071 manually refined frames, and even taking 30 frames per video gives 1530, not 7830. The number 7830 is close to 375 available views times 21 frames, suggesting that all views rather than three are annotated, which conflicts with Section I. Please reconcile the count and specify, for each released mask, whether it is manually refined, spatially propagated, or temporally propagated.
- [Section III.C; Section IV.C; Fig. 7] Ground-truth mask accuracy is not validated, and the baseline evaluation is at risk of self-agreement. In the annotation pipeline, only keyframes of the initial camera cini are manually refined; non-initial camera keyframes and all intermediate frames are produced by propagating masks with XMem and SAM. The baseline method in Section IV.C is exactly XMem applied spatially and temporally. Systematic XMem errors, such as label switches or boundary drift, can therefore be embedded in the 'ground truth' and inflate the scores reported in Tables III and IV; Fig. 7 documents that masks can be applied to nearby objects of the same class, and no evidence is given that the ground-truth column is not affected by the same failure mode. No inter-annotator agreement, per-frame error rate, or independent quality check is reported. Please add a quantitative validation study, for example human correction rates on a random sample of frames or a comparison against independently produced masks, and report how much of the baseline score is self-agreement with the automatically propagated labels.
- [Section V.B; Section V.C; Table VIII] The 3D object segmentation benchmark's per-condition conclusions rest on very few objects. Table VI shows that occluded, small-scale, and complex-structure categories each contain only 3 objects, and Table VIII reports mean IoU for these conditions without variance or per-object results. With n=3 per condition, the paper's claim of a 'more comprehensive analysis' of method strengths and limitations is stronger than the evidence supports. The final paragraph of Section V.C acknowledges the small subset, but the quantitative tables should also report per-object IoU or confidence intervals, and the conclusions should be correspondingly qualified.
minor comments (5)
- [Abstract] The first sentence has a subject-verb agreement error: 'The application of methods ... have steadily gained popularity' should be 'has steadily gained popularity'.
- [Table I] The caption contains a typo: 'SPATIO-TEMPORAL COSISTENCE' should read 'CONSISTENCE'.
- [Section II.B; Section V.A] The name of the 3D segmentation method by Ye et al. is rendered inconsistently as 'LeRF-Mask' in Section V.A and 'LERF-Mask' in Section II.B; please use one spelling throughout.
- [Figure 5] The axis labels and tick labels in Figure 5 appear as garbled placeholder glyphs in the manuscript, making the figure difficult to read; please regenerate it with standard fonts.
- [Section V.B] The benchmark description does not mention whether the evaluation prompts (points or strokes) are provided under the same protocol across the four methods, which matters for fairness; please specify the prompting procedure.
Circularity Check
The benchmark's ground truth is produced by the same XMem propagation used as the baseline, so the reported J&F scores largely measure XMem against itself.
-
other
[Section III.C (Annotation Method, Keyframes spatial propagation) and Section IV.C (Baseline Evaluation)]
"The annotated masks from the keyframes of the initial camera cini are propagated to other camera views using the semi-supervised video object segmentation model XMem [57]. ... we adopt a straightforward two-stage baseline by applying XMem [57] both spatially and temporally across multi-view video sequences."
The masks used as ground truth for evaluating the baseline are generated by the same XMem model that the baseline applies. Non-initial-camera keyframes are produced by spatial XMem propagation from cini keyframes, and non-keyframe temporal masks are propagated from manually corrected keyframes; the baseline then runs XMem spatially and temporally over the same input. The J&F3 scores in Tables III, IV and V therefore measure XMem's agreement with its own propagation output rather than independent annotation accuracy. No independent verification (e.g., inter-annotator agreement or error-rate analysis) of the propagated masks is reported, so the benchmark numbers are not an external check on the dataset.
full rationale
MUVOD is a dataset and benchmark paper rather than a theoretical derivation, and most of its content is not circular: the category taxonomy, scene selection, metric definition, and 3D-object-segmentation comparison are independent of the paper's claims. The one load-bearing circular step is the overlap between annotation generation and baseline evaluation. Section III.C states that masks from the initial camera's keyframes are propagated to other views with XMem, and Section IV.C defines the baseline as applying XMem both spatially and temporally. Consequently the ground-truth masks scored in Tables III-V are largely XMem outputs, so the reported 79.4% and 75.6% J&F3 values characterize self-consistency of the propagation pipeline, not accuracy against human-verified labels. The paper also reports no quantitative mask-quality check or inter-annotator agreement, and the abstract's 7830 RGB images (30 frames per video) is hard to reconcile with the protocol of three annotated views per scene (17 x 3 x 30 = 1530 frames), which further undermines the reliability claim. These are dataset-validity concerns; they do not involve a self-citation chain or fitted parameters, but they do make the benchmark's headline numbers circular by construction. I therefore assign a partial circularity score of 6.
Assumptions & free parameters
assumptions (4)
- domain assumption The source multi-view videos are time-synchronized and geometrically calibrated, allowing cross-view propagation.
- domain assumption The depth-layer ordering assigned by annotators resolves occlusions deterministically.
- domain assumption XMem and SAM produce masks accurate enough for propagation, and manual refinement of keyframes corrects their errors in all propagated frames.
- domain assumption The J&F_N metric (average of per-camera J&F) validly measures multi-view video object segmentation performance.
Cite this review
Pith. "Pith review of MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation." pith.science (2026). https://pith.science/paper/MPT5NHSN
@misc{pith2026250707519,
author = {Pith},
title = {Pith review of: MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/MPT5NHSN}},
note = {Machine review of arXiv:2507.07519}
}
read the original abstract
The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate efficacy in a range of 3D scene understanding and editing tasks. Nevertheless, the 4D object segmentation of dynamic scenes remains an underexplored field due to the absence of a sufficiently extensive and accurately labelled multi-view video dataset. In this paper, we present MUVOD, a new multi-view video dataset for training and evaluating object segmentation in reconstructed real-world scenarios. The 17 selected scenes, describing various indoor or outdoor activities, are collected from different sources of datasets originating from various types of camera rigs. Each scene contains a minimum of 9 views and a maximum of 46 views. We provide 7830 RGB images (30 frames per video) with their corresponding segmentation mask in 4D motion, meaning that any object of interest in the scene could be tracked across temporal frames of a given view or across different views belonging to the same camera rig. This dataset, which contains 459 instances of 73 categories, is intended as a basic benchmark for the evaluation of multi-view video segmentation methods. We also present an evaluation metric and a baseline segmentation approach to encourage and evaluate progress in this evolving field. Additionally, we propose a new benchmark for 3D object segmentation task with a subset of annotated multi-view images selected from our MUVOD dataset. This subset contains 50 objects of different conditions in different scenarios, providing a more comprehensive analysis of state-of-the-art 3D object segmentation methods. Our proposed MUVOD dataset is available at https://volumetric-repository.labs.b-com.com/#/muvod.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
3d face reconstruction and gaze tracking in the hmd for virtual interaction,
S.-Y . Chen, Y .-K. Lai, S. Xia, P. L. Rosin, and L. Gao, “3d face reconstruction and gaze tracking in the hmd for virtual interaction,” IEEE Transactions on Multimedia , vol. 25, pp. 3166–3179, 2022
2022
-
[2]
Real-time visual–inertial slam based on adap- tive keyframe selection for mobile ar applications,
J.-C. Piao and S.-D. Kim, “Real-time visual–inertial slam based on adap- tive keyframe selection for mobile ar applications,” IEEE Transactions on Multimedia, vol. 21, no. 11, pp. 2827–2836, 2019
2019
-
[3]
Learning a multi-view stereo machine,
A. Kar, C. H ¨ane, and J. Malik, “Learning a multi-view stereo machine,” Advances in neural information processing systems , vol. 30, 2017. 14
2017
-
[4]
Soft 3d reconstruction for view synthesis,
E. Penner and L. Zhang, “Soft 3d reconstruction for view synthesis,” ACM Transactions on Graphics (TOG) , vol. 36, no. 6, pp. 1–11, 2017
2017
-
[5]
High-quality streamable free- viewpoint video,
A. Collet, M. Chuang, P. Sweeney, D. Gillett, D. Evseev, D. Calabrese, H. Hoppe, A. Kirk, and S. Sullivan, “High-quality streamable free- viewpoint video,” ACM Transactions on Graphics (ToG), vol. 34, no. 4, pp. 1–13, 2015
2015
-
[6]
Fusion4d: Real-time performance capture of challenging scenes,
M. Dou, S. Khamis, Y . Degtyarev, P. Davidson, S. R. Fanello, A. Kow- dle, S. O. Escolano, C. Rhemann, D. Kim, J. Taylor et al., “Fusion4d: Real-time performance capture of challenging scenes,” ACM Transac- tions on Graphics (ToG) , vol. 35, no. 4, pp. 1–13, 2016
2016
-
[7]
V oxel structure-based mesh reconstruction from a 3d point cloud,
C. Lv, W. Lin, and B. Zhao, “V oxel structure-based mesh reconstruction from a 3d point cloud,” IEEE Transactions on Multimedia , vol. 24, pp. 1815–1829, 2021
2021
-
[8]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
Show all 87 references
-
[9]
Plenoxels: Radiance fields without neural networks,
S. Fridovich-Keil, A. Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa, “Plenoxels: Radiance fields without neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5501–5510
2022
-
[10]
Instant neural graphics primitives with a multiresolution hash encoding,
T. M ¨uller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM Transactions on Graphics (ToG), vol. 41, no. 4, pp. 1–15, 2022
2022
-
[11]
Mip-nerf: A multiscale representation for anti- aliasing neural radiance fields,
J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan, “Mip-nerf: A multiscale representation for anti- aliasing neural radiance fields,” in Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision , 2021, pp. 5855–5864
2021
-
[12]
Point-nerf: Point-based neural radiance fields,
Q. Xu, Z. Xu, J. Philip, S. Bi, Z. Shu, K. Sunkavalli, and U. Neumann, “Point-nerf: Point-based neural radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5438–5448
2022
-
[13]
Sd-nerf: Towards lifelike talking head animation via spatially-adaptive dual-driven nerfs,
S. Shen, W. Li, X. Huang, Z. Zhu, J. Zhou, and J. Lu, “Sd-nerf: Towards lifelike talking head animation via spatially-adaptive dual-driven nerfs,” IEEE Transactions on Multimedia , 2023
2023
-
[14]
Gaussian grouping: Segment and edit anything in 3d scenes,
M. Ye, M. Danelljan, F. Yu, and L. Ke, “Gaussian grouping: Segment and edit anything in 3d scenes,” arXiv preprint arXiv:2312.00732, 2023
2023 arXiv
-
[15]
Interactive segmen- tation of radiance fields,
R. Goel, D. Sirikonda, S. Saini, and P. Narayanan, “Interactive segmen- tation of radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 4201–4211
2023
-
[16]
Segment anything in 3d with nerfs,
J. Cen, Z. Zhou, J. Fang, W. Shen, L. Xie, D. Jiang, X. Zhang, Q. Tian et al., “Segment anything in 3d with nerfs,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[17]
Segment any 3d gaussians,
J. Cen, J. Fang, C. Yang, L. Xie, X. Zhang, W. Shen, and Q. Tian, “Segment any 3d gaussians,” arXiv preprint arXiv:2312.00860 , 2023
2023 arXiv
-
[18]
Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields,
S. Zhou, H. Chang, S. Jiang, Z. Fan, Z. Zhu, D. Xu, P. Chari, S. You, Z. Wang, and A. Kadambi, “Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp...
2024
-
[19]
Pointnet++: Deep hierarchical feature learning on point sets in a metric space,
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[20]
Geometric back-projection net- work for point cloud classification,
S. Qiu, S. Anwar, and N. Barnes, “Geometric back-projection net- work for point cloud classification,” IEEE Transactions on Multimedia , vol. 24, pp. 1943–1955, 2021
1943
-
[21]
3d object segmentation using cross-window point transformer with latent semantic boundary guidance,
Q. Wang, D. Liu, Z. Liu, J. Xu, and J. Tan, “3d object segmentation using cross-window point transformer with latent semantic boundary guidance,” IEEE Transactions on Multimedia , 2023
2023
-
[22]
Pcl: Point contrast and labeling for weakly supervised point cloud semantic segmentation,
A. Du, T. Zhou, S. Pang, Q. Wu, and J. Zhang, “Pcl: Point contrast and labeling for weakly supervised point cloud semantic segmentation,” IEEE Transactions on Multimedia , 2024
2024
-
[23]
Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields,
L. Song, A. Chen, Z. Li, Z. Chen, L. Chen, J. Yuan, Y . Xu, and A. Geiger, “Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields,” IEEE Transactions on Visualization and Computer Graphics , vol. 29, no. 5, pp. 2732–2742, 2023
2023
-
[24]
Robust dynamic radiance fields,
Y .-L. Liu, C. Gao, A. Meuleman, H.-Y . Tseng, A. Saraf, C. Kim, Y .-Y . Chuang, J. Kopf, and J.-B. Huang, “Robust dynamic radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 13–23
2023
-
[25]
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis,
J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis,” in 3DV, 2024
2024
-
[26]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,
Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 331–20 341
2024
-
[27]
Spacetime gaussian feature splatting for real-time dynamic view synthesis,
Z. Li, Z. Chen, Z. Li, and Y . Xu, “Spacetime gaussian feature splatting for real-time dynamic view synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 8508–8520
2024
-
[28]
4d gaussian splatting for real-time dynamic scene rendering,
G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 310–20 320
2024
-
[29]
How personalized and effective is immersive virtual reality in education? a systematic literature review for the last decade,
A. Marougkas, C. Troussas, A. Krouska, and C. Sgouropoulou, “How personalized and effective is immersive virtual reality in education? a systematic literature review for the last decade,” Multimedia Tools and Applications, vol. 83, no. 6, pp. 18 185–18 233, 2024
2024
-
[30]
SemanticKITTI: A Dataset for Semantic Scene Under- standing of LiDAR Sequences,
J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, C. Stachniss, and J. Gall, “SemanticKITTI: A Dataset for Semantic Scene Under- standing of LiDAR Sequences,” in Proc. of the IEEE/CVF International Conf. on Computer Vision (ICCV) , 2019
2019
-
[31]
Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,
Y . Liao, J. Xie, and A. Geiger, “Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 3, pp. 3292– 3310, 2022
2022
-
[32]
The ikea asm dataset: Understanding people assembling furniture through actions, objects and pose,
Y . Ben-Shabat, X. Yu, F. Saleh, D. Campbell, C. Rodriguez-Opazo, H. Li, and S. Gould, “The ikea asm dataset: Understanding people assembling furniture through actions, objects and pose,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021...
2021
-
[33]
Hi4d: 4d instance segmentation of close human interaction,
Y . Yin, C. Guo, M. Kaufmann, J. J. Zarate, J. Song, and O. Hilliges, “Hi4d: 4d instance segmentation of close human interaction,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17 016–17 027
2023
-
[34]
4d temporally coherent light-field video,
A. Mustafa, M. V olino, J.-Y . Guillemaut, and A. Hilton, “4d temporally coherent light-field video,” in 2017 International Conference on 3D Vision (3DV). IEEE, 2017, pp. 29–37
2017
-
[35]
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,
C. Ionescu, D. Papava, V . Olaru, and C. Sminchisescu, “Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,” IEEE transactions on pattern analysis and machine intelligence, vol. 36, no. 7, pp. 1325–1339, 2013
2013
-
[36]
Outdoor dynamic 3-d scene reconstruction,
H. Kim, J.-Y . Guillemaut, T. Takai, M. Sarim, and A. Hilton, “Outdoor dynamic 3-d scene reconstruction,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 11, pp. 1611–1622, 2012
2012
-
[37]
High-quality video view interpolation using a layered representation,
C. L. Zitnick, S. B. Kang, M. Uyttendaele, S. Winder, and R. Szeliski, “High-quality video view interpolation using a layered representation,” ACM transactions on graphics (TOG), vol. 23, no. 3, pp. 600–608, 2004
2004
-
[38]
Unstructured video-based rendering: Interactive exploration of casually captured videos,
L. Ballan, G. J. Brostow, J. Puwein, and M. Pollefeys, “Unstructured video-based rendering: Interactive exploration of casually captured videos,” in ACM SIGGRAPH 2010 papers , 2010, pp. 1–11
2010
-
[39]
Mose: A new dataset for video object segmentation in complex scenes,
H. Ding, C. Liu, S. He, X. Jiang, P. H. Torr, and S. Bai, “Mose: A new dataset for video object segmentation in complex scenes,” inProceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 20 224–20 234
2023
-
[40]
Video panoptic segmenta- tion,
D. Kim, S. Woo, J.-Y . Lee, and I. S. Kweon, “Video panoptic segmenta- tion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 9859–9868
2020
-
[41]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields,
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, “Mip-nerf 360: Unbounded anti-aliased neural radiance fields,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5470–5479
2022
-
[42]
Tensorf: Tensorial radiance fields,
A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su, “Tensorf: Tensorial radiance fields,” in European Conference on Computer Vision . Springer, 2022, pp. 333–350
2022
-
[43]
3d gaussian splatting for real-time radiance field rendering,
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics (ToG), vol. 42, no. 4, pp. 1–14, 2023
2023
-
[44]
Garfield: Group anything with radiance fields,
C. M. Kim, M. Wu, J. Kerr, M. Tancik, K. Goldberg, and A. Kanazawa, “Garfield: Group anything with radiance fields,” in arXiv, 2024
2024
-
[45]
Nerfshop: Interactive editing of neural radiance fields,
C. Jambon, B. Kerbl, G. Kopanas, S. Diolatzis, T. Leimk ¨uhler, and G. Drettakis, “Nerfshop: Interactive editing of neural radiance fields,” Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 6, no. 1, May 2023. [Online]. Available: https://repo-sam.i...
2023
-
[46]
In-place scene labelling and understanding with implicit scene representation,
S. Zhi, T. Laidlow, S. Leutenegger, and A. J. Davison, “In-place scene labelling and understanding with implicit scene representation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 15 838–15 847
2021
-
[47]
Panoptic lifting for 3d scene understanding with neural fields,
Y . Siddiqui, L. Porzi, S. R. Bul `o, N. M ¨uller, M. Nießner, A. Dai, and P. Kontschieder, “Panoptic lifting for 3d scene understanding with neural fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9043–9052. 15
2023
-
[48]
Panoptic neural fields: A semantic object-aware neural scene representation,
A. Kundu, K. Genova, X. Yin, A. Fathi, C. Pantofaru, L. J. Guibas, A. Tagliasacchi, F. Dellaert, and T. Funkhouser, “Panoptic neural fields: A semantic object-aware neural scene representation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2022
-
[49]
The replica dataset: A digital replica of indoor spaces,
J. Straub, T. Whelan, L. Ma, Y . Chen, E. Wijmans, S. Green, J. J. Engel, R. Mur-Artal, C. Ren, S. Verma et al. , “The replica dataset: A digital replica of indoor spaces,” arXiv preprint arXiv:1906.05797 , 2019
1906 arXiv
-
[50]
Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding,
M. Roberts, J. Ramapuram, A. Ranjan, A. Kumar, M. A. Bautista, N. Paczan, R. Webb, and J. M. Susskind, “Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 1...
2021
-
[51]
Are we ready for autonomous driving? the kitti vision benchmark suite,
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in 2012 IEEE conference on computer vision and pattern recognition . IEEE, 2012, pp. 3354–3361
2012
-
[52]
Neural volumetric object selection,
Z. Ren, A. Agarwala, B. Russell, A. G. Schwing, and O. Wang, “Neural volumetric object selection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 6133–6142
2022
-
[53]
Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,
B. Mildenhall, P. P. Srinivasan, R. Ortiz-Cayon, N. K. Kalantari, R. Ra- mamoorthi, R. Ng, and A. Kar, “Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,” ACM Transactions on Graphics (TOG), vol. 38, no. 4, pp. 1–14, 2019
2019
-
[54]
Spin-nerf: Multiview segmentation and perceptual inpainting with neural radiance fields,
A. Mirzaei, T. Aumentado-Armstrong, K. G. Derpanis, J. Kelly, M. A. Brubaker, I. Gilitschenski, and A. Levinshtein, “Spin-nerf: Multiview segmentation and perceptual inpainting with neural radiance fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Patte...
2023
-
[55]
Fast online object tracking and segmentation: A unifying approach,
Q. Wang, L. Zhang, L. Bertinetto, W. Hu, and P. H. Torr, “Fast online object tracking and segmentation: A unifying approach,” in Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recogni- tion, 2019, pp. 1328–1338
2019
-
[56]
State-aware tracker for real-time video object segmentation,
X. Chen, Z. Li, Y . Yuan, G. Yu, J. Shen, and D. Qi, “State-aware tracker for real-time video object segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9384–9393
2020
-
[57]
Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,
H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in European Conference on Computer Vision . Springer, 2022, pp. 640–658
2022
-
[58]
Semantic video cnns through representation warping,
R. Gadde, V . Jampani, and P. V . Gehler, “Semantic video cnns through representation warping,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 4453–4462
2017
-
[59]
Video instance segmentation,
L. Yang, Y . Fan, and N. Xu, “Video instance segmentation,” in Proceed- ings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 5188–5197
2019
-
[60]
Vip-deeplab: Learning visual perception with depth-aware video panoptic segmenta- tion,
S. Qiao, Y . Zhu, H. Adam, A. Yuille, and L.-C. Chen, “Vip-deeplab: Learning visual perception with depth-aware video panoptic segmenta- tion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3997–4008
2021
-
[61]
Amanet: Adaptive multi-path aggregation for learning human 2d-3d correspondences,
X. Wang, Y . Guo, J. Song, L. Gao, and H. T. Shen, “Amanet: Adaptive multi-path aggregation for learning human 2d-3d correspondences,” IEEE Transactions on Multimedia , vol. 25, pp. 979–992, 2021
2021
-
[62]
Cpi-parser: Integrating causal properties into multiple human parsing,
X. Wang, X. Chen, L. Gao, J. Song, and H. T. Shen, “Cpi-parser: Integrating causal properties into multiple human parsing,” IEEE Trans- actions on Image Processing , 2024
2024
-
[63]
Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera,
J. S. Yoon, K. Kim, O. Gallo, H. S. Park, and J. Kautz, “Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 5336–5345
2020
-
[64]
Consistent video depth estimation,
X. Luo, J.-B. Huang, R. Szeliski, K. Matzen, and J. Kopf, “Consistent video depth estimation,” ACM Transactions on Graphics (ToG), vol. 39, no. 4, pp. 71–1, 2020
2020
-
[65]
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,
S. Peng, Y . Zhang, Y . Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” in CVPR, 2021
2021
-
[66]
Reconstructing 3d human pose by watching humans in the mirror,
Q. Fang, Q. Shuai, J. Dong, H. Bao, and X. Zhou, “Reconstructing 3d human pose by watching humans in the mirror,” in CVPR, 2021
2021
-
[67]
Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,
L. Sigal, A. O. Balan, and M. J. Black, “Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,” International journal of computer vision , vol. 87, no. 1-2, pp. 4–27, 2010
2010
-
[68]
Semantically coherent 4d scene flow of dynamic scenes,
A. Mustafa and A. Hilton, “Semantically coherent 4d scene flow of dynamic scenes,” International Journal of Computer Vision , vol. 128, no. 2, pp. 319–335, 2020
2020
-
[69]
Light field content from 16-camera rig,
D. Doyen, T. Langlois, B. Vandame, F. Babon, G. Boisson, N. Sabater, R. Gendrot, and A. Schubert, “Light field content from 16-camera rig,” ISO/IEC JTC1/SC29 WG11 Doc. m40010. Geneva , 2017
2017
-
[70]
Barn new natural content proposal for miv,
T. Tapie, A. Schubert, R. Gendrot, G. Briand, F. Thudor, and R. Dor ´e, “Barn new natural content proposal for miv,” ISO/IEC JTC 1/SC 29/WG 4 m56632, Online, April 2021
2021
-
[71]
Breakfast new natural content proposal for miv,
——, “Breakfast new natural content proposal for miv,” ISO/IEC JTC 1/SC 29/WG 4 m56730, Online, April 2021
2021
-
[72]
Kermit test sequence for windowed 6dof activities,
B. Salahieh, B. Marva, M. Nentedem, A. Kumar, V . Popvic, K. Seshadri- nathan, O. Nestares, and J. Boyce, “Kermit test sequence for windowed 6dof activities,” ISO/IEC JTC1/SC29/WG11 MPEG M43748, July 2018
2018
-
[73]
[MPEG-I Vi- sual] Natural outdoor test sequences,
D. Mieloch, A. Dziembowski, and M. Doma ´nski, “[MPEG-I Vi- sual] Natural outdoor test sequences,” ISO/IEC JTC1/SC29/WG11 MPEG2020/M51598, January 2020
2020
-
[74]
Multiview test video sequences for free navigation exploration obtained using pairs of cameras,
M. Doma ´nski, A. Dziembowski, A. Grzelka, D. Mieloch, O. Stankiewicz, and K. Wegner, “Multiview test video sequences for free navigation exploration obtained using pairs of cameras,” ISO/IEC JTC1/SC29/WG11, Doc. MPEG M , vol. 38247, 2016
2016
-
[75]
New test sequence for mpeg-i visual,
X. Sheng, “New test sequence for mpeg-i visual,” ISO/IEC JTC1/SC29/WG04 MPEG136/M58275, Online, October 2021
2021
-
[76]
[MIV] new natural content - martialarts,
D. Mieloch, A. Dziembowski, B. Szydełko, D. Kl ´oska, A. Grzelka, J. Stankowski, M. Doma ´nski, G. Lee, and J. Jeong, “[MIV] new natural content - martialarts,” ISO/IEC JTC1/SC29/WG4 MPEG 141, M61949, January 2023
2023
-
[77]
Neural 3d video synthesis from multi-view video,
T. Li, M. Slavcheva, M. Zollhoefer, S. Green, C. Lassner, C. Kim, T. Schmidt, S. Lovegrove, M. Goesele, R. Newcombe et al., “Neural 3d video synthesis from multi-view video,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5521–5531
2022
-
[78]
Immersive light field video with a layered mesh representation,
M. Broxton, J. Flynn, R. Overbeck, D. Erickson, P. Hedman, M. Duvall, J. Dourgarian, J. Busch, M. Whalen, and P. Debevec, “Immersive light field video with a layered mesh representation,” ACM Transactions on Graphics (TOG), vol. 39, no. 4, pp. 86–1, 2020
2020
-
[79]
A multi-view stereoscopic video database with green screen (mtf) for video transition quality- of-experience assessment,
N. Hobloss, L. Zhang, and M. Cagnazzo, “A multi-view stereoscopic video database with green screen (mtf) for video transition quality- of-experience assessment,” in 2021 13th International Conference on Quality of Multimedia Experience (QoMEX) . IEEE, 2021, pp. 201– 206
2021
-
[80]
Poznan blocks-a multiview video test sequence and camera parameters for free viewpoint television,
M. Domanski, A. Dziembowski, A. Kuehn, M. Kurc, A. Luczak, D. Mieloch, J. Siast, O. Stankiewicz, and K. Wegner, “Poznan blocks-a multiview video test sequence and camera parameters for free viewpoint television,” ISO/IEC JTC1/SC29/WG11 MPEG2014 M , vol. 32243, pp. 13–17, 2014
2014
-
[81]
Large-scale video panoptic segmentation in the wild: A benchmark,
J. Miao, X. Wang, Y . Wu, W. Li, X. Zhang, Y . Wei, and Y . Yang, “Large-scale video panoptic segmentation in the wild: A benchmark,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 21 033–21 043
2022
-
[82]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” arXiv preprint arXiv:2304.02643 , 2023
2023 arXiv
-
[83]
Similarity measures for inter- section of camera view frustums,
Y . Zamani, H. Shirzad, and S. Kasaei, “Similarity measures for inter- section of camera view frustums,” in 2017 10th Iranian Conference on Machine Vision and Image Processing (MVIP) . IEEE, 2017, pp. 171– 175
2017
-
[84]
Lightglue: Local feature matching at light speed,
P. Lindenberger, P.-E. Sarlin, and M. Pollefeys, “Lightglue: Local feature matching at light speed,” arXiv preprint arXiv:2306.13643 , 2023
2023 arXiv
-
[85]
Scene- generalizable interactive segmentation of radiance fields,
S. Tang, W. Pei, X. Tao, T. Jia, G. Lu, and Y .-W. Tai, “Scene- generalizable interactive segmentation of radiance fields,” inProceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 6744–6755
2023
-
[86]
Tracking anything with decoupled video segmentation,
H. K. Cheng, S. W. Oh, B. Price, A. Schwing, and J.-Y . Lee, “Tracking anything with decoupled video segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 1316–1326
2023
-
[87]
Emerging properties in self-supervised vision transformers,
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 9650–9660
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.