{"id":"b176547c-6d37-4d89-918e-ddcf0eb87ffc","arxiv_id":"2502.02334","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fusing event-camera data during the 2D-to-3D lifting step improves semantic scene completion accuracy and robustness on a new real-world benchmark and on corrupted SemanticKITTI.","lead":"This paper introduces DSEC-SSC, a new real-world driving dataset pairing ordinary camera images with event-camera data for 3D scene completion, plus a model called EvSSC that fuses both signals. The goal is to make self-driving perception more reliable under motion blur, darkness, fog, and partial camera sensor failure.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DSEC-SSC's unvalidated 4D labels are the only real-event evidence for EvSSC, so if dynamic-object labels are unreliable the real-world support for the central claim collapses.","rationale":"I read the paper as a method plus benchmark contribution. The method's benefit is claimed to be real-world robustness, and the only real-world measurements are the DSEC-SSC rows. Those rows depend on ground-truth labels that the authors themselves describe as involving manual BEV annotation and interpolation for dynamic objects, with no quality check. This is the most load-bearing weakness because it is upstream of all real-event conclusions: if the labels are noisy, the benchmark cannot rank methods, and the paper's only non-simulated evidence for event-aided lifting disappears. The simulation concern is real but secondary: the shot-noise improvement (52.5%) is physically plausible because an independent event sensor would be unaffected by image-sensor shot noise, and the simulation concern mainly affects fog/darkness cells. Label trustworthiness, by contrast, threatens the entire real-world portion of the claim. I agree with the reader's weakest_assumption. A conditional verdict is appropriate pending the label audit; I would keep the reader's CONDITIONAL rather than moving to reject, because the method has consistent (if small) gains, a clear ablation, and a plausible mechanism, but the evidence is not yet solid. The proposed audit would settle whether the labels are trustworthy. If the labels fail, the verdict should move to REJECT or UNVERDICTED until real-event validation is provided.","tokens_in":19256,"tokens_out":9104,"duration_ms":99013,"concrete_test":"Audit the DSEC-SSC dynamic-object labels: select, say, 3 of the 6 validation sequences, regenerate dynamic-object voxels from the raw LiDAR using an independent off-the-shelf 3D tracker or manual 3D annotation (not the paper's BEV pipeline), and compute per-class IoU between the two label sets. Also recompute Table II mIoU restricted to static classes only (road, sidewalk, building, vegetation, terrain, fence). If dynamic-label IoU is below ~0.7 or the static-only EvSSC gain vanishes, the real-event support for EvSSC is not established and the benchmark needs re-labeling before the central claim can be evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that event-aided lifting improves SSC accuracy and robustness. The only real-event support is Table II on DSEC-SSC (gains of +0.72 and +0.49 mIoU). These scores are computed against ground truth produced by the Sec. III pipeline: static labels come from projected 2D semantic maps plus clustering, while dynamic objects are reconstructed by manually annotating 'pipeline' instances in BEV and linearly interpolating position/orientation when LiDAR misses them (Sec. III-B). No annotation agreement, error analysis, or independent LiDAR cross-check is reported; the Limitations section even concedes that more accurate pose estimation is needed. If dynamic-object labels are systematically misplaced or missing, the small real-event gains in Table II could be within label noise, and the paper is left with no real-event validation of the core claim. The SemanticKITTI-C robustness results (including the 52.5% shot-noise gain) are entirely based on DVS-Voltmeter simulated events, which cannot serve as real-world confirmation. Thus the real-world leg of the argument is unsecured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DSEC-SSC, a real-world event-aided semantic scene completion (SSC) benchmark derived from the DSEC dataset through a semi-automatic 4D labeling pipeline, and proposes EvSSC, a fusion framework whose Event-aided Lifting Module (ELM) injects event features into the 2D-to-3D lifting stage. The method is evaluated on DSEC-SSC and on simulated SemanticKITTI-E/C datasets using VoxFormer and SGN backbones. The central claim is that event-aided lifting yields consistent improvements in mIoU across five degradation modes and both in-domain and out-of-domain settings, with a headline 52.5% relative improvement in the in-domain shot-noise setting.","tokens_in":19509,"tokens_out":5738,"duration_ms":57230,"significance":"If the results hold up, the paper makes a useful contribution: it provides the first real-world event-camera benchmark for SSC, proposes a simple and architecture-agnostic fusion point (during lifting) that is ablated against two alternative fusion paradigms, and demonstrates that event cues can help under image degradation. The commitment to release the dataset and code, the benchmarking of multiple baselines, and the explicit ablation of design choices are strengths. The main significance risk is that the only real-event evidence (Table II) rests on a labeling pipeline whose accuracy is not validated, and the headline robustness results are entirely based on simulated events; both issues are addressable in revision.","major_comments":[{"comment":"The DSEC-SSC dynamic-object labels are not validated. Section III-B describes a pipeline that manually annotates 'pipeline' instances in BEV and uses linear interpolation to place dynamic objects when LiDAR misses them, but no inter-annotator agreement, no error analysis, and no independent LiDAR cross-check are reported; the Limitations section (Sec. VI) also concedes that pose-estimation accuracy needs improvement. This is load-bearing because the only real-event support for the central claim is Table II, where the EvSSC gains over RGB-only baselines are +0.72 and +0.49 mIoU. If dynamic-object labels are systematically displaced or missing, these differences could lie within label noise. The authors should validate the labels, for example by comparing a subset against dense manual annotation from the LiDAR sweeps or against an independent sensor, and report per-class label agreement.","section":"Section III-B, Table II, Section VI"},{"comment":"No uncertainty estimates or significance tests are reported. Many of the claimed gains are small: on DSEC-SSC the gains are +0.72 and +0.49 mIoU, and in Table IV most in-domain corruption gains are below 0.9 mIoU. Without repeated runs or paired significance testing, the 'consistently improved prediction accuracy' claim is not statistically grounded. Please report mean and standard deviation over at least three seeds for the key comparisons, or provide a significance test on the validation-set metric.","section":"Tables II-IV"},{"comment":"The SemanticKITTI-C robustness protocol does not state whether the event streams are generated from the clean RGB images or from the corrupted RGB images. If the events are generated from clean images, the event branch is never exposed to the degradation, which would make part of the robustness gain a by-construction effect. If the events are generated from corrupted images, the fidelity of DVS-Voltmeter simulation under strong blur, noise, or fog needs at least a discussion, since most quantitative robustness claims in Table IV come from this simulated setting. The authors should specify the exact protocol and, ideally, report results for both clean-generated and corrupted-generated events.","section":"Section V.A, Section V.C, Table IV"},{"comment":"The abstract's 52.5% relative improvement is drawn from a single cell of Table IV (in-domain shot noise, VoxFormer-S: 8.29 to 12.64). All other corruption cells show much smaller gains, and the SGN-S cell for the same shot-noise condition is only +0.70 mIoU. While 'up to 52.5%' is literally correct, the claim would be better supported if the text reported the range or median improvement across the ten corruption cells and noted that this large gain is an outlier. Please recontextualize the headline number or provide an explanation for why this particular cell is so much larger than the others.","section":"Abstract, Table IV, Section V-C"}],"minor_comments":[{"comment":"The text says 'EvSSC (SGN) boots mIoU by 0.31 compared to its RGB baseline (29.37 vs. 29.06)', but Table II lists the EvSSC (SGN-S) mIoU as 29.55, not 29.37. The word 'boots' should be 'boosts', and the inconsistent SGN/SGN-S naming should be harmonized.","section":"Section V.B, Table II"},{"comment":"Equation (2) is notationally unclear: 'Mg = arg max N X i=1 1(di < eps)' does not define a well-posed optimization over the plane parameters, and the sentence 'Here, 1 is an indicator function... and MS∈MS' is garbled. The plane-fitting objective and the role of the K-means cluster count in Eq. (3) should be stated explicitly, including the values of epsilon and the cluster count used.","section":"Section III-C, Eqs. (1)-(3)"},{"comment":"The HATS entry in Table V(b) cites reference [98], but [98] is 'Recovering accurate 3D human pose in the wild using IMUs and a moving camera', which is not the HATS event representation. The HATS method (Sironi et al., CVPR 2018) should be cited instead.","section":"Table V(b), Reference [98]"},{"comment":"The sentence 'we report training results using both 2D images (xrgb) and event data (xevent) as inputs' is ambiguous; the tables actually distinguish RGB-only, event-only, and EvSSC (RGB+event) rows. Please rephrase to describe the three input configurations.","section":"Section V.A"},{"comment":"The severity levels used to create SemanticKITTI-C (e.g., the OpenCV or image-processing parameters for motion blur, fog, brightness, darkness, and shot noise) are not specified. These settings should be reported in the paper or in the supplementary material to make the robustness results reproducible.","section":"Section V.A, Table IV"}],"recommendation":"major_revision","confidential_remarks":"The central idea is promising and the paper is generally honest about its limitations, but the real-world leg of the evidence is currently thin: the DSEC-SSC label pipeline has no validation, and the robustness headline rests on one simulated-event outlier. These are fixable with additional experiments and analysis, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my take on the Event-aided SSC paper. The two things worth knowing: first, it ships the first real-world event+RGB SSC dataset (DSEC-SSC) with a genuinely clever labeling pipeline that deals with dynamic objects in 4D. Second, the fusion-in-lifting idea (ELM) is a real and sensible contribution. But the evidence for the central robustness claim is thinner than the abstract suggests. The 52.5% number is one cherry-picked cell from SemanticKITTI-C shot noise; most other gains are a few tenths of mIoU.\n\nWhat the paper does well: the DSEC-SSC benchmark is a real gap-filler. The labeling pipeline—project 2D semantic maps to point clouds, separate static/dynamic, use the 'pipeline' effect plus manual BEV annotation and linear interpolation for missing dynamic objects—is a practical approach that others can copy. The ELM module is a clean way to inject event features into the lifting stage, and the ablations (fusion-then-lifting, decode-then-fusion) do isolate the benefit. The gains are consistent across two architectures (VoxFormer, SGN) and two datasets, which gives me some confidence that the effect is real.\n\nWhere it is soft: the DSEC-SSC labels are unvalidated. No annotation agreement, no error analysis, no independent LiDAR cross-check. The paper itself concedes pose estimation needs improvement. If the dynamic-object labels are systematically off, the real-world gains (which are only +0.72 and +0.31 mIoU) could be label noise. Second, most robustness evidence comes from DVS-Voltmeter simulated events, not a real event camera. Third, no error bars anywhere.\n\nThese are proportionately serious: the core architectural claim (event features help in lifting) is plausibly true and the simulated evidence supports it, but the headline numbers are not as robust as they look. The paper is honest about limitations, and the contributions are real enough to deserve referee time. My recommendation: accept for peer review, but the authors should be pushed to validate the DSEC-SSC labels, report variance, and soften the 52.5% framing.\n\nThis is a paper for people working on occupancy prediction or event-based perception. I'd bring it to a reading group.\n\nBest,","headline":"New dataset and a sensible fusion idea, but the real-world legs are weak and the headline gain is a single cell.","tokens_in":20027,"tokens_out":2462,"would_cite":true,"duration_ms":24087,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fusing event data into the 2D-to-3D lifting step makes camera-based semantic scene completion more accurate and robust under degraded imaging.","keywords":["semantic scene completion","event camera","RGB-event fusion","3D occupancy prediction","view transformation","autonomous driving","robustness to corruption","DSEC-SSC benchmark"],"falsifier":"Run the SemanticKITTI-C shot-noise protocol with real event streams recorded by an event camera synchronized to the RGB image, instead of DVS-Voltmeter-simulated events. If EvSSC's 52.5% relative mIoU gain over VoxFormer-S collapses or becomes negative, then the simulated-event premise, not the fusion mechanism, produced the headline robustness.","tokens_in":19067,"feed_emoji":"🚗","tokens_out":6931,"duration_ms":58160,"temperature":0.7,"pith_summary":"This paper is trying to establish that event cameras, which report brightness changes asynchronously, can make camera-based semantic scene completion—predicting a full 3D occupancy grid with semantic labels from images—more accurate and much more robust when ordinary images degrade. To test this, it builds DSEC-SSC, a real-world benchmark with event, RGB, and LiDAR data, and proposes EvSSC, a fusion framework whose Event-aided Lifting Module (ELM) mixes RGB and event features during the 2D-to-3D view-transformation step. Across transformer-based and LSS-based architectures, EvSSC improves mean IoU on clean benchmarks and, on corrupted SemanticKITTI-C, consistently beats RGB-only baselines under motion blur, fog, brightness, darkness, and shot noise, with the largest gain when the image sensor partially fails. If these results hold, event-aided lifting is a low-cost way to harden occupancy prediction for autonomous driving against exactly the conditions that break RGB-only systems.","feed_headline":"Event camera fusion lifts 3D scene completion accuracy by up to 52.5%","feed_subtitle":"Fusing events with RGB during 2D-to-3D lifting keeps occupancy stable under blur, fog, darkness, and shot noise.","key_machinery":"The load-bearing mechanism is the Event-aided Lifting Module (ELM), a fusion-based lifting block inserted during 2D-to-3D view transformation. ELM takes separately encoded image and event features, sums their key and value pairs, applies self-attention to the aggregated features to compute a per-location gating weight $w$, blends the two modalities as $(1-w)F_{img}+w F_{event}$, and then uses deformable attention to turn those fused key/value features into 3D voxel queries. The module's job is to do cross-modal fusion at the exact step where 2D features are projected into 3D, preserving spatial fidelity and keeping temporal event cues available to the 3D volume construction.","core_discovery":"The paper's central discovery is that the right place to fuse event data with RGB in an SSC network is inside the lifting step, not before or after it. EvSSC builds two encoders, one for images and one for rasterized events; in ELM, the key and value features of both modalities are added, fed through self-attention to produce a scalar gating weight, and then blended into fused key and value features; a deformable-attention stage uses those fused values to query voxel features for the 3D volume. This 'fusion-based lifting' yields higher mIoU than early 2D fusion or late 3D fusion in ablations, and it plugs into both transformer-based (VoxFormer) and LSS-based (SGN) occupancy models. The paper also introduces DSEC-SSC, the first real-world event-aided SSC dataset, with a semi-automatic 4D labeling pipeline that separates static and dynamic objects and reconstructs dynamic instances in time so labels follow moving vehicles and people.","pith_inferences":["A direct extension the paper does not run is testing the corruption protocol with real event streams rather than DVS-Voltmeter simulation; the headline 52.5% gain comes from simulated events, so a real-event test under matched motion blur and low light would show whether the gain persists under actual sensor noise.","The gating weight $w$ in ELM is a soft per-location choice between RGB and event evidence, suggesting the same fusion-based lifting design could transfer to other complementary sensors, such as radar or thermal cameras, whenever their features can be expressed as key/value pairs.","The DSEC-SSC labeling pipeline is sensor-agnostic and low-cost, so it could be reused to create event-aided SSC labels from other sparse-LiDAR datasets; a useful validation would be to compare DSEC-SSC dynamic-object labels against a dense or independently annotated reference."],"forward_implications":["On the simulated SemanticKITTI-E benchmark, EvSSC raises VoxFormer mIoU from 12.86 to 13.61 and SGN from 14.55 to 15.15, with only a millisecond-level latency increase.","On corrupted SemanticKITTI-C, event fusion improves mIoU over RGB-only baselines in all five degradations, in both out-of-domain and in-domain settings; VoxFormer-S goes from 8.29 to 12.64 mIoU under in-domain shot noise, a 52.5% relative gain.","On real-world DSEC-SSC, adding events raises VoxFormer mIoU from 25.62 to 26.34 and SGN from 29.06 to 29.55, and helps recover vehicles and traffic signs in low-light frames.","Because ELM works on both transformer-based and LSS-based occupancy models, event-aided lifting can likely be added to other camera-based 3D occupancy architectures without redesigning the rest of the network."],"supporting_citations":[{"why":"Supplies the real-world RGB, event, and LiDAR data from which DSEC-SSC is built.","marker":"[90]"},{"why":"Provides the voxelized occupancy labels and train/val splits used for SemanticKITTI-E and SemanticKITTI-C.","marker":"[7]"},{"why":"VoxFormer is the transformer-based SSC baseline into which ELM is inserted and which anchors the robustness comparisons.","marker":"[21]"},{"why":"SGN supplies the LSS-based SSC baseline showing that ELM generalizes across architecture families.","marker":"[22]"},{"why":"DVS-Voltmeter generates the simulated event streams used to build SemanticKITTI-E and all corrupted-condition experiments.","marker":"[96]"},{"why":"MonoScene is one of the RGB-only baselines benchmarked on DSEC-SSC, used to show that event-only inputs underperform RGB and that EvSSC improves on RGB.","marker":"[37]"}],"fun_headline_variants":["Fusion in the lift: events boost 3D scene completion by 52.5%","When RGB fails, events pick up the slack: up to 52.5% better","Event cameras help cars see through blur and fog, says new study","New dataset and fusion method make SSC robust to bad weather","Blurred? Foggy? Events still guide 3D scene completion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The DSEC-SSC ground-truth labels come from a semi-automatic pipeline that uses manual BEV annotation and linear interpolation for dynamic objects, and most robustness numbers use simulated rather than real events; if either source is unreliable, the reported gains may not transfer to real conditions.","fun_headline_variants_meta":{"raw":{"variants":["Fusion in the lift: events boost 3D scene completion by 52.5%","When RGB fails, events pick up the slack: up to 52.5% better","Event cameras help cars see through blur and fog, says new study","New dataset and fusion method make SSC robust to bad weather","Blurred? Foggy? Events still guide 3D scene completion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000635,"raw_usage":{"total_tokens":2973,"prompt_tokens":1033,"completion_tokens":1940,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":1838}},"tokens_in":649,"tokens_out":1940,"duration_ms":12246,"temperature":1.0,"reasoning_tokens":1838,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T12:29:42.351595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the SemanticKITTI-C shot-noise protocol with real event streams recorded by an event camera synchronized to the RGB image, instead of DVS-Voltmeter-simulated events. If EvSSC's 52.5% relative mIoU gain over VoxFormer-S collapses or becomes negative, then the simulated-event premise, not the fusion mechanism, produced the headline robustness.","supporting_citations":[{"cited_title":"DSEC: A stereo event camera dataset for driving scenarios,","cited_arxiv_id":null,"evidence_quote":"Supplies the real-world RGB, event, and LiDAR data from which DSEC-SSC is built."},{"cited_title":"DVS-V oltmeter: Stochastic process- based event simulator for dynamic vision sensors,","cited_arxiv_id":null,"evidence_quote":"DVS-Voltmeter generates the simulated event streams used to build SemanticKITTI-E and all corrupted-condition experiments."},{"cited_title":"MonoScene: Monocular 3D semantic scene completion,","cited_arxiv_id":null,"evidence_quote":"MonoScene is one of the RGB-only baselines benchmarked on DSEC-SSC, used to show that event-only inputs underperform RGB and that EvSSC improves on RGB."}],"review_version":1}