{"id":"e9e5358a-7308-4200-9b8b-fd5396ead284","arxiv_id":"2607.07374","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":9,"one_line_summary":"PLED-VINS improves event-camera SLAM in dynamic environments by fusing an entropy-recency temporal reliability score with geometric bundle-adjustment weights for point and line features.","lead":"This paper presents a SLAM system for event cameras that handles moving objects by combining temporal reliability (from event statistics) with geometric reliability (from bundle adjustment) to weight features. A smart generalist might read it to understand how event cameras can improve robot navigation in dynamic real-world scenes like streets with moving cars.","discovery_kind":"new_method","skeptic_critique":{"model":"glm-5.2","headline":"Single-sequence ablation cannot isolate whether the entropy-recency score map or the addition of line features drives the reported improvements across the full evaluation.","rationale":"The reader correctly identifies several conditions that prevent full acceptance, and the rigid-motion hypothesis is a legitimate scope limitation. However, I rank the attribution problem as more load-bearing because it directly threatens the core novelty claim. The paper's central contribution is the entropy-recency score map and adaptive fusion, not line feature integration (which is well-established in prior work like PL-VINS and PL-EVIO). Yet the strongest experimental evidence (Tables I–II) does not control for line feature addition, and the only controlled ablation covers one of five evaluated dynamic sequences. This is a standard experimental design gap, not an internal inconsistency or a disagreement with consensus. The method itself is well-motivated and the formulation (Eqs. 2–5, 12–17) is internally consistent. The per-dataset hyperparameter variation compounds the attribution concern: if the method requires different λw/λm settings per dataset, and the ablation only covers one dataset with one parameter setting, the generalization of the temporal reliability contribution is underestablished. The verdict should remain CONDITIONAL; the reader's assessment is sound but the emphasis should shift from the rigid-motion assumption (which is a known, acknowledged scope limitation) to the experimental attribution gap (which is addressable and directly bears on the central claim).","tokens_in":13227,"tokens_out":2920,"duration_ms":149424,"concrete_test":"Run the Table III ablation (geometric-only vs. geometric+temporal) on all three VIODE high sequences and both DAVIS 240C dynamic sequences, using a single fixed set of hyperparameters (λw, λm) across all sequences. If the geometric+temporal configuration does not consistently reduce MAE and increase w_ratio relative to geometric-only across all five sequences, the claim that the entropy-recency score map is the key enabler of robustness weakens, and the improvements may be attributable to line features or per-dataset tuning instead.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim attributes improved state estimation to two novel components: (1) the entropy-recency score map for temporal reliability and (2) the adaptive fusion strategy. However, the experimental design cannot cleanly isolate these contributions. The full trajectory comparisons in Tables I–II compare PLED-VINS (point+line, temporal+geometric) against DynaVINS (point-only, geometric-only), conflating the novel temporal reliability with the addition of line features. The only controlled ablation (Table III, geometric-only vs. geometric+temporal, both with line features) is run on a single sequence (parking_lot_high). Without ablations on the other high-dynamic sequences (city_day_high, city_night_high) and the DAVIS 240C sequences, it remains unclear whether the entropy-recency score map generalizes or whether the trajectory improvements stem primarily from line feature augmentation and per-dataset hyperparameter tuning (λw varies 4×, λm varies 10× across datasets). The reader's identified concern about the rigid-motion hypothesis is a genuine scope limitation acknowledged by the authors, but it is less load-bearing than this attribution problem: even within the rigid-motion regime, the evidence does not establish that the proposed temporal reliability signal is the primary driver of improvement.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The paper proposes PLED-VINS, a monocular event-camera visual-inertial SLAM framework for dynamic environments. The core technical contributions are: (1) an entropy-recency score map that quantifies temporal reliability of features from motion-compensated event streams, and (2) an adaptive fusion strategy combining temporal and geometric reliability cues (including motion-conditioned modeling for line features) within a unified point-line robust bundle adjustment. The system is evaluated on the VIODE, DAVIS 240C, and DSEC datasets against five baselines, showing consistent ATE/MPE improvements on high-dynamic sequences. The approach is well-motivated and the individual technical components are clearly described.","tokens_in":14141,"tokens_out":1159,"duration_ms":247735,"significance":"The paper addresses a genuine gap: most event-based SLAM frameworks assume static scenes, and existing dynamic-environment methods either rely on expensive segmentation or compress temporal information into single statistics. The entropy-recency score map is a principled contribution that preserves distributional temporal information. The motion-conditioned fusion coefficient for line features (Eqs. 15–17) is a novel and sensible design that accounts for geometric observability. The system achieves real-time performance (21.25 Hz at 240×180). The ablation study (Table III) provides weight-quality metrics (MAE, w_ratio) against ground-truth segmentation labels, which is a more informative evaluation than trajectory error alone. The explicit acknowledgment of the rigid-motion hypothesis limitation in the Conclusion is appropriate. However, the experimental design does not fully isolate the proposed temporal reliability contributions from the addition of line features, and per-dataset hyperparameter variation raises generalization concerns.","major_comments":[{"comment":"Table III (ablation study) and Tables I–II (main results): The central claim attributes improvement to two novel components—temporal reliability and adaptive fusion—but the experimental design cannot cleanly isolate these from the addition of line features. Tables I–II compare PLED-VINS (point+line, temporal+geometric) against DynaVINS (point-only, geometric-only), conflating the temporal reliability contribution with line feature augmentation. The only controlled ablation (Table III, geometric-only vs. geometric+temporal, both with line features) is run on a single VIODE sequence (parking_lot_high). Without ablations on additional high-dynamic sequences (city_day_high, city_night_high) and the DAVIS 240C sequences, it remains unclear whether the entropy-recency score map generalizes or whether trajectory improvements stem primarily from line features. Adding at least 2–3 more sequences'","section":null}],"minor_comments":[{"comment":"Section III.B, Eq. (2): The ε appears both inside and outside the logarithm. Clarify whether ε is added before or after taking log, as this affects numerical behavior near zero-probability bins.","section":null},{"comment":"Section III.E, Eq. (15): The relationship between α_line and the text description could be clearer. The text states α_line 'increases when r indicates strong alignment and remains high under tangential motion,' but Eq. (15) defines α_line = 1 − p_n(1−r). A brief derivation or intermediate step showing how p_n and r interact to produce the described behavior would improve clarity.","section":null},{"comment":"Section IV.B: The use of v2e [36] to generate synthetic events for VIODE is mentioned briefly. A sentence noting the known limitations of v2e (e.g., noise model fidelity) and why they do not undermine the VIODE results would strengthen the evaluation.","section":null},{"comment":"Table I: PL-VINS results for parking_lot mid and high are missing (shown as '—'). A footnote explaining why (e.g., tracking failure) would be helpful.","section":null},{"comment":"Fig. 6: The y-axis starts at 10^{-1}, which compresses the visual differences. Consider adjusting the axis range or using a log scale to better visualize relative differences.","section":null},{"comment":"Section III.C, Eq. (6): The line band width W is listed as a free parameter but its value is not reported in the experimental settings (Section IV.A). Stating the value used would aid reproducibility.","section":null},{"comment":"Reference [7] (E2-VINS): The venue is listed as 'Appl. Sci., vol. 15, no. 3, p. 1314, 2025.' Please verify this reference is correctly cited and accessible, as it appears to be a very recent publication.","section":null},{"comment":"The paper would benefit from a brief discussion of failure modes beyond the rigid-motion assumption, such as scenarios where the entropy-recency map may produce false positives (e.g., flickering textures).","section":null}],"recommendation":"major_revision","confidential_remarks":"The reader's concern about the recurrent fusion loop (Eqs. 12–13 feeding back into Eq. 8) is worth noting but is not necessarily a circularity problem: the temporal reliability w_EV is computed independently from event statistics and does not depend on BA outputs, so the recurrence is a standard iterative refinement of geometric weights informed by an external signal. The more substantive concern is the attribution problem: the single-sequence ablation is insufficient to support the central claim that temporal reliability (not line features) drives the improvements. This is fixable with additional ablation runs and should be required before acceptance. The per-dataset hyperparameter variation (λw: 2.0/8.0/2.0, λm: 0.2/1.0/2.0) is also worth flagging to the authors as a generalization concern, though some variation across sensor configurations is expected."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful and constructive review. The referee correctly identifies that our central claim—improvement from temporal reliability and adaptive fusion—is not cleanly isolated from the contribution of line features in the main results tables, and that the ablation study is limited to a single VIODE sequence. We agree this is a legitimate concern and will address it in revision.","responses":[{"response":"The referee is correct that the current experimental design does not cleanly isolate the temporal reliability contribution from the line feature augmentation. We acknowledge this conflation in Tables I–II, where PLED-VINS (point+line, temporal+geometric) is compared against DynaVINS (point-only, geometric-only). The referee is also correct that the ablation in Table III is limited to a single sequence (parking_lot_high), which is insufficient to demonstrate generalization of the entropy-recency score map. We will address this in the revised manuscript by expanding the ablation study to include at least three additional high-dynamic sequences: city_day_high and city_night_high from VIODE, and the dynamic_6dof sequence from DAVIS 240C. For each sequence, we will report four configurations in a controlled factorial design: (A) point-only + geometric-only (replicating DynaVINS), (B) point+line + geometric-only (isolating the line feature contribution), (C) point-only + geometric+temporal (isolating the temporal reliability contribution), and (D) point+line + geometric+temporal (full model). This design will allow clean attribution of trajectory improvements to each component. We will also extend the weight-quality metrics (MAE, w_ratio) from Table III to the additional VIODE sequences where ground-truth segmentation labels are available. We agree that without these additional ablations, the generalization claim for the entropy-recency score map is not adequately supported.","revision_made":"yes","referee_comment":"Table III (ablation study) and Tables I–II (main results): The central claim attributes improvement to two novel components—temporal reliability and adaptive fusion—but the experimental design cannot cleanly isolate these from the addition of line features. Tables I–II compare PLED-VINS (point+line, temporal+geometric) against DynaVINS (point-only, geometric-only), conflating the temporal reliability contribution with line feature augmentation. The only controlled ablation (Table III, geometric-only vs. geometric+temporal, both with line features) is run on a single VIODE sequence (parking_lot_high). Without ablations on additional high-dynamic sequences (city_day_high, city_night_high) and the DAVIS 240C sequences, it remains unclear whether the entropy-recency score map generalizes or whether trajectory improvements stem primarily from line features. Adding at least 2–3 more sequences'"}],"tokens_in":12905,"tokens_out":593,"duration_ms":80131,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Bottom line: the entropy-recency score map (Eqs. 2–5) is a genuinely new idea for event-based dynamic SLAM, and the motion-conditioned line fusion (Eqs. 14–17) is a thoughtful extension. But the experimental design can't cleanly attribute the reported improvements to the temporal reliability signal versus the addition of line features, and that's the main thing you should know before deciding how to handle this.","headline":"Entropy-recency score map is a genuine new idea for event-based dynamic SLAM, but the ablation is too thin to prove it drives the gains.","tokens_in":14032,"tokens_out":160,"would_cite":false,"duration_ms":60600,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Event-camera SLAM survives dynamic scenes via entropy-recency scoring","keywords":["event camera","visual-inertial SLAM","dynamic environments","entropy","temporal reliability","bundle adjustment","point-line features","motion compensation"],"falsifier":"If one constructed a scene where static structures produce late-biased, high-entropy event distributions after motion compensation — for instance, through rapid camera translation past nearby high-texture surfaces — the entropy-recency score would incorrectly flag them as dynamic, degrading pose estimation rather than improving it.","tokens_in":13328,"feed_emoji":"📸","tokens_out":845,"duration_ms":138575,"temperature":0.7,"pith_summary":"Event cameras asynchronously record per-pixel brightness changes, giving them an edge over conventional cameras in fast-motion and low-light scenes. But most event-camera SLAM systems still assume the world is static, so moving objects inject false geometric constraints and corrupt pose estimates. PLED-VINS tackles this by building an entropy-recency score map: after warping raw events to a reference frame using IMU data, it bins events at each pixel into time intervals, computes the Shannon entropy of that temporal distribution, and combines it with a recency measure that flags late-biased activations. The core idea is that static structures, once motion-compensated, produce temporally structured event patterns, while independently moving objects leave dispersed, recent-biased temporal signatures even after compensation. These temporal reliability scores are then fused with geometric reliability from a robust bundle adjustment that jointly handles point and line features, with an adaptive scheme that shifts weight between temporal and geometric cues depending on motion conditions. The result is a monocular event-camera visual-inertial SLAM system that suppresses dynamic observations without expensive object-level segmentation.","feed_headline":"Entropy-Recency Map Makes Event-Camera SLAM Work in Crowded Scenes","feed_subtitle":"By scoring how event timestamps cluster after motion compensation, PLED-VINS suppresses moving-object noise without segmentation, beating","key_machinery":"entropy-recency score map","core_discovery":"The entropy-recency score map is the central object. It quantifies temporal reliability per pixel by combining two complementary signals: normalized Shannon entropy of motion-compensated event timestamps across temporal bins (capturing dispersion), and a recency score weighted toward later bins (capturing late-biased activation). Their nonlinear symmetric fusion yields a per-pixel score that highlights independently moving objects while suppressing static background and noise. When integrated with geometric reliability from robust bundle adjustment through an adaptive, motion-conditioned weighting strategy, this score map enables a monocular event-camera SLAM system to achieve the lowest ATE","pith_inferences":["The entropy-recency score map implicitly performs a soft, per-pixel motion segmentation without requiring object classes or semantic labels, which could make it more generalizable than learning-based dynamic SLAM methods.","The 21.25 Hz processing frequency at 240x180 resolution suggests the method is borderline for real-time deployment on higher-resolution event cameras, and the entropy computation over temporal bins may become a bottleneck at scale.","The reliance on a single dominant rigid-motion hypothesis means the system should degrade gracefully (falling back to geometric-only weighting) rather than catastrophically when multiple moving objects with comparable event signatures are present, but this remains untested."],"forward_implications":["Event cameras could become the default sensor for SLAM in dynamic environments, replacing frame-based approaches that fail under motion blur and aggressive dynamics.","The entropy-recency formulation could extend to multi-motion hypotheses, enabling segmentation of several independently moving objects without explicit object detection.","Line-feature motion conditioning could generalize to other geometric primitives (planes, curves) where observability depends on motion direction relative to the primitive's geometry.","The approach could integrate with stereo or depth-aware event systems to address the depth-dependent parallax limitation identified by the authors."],"fun_headline_variants":["Entropy-Recency Map Tells Event SLAM Which Pixels to Distrust","No Segmentation Needed: Entropy-Recency Map Cuts Moving-Object Drift in Event SLAM","Motion-Compensated Event Timestamps Flag Unreliable Pixels for SLAM","Entropy of Compensated Timestamps Sharpens Event SLAM in Dynamic Scenes","Per-Pixel Temporal Reliability Score Stabilizes Monocular Event SLAM"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The entropy-recency score map depends on IMU-based motion compensation correctly aligning events from static structures to consistent pixels. If compensation is imperfect — due to depth-dependent parallax, IMU bias drift, or complex multi-body dynamics — static structures produce dispersed temporal distributions indistinguishable from dynamic ones, corrupting the temporal reliability signal.","fun_headline_variants_meta":{"raw":{"variants":["Entropy-Recency Map Tells Event SLAM Which Pixels to Distrust","No Segmentation Needed: Entropy-Recency Map Cuts Moving-Object Drift in Event SLAM","Motion-Compensated Event Timestamps Flag Unreliable Pixels for SLAM","Entropy of Compensated Timestamps Sharpens Event SLAM in Dynamic Scenes","Per-Pixel Temporal Reliability Score Stabilizes Monocular Event SLAM"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":621,"prompt_tokens":512,"completion_tokens":109,"prompt_tokens_details":null},"tokens_in":512,"tokens_out":109,"duration_ms":33664,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T12:39:51.287387+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If one constructed a scene where static structures produce late-biased, high-entropy event distributions after motion compensation — for instance, through rapid camera translation past nearby high-texture surfaces — the entropy-recency score would incorrectly flag them as dynamic, degrading pose estimation rather than improving it.","supporting_citations":[],"review_version":1}