{"id":"b02a3e32-015d-4bb6-828b-5fc6b84fa177","arxiv_id":"2608.10479","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Event-guided adapters using IWEs and bidirectional flow, injected into a pre-trained Wan2.1 FLF2V diffusion model, improve perceptual quality and temporal coherence of generated intermediate frames.","lead":"The authors adapt a pre-trained video diffusion model to interpolate frames using event camera data, by injecting event-derived edge maps and optical flow through lightweight adapters. The method reports better perceptual quality and temporal consistency than existing baselines on several event video benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No deduplication between EvPexels training videos and Pexels test clips is reported; if they overlap, the headline Pexels results are invalid.","rationale":"The reader's stated weakest assumption was the accuracy of the event-only optical flow used for warping. That is a legitimate concern, but I do not think it is the most load-bearing one. Section 4.3.1 shows that removing flow warping degrades PSNR/SSIM/LPIPS/FID/FVD (Tab. 3, rows 'w/o flows warping' vs 'Full model'), so the flow guidance is empirically beneficial even if the flow fields are imperfect; moreover, the paper's Sec. 4.4 provides only a qualitative visualization, not a quantitative flow-accuracy check. In contrast, the Pexels training/test overlap issue directly threatens the numeric basis of the headline claim. The paper explicitly builds a Pexels-derived training set and then evaluates on Pexels clips without stating any deduplication procedure. This is not a subtle modeling assumption; it is a standard benchmark-integrity condition that can be checked by file/source-ID comparison or perceptual hashing. Because the reader already assigned a CONDITIONAL verdict and explicitly noted the Pexels overlap concern in the rationale, I keep the verdict unchanged: the paper should be accepted only after the overlap check is performed and reported. If overlap is found, the Pexels results and the 'consistently outperforms' wording must be corrected, and the remaining claims should be re-evaluated on BS-ERGB and DAVIS alone.","tokens_in":12085,"tokens_out":4645,"duration_ms":45105,"concrete_test":"Obtain or reconstruct the source-video identifiers for the 30 Pexels test clips and the 1,100 EvPexels training videos; compute exact source-ID matches and near-duplicate content overlap using perceptual hashing on sampled frames. If any test clip shares a source video or an overlapping temporal segment with the training set, remove those test clips and recompute the Tab. 2 Pexels metrics. If the margin over VDM-EVFI-Wan2.1 disappears on the disjoint subset, the 'consistently outperforms' claim must be revised; if no overlap exists, the conditional verdict can stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of consistent state-of-the-art performance rests heavily on Tab. 2, where the method wins all five metrics on Pexels. Section 4.1.1 states that the EvPexels training set is 'constructed from videos collected via Pexels' (1,100 sequences), while Section 4.1.2 states that the test set includes '30 clips from Pexels'. Nowhere do the authors state that these 30 test clips are disjoint from the 1,100 EvPexels training videos, nor do they report any source-video ID check, frame-level deduplication, or temporal-segment overlap analysis. Because both training and test data come from the same platform, the possibility of exact or near-duplicate content is real and unaddressed. If even a subset of the Pexels test clips overlaps with EvPexels training sequences, the Pexels columns in Tab. 2 are inflated by training/test leakage, and the abstract's claim that the method 'consistently outperforms existing state-of-the-art approaches' is unsupported on one of the three benchmarks. The concern is decisive because it is about the validity of the empirical evidence itself, not about an ablation or a modeling choice. If the overlap is zero, the concern is fully resolved; if it is nonzero, the Pexels results must be recomputed on a disjoint subset and the headline claim revised. This is a concrete, checkable condition, and the paper currently provides no information to verify it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an adapter-based framework for event-guided video frame interpolation built on a frozen DiT-based image-to-video diffusion model (Wan2.1 FLF2V). Event streams are converted into Image Warped Events (IWEs) and bidirectional sparse optical flow by contrast maximization; an IWE encoder injects edge-aware structural features into the latent input, and a flow-based alignment-and-fusion adapter warps DiT features from neighboring latent frames before selected DiT blocks. The model is fine-tuned with LoRA on a newly introduced synthetic dataset, EvPexels, plus the BS-ERGB training split. Experiments at x24 interpolation on BS-ERGB, DAVIS, and Pexels report state-of-the-art perceptual metrics on BS-ERGB and best all-metric results on DAVIS and Pexels.","tokens_in":12487,"tokens_out":3252,"duration_ms":27921,"significance":"If the empirical claims hold, the paper offers a practically valuable recipe for incorporating event data into large pre-trained I2V diffusion models without retraining from scratch, and the EvPexels dataset is a potentially useful community resource. The design is clean: event cues are converted into modalities compatible with standard diffusion control, and the ablations in Tables 3-5 support the contribution of each component. However, the headline claim of consistent state-of-the-art performance currently rests on results that are not fully verified due to a possible train/test overlap on the Pexels benchmark, as well as an abstract that overstates the BS-ERGB numbers.","major_comments":[{"comment":"The Pexels test set is drawn from the same platform used to construct the EvPexels training set, but the paper reports no source-video ID check, frame-level deduplication, or temporal-segment overlap analysis between the 30 Pexels test clips and the 1,100 EvPexels training sequences. If any of the test clips overlap with training content, the Pexels columns in Table 2 are inflated and the headline claim of consistent state-of-the-art performance is unsupported on one of the three benchmarks. Please report the exact source video IDs/URLs for both sets, verify disjointness, and if any overlap exists recompute the Pexels metrics on a strictly disjoint subset.","section":"4.1.1, 4.1.2, Table 2"},{"comment":"The abstract's statement that the method 'consistently outperforms existing state-of-the-art approaches' is not supported by Table 1: on BS-ERGB, CBMNet-Large achieves higher PSNR (25.306 vs. 23.261) and SSIM (0.7120 vs. 0.704), and TimeLens achieves higher PSNR (24.704 vs. 23.261). The body text at the end of Section 4.2.1 correctly qualifies the result as state-of-the-art on perceptual metrics, so the abstract and conclusion should be revised to match, e.g., 'state-of-the-art perceptual quality on BS-ERGB and best all-metric results on DAVIS and Pexels.'","section":"Abstract, 4.2.1, Table 1"},{"comment":"The whole pipeline depends on event-only sparse bidirectional optical flow to warp DiT features across the temporally compressed latent space, but the paper provides no quantitative accuracy check of these flow estimates against ground truth; Section 4.4 offers only visual inspection. If the flows are noisy or incomplete in occluded, low-event-density, or large-appearance-change regions, the warping in Eq. (7) could corrupt features rather than align them. Please add a quantitative flow-error evaluation, e.g., endpoint error on a synthetic benchmark with ground-truth flow, and discuss failure modes for the event-only flow estimates.","section":"3.2.1, Eq. (6)-(8), 4.4"}],"minor_comments":[{"comment":"The heading 'Ablation Studys' contains a typo and should read 'Ablation Study.'","section":"4.3"},{"comment":"The header 'DA VIS' has an erroneous space and should read 'DAVIS.'","section":"Table 2"},{"comment":"Reference [12] contains the typo 'W ACV' in the venue name; it should be 'WACV.'","section":"References"},{"comment":"The notation in Figure 2, especially the arrow labels such as 'w' and the relationship between IWEs, flows, and the alignment adapters, is dense and not fully explained in the caption; a more detailed caption or a legend would aid reproducibility.","section":"Figure 2"},{"comment":"The main fine-tuning uses 4,000 steps on 8 NVIDIA A800 GPUs, while the ablations are trained for 5,400 steps on a single GPU; this difference in training budget should be stated clearly so that the ablation numbers in Table 3 are not directly compared to Table 1 without this caveat.","section":"4.1.3, 4.3.1"}],"recommendation":"major_revision","confidential_remarks":"The data-leakage concern on the Pexels benchmark is the most serious issue and is likely decisive for the paper's central empirical claim. The authors need to provide concrete evidence of disjointness between the EvPexels training set and the Pexels test clips. If the overlap is nonzero, the headline results must be recomputed and the abstract revised. The technical approach is otherwise plausible and the ablations are informative, so the paper can be made suitable for publication after this verification and the abstract correction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my read on arXiv:2608.10479. The thing to know: this is a genuinely new adapter recipe for event-guided video interpolation in a DiT-based I2V model. Prior work (VDM-EVFI) used a ControlNet-style voxel grid on a U-Net; this paper instead extracts IWEs and bidirectional sparse flow via contrast maximization and injects them as lightweight adapters into Wan2.1's FLF2V. That is a sensible design, and the ablations show each piece contributes. The paper also introduces EvPexels, a large synthetic event-video dataset, which the community could reuse.\n\nThe experimental work is broad: five baselines, three benchmarks, x24 interpolation, and the perceptual metrics (LPIPS, FID, FVD) are consistently strong. The text is honest about the PSNR/SSIM tradeoff on BS-ERGB, where CBMNet-Large beats them. But the abstract's \"consistently outperforms\" is directly contradicted by Table 1, and that should be fixed before publication.\n\nThe bigger soft spot is the data-leakage risk. The EvPexels training set is built from Pexels videos (1,100 sequences), and the Pexels test set is 30 clips from the same platform. The paper never reports a deduplication check. That is a concrete, checkable issue. If the test clips overlap with training sequences, the Pexels columns in Table 2 are inflated. The authors should either show disjoint source-video IDs and frame-level dedup, or recompute on a clearly disjoint subset. This does not sink the whole paper, since DAVIS and BS-ERGB stand independently, but the Pexels claim is currently unverified.\n\nAlso minor: the paper relies on event-only optical flow from contrast maximization, but provides only a qualitative visual check (Fig. 6) that these flows are accurate. A quantitative accuracy number would strengthen trust in the warping adapter. And there is no code or data release, so independent verification is not yet possible.\n\nWho this is for: researchers in event-based vision and diffusion-based video interpolation. The adapter idea is reusable beyond VFI.\n\nMy bottom line: worth a serious referee. Send it to peer review with a required deduplication analysis and an abstract rewrite. If the overlap turns out to be zero, the paper is in good shape; if not, the Pexels results need to be re-run.","headline":"A well-ablated adapter recipe for event-guided DiT interpolation, but the 'consistently outperforms' claim is too strong and the Pexels test set may overlap the EvPexels training data.","tokens_in":12938,"tokens_out":2508,"would_cite":true,"duration_ms":20774,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Event motion cues beat frame-only baselines at x24 video interpolation.","keywords":["event cameras","video frame interpolation","diffusion transformer","image-to-video diffusion","optical flow","image warped events","contrast maximization","adapter fine-tuning"],"falsifier":"Measure the estimated bidirectional flows against ground-truth optical flow on a synthetic scene with known dense motion; if the event-only flows are systematically wrong in occluded or textureless regions and those errors propagate to the interpolated frames (visible as warping artifacts or drops in FVD/LPIPS), the core assumption fails.","tokens_in":11868,"feed_emoji":"🎬","tokens_out":7018,"duration_ms":55930,"temperature":0.7,"pith_summary":"This paper claims that sparse, asynchronous event-camera data—converted into image warped events (IWEs) and bidirectional sparse optical flow via contrast maximization—can be injected into a pre-trained DiT-based image-to-video diffusion model through two lightweight adapters, improving video frame interpolation across large temporal gaps. The proposal matters because event cameras capture high-temporal-resolution motion that frame-only methods miss, and the adapter approach avoids training an event-assisted generative model from scratch. On the real BS-ERGB benchmark the method sets state-of-the-art perceptual metrics (LPIPS, FID, FVD), and on the synthetic DAVIS and Pexels benchmarks it leads on all five reported metrics at x24 interpolation. The paper also contributes EvPexels, a synthetic event-video dataset of about 390,000 frames, to support training and benchmarking.","feed_headline":"Event motion cues beat frame-only baselines at x24 video interpolation","feed_subtitle":"Event-derived edges and optical flow, added via adapters, raise interpolation fidelity and temporal coherence.","key_machinery":"The engine of the method is the pairing of two event-derived signals—Image Warped Events and bidirectional sparse optical flow—both computed from the raw event stream via contrast maximization between consecutive latent frames. The IWE encoder is a compact 3D-convolution network whose features are added element-wise to the input latents, supplying edge-like structural cues. The flow-based alignment-and-fusion adapter warps DiT features from the previous and next latent frames toward the current frame using the estimated flows, aggregates them with the current frame's features, and adds the fused result as a residual correction before selected DiT blocks. This explicit motion-guided feature warping is the mechanism that enforces temporal coherence, supported by LoRA fine-tuning of the frozen backbone.","core_discovery":"The central claim is that event streams, when expressed as IWEs (edge-aligned images obtained by warping events along flow) and bidirectional optical flow, provide spatially and temporally aligned guidance that a frozen DiT-based image-to-video model can consume with minimal architectural change. The paper shows that adding an IWE encoder that injects edge features into the latent input, plus flow-based alignment-and-fusion adapters that warp DiT features from neighboring latent frames toward the current frame before selected blocks, reduces motion blur, structural distortion, and temporal inconsistency relative to frame-only interpolation. The authors demonstrate this on real and synthetic benchmarks, reporting that the method outperforms existing state-of-the-art approaches: state-of-the-art perceptual quality on BS-ERGB and best all-metric results on DAVIS and Pexels at x24 interpolation. They further argue that this proves event-derived flow and IWE can serve as drop-in guidance for large generative interpolators, circumventing the need to adapt sparse event streams directly into dense grid-based representations.","pith_inferences":["Because the flow-warping adapter is agnostic to how the flow was produced, the same architecture could in principle be fed optical flow estimated from RGB frames or from a hybrid event-frame estimator, making the approach portable to frame-only interpolation settings.","The accuracy of the event-only flow is the ceiling on the gains: in occluded or low-event-density regions, noisy warps could actively corrupt features, so a learned flow-refinement step or confidence-weighted fusion would be a natural, testable extension.","The benchmark pattern—perceptual metrics winning while PSNR trails traditional methods on BS-ERGB—suggests the model trades pixel-exactness for perceived realism; a user study or task-based evaluation would quantify whether that trade-off is preferred.","The ablation showing injection into the first DiT blocks favors PSNR/SSIM while last blocks favor LPIPS/FID/FVD hints at a complementary schedule: injecting at both early and late blocks might combine reconstruction and perceptual strengths."],"forward_implications":["Event cameras become a practical conditioning signal for large generative video models, enabling high-quality interpolation at x24 temporal gaps with only adapter-level fine-tuning.","The same IWE-plus-flow adapter recipe could be transplanted to other DiT-based image-to-video backbones, since it leaves the base denoiser frozen.","The released synthetic event-video dataset EvPexels allows other researchers to train and evaluate event-guided interpolation without collecting real event data.","The ablation results imply that explicit motion-guided feature warping is more effective than simply concatenating event features into the latent input, guiding future designs for temporal conditioning."],"supporting_citations":[{"why":"Supplies the off-the-shelf contrast-maximization method used to compute bidirectional sparse optical flow and the warp-based IWE representations from raw event streams.","marker":"[24]"},{"why":"Establishes the contrast maximization principle that the event-to-flow/IWE extraction builds on.","marker":"[26]"},{"why":"Provides the pre-trained DiT-based image-to-video (FLF2V) model that the adapter framework fine-tunes, serving as the frozen backbone.","marker":"[31]"},{"why":"Provides the Vid2e simulator used to synthesize event streams from RGB videos when constructing the EvPexels training dataset.","marker":"[7]"},{"why":"Supplies the BS-ERGB real high-speed event-image dataset used for training and the primary real-benchmark test set.","marker":"[28]"},{"why":"Serves as the main event-conditioned diffusion baseline (VDM-EVFI) that the method is compared against on the Wan2.1 backbone.","marker":"[4]"},{"why":"A strong event-based interpolation baseline (CBMNet) that the method must beat on distortion and perceptual metrics.","marker":"[13]"},{"why":"Motivates the IWE encoder by showing that image warped events correlate strongly with scene edges and object boundaries.","marker":"[12]"}],"fun_headline_variants":["Event adapters steer frozen DiT to sharper x24 interpolation","Frozen diffusion learns from events: state-of-the-art video interpolation","Event-guided X24 interpolation: adapters beat frame-only baselines","IWEs plus flow adapters: minimal changes, maximal interpolation gains","Event streams plug into DiT: sharper, faster video interpolation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the event-only optical flow estimates, produced by contrast maximization in temporal segments between latent frames, stay accurate enough to warp DiT features across the 4x temporal compression, even in regions with occlusions, low event density, or large appearance changes.","fun_headline_variants_meta":{"raw":{"variants":["Event adapters steer frozen DiT to sharper x24 interpolation","Frozen diffusion learns from events: state-of-the-art video interpolation","Event-guided X24 interpolation: adapters beat frame-only baselines","IWEs plus flow adapters: minimal changes, maximal interpolation gains","Event streams plug into DiT: sharper, faster video interpolation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000689,"raw_usage":{"total_tokens":3112,"prompt_tokens":928,"completion_tokens":2184,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":2095}},"tokens_in":544,"tokens_out":2184,"duration_ms":14351,"temperature":1.0,"reasoning_tokens":2095,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:18:58.281106+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the estimated bidirectional flows against ground-truth optical flow on a synthetic scene with known dense motion; if the event-only flows are systematically wrong in occluded or textureless regions and those errors propagate to the interpolated frames (visible as warping artifacts or drops in FVD/LPIPS), the core assumption fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the BS-ERGB real high-speed event-image dataset used for training and the primary real-benchmark test set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the contrast maximization principle that the event-to-flow/IWE extraction builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Vid2e simulator used to synthesize event streams from RGB videos when constructing the EvPexels training dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as the main event-conditioned diffusion baseline (VDM-EVFI) that the method is compared against on the Wan2.1 backbone."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A strong event-based interpolation baseline (CBMNet) that the method must beat on distortion and perceptual metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the IWE encoder by showing that image warped events correlate strongly with scene edges and object boundaries."}],"review_version":1}