{"id":"ec387e77-10c5-45bf-9112-346be797e3d8","arxiv_id":"2502.06734","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A 2M-pair instruction-based video editing dataset built from real videos and specialist models, demonstrated to train editors that beat prior methods.","lead":"The authors built a dataset of about two million video editing pairs by training four specialist models and filtering the results. The dataset is meant to give end-to-end video editing models the high-quality training data they currently lack.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Filtering pipeline is validated only on a tiny task-mixed set with hand-set thresholds; without per-task validation, the 'high-quality' label for all 18 tasks is not established.","rationale":"The reader's weakest assumption correctly identifies the quality-filtering classifier and hand-tuned thresholds as the load-bearing point. The paper's own Appendix D.1 documents the tiny validation set and the fixed thresholds, which is a genuine external-validity risk rather than an internal inconsistency. The concern is not that the pipeline is obviously broken, but that the evidence provided is too weak to certify 'high-quality' across 18 heterogeneous tasks, especially because object removal is explicitly exempted from text-alignment filtering and local edits are scored over full frames. I independently considered the CLIP circularity issue (CLIP used both as a curation filter and as an evaluation metric), but that is secondary: even if CLIPScore were partially inflated, the Ewarp and Temporal Consistency gains and the user-preference result provide some independent signal. The most decisive missing check is a per-task human evaluation of the retained pairs, which would directly test the filter's generalization. The controlled ablation in Table 3 (same sample count and epochs for InsV2V vs. Señorita-2M) is good support for the dataset's utility, and the expert-model comparisons in Tables 5-8 are useful, but neither isolates the filtering pipeline's quality. Because the concern strengthens the need for release and external validation rather than overturning the central contribution, the existing CONDITIONAL verdict remains appropriate; I mark this as UNCHANGED to indicate no verdict movement beyond what the reader already assigned.","tokens_in":31594,"tokens_out":4492,"duration_ms":44605,"concrete_test":"Ask the authors to release (or provide) the quality classifier scores and labels, or a random 10K filtered subset with human ratings. Have independent annotators label a task-stratified sample of at least 200 edited pairs per task (3,600 pairs) as pass/fail. Compute precision and recall at the fixed thresholds (0.6, 0.2/0.22, 0.95) and compare with thresholds tuned on this held-out set. If per-task precision at 0.6 falls below roughly 90% for any major task, or if optimal per-task thresholds deviate by more than about 0.1 from the fixed values, the 'high-quality' assertion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Señorita-2M is 'high-quality' rests on the filtering pipeline in §3.2.4 and Appendix D.1, but that pipeline is calibrated on a very small, task-mixed sample: 5,000 manually annotated edited videos (~1,000 failures), a validation set of only 120 successful and 60 failed samples, and fixed thresholds (0.6 for quality, 0.2/0.22 for text alignment, 0.95 for unchanged-pair removal) chosen without an independent task-stratified sweep. The 18 tasks have very different failure modes: object removal has no text-alignment filter at all, local stylization edits only a masked region, and global stylization compares long style prompts. A single ensemble classifier trained on features from 17 frames via CLIP ViT-H cannot be assumed to transfer uniformly to all tasks. If the classifier is miscalibrated on, say, object swap, inpainting, or outpainting, a substantial fraction of the retained 2M pairs may be low quality, and the downstream gains in Table 3 could reflect dataset scale or architecture choices rather than dataset quality. The paper reports no per-task precision/recall, no threshold sensitivity analysis, and the data and classifiers are not released, so this assumption is currently unfalsifiable from the paper alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Señorita-2M, a dataset of approximately two million instruction-based video editing pairs built from 388,909 real videos crawled from Pexels. The construction pipeline trains four expert models (global stylizer, local stylizer, inpainter, remover) on a WebVid-10M-derived annotation set, uses these experts plus vision tools (SAM2, depth/HED/Canny detectors) to generate edited pairs, and applies a cascade of quality, text-alignment, and visual-similarity filters. The authors also train several end-to-end editing architectures on the dataset and report that their best model outperforms prior methods on Ewarp, CLIPScore, temporal consistency, and user preference (Table 2). An ablation (Table 3) shows that training on Señorita-2M improves over training on InsV2V with the same sample budget.","tokens_in":31877,"tokens_out":3192,"duration_ms":29506,"significance":"If the quality claims are substantiated, this is a useful contribution to the video-editing community: it is among the first large-scale, instruction-based video editing datasets with both local and global edit types, and the expert-model design and filtering pipeline are a serious attempt at curating training data for end-to-end editors. The paper also reports a systematic comparison of editing architectures (Table 4), and the authors commit to open-sourcing the dataset and models, which would enable reproducibility and follow-up work. However, the central claim that Señorita-2M is 'high-quality' is currently supported mainly by the authors' own filtering pipeline, whose validation is thin, and by evaluations that lack statistical rigor. A careful revision that adds per-task filter validation, an independent text-alignment metric, and uncertainty measures would substantially strengthen the contribution.","major_comments":[{"comment":"The quality filter is the load-bearing component for the 'high-quality' claim, but its validation is reported only as 5,000 annotated videos (~1,000 failures) and a validation set of 120 successful and 60 failed samples, with no per-task precision, recall, or retention rates. Because the 18 tasks have very different failure modes (object removal has no text-alignment filter, local stylization compares masked regions, global stylization compares long style prompts), a single ensemble classifier with a fixed 0.6 threshold cannot be assumed to transfer uniformly. The paper should report per-task filter performance (precision/recall on a held-out stratified sample), the number of pairs retained per task, and a threshold sensitivity analysis. Without this, the 'high-quality' label for all 2M pairs is not established.","section":"§3.2.4 / Appendix D.1"},{"comment":"The text-alignment filter in Appendix D.2 uses CLIP text-video similarity to accept edited pairs, and the main evaluation in §4.3.1 uses CLIPScore to measure text-video alignment. Consequently, the reported text-alignment improvement (Table 2, CLIPScore 0.2895 vs. 0.2723) may be partly inherited from the curation criterion rather than reflecting a genuinely better editor. To make the claim convincing, the authors should evaluate text alignment with an independent metric (e.g., a different vision-language model or a human judgment study on the DAVIS outputs) or at least show that the improvement persists when the evaluation prompts are outside the distribution used for filtering.","section":"§4.3.1 and §D.2"},{"comment":"The quantitative comparison reports no error bars, no confidence intervals, and no significance tests. The differences in Ewarp and CLIPScore between methods are small in absolute terms (e.g., CLIPScore 0.2895 vs. 0.2723), and the user study is described only by a single preference percentage (53.17%) with no information about the number of participants, the number of videos rated, or the rating protocol. Without this information, the claim of state-of-the-art performance is not statistically supported. The authors should provide variance estimates (e.g., bootstrap intervals) across DAVIS videos and full details of the user study.","section":"Table 2 and §4.3.2"},{"comment":"The paper claims 18 editing tasks and approximately 2M pairs, but it does not report the number of pairs per task after filtering. Since tasks such as inpainting/outpainting contribute roughly 60,000 pairs (Appendix C.4.4) and conditional generation tasks (depth, HED, etc.) are likely much larger, the distribution is highly uneven. If some tasks have very few retained pairs after filtering, the dataset's usefulness for 'general' video editing is unclear. Reporting per-task counts before and after filtering is necessary to assess coverage.","section":"§3.2.2 and Table 1"},{"comment":"The expert models' quantitative comparisons (Tables 5–8) are computed on datasets that appear to be the same or similar to the expert training/evaluation data, and they are not accompanied by any error bars either. For instance, Table 7 reports that InsV2V has the lowest Ewarp on object swap, which the authors explain as a failure mode (no actual swap), but this interpretation should be validated by a human study or by reporting per-sample statistics. More generally, the lack of uncertainty quantification in all experimental tables makes it difficult to judge whether the reported differences are meaningful.","section":"§B.4 (Table 7) and §3.2.4"}],"minor_comments":[{"comment":"There are numerous typos and formatting errors, e.g., 'blodfaced' in Tables 2 and 8, the accented title 'Se\\~norita' appearing inconsistently in the text, and several broken bibliography entries (e.g., '?;' in the first paragraph of the introduction and the Pexels URL formatting).","section":"General"},{"comment":"The text says 'Different thresholds were applied for different tasks' in the quality filter, but §4.1 only reports a single threshold of 0.6 and mentions a lower threshold for object addition. The threshold values for each task should be listed explicitly in one place.","section":"§3.2.4 / §D.1"},{"comment":"The evaluation uses 'randomly generated editing prompts' on DAVIS, but it is unclear how these prompts were generated and whether they are appropriate for all compared methods, especially inversion-based methods that may require specific prompt formats. A description of prompt generation and validation is missing.","section":"§4.3.1"},{"comment":"The text says 'we set thresholds of 0.2 and 0.22, respectively' for object swap and local stylization, but the preceding sentence says 'we set a lower threshold of 0.2' for global stylization and later §4.1 says global stylization and object addition use 0.2. Please clarify the exact thresholds for each task in the main text to avoid confusion.","section":"Appendix D.2"},{"comment":"The training details for the editing model report two stages (first stage at 336×592, second stage at 448×768), but the second-stage data and any potential domain shift are not described. A brief note on the resolution adaptation would help reproducibility.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The central idea—a large, instruction-based video editing dataset built by expert models and filtered by a quality pipeline—is timely and likely to be of interest to the community. The main concern is that the paper overclaims the quality and the evaluation. The missing per-task filter validation, the CLIP-in-filter/CLIP-in-evaluation circularity, and the absence of error bars and user-study details are fixable with additional experiments and reporting, so I do not recommend rejection. However, the authors should be required to provide the per-task breakdowns and a more rigorous evaluation before publication. I also note that the paper says the dataset will be open-sourced 'upon acceptance'; if the data and filtering code are not released, some of the validation we request cannot be independently checked, so the availability of these artifacts should be confirmed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Señorita-2M is the first genuinely large instruction-based video editing dataset built from real videos, and that is the main story. Roughly two million pairs across 18 tasks, with four trained specialist models doing the editing and a three-stage filter cleaning the result. If the release matches the description, this becomes a standard training resource for end-to-end video editors, similar to what InstructPix2Pix and UltraEdit did for images. That is the thing to know before reading.\n\nWhat is actually new: no existing open dataset comes close on scale and task diversity. VIVID-10M is not released and is local-only; InsV2V is synthetic and small. The four expert models—global stylizer, local stylizer, inpainter, remover—are trained on CogVideoX with a fair amount of care, and the appendix gives real implementation detail. The ablation against InsV2V with matched 60K training samples, and the architecture comparison, are the right kind of evidence. The filtering pipeline—quality classifier, CLIP text-alignment, CLIP visual-similarity—is a sensible design.\n\nNow the soft spots. The quality filtering rests on a small validation set: roughly 120 successful and 60 failed videos, with hand-set thresholds for different tasks, and no per-task precision/recall or threshold sensitivity analysis. The paper itself notes that object removal gets no text-alignment filter and that thresholds differ by task; that is honest, but it also means the 'high-quality' label for all 18 tasks is not actually established. The CLIP circularity is real but partial: CLIP is used in curation and in the CLIPScore metric, so the text-alignment gains in Table 2 are partly inherited. The user study helps, but participant and video counts are missing. There are no error bars anywhere, and the DAVIS benchmark with randomly generated prompts is not a standard evaluation. Data and models are not released.\n\nNone of this kills the resource. The ablation comparing Señorita-2M against InsV2V on the same 60K budget shows gains in temporal consistency and CLIPScore, and the user preference numbers point the same way. The central claim—a large real-video dataset helps train better end-to-end editors—holds up. It is just weaker than the paper's tone suggests.\n\nBottom line: this deserves peer review, not desk rejection. A serious referee should ask for the release and for per-task filtering validation, or at least threshold sensitivity, plus basic statistics on the user study. I would read it carefully if I were building video editors and bring it to a reading group, but I would not cite it as a released resource until the data is actually out.","headline":"A genuinely new, large-scale real-video editing dataset with solid engineering, but the filtering quality and evaluation are less rigorous than the claims require.","tokens_in":32431,"tokens_out":2694,"would_cite":false,"duration_ms":25694,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-million-pair instruction dataset, produced by four specialist editors and filtered with a learned quality classifier, trains a first-frame-guided ControlNet editor that outperforms prior video editing methods on frame consistency…","keywords":["video editing dataset","instruction-based editing","end-to-end video editing","diffusion models","data filtering","video inpainting","ControlNet","text-video alignment"],"falsifier":"Take a random sample of pairs from each of the 18 task categories, run the same quality classifier and CLIP thresholds, and have annotators judge success and failure; if the annotators' failure rate is far above the pipeline's implied rate, or if many pairs judged successful show no visible change, the filtering claim fails. A simpler direct check is to train the identical editor on a 120K pre-filter subset versus a 120K post-filter subset; if metrics do not improve, the filter is not doing the work.","tokens_in":31378,"feed_emoji":"🎬","tokens_out":8963,"duration_ms":73654,"temperature":0.7,"pith_summary":"Señorita-2M sets out to close the data gap that holds back end-to-end video editing: models trained on (source video, edited video, instruction) triples edit fast, but until now no large set of clean training pairs existed. The paper builds roughly two million such triples by running four specialized editing models over real Internet videos, rewriting edit prompts into natural instructions with a language model, and discarding failures through a learned quality classifier plus CLIP-based thresholds for text alignment and amount of change. A first-frame-guided ControlNet editor trained on this dataset reports lower warp error, higher text alignment, and better temporal consistency than TokenFlow, Flatten, AnyV2V, and InsV2V, and wins 53.17% of user preferences. The ablation shows that Señorita-2M pairs improve scores over InsV2V pairs even when the sample count is held at 60K, so the dataset's quality, not just its size, is the point.","feed_headline":"2M pairs beat inversion video editors","feed_subtitle":"Señorita-2M adds ~2 million filtered instruction-video pairs; the trained editor leads on consistency and text alignment.","key_machinery":"The load-bearing object is the dataset construction and filtering pipeline rather than any single network. Candidate edits are generated by four specialists built on a common video diffusion base; a language model converts object names and style prompts into clear instructions; and a three-stage filter keeps only pairs that look edited (ensemble MLP classifiers on CLIP frame features, threshold 0.6), match their instruction (CLIP text–video similarity, thresholds 0.2 to 0.22 depending on task), and differ enough from the source (vision-similarity threshold 0.95). On the training side, the architecture that carries the result is a ControlNet-conditioned video diffusion transformer in which the edited first frame is concatenated into the main branch, letting a single edit anchor the whole clip. The remover also relies on a deliberate training trick: 90% of its masks come from unrelated videos, which breaks the usual correlation between mask shape and content and lets classifier-free guidance erase the target object.","core_discovery":"The central claim is that specialist-generated, heavily filtered training pairs can make end-to-end instruction-based video editing the strongest option, not merely a fast one. Four expert editors—a global stylizer, a local stylizer, a text-guided inpainter, and a remover—are trained on a common video diffusion base and applied to 388,909 crawled videos, producing candidates across 18 tasks including style transfer, object swap, removal, addition, inpainting, outpainting, grounding, and conditional generation. The candidates pass through a cascade that removes low-quality edits (quality classifier, threshold 0.6), text-misaligned edits (CLIP similarity thresholds 0.2 and 0.22), and near-identical pairs (similarity above 0.95). Training the final editor with a first-frame-guided ControlNet on the surviving pairs yields Ewarp 9.42, CLIPScore 0.2895, temporal consistency 0.9775, and 53.17% user preference against the four baselines, and the equal-sample ablation attributes a clear part of that gain to the dataset itself.","pith_inferences":["The quality label would benefit from an external audit: randomly sample pairs from all 18 tasks and compare human failure judgments with the quality classifier's decisions, something the paper does not report.","The user-preference figure is reported without participant counts or protocol details in the main text, so it is best read as a directional signal until the study design is available.","A direct way to isolate the filter's contribution would be to train the same editor on random pre-filter pairs versus post-filter pairs at equal size; the metric gap would quantify how much of the gain comes from filtering rather than from the specialist generators.","The remover's unrelated-mask trick is a transferable idea: any inpainting-style system that should erase rather than regenerate can break the mask–content correlation the same way."],"forward_implications":["An end-to-end editor trained on Señorita-2M inherits the fast single-pass inference of supervised methods while beating inversion-based baselines on frame consistency and text alignment.","Holding training samples at 60K, Señorita-2M still improves CLIPScore and temporal consistency over InsV2V data, so the dataset can benefit researchers who cannot reproduce the four expert generators.","The best architecture found, first-frame-guided ControlNet, points to edited-key-frame conditioning as the productive design for learning video edits from pairs.","The 18 task types, including grounding and conditional generation, make the dataset a plausible starting point for a single multi-task video editor rather than separate per-task models.","Releasing the dataset and trained models, as the paper says it will, would let others reproduce the reported numbers and train editors on new base models."],"supporting_citations":[{"why":"Supplies the CogVideoX video diffusion transformer used as the base of all four expert editors and of the final editing model.","marker":"Yang et al., 2024"},{"why":"Provides WebVid-10M, the training corpus on which the four expert editors are trained.","marker":"Bain et al., 2021"},{"why":"CogVLM2 captions the videos and lists object names used for mask and phrase generation.","marker":"Hong et al., 2024"},{"why":"Grounded-SAM2 produces the object masks and tracked segments used for local edits and expert training.","marker":"Liu et al., 2023a; Ravi et al., 2024"},{"why":"Flux-Fill edits the first frame that guides object swap and inpainting.","marker":"Black Forest Labs, 2024"},{"why":"LLaMA-3 rewrites object names and style prompts into the final editing instructions.","marker":"Dubey et al., 2024"},{"why":"InsV2V is the trained baseline and the data source replaced in the equal-sample ablation.","marker":"Cheng et al., 2024"},{"why":"InstructPix2Pix provides the architecture compared as Ins-Edit and the precedent for prompt-to-prompt filtering.","marker":"Brooks et al., 2023"},{"why":"ControlNet provides the conditional control branch used in the final editor and in the stylizers.","marker":"Zhang et al., 2023a"},{"why":"CLIP features drive the quality classifier, text-alignment filter, similarity filter, and evaluation metrics.","marker":"Radford et al., 2021"}],"fun_headline_variants":["2M filtered pairs make end-to-end video editing win","Señorita-2M: 2M pairs beat inversion editors","Specialist-built 2M pairs improve end-to-end video editing","High-quality 2M pairs boost end-to-end video editing","Filtered 2M pairs make end-to-end editors outperform inversion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole dataset's 'high quality' label rests on a failure detector trained from 5,000 hand-labeled videos and on three similarity thresholds chosen without a separate validation set; if that detector and those thresholds do not transfer to all 18 tasks and all two million pairs, the dataset quality claim is not established.","fun_headline_variants_meta":{"raw":{"variants":["2M filtered pairs make end-to-end video editing win","Señorita-2M: 2M pairs beat inversion editors","Specialist-built 2M pairs improve end-to-end video editing","High-quality 2M pairs boost end-to-end video editing","Filtered 2M pairs make end-to-end editors outperform inversion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000897,"raw_usage":{"total_tokens":3902,"prompt_tokens":1017,"completion_tokens":2885,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":2795}},"tokens_in":633,"tokens_out":2885,"duration_ms":17350,"temperature":1.0,"reasoning_tokens":2795,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:29:47.579917+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of pairs from each of the 18 task categories, run the same quality classifier and CLIP thresholds, and have annotators judge success and failure; if the annotators' failure rate is far above the pipeline's implied rate, or if many pairs judged successful show no visible change, the filtering claim fails. A simpler direct check is to train the identical editor on a 120K pre-filter subset versus a 120K post-filter subset; if metrics do not improve, the filter is not doing the work.","supporting_citations":[],"review_version":1}