{"id":"6c17467c-b2b5-417c-8e9f-7c301b973905","arxiv_id":"2604.24947","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A new large-scale subjective database for video portrait region cropping with temporal smoothing, benchmarked using existing models and compared to saliency predictions.","lead":"The paper introduces the LIVE-YT VC database of 1800 videos with subjective human annotations for cropping important portrait regions from landscape videos, plus a temporally smoothed version called LIVE-YT VC++. This resource supports development of better video aspect-ratio adaptation methods for mobile viewing that avoid distortion while keeping essential content.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Claim that the database 'is the largest publicly-available' is unsupported because the abstract only says the authors 'plan to open source' it.","rationale":"Reader's weakest assumption concerns annotation reliability and generalizability. The load-bearing issue for the stated strongest claim is instead the mismatch between 'is publicly-available' and the future-tense release plan; these are distinct failure modes. The concern is concrete and falsifiable by checking for an actual release.","tokens_in":1777,"tokens_out":272,"duration_ms":49054,"concrete_test":"Scan the full manuscript (including any data availability statement or footnote) for an actual release URL, GitHub repository, or explicit statement that the dataset is already downloadable; if only 'plan to' language remains, the headline claim must be qualified or retracted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that LIVE-YT VC 'is the largest publicly-available subjective video portrait region cropping database.' The abstract ends by stating 'we plan to open source the project,' using future tense. No release link, DOI, or confirmation of current public availability appears in the provided text. This directly falsifies the present-tense 'is publicly-available' predicate required for the claim to hold, independent of annotation quality or size comparisons.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces the LIVE-YouTube Video Cropping (LIVE-YT VC) database of 1800 videos annotated by 90 human subjects for subjective portrait region cropping, drawn from YouTube-UGC and LSVQ sources. It also describes LIVE-YT VC++, a smoothed variant created via a novel intra-frame temporal filter. The work claims this is the largest publicly-available database for the task, demonstrates its utility via the SmartVidCrop algorithm and fine-tuned video grounding models, compares annotations to saliency predictions, and states plans to open source the resource as a benchmark for aspect-ratio adaptation in mobile video.","tokens_in":1873,"tokens_out":448,"duration_ms":66873,"significance":"If the database is released and the annotations are shown to be reliable, this would supply a valuable large-scale subjective resource for video cropping research, filling a gap for methods that adapt landscape video to portrait/mobile formats while preserving content. The temporal smoothing idea directly targets annotation variability, and the empirical benchmarking against existing models provides usable baselines. The scale (1800 videos, 90 annotators) and connection to saliency tasks strengthen its potential as a community benchmark.","major_comments":[{"comment":"Abstract: the statement that 'this new resource is the largest publicly-available subjective video portrait region cropping database' is unsupported. The abstract ends by stating 'we plan to open source the project' (future tense) with no release link, DOI, or confirmation of current public availability, directly falsifying the present-tense claim that is central to the paper's contribution.","section":"Abstract"},{"comment":"Database construction and LIVE-YT VC++ sections: the manuscript provides no details on inter-annotator agreement among the 90 subjects, the exact equations or parameters of the intra-frame temporal filter, or statistical validation of the smoothed ++ version relative to the raw annotations. These omissions prevent assessment of whether the resource reliably captures 'significant regions' and can serve as a reproducible benchmark.","section":"Database construction and LIVE-YT VC++"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments on our manuscript. We address each major comment below, indicating where revisions will be made to strengthen the presentation and ensure accuracy.","responses":[{"response":"We agree that the abstract contains an inconsistency in tense that requires correction. The database has been fully collected and processed, but public release is planned upon acceptance to permit any final refinements. We will revise the abstract to state that the resource 'will be the largest publicly-available' upon release and will add a footnote or dedicated section providing a planned release link or repository identifier. This change ensures the claim is accurately supported without misrepresentation.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the statement that 'this new resource is the largest publicly-available subjective video portrait region cropping database' is unsupported. The abstract ends by stating 'we plan to open source the project' (future tense) with no release link, DOI, or confirmation of current public availability, directly falsifying the present-tense claim that is central to the paper's contribution."},{"response":"We concur that these details are necessary to demonstrate reliability and enable reproducibility. In the revised manuscript we will expand the relevant sections to report: quantitative inter-annotator agreement metrics (e.g., mean IoU and Fleiss' kappa across the 90 subjects); the exact mathematical formulation and all parameter values of the intra-frame temporal filter (including smoothing coefficients and window sizes); and statistical comparisons between the raw and smoothed annotations (e.g., variance reduction and temporal consistency measures). These additions will allow readers to evaluate the resource's suitability as a benchmark.","revision_made":"yes","referee_comment":"[Database construction and LIVE-YT VC++] Database construction and LIVE-YT VC++ sections: the manuscript provides no details on inter-annotator agreement among the 90 subjects, the exact equations or parameters of the intra-frame temporal filter, or statistical validation of the smoothed ++ version relative to the raw annotations. These omissions prevent assessment of whether the resource reliably captures 'significant regions' and can serve as a reproducible benchmark."}],"tokens_in":1521,"tokens_out":455,"duration_ms":35431,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is a new collection of subjective labels for cropping important regions in landscape videos to fit portrait displays, plus a post-processing filter to smooth those labels over time. They took 1800 videos from YouTube-UGC and LSVQ, had 90 people annotate them, and created a smoothed LIVE-YT VC++ version. They then ran SmartVidCrop and some video grounding models on it, and checked overlap with saliency predictions. That gives a concrete resource for a common mobile-video problem where static crops or warps look bad.","headline":"The paper assembles a large set of human annotations for video portrait cropping and adds temporal smoothing, but the data is not released and key methodological details are missing.","tokens_in":2352,"tokens_out":190,"would_cite":false,"duration_ms":39517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A new database collects human annotations for cropping significant regions from landscape videos into portrait format.","keywords":["video cropping","aspect ratio","subjective annotations","temporal smoothing","portrait region","video database","mobile video consumption","video saliency"],"falsifier":"If independent human evaluators consistently prefer crops generated by models not using this dataset over those trained on it, when judging preservation of important content and visual quality.","tokens_in":2691,"feed_emoji":"📹","tokens_out":555,"duration_ms":52315,"temperature":0.7,"pith_summary":"The paper establishes a large-scale resource to support cropping videos for different aspect ratios without losing key content or introducing distortion. It gathers annotations from 90 subjects across 1800 videos drawn from YouTube-UGC and LSVQ collections, creating the largest such public dataset. A post-processed version applies a novel temporal filter to smooth the annotations frame by frame. The authors show the data's value by applying it to the SmartVidCrop method and fine-tuning video grounding models, while also comparing the labels to saliency predictions.","feed_headline":"New database labels 1800 videos for portrait cropping","feed_subtitle":"Annotations from 90 subjects on landscape videos create a resource for aspect ratio adaptation in mobile viewing.","key_machinery":"The LIVE-YT VC database of subjective portrait region annotations and the intra-frame temporal filter for smoothing annotations within each video.","core_discovery":"We introduce the LIVE-YT VC database featuring 1800 videos annotated by 90 human subjects for portrait region cropping, along with a temporally smoothed version called LIVE-YT VC++ using an intra-frame temporal filter, and demonstrate its usefulness for video aspect ratio transformation using SmartVidCrop and state-of-the-art models.","pith_inferences":["If the annotations reflect broad human judgments of significance, automated systems could use them to retain narrative elements when resizing videos for different platforms.","The temporal smoothing method could be tested on other types of subjective video labels to improve consistency.","Expanding the database to include more diverse video sources might test how well the annotations generalize beyond the original collection."],"forward_implications":["Models trained or fine-tuned on this dataset can better adapt videos to mobile display aspect ratios while preserving meaning.","The dataset provides a benchmark for future research on subjective video cropping and grounding.","Labels resembling saliency annotations allow exploration of similarities between cropping regions and saliency predictions.","Opening the data to the community will advance development of quality-preserving video reshaping techniques."],"fun_headline_variants":["1800 Videos in LIVE-YT VC Database for Portrait Cropping","Temporal Smoothing of Subjective Annotations for Video Cropping","LIVE-YT VC++ Applies Intra-Frame Filter to Annotations","Resource for Aspect Ratio Adaptation Using Cropping Labels"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The annotations provided by the 90 subjects accurately and reliably identify the significant regions that should be kept when cropping videos for different viewers and contexts.","fun_headline_variants_meta":{"raw":{"variants":["1800 Videos in LIVE-YT VC Database for Portrait Cropping","Temporal Smoothing of Subjective Annotations for Video Cropping","LIVE-YT VC++ Applies Intra-Frame Filter to Annotations","Resource for Aspect Ratio Adaptation Using Cropping Labels"]},"model":"grok-4.3","cost_usd":0.008074,"raw_usage":{"total_tokens":3627,"prompt_tokens":741,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":80740500,"prompt_tokens_details":{"text_tokens":741,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2820,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":741,"tokens_out":66,"duration_ms":18251,"temperature":1.0,"reasoning_tokens":2820,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-08T04:20:29.396192+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If independent human evaluators consistently prefer crops generated by models not using this dataset over those trained on it, when judging preservation of important content and visual quality.","supporting_citations":[],"review_version":1}