{"id":"21f648e3-d77a-4da3-bafe-66d086fe264e","arxiv_id":"2507.21837","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"VeasyGuide automatically highlights instructor actions in presentation videos, raising low-vision learners' action detection from 61% to 88% and reducing reported cognitive load.","lead":"VeasyGuide is a tool that detects an instructor's pointing, marking, and sketching in presentation videos and highlights or magnifies those actions to help low-vision learners see them. In a study with 8 low-vision viewers, detecting such actions rose from 61% to 88% with the tool, and viewers reported less mental effort.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task 1 split the four videos between baseline and VeasyGuide conditions, so the unpaired Welch comparison conflates the highlight effect with per-video difficulty; the 61% to 88% gain is not yet attributable to VeasyGuide.","rationale":"I read the paper as claiming that VeasyGuide causally improves low-vision learners' detection of instructor actions on top of whatever content appears in the videos. For that causal claim, the most load-bearing assumption is not that the hand-tuned detection thresholds generalize to new videos, but that the measured 61% to 88% difference is actually caused by the highlight intervention. The reported experiment makes this uncertain because Task 1 assigned different videos to the two conditions: Table 7's per-participant activity totals are unequal between baseline and VeasyGuide (e.g., 11 vs 18), so the condition effect is entangled with video identity and difficulty. The paper's unpaired analysis does not adjust for this, and no video-stratified result is reported. I do not think this is fatal: a rough paired comparison of the Table 7 participant-level rates still shows a significant difference, which is evidence that the effect is not solely a participant-ability artifact. The correct response is to require a reanalysis that controls for video identity, not to reject the paper outright. The reader's conditional verdict is therefore the right one, but for a more pointed reason than the detection-threshold generalization concern. The co-design process, the qualitative reports, and the sighted-user results remain valuable independent support; the unresolved piece is the quantitative attribution of the main success-rate gain to VeasyGuide rather than to the particular videos used in each condition. My concrete test would settle this with data the authors already possess.","tokens_in":26121,"tokens_out":13977,"duration_ms":169691,"concrete_test":"Request the raw per-trial localization data and re-run the analysis with condition as a fixed effect and participant and video as random effects in a mixed-effects logistic regression. As a minimal check, compute the baseline-versus-VeasyGuide success difference separately within each of V1–V4, since each video was seen by different subsets of participants in each condition, and test for a consistent within-video direction. If V4 drives the effect, or if the adjusted effect shrinks below significance (p >= 0.05 or d < 0.8), the headline claim is not supported; if the within-video differences are consistently positive and the mixed-model effect remains significant with a similar magnitude, the video-confound concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest quantitative claim (RQ1) rests on comparing per-participant success rates between baseline and VeasyGuide conditions. Table 7 shows that the number of cued activities differs across conditions for the same participant (e.g., L1: 11 vs 18; L7: 18 vs 11), confirming that each participant did not view the same videos in both conditions; V1–V4 were divided, with one subset in baseline and the complementary subset in VeasyGuide. The paper acknowledges this by performing an unpaired Welch test 'since different videos are used.' This makes the condition contrast a between-video comparison. Appendix E shows large per-video difficulty differences (V4 had <25% baseline success), so if harder or easier videos were unevenly placed across conditions, the 27-point mean difference could be inflated or entirely produced by video content rather than by highlighting. Counterbalancing mitigates but does not eliminate this risk in a sample of 8 participants and 4 videos, and no stratified per-video analysis, mixed-effects model with video as a random effect, or raw trial data is reported. A paired by-participant comparison of the numbers in Table 7 is suggestive (mean difference roughly 27 points, p approximately 0.03 by paired t-test), so the effect may survive, but the paper's reported analysis does not establish that the effect is attributable to VeasyGuide rather than to which videos appeared in each arm. The central 88% claim is therefore not yet secured.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VeasyGuide, a web-based tool that uses motion detection to identify instructor actions (pointing, marking, sketching) in educational presentation videos and augments playback with personalized visual highlights and magnification. The system was developed through a co-design study with three low-vision (LV) participants. The evaluation involved 8 LV and 8 sighted participants in two tasks: a localization task (Task 1, videos V1–V4) measuring detection success and response time, and a viewing task (Task 2, videos V5–V6) measuring cognitive load and experience. The paper reports that VeasyGuide significantly improved LV users' detection of visual activities (61% to 88%), reduced NASA-TLX workload across all dimensions, and received positive feedback from sighted users, while detection time differences were not statistically significant. The main quantitative claim (RQ1) is based on an unpaired comparison of participants' success rates between conditions that used different video subsets, a design confound that the paper acknowledges but does not resolve.","tokens_in":26450,"tokens_out":6453,"duration_ms":76443,"significance":"If the reported effects hold, VeasyGuide addresses a real and under-explored accessibility barrier: LV learners missing instructor actions in slide-based videos. The co-design process and the derived design implications (DI1–DI4) are valuable and actionable, and the public web application makes the tool usable beyond the paper. The paper is honest about several limitations (Section 7), including the small sample and the lack of a systematic personalization evaluation in Task 1. The per-participant data in Table 7 are detailed and enable reanalysis, which is a strength. However, the headline success-rate claim currently rests on a between-video comparison, and the paper overstates the response-time finding, so the central contribution is not yet fully secured.","major_comments":[{"comment":"The RQ1 success-rate analysis uses an unpaired Welch t-test on data that are paired at the participant level, but the baseline and VeasyGuide conditions used complementary subsets of videos V1–V4. Table 7 shows that the number of cued activities per condition differs within participants (e.g., L1: 11 vs 18; L7: 18 vs 11), and Appendix E reveals large per-video difficulty differences (V4 baseline success <25%). Consequently, the 61% vs 88% mean difference may be inflated or produced by which videos happened to appear in each condition rather than by the highlighting itself. Please reanalyze with a mixed-effects model that includes participant and video as random effects, report a stratified per-video comparison, or at minimum provide the paired per-participant difference and its confidence interval using the data in Table 7, and discuss how video difficulty is accounted for.","section":"§5.1.4, Table 7, Appendix E"},{"comment":"The abstract and Section 1 claim that VeasyGuide yields 'faster response times' and 'improves the speed of detection,' but Section 5.2.2 states that the Mann-Whitney U test found no statistically significant difference (p > 0.05). Presenting the non-significant mean/median speed improvements as a demonstrated benefit overstates the evidence. Please either soften the speed claim to a non-significant trend or report a test that properly accounts for the paired design and video-level variation.","section":"Abstract, §1, §5.2.2"},{"comment":"The activity recognition pipeline uses six hand-chosen thresholds (§4.1.1: RoC area 0.01% of frame, temporal closeness 3s, spatial closeness 5% of frame diagonal, Hu-moment merge threshold 0.5, minimum activity duration 5s, and 1.5s pre-trigger), but the paper reports no precision/recall evaluation of the detector on the six study videos. The user-success outcome in the VeasyGuide condition depends on the pipeline actually highlighting the cued activities; if recall is imperfect or the thresholds are implicitly tuned to the study videos, the 88% success rate is not a general property of the system. Please add a detection-accuracy evaluation (per-video recall/precision against the gold-standard activities) or a sensitivity analysis of the thresholds, and discuss the generalizability of these parameters to other presentation video styles.","section":"§4.1, §5.2.1"}],"minor_comments":[{"comment":"The statement that 'detection success gap reduced by 80.5%, and detection speed gap reduced by 27.3%' is not derived from the results in Section 5; please specify the formula (e.g., based on means or medians) or remove these numbers.","section":"§6.1.4"},{"comment":"Please clarify how the total number of cued activities per condition in Table 7 was determined, including how V1–V4 were assigned to conditions for each participant and how activities shorter than one second were treated in the totals.","section":"§5.1.3, Table 7"},{"comment":"Two of the eight evaluation participants (L3 and L7) also participated in the co-design study; please acknowledge this overlap as a potential familiarity bias in the evaluation.","section":"§5.1.1, Table 5"},{"comment":"The horizontal axis of Figure 5 appears to be logarithmic; please label it as such or use a linear scale with clear units.","section":"Figure 5"},{"comment":"The 1.5-second pre-activity highlight trigger is an important design element for predictability (DI3), but its isolated effect is not evaluated; please mention this explicitly in the limitations or future work.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of ASSETS and the overall contribution is promising, but the main statistical analysis needs to be fixed before I can support acceptance. The unpaired test on paired data is the central issue; if a reanalysis with video as a random effect still shows a significant benefit, the paper would be much stronger. The overlap between co-design and evaluation participants is another concern to watch. No citation issues detected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is good: use motion detection to highlight and magnify instructor actions in presentation videos for low-vision learners, co-designed with the population it serves. The paper is honest about its lineage—the recognition pipeline is close to Ram et al., and the novelty lies in the LV co-design, the personalization layer, and the evaluation. That framing is fair, and the qualitative material is rich. The design implications (familiarity, predictability, personalization, agency) are clearly grounded in what participants said, not bolted on after the fact.\n\nThe weak spot is the RQ1 analysis, and the stress-test note lands. Task 1 splits four videos between baseline and VeasyGuide, then compares unpaired per-participant success rates with a Welch t-test. Appendix E shows big per-video difficulty differences—V4 was brutal in baseline. So the 61% vs. 88% gap is confounded with which videos landed in each arm. The paper acknowledges the unpaired test, but that does not fix the attribution problem. A paired by-participant comparison of Table 7 suggests the effect may survive, but that is not the reported analysis. The fix is straightforward: report a mixed-effects model with video as a random effect, or at least a per-video stratified breakdown. Without that, the headline claim is not yet secured.\n\nThe secondary overclaim is the speed benefit. The Mann-Whitney U finds no significant difference, yet the abstract and intro present faster detection as a result. That should be softened. The cognitive-load reductions in Task 2 are paired and do look solid, so the paper's qualitative and secondary results carry more weight than the main success-rate claim.\n\nNo code or data are shipped, which is a limitation for a tool paper, but not a fatal one. The sample sizes are small, as the paper itself notes.\n\nVerdict: this deserves a serious referee and likely acceptance after revision. The methodological gap is addressable, and the co-design contribution is genuine. I would send it to review, but I would insist the authors reanalyze Task 1 properly. For my own work, I would not cite the 88% number yet, but I would point to the design process and implications.","headline":"A thoughtful cdesigned accessibility tool with real promise, but the headline detection gain rests on a between-video comparison that the paper's own analysis does not yet secure.","tokens_in":26947,"tokens_out":2464,"would_cite":false,"duration_ms":35118,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that VeasyGuide, a real-time motion-detection overlay that highlights and magnifies instructor pointing, marking, and sketching, lifts low-vision learners' action detection from 61% to 88% and cuts self-reported cognitive…","keywords":["low vision","presentation videos","video accessibility","visual accessibility","motion detection","co-design","magnification","universal design"],"falsifier":"Annotate a held-out set of presentation videos with the true locations of pointing, marking, and sketching actions, run VeasyGuide's pipeline on them, and compare the overlap of its activity boxes with the annotations; if low-vision users' success rate on those videos fails to beat the 61% baseline by a similar margin, the generalizability claim is not supported.","tokens_in":25952,"feed_emoji":"🔍","tokens_out":7713,"duration_ms":91478,"temperature":0.7,"pith_summary":"VeasyGuide addresses a specific gap in educational presentation videos: instructors communicate by pointing, marking, and sketching on slides, but these visual actions are rarely described verbally, so low-vision learners must hunt for them or miss them. The paper claims that a real-time system which detects these actions from motion and renders a consistent, personalized highlight plus an optional auto-following zoom lets low-vision learners find what they are looking for, raising mean detection success from 61% to 88% and reducing self-reported cognitive load on every NASA-TLX dimension. The design was shaped by a co-design study with three low-vision participants and tested with eight low-vision and eight sighted viewers. The authors' position is that enhancing residual vision with guidance beats substituting for it with audio description, and that the same guidance can help sighted viewers stay attentive.","feed_headline":"Low-vision learners spot 88% of instructor actions with VeasyGuide","feed_subtitle":"A motion-detection overlay lifts detection from 61% to 88% and cuts cognitive load.","key_machinery":"The load-bearing mechanism is a two-stage activity recognition pipeline followed by a visualization module. The pipeline splits the video into shots with off-the-shelf shot detection, computes frame differences on one-third-second segments, keeps regions of change whose area exceeds $0.01\\%$ of the frame, and builds a graph whose nodes are those regions, connected when they occur within 3 seconds and within 5% of the frame diagonal of each other. Edge weights come from Hu moments, seven shape-descriptor numbers that measure visual difference between regions, which lets the system merge or reject transient pointer trails. Connected components become activities, each rendered as a stationary, box-shaped, personalized highlight with a pointer icon and an optional auto-following zoom toggled with the Z key; a 1.5-second pre-activity trigger gives learners advance notice of where to look. The graph representation is what turns noisy per-frame motion into discrete, stable activities that can be highlighted and magnified.","core_discovery":"The central claim is that the visual-search bottleneck for low-vision learners is not just the size of the content but knowing what to look for and where, and that an overlay can supply that knowledge in real time. On the paper's evidence, detection success rose from 61% to 88% (Welch's t-test $p<0.05$, Cohen's $d=1.30$), mean search time fell from 2.57s to 1.53s without reaching significance, and all NASA-TLX workload dimensions dropped significantly. In the guided condition no low-vision participant paused playback and fewer used screen magnification, suggesting the tool replaces part of the stop-and-search behavior of baseline viewing. The paper also reports that sighted viewers, already at ceiling on objective measures, subjectively preferred VeasyGuide for focus and comprehension, so the authors propose the mechanism as a broadly useful attention aid rather than only an accessibility tool.","pith_inferences":["The localization task fixed highlight style to system defaults, so the 88% figure is a default-settings result; testing with participants' own personalized styles might change detection rates in either direction and would separate the effect of personalization from the effect of highlighting per se.","The pre-activity trigger likely explains part of the benefit; an ablation that removes the 1.5-second warning would isolate how much of the gain comes from anticipation versus visibility.","The graph representation is blind to semantic content, so the same machinery could be pointed at other moving targets, such as cursors in coding screencasts or moving regions in sports and remote collaboration, by retuning only the thresholds.","Since sighted viewers reported less clutter and better focus, the system may be a test bed for attention-as-accessibility design, where the same cue serves both perceptual and attentional functions."],"forward_implications":["A learner who misses roughly four in ten instructor actions could miss roughly one in ten with the default settings, a per-participant mean improvement of 74.5%.","The tool changes search strategy: in the guided condition no low-vision participant paused playback during the localization task, and screen-magnifier reliance dropped from six users to two.","The design implications—consistent familiar visuals, predictable spatial context including a 1.5-second pre-trigger, real-time personalization with immediate feedback, and preserving user agency—can guide future accessibility tools for visual search in video.","Because the pipeline uses lightweight motion detection and graph operations, the paper expects it to run on-device and to extend to live lectures and mainstream video platforms."],"supporting_citations":[{"why":"supplies the motion-detection technique the activity pipeline is built on.","marker":"[56]"},{"why":"provides the graph-based activity representation that the pipeline adapts from blackboard-style videos.","marker":"[63]"},{"why":"shot detection that splits the video so slide transitions do not generate false activities.","marker":"[14]"},{"why":"Hu-moment shape descriptors used to decide which motion regions merge into one node.","marker":"[31]"},{"why":"identifies pointer trails as transient motion noise the merging step is meant to suppress.","marker":"[17]"},{"why":"the earlier highlight-based accessibility approach that the base prototype extends.","marker":"[66]"},{"why":"prior saliency-based magnification for low-vision video viewers that motivates the auto-zoom design.","marker":"[4]"},{"why":"prior presentation-video accessibility system whose non-visual exploration findings frame why instructor actions are under-served.","marker":"[58]"},{"why":"evidence that visual cues aid low-vision visual search, the core mechanism being evaluated.","marker":"[87]"},{"why":"shows video accessibility preferences depend on context, justifying the personalization requirements.","marker":"[37]"}],"fun_headline_variants":["VeasyGuide raises action detection to 88% for low-vision","Overlay lifts low-vision action detection to 88%","VeasyGuide: 88% detection, lower cognitive load","From 61% to 88%: VeasyGuide sharpens action detection","VeasyGuide reduces search effort and lifts detection to 88%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fixed detection settings—how much motion counts as a region, how close regions must be in space and time to merge, and how short an activity can be—work on presentation videos beyond the six used in the study.","fun_headline_variants_meta":{"raw":{"variants":["VeasyGuide raises action detection to 88% for low-vision","Overlay lifts low-vision action detection to 88%","VeasyGuide: 88% detection, lower cognitive load","From 61% to 88%: VeasyGuide sharpens action detection","VeasyGuide reduces search effort and lifts detection to 88%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001313,"raw_usage":{"total_tokens":5336,"prompt_tokens":915,"completion_tokens":4421,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":4328}},"tokens_in":531,"tokens_out":4421,"duration_ms":36212,"temperature":1.0,"reasoning_tokens":4328,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:17:33.558349+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate a held-out set of presentation videos with the true locations of pointing, marking, and sketching actions, run VeasyGuide's pipeline on them, and compare the overlap of its activity boxes with the annotations; if low-vision users' success rate on those videos fails to beat the 61% baseline by a similar margin, the generalizability claim is not supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the motion-detection technique the activity pipeline is built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the graph-based activity representation that the pipeline adapts from blackboard-style videos."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Hu-moment shape descriptors used to decide which motion regions merge into one node."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the earlier highlight-based accessibility approach that the base prototype extends."},{"cited_title":"Touhidul Islam and Syed Masum Billah","cited_arxiv_id":null,"evidence_quote":"shows video accessibility preferences depend on context, justifying the personalization requirements."}],"review_version":1}