{"id":"9df36e47-4794-4806-9f2b-d705c14970f8","arxiv_id":"2506.02676","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A wearable assistive system for blind users solved 95.7% of eight Cybathlon VIS tasks in its authors' training environment and won the qualification run, but hardware failures cost it in the final.","lead":"Sight Guide is a wearable camera-and-sensor system that helps blind people navigate and understand scenes by vibrating a belt and speaking instructions. The authors tested it on eight competition tasks and report a 95.7% success rate on their own course, then placed third at the Cybathlon 2024 Vision Assistance Race.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 95.7% headline is a device-only mean over seven selected tasks; recomputing actual task success from Table I gives about 90%, so the central quantitative claim needs re-scoping.","rationale":"Good-faith reading: this is a competition engineering report, and its success claim is scoped to the Cybathlon VIS training environment. The rectangular-boundary assumption that the reader flagged is real but is an acknowledged competition-specific constraint; the paper explicitly limits real-world applicability, so I do not see it as the primary threat to the central claim. The more direct threat is the mismatch between the 95.7% number quoted in the abstract/conclusion and what Table I actually supports. The paper does disclose both device and pilot rates, but the central claim selects the higher device-only rate and calls it 'task success.' That is not a mere presentation preference: it changes the headline from 95.7% to 90.0% when actual task completion is counted, and the Grocery sentence contradicts the same table. The lack of per-run logs is why this cannot be resolved by inspection. I therefore partially agree with the reader: their weakest_assumption (navigation geometry) is not the one I would put first; their rationale's mention of raw data and task exclusion is closer. The right resolution is to keep the paper CONDITIONAL, with conditions being a re-stated metric and release of per-run outcome logs, not to reject it. Hence verdict_should_be is UNCHANGED (still CONDITIONAL).","tokens_in":11711,"tokens_out":7056,"duration_ms":70767,"concrete_test":"Recompute the headline from Table I using the pilot-success column as the task success rate over the same eight tasks: (100+100+90+100+70+80+90)/7 = 90.0%; including Serving gives 91.3%. If this is the relevant 'task success rate,' the abstract's 95.7% is incorrect as stated. Separately, inspect the per-run logs for the single Grocery failure: if the device emitted the wrong shelf cell, the Grocery device success should be 90% and the text 'In all runs ... Grocery' is wrong; if the pilot selected the wrong cell, then the device success rate should not be 90% without a definition of what 'device success' means when the final action is wrong. Releasing a per-task, per-run table with device/pilot labels would settle both points.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-B and Table I define two numbers: 'Success Device' and 'Success Pilot.' The abstract and conclusion quote 95.7% as 'task success rate,' but 95.7% is the average of the device-success column over the seven non-Serving tasks. The pilot-success column, which is the actual task-completion rate, averages (100+100+90+100+70+80+90)/7 = 90.0% over those same tasks (91.3% if Serving is included). The claim as worded therefore overstates the probability that a task was actually completed. The Table's own text is internally inconsistent: Section III-B says 'In all runs, the device successfully solved ... Grocery,' then immediately says 'In one trial, the system identified the wrong target cell during the Grocery task,' while Table I lists Grocery device success as 90%. The eight-task average also excludes Footpath and Finder, two developed tasks that were not attempted in competition, so the headline cannot be read as a success rate over the full VIS task set. What is missing is a clear definition of 'task success' and per-run logs that map each of the 80 attempts to device, pilot, and overall outcomes.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Sight Guide, a wearable assistive system developed for the Vision Assistance Race (VIS) at Cybathlon 2024. The hardware combines chest-mounted stereo and depth cameras, a handheld RGB camera, an NVIDIA Jetson Orin NX, a vibration belt, and audio feedback. The software integrates VIO, volumetric mapping, task-boundary detection, A* planning, and task-specific modules for OCR, object detection, semantic mapping, and touchscreen interaction. The system was evaluated in the authors' training environment with ten runs per task across eight tasks, and the paper reports a 95.7% device success rate and a 91.3% pilot success rate. The authors also describe their Cybathlon results, including first place in qualification and third place overall, and conclude that the device achieved a 95.7% task success rate and that their results validate the approach.","tokens_in":11973,"tokens_out":6693,"duration_ms":66300,"significance":"If the quantitative claims are properly scoped, this is a useful systems contribution: it demonstrates an integrated wearable assistive platform operating with a blind pilot in a realistic competition, it separates device errors from pilot errors, which is good evaluation practice, and it records detailed lessons about depth sensing, feedback design, and real-world deployment. The paper does not provide code, data, or formal derivations, and the headline success rate is not the true task-completion rate. Nevertheless, the architecture description and the honest discussion of limitations provide a valuable reference for the assistive robotics and human-robot interaction communities.","major_comments":[{"comment":"The 95.7% figure quoted in the Abstract and Conclusion is the average of the 'Success Device' column over the seven non-Serving tasks, not the task success rate. The 'Success Pilot' column, which reflects actual task completion, averages 91.3% if Serving is included or 90.0% if it is excluded, with per-task values ranging from 70% to 100%. The claim '95.7% task success rate' therefore overstates the probability that a task was completed. Please define 'task success' explicitly, report per-run outcomes for all 80 attempts, and label the headline number as device-only success in the Abstract, Table I, and Conclusion.","section":"Section III-B, Table I, Abstract, Section V"},{"comment":"The prose contradicts Table I: the text states that 'In all runs, the device successfully solved Doorbell, Free Seats, Grocery, Sidewalk, and Touchscreen,' immediately adds that 'In one trial, the system identified the wrong target cell during the Grocery task,' and Table I reports Grocery device success as 90%. This internal inconsistency must be corrected, and the authors should indicate whether the Grocery failure was classified as a device error or a pilot error.","section":"Section III-B"},{"comment":"The evaluation uses ten runs per task without confidence intervals, and the criteria for distinguishing device success from pilot success are not defined. With ten runs, a single failure changes the reported rate by 10 percentage points; for example, Colours shows device success 90% and pilot success 70% (i.e., 7/10). Because the system was iteratively tuned in the same training environment, the reported rates are in-sample measurements. Please provide per-run logs, define the error-classification criteria, and explicitly state the in-sample nature of the headline result as a limitation.","section":"Section III-B, Table I"}],"minor_comments":[{"comment":"There are typographical errors: 'quantitavely' should be 'quantitatively,' and 'detailled,' 'taylored,' and 'Coulors' appear in the text.","section":"Section III-B"},{"comment":"Task names are inconsistent: the Introduction uses 'Free Chairs,' Section III-B and Table I use 'Seatfinder' and 'Free Seats,' and 'Dish Up' appears interchangeably with 'Serving' and 'Tablet' with 'Touchscreen.' Please unify the terminology.","section":"Introduction, Section III-B, Table I"},{"comment":"The navigation module assumes a known-size rectangular task boundary and projects Canny edges onto the ground plane; this competition-specific assumption is not listed among the limitations in Section IV-A. Please state explicitly in the limitations discussion that the boundary detection would need to be generalized for non-rectangular or non-planar environments.","section":"Section II-B1, Section IV-A"},{"comment":"The conclusion that 'our results in the Cybathlon 2024 ... validate the effectiveness' is stronger than the evidence, which consists of one successful qualification run and a final run cut short by hardware failure. Please temper this claim to reflect the qualitative nature of the competition data.","section":"Section III-C, Section V"},{"comment":"The 'All' column mixes definitions: for Device success it averages seven tasks, for Pilot success it includes eight tasks, and for Time it sums eight tasks. Please clarify which tasks are included in each aggregate and avoid comparing numbers computed over different task sets.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid systems description with a clear real-world evaluation context, but the central quantitative claim is currently overstated and internally inconsistent. The authors can fix this by re-scoping the headline number, adding per-run data, and clarifying the device/pilot error split. I see no citation or novelty concerns beyond the need to make the task-success definition explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a solid, honest engineering report. What is genuinely new is the integration: combining VIO, Wavemap, nvblox, YOLO, OCR, LoFTR, a vibration belt, and audio feedback into one wearable platform tuned for the Cybathlon VIS tasks. The design decisions are described in enough detail to be reimplemented, and the authors are unusually candid about hardware failures, pilot errors, and tasks they chose to skip. Separating device success from pilot success is good practice and should be standard.\n\nThat said, the headline number needs re-scoping. The abstract and conclusion say \"95.7% task success rate,\" but Table I shows that 95.7% is the average of the device-success column over seven tasks, excluding Serving. The pilot-success column, which is the actual task-completion rate, averages 90.0% over those same seven tasks (91.3% with Serving). So as written, the central claim overstates how often a task was actually completed. There is also an internal contradiction: Section III-B says the device solved Grocery in all runs, then says it identified the wrong target cell in one trial, while Table I lists Grocery device success at 90%. The paper should fix this, define \"task success\" explicitly, and report per-run outcomes or confidence intervals, especially since the evaluation is in-sample with only ten runs per task.\n\nThe exclusion of Footpath and Finder from the headline is disclosed, which is good, but it means 95.7% cannot be read as a success rate over the full VIS task set. The navigation stack also assumes a known rectangular ground-plane boundary, so the system is scoped to competition-like environments; the authors acknowledge this in the discussion. None of these issues are fatal, but they are real and should be corrected.\n\nThis paper is for robotics practitioners building assistive wearables or preparing for Cybathlon. It is not a new perception algorithm paper, and it does not claim to be. The value is in the integration blueprint, the evaluation methodology (with the caveats above), and the lessons learned.\n\nI would send this to peer review with a request to correct the headline and add the missing detail. The engineering is sound, the transparency is refreshing, and the community benefits from having this on record.","headline":"A transparent, useful engineering report on a Cybathlon assistive system, but the 95.7% headline is a device-only mean over seven selected tasks and should be rescoped before publication.","tokens_in":12507,"tokens_out":1769,"would_cite":false,"duration_ms":18870,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sight Guide is a wearable, multi-camera assistive system that guided a blind pilot through eight Cybathlon 2024 Vision Assistance Race tasks, achieving a 95.7% device success rate in training and a first-place qualification run.","keywords":["wearable assistive technology","vision assistance race","Cybathlon 2024","visual-inertial odometry","obstacle avoidance","semantic scene understanding","vibration feedback","RGB-D perception"],"falsifier":"Take the same navigation stack into a flat-floor environment whose boundary is not a rectangle of known size, for example an L-shaped or circular room, and remove any pre-programmed boundary dimensions. If boundary detection still returns a confident rectangle and the planner still guides the pilot successfully, the geometric assumption is not load-bearing; if it fails or guides into walls, that assumption is confirmed as the system's limit.","tokens_in":11525,"feed_emoji":"🦯","tokens_out":11119,"duration_ms":101850,"temperature":0.7,"pith_summary":"The paper presents Sight Guide, a wearable assistive system built for the Vision Assistance Race at Cybathlon 2024, and it aims to show that a single integrated device can guide a blind pilot through both obstacle-avoidance and scene-understanding tasks. The hardware is a backpack computer with chest- and hand-mounted RGB-depth cameras, a 16-motor vibration belt, and audio output; the software is a modular stack combining visual-inertial odometry, occupancy mapping, planning, and task-specific perception. In a training replica of the competition, the device completed 95.7% of the seven device-supported tasks across ten randomized runs, while the pilot completed 91.3%, and the best official qualification run solved seven of eight tasks. The paper also argues that the remaining gap to real-world use is not the perception approach itself but environmental assumptions, hardware robustness, and feedback precision.","feed_headline":"Wearable vision-belt system solves 95.7% of assistive tasks","feed_subtitle":"Backpack computer, chest cameras, and a 16-point vibration belt carried a blind pilot through Cybathlon 2024.","key_machinery":"The load-bearing mechanism is the coupling between a known-rectangle boundary model and a task-state machine. Boundary detection takes edge points from the left camera, projects them onto a fitted ground plane, and uses RANSAC to fit a rectangle of known dimensions, producing the four corners that anchor planning, finish-line detection, seat-row cropping, and shelf analysis. The state machine then switches between navigation and one scene-understanding module at a time, so the embedded computer never runs all networks simultaneously. The same loop, perceive, localize, command a heading, confirm by audio, repeats across all tasks: vibration for direction and speech for semantic instructions such as \"row 1, cell 3\" or \"left, position 2.\"","core_discovery":"The central claim is that the eight VIS tasks do not require a separate device per task: a wearable multi-camera rig with a vibration belt can carry a blind user through them end to end. The navigation stack estimates pose with visual-inertial odometry, builds a wavelet-compressed occupancy map, detects the task area by fitting a known-size rectangle to ground-plane edge projections, plans with A*, and steers the pilot through 1 Hz vibration commands. Around this core, task-specific modules use text recognition with fuzzy matching, object detection with grid fitting, semantic 3D mapping for free-seat counting, and homography-rectified finger tracking for the touchscreen, all coordinated by a state machine that loads one module at a time. The evidence is ten randomized runs in a competition replica, giving a 95.7% device success rate across seven scored tasks, and a Cybathlon qualification run in which seven of eight attempted tasks were completed within the time limit.","pith_inferences":["Beyond the paper: the known-rectangle assumption is the most testable single point of failure; replacing it with free-space segmentation or room-boundary reasoning would likely let the same belt-feedback loop work in offices and homes.","Beyond the paper: the depth-fusion strategy, where stereo reprojection fills pixels where time-of-flight depth drops out on black or thin objects, could be evaluated as a general sensor-fusion recipe on a benchmark of reflective and low-texture objects.","Beyond the paper: the interactive pattern of pointing a handheld camera and receiving an audio proximity cue, used in the Finder and Touchscreen tasks, could be generalized into a single \"what am I pointing at?\" query interface, especially if paired with a vision-language model.","Beyond the paper: the authors report that a second blind user reached comparable performance after 20 minutes of instruction; if replicated with more users, that would separate the system's usability from the specific pilot's long training."],"forward_implications":["If the training results are representative, an integrated wearable system can finish the full eight-task VIS battery in an average of 469 seconds, under the 480-second competition limit.","The 100% device-success runs on Doorbell, Free Seats, Sidewalk, and Tablet indicate that combining OCR with fuzzy matching, semantic mapping, and vibration-belt navigation is reliable in controlled settings.","Since pilot success (91.3%) trails device success (95.7%), interface-level mistakes, such as shifting a finger while lifting it from the touchscreen, are a comparable source of failure to perception errors, so better feedback could raise end-to-end performance without new sensors.","The paper's own limitation discussion implies that the same stack will not transfer as-is to arbitrary environments, because its scene modules assume predefined object classes and task geometry; generalizing would require replacing those assumptions rather than tuning the current modules."],"supporting_citations":[{"why":"Defines the Cybathlon VIS competition and the task set; the paper's system and evaluation are organized around these tasks.","marker":"[7]"},{"why":"Supplies the semidirect visual-odometry front end that estimates camera pose, the backbone of navigation, mapping, and seat analysis.","marker":"[9]"},{"why":"Supplies the keyframe-based nonlinear-optimization back end that refines the pose estimates used across the navigation stack.","marker":"[10]"},{"why":"Provides the wavelet-compressed occupancy map from which the 2D cost map and A* planner are built.","marker":"[11]"},{"why":"Provides the text-detection stage of the OCR pipeline used in the doorbell and grocery tasks.","marker":"[13]"},{"why":"Provides text recognition, converting detected text regions into words that are matched to known names or keywords.","marker":"[14]"},{"why":"Supplies the string-distance threshold that makes name and keyword matching tolerant to small OCR errors.","marker":"[15]"},{"why":"Provides the object-detection network used for seats, groceries, finder objects, and touchscreen finger and tablet detection.","marker":"[16]"},{"why":"Supplies the GPU-accelerated semantic 3D mapping used to accumulate chair, person, and backpack detections for free-seat classification.","marker":"[17]"},{"why":"Supplies the dense feature matcher that finds the target item on the cropped touchscreen image.","marker":"[19]"}],"fun_headline_variants":["Wearable vision belt hits 95.7% success in assistive tasks","Sight Guide wearable aids blind users through Cybathlon 2024","Multi-camera wearable navigates and identifies objects for blind","Vibration-guided wearable achieves 95.7% task success for blind","Cybathlon 2024: wearable system completes 95.7% of vision tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything that needs a spatial frame assumes the task area is a rectangle of known size lying in a flat ground plane, detected once from projected edge points; in a space without a visible rectangular boundary or with a non-planar floor, the navigation goal and the seat and shelf crops have no reliable anchor.","fun_headline_variants_meta":{"raw":{"variants":["Wearable vision belt hits 95.7% success in assistive tasks","Sight Guide wearable aids blind users through Cybathlon 2024","Multi-camera wearable navigates and identifies objects for blind","Vibration-guided wearable achieves 95.7% task success for blind","Cybathlon 2024: wearable system completes 95.7% of vision tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000614,"raw_usage":{"total_tokens":2848,"prompt_tokens":933,"completion_tokens":1915,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":1816}},"tokens_in":549,"tokens_out":1915,"duration_ms":12420,"temperature":1.0,"reasoning_tokens":1816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:19:13.359474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same navigation stack into a flat-floor environment whose boundary is not a rectangle of known size, for example an L-shaped or circular room, and remove any pre-programmed boundary dimensions. If boundary detection still returns a confident rectangle and the planner still guides the pilot successfully, the geometric assumption is not load-bearing; if it fails or guides into walls, that assumption is confirmed as the system's limit.","supporting_citations":[{"cited_title":"Cybathlon 2024 the third edition: What’s new and what’s dif- ferent?[competitions],","cited_arxiv_id":null,"evidence_quote":"Defines the Cybathlon VIS competition and the task set; the paper's system and evaluation are organized around these tasks."},{"cited_title":"SVO: Semidirect visual odometry for monocular and multicamera systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the semidirect visual-odometry front end that estimates camera pose, the backbone of navigation, mapping, and seat analysis."},{"cited_title":"Keyframe-based visual–inertial odometry using nonlinear optimiza- tion,","cited_arxiv_id":null,"evidence_quote":"Supplies the keyframe-based nonlinear-optimization back end that refines the pose estimates used across the navigation stack."},{"cited_title":"Efficient volumetric mapping of multi-scale environments using wavelet-based compression,","cited_arxiv_id":null,"evidence_quote":"Provides the wavelet-compressed occupancy map from which the 2D cost map and A* planner are built."},{"cited_title":"Vision transformer for fast and efficient scene text recogni- tion,","cited_arxiv_id":null,"evidence_quote":"Provides text recognition, converting detected text regions into words that are matched to known names or keywords."},{"cited_title":"Binary codes capable of correcting deletions, insertions, and reversals,","cited_arxiv_id":null,"evidence_quote":"Supplies the string-distance threshold that makes name and keyword matching tolerant to small OCR errors."},{"cited_title":"nvblox: Gpu-accelerated incremental signed distance field mapping,","cited_arxiv_id":null,"evidence_quote":"Supplies the GPU-accelerated semantic 3D mapping used to accumulate chair, person, and backpack detections for free-seat classification."},{"cited_title":"Efficient loftr: Semi- dense local feature matching with sparse-like speed,","cited_arxiv_id":null,"evidence_quote":"Supplies the dense feature matcher that finds the target item on the cropped touchscreen image."}],"review_version":1}