{"id":"95e1dcc8-dad3-4f16-92f7-4e7cb7720254","arxiv_id":"2508.13319","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":2.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A voice-controlled mobile surveillance robot built from Raspberry Pi, YOLOv3, and a Kinect sensor is shown to stream video and execute spoken navigation commands in indoor tests.","lead":"This paper builds a wheeled surveillance robot from two Raspberry Pi computers, a Kinect depth camera, and open-source software for object detection and voice control. It is a system demonstration that may interest hobbyists or educators rather than researchers seeking new algorithms.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim's 'interactive frame rates on CPU' and 'recognises commands reliably' are unquantified; with the full text unreadable, there is no verifiable evidence that YOLOv3 plus a Raspberry Pi 4 meets the throughput needed for the claimed autonomous action loop.","rationale":"The reader's verdict of UNVERDICTED is appropriate because the full text is corrupted and the abstract alone does not supply the evidence needed to evaluate the central claim. My stress-test agrees that the evidence is missing, but I locate the most load-bearing gap more specifically in the unquantified performance claim: 'interactive frame rates on CPU' combined with 'YOLOv3' and 'Raspberry Pi 4' is technically suspicious, and the absence of measured FPS, model variant, input resolution, and command accuracy prevents the claim from being checked. This is not a claim of fraud or even of falsehood; it is a claim that the central assertion is currently unsupported. The reader focused on environmental robustness of depth/detector; I focus on whether the stated hardware and software can plausibly meet the real-time requirement at all, which is a more fundamental precondition. Thus agreement is partial. Since both the reader's concern and mine reduce to the same conclusion (evidence is unavailable and the verdict must stay UNVERDICTED until the full text or artifacts are inspectable), no verdict change is warranted. The concrete test I propose would settle the concern either by recovering the needed numbers from the manuscript or by an independent CPU benchmark of the described YOLOv3 configuration.","tokens_in":6298,"tokens_out":3106,"duration_ms":35815,"concrete_test":"Obtain a readable version of the full manuscript or the authors' repository, then locate the evaluation section and extract: (a) YOLOv3 variant, input resolution, and CPU model; (b) measured detection FPS and total end-to-end latency; (c) number of command trials and recognition accuracy; (d) whether obstacle avoidance and command-to-action execution were tested in varied lighting/clutter conditions. If no such metrics exist, run an independent benchmark: deploy the described YOLOv3 configuration on a Raspberry Pi 4 with the Kinect stream and measure detection FPS over at least 100 frames. If the measured FPS is below the threshold needed for the navigation/action loop (e.g., below the command-response rate), the central claim 'without manual control' is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract asserts that 'in indoor tests the robot detects common objects at interactive frame rates on CPU, recognises commands reliably, and translates them to actions without manual control.' For this to hold, the perception stack must produce detections fast enough to close the loop between speech, obstacle avoidance, and motion commands. The manuscript does not report model variant, input resolution, achieved FPS, command accuracy, trial counts, or environment conditions, and the full text is corrupted to the point of being unreadable. The specific technical risk is that YOLOv3 at standard resolutions on a CPU-only Raspberry Pi 4 is widely reported to run at well below real-time rates; if the actual settings use a tiny model, a low resolution, or a faster machine than stated, the reproducibility claim in the abstract ('off-the-shelf hardware and open software') is materially weakened. The abstract also claims 'We discuss limits' but gives no actual limits, and no code repository or artifact is mentioned, so the 'easy to reproduce' claim cannot be checked. The load-bearing concern is therefore not an internal contradiction but an unsupported correctness and reproducibility claim: without quantitative evaluation, the central claim is unfalsifiable as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a mobile surveillance robot built from two Raspberry Pi 4 units on a differential-drive base, with a camera, microphone, speaker, and Kinect RGB-D sensor. The system streams video with FFmpeg, detects objects using YOLOv3, and accepts spoken commands through Python speech-recognition, translation, and text-to-speech libraries. The abstract claims that indoor tests demonstrate interactive frame rates on CPU, reliable command recognition, and translation of commands to actions without manual control, and it further claims that the design is easy to reproduce using off-the-shelf hardware and open software. The provided full text is heavily corrupted and unreadable, so none of the experimental details, system diagrams, or the promised discussion of limits can be verified.","tokens_in":6465,"tokens_out":3979,"duration_ms":42974,"significance":"If the claimed results hold and are properly quantified, the work would be a modest but useful integration demonstration: a low-cost, voice-controlled surveillance robot assembled from accessible components. I credit the authors for choosing widely available hardware and open-source software, properties that could make the system a useful reference for hobbyists or teaching labs. However, the contribution is empirical in nature, and the submitted manuscript provides no quantitative evidence for its central claims. There are no code artifacts, machine-checked proofs, or parameter-free derivations to offset the missing evaluation. As it stands, the paper cannot be assessed beyond the abstract, and the significance of the claimed system remains unverified.","major_comments":[{"comment":"The central claim, \"In indoor tests the robot detects common objects at interactive frame rates on CPU, recognises commands reliably, and translates them to actions without manual control,\" is unquantified. The paper reports no frame rate in frames per second, no detection accuracy or command-recognition accuracy, no trial counts, no latency measurements, and no description of the test environment. Because this is an empirical systems paper, these measurements are the load-bearing evidence and must be reported.","section":"Abstract"},{"comment":"The supplied full text is corrupted to the point of being unreadable: practically every line consists of garbled characters, so the system architecture, the evaluation protocol, the results, and the advertised discussion of limits cannot be inspected. The authors must resubmit a cleanly encoded version of the manuscript. Without this, the paper cannot be meaningfully reviewed.","section":"Full text"},{"comment":"The claim that the system is \"easy to reproduce\" because it relies on off-the-shelf hardware and open software is not supported by any accessible artifact or instruction. The abstract gives no code repository, no software versions, no wiring or assembly details, and no configuration parameters, and none of these are readable in the corrupted full text. A reproducibility claim without any of this information is unverifiable.","section":"Abstract"},{"comment":"The phrase \"interactive frame rates on CPU\" is not tied to any model variant or input resolution. Given the stated platform, a Raspberry Pi 4 running YOLOv3 on CPU, the qualitative wording does not establish that detection runs fast enough to support the closed loop implied by \"translates them to actions without manual control.\" The authors need to state the YOLO variant, input resolution, measured FPS, and, if applicable, the latency of the voice-command-to-action pipeline.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract promises that the authors \"discuss limits,\" but no specific limits are stated in the abstract and none can be read in the full text. The resubmission should name concrete operating constraints, such as lighting conditions, obstacle types, network latency, or command vocabulary size.","section":"Abstract"},{"comment":"The claim of \"multilingual translation\" is unsupported: no languages are named, and no evaluation of translation quality is mentioned. Please specify which languages were tested and how translation was assessed.","section":"Abstract"},{"comment":"The rendering problem appears to affect the entire document, including equations and figure captions if any exist. The authors should verify that the PDF and the source produce identical, readable text before resubmission.","section":"Full text"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unreadable in its current form, and the abstract alone does not provide the quantitative support needed for a systems paper. I suggest the editor ask the authors to resubmit a cleanly encoded version and to add a proper evaluation section before the paper is sent out for review again. The novelty appears limited, but that is secondary to the missing evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a straightforward integration of off-the-shelf parts—two Raspberry Pi 4 boards, YOLOv3, a Kinect, FFmpeg, and standard speech libraries—with no claim of a new algorithm, dataset, or measurement result. The abstract says the robot detects objects at \"interactive frame rates on CPU\" and \"recognises commands reliably,\" but gives no FPS, accuracy percentages, trial counts, or environment conditions. To make matters worse, the full text I received is corrupted beyond use, so I cannot check whether the actual paper has tables or details that support these claims.\n\nThe one thing the paper does well is present a clean system architecture with sensible component choices. The idea of using a central Pi for perception and a front Pi for the drive base is reasonable, and the listed extensions (ultrasonic ranging, GPU acceleration, face and text recognition) are practical next steps rather than moonshots. If the full text is intact in the arXiv version, it may serve as a useful build guide for hobbyists or students.\n\nThe soft spots are in the claims, not the design. There is no quantitative evaluation anywhere in the abstract. The stress-test point about YOLOv3 on a CPU-only Raspberry Pi 4 is worth taking seriously: at typical settings, that combination often runs at a few frames per second, which may not be enough for the closed-loop obstacle avoidance and speech-to-action loop described. It is possible the authors used a small model or low input resolution, but they do not say, and I cannot verify. The abstract also says \"We discuss limits\" without listing any, and no code repository or configuration files are mentioned, so the \"easy to reproduce\" claim is unverifiable as stated.\n\nNone of this amounts to an internal contradiction. The system likely works in a limited indoor setting, and the engineering appears coherent. But as a research paper it is unfalsifiable from the available text. This is the kind of work that belongs in a workshop or demo session, not a peer-reviewed archival venue, unless the authors add real measurements and ship the code.\n\nMy recommendation: desk reject as submitted. If the authors restore the full text and include basic performance numbers—FPS, command accuracy over N trials, latency—it becomes a reasonable workshop submission. For now, it is not referee material.","headline":"A clean off-the-shelf integration whose key performance claims are unquantified and unverifiable from the corrupted full text; workshop material at best, not a research paper.","tokens_in":7019,"tokens_out":3014,"would_cite":false,"duration_ms":28979,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports a mobile surveillance robot built from two Raspberry Pi 4 boards, a Kinect RGB-D sensor, and YOLOv3 object detection that, in indoor tests, detects common objects at interactive CPU frame rates and translates voice…","keywords":["surveillance robot","Raspberry Pi","YOLOv3","object detection","speech recognition","voice control","Kinect RGB-D"],"falsifier":"Run the same robot in a room with strong window light and a dark floor, or with low-reflectivity obstacles such as a glass coffee table, and count how often the Kinect depth map misses the obstacle or YOLOv3 fails to detect a person; frequent collisions or misdetections in these unreported conditions would falsify the claim of reliable no-manual-control operation.","tokens_in":6068,"feed_emoji":"🤖","tokens_out":5398,"duration_ms":45590,"temperature":0.7,"pith_summary":"This paper tries to establish that a useful mobile surveillance robot can be assembled from cheap off-the-shelf hardware and open software, with no GPU and no manual joystick. The robot carries two Raspberry Pi 4 computers: one on a two-wheeled base handles camera, microphone, and speaker, and the second streams live video with FFmpeg and runs perception and speech. A Kinect sensor supplies color and depth frames, and a pretrained YOLOv3 detector finds common objects to inform navigation and event awareness. The paper reports that in indoor tests the robot identifies everyday objects at interactive frame rates on CPU and turns spoken commands into actions reliably. If true, the design offers a reproducible, low-cost baseline for voice-driven surveillance robotics.","feed_headline":"Voice commands steer a low-cost surveillance robot","feed_subtitle":"Two Raspberry Pi boards, a Kinect, and YOLOv3 make a robot that detects objects and replies to speech.","key_machinery":"The load-bearing mechanism is the split-compute architecture: a front Raspberry Pi on the drive base handles the camera, microphone, speaker, and motors, while a second Raspberry Pi serves the video stream and runs perception and speech. The central perception object is YOLOv3, a pretrained convolutional object detector that localizes common objects in a single pass, running on CPU and informed by the Kinect RGB-D stream for depth-based obstacle cues. The speech pipeline, built from speech recognition, multilingual translation, and text-to-speech libraries, is what converts spoken commands into robot actions and replies.","core_discovery":"The central discovery is that a two-board Raspberry Pi architecture can keep the live video streaming path separate from the perception and speech path, and this division lets a CPU-only YOLOv3 detector run at interactive frame rates while a Kinect supplies depth cues for obstacle awareness. Voice commands are recognized, translated when needed, and mapped to drive actions, while the robot reads back responses in the requested language. In indoor tests, the combined system detects common objects, recognizes commands, and acts on them without manual control.","pith_inferences":["A natural next experiment, not reported in the paper, would measure mission success under changing daylight, cluttered floors, or Wi-Fi latency; those are exactly the conditions where a single Kinect depth stream and CPU-only YOLOv3 would be stressed.","The split-compute design suggests a scaling path: offloading YOLOv3 to a GPU-equipped server or central unit could make the robot's onboard unit lighter and cheaper while keeping the same voice and streaming interfaces.","Adding the ultrasonic range sensor that the paper lists as an extension could provide a collision-avoidance layer independent of the object detector, making 'without manual control' more robust than depth data alone allows.","If the speech-to-action path is direct command mapping, the same pipeline could be extended to event-triggered speech, such as announcing when YOLOv3 detects a person, without changing the core architecture."],"forward_implications":["If the indoor results hold, a voice-controlled surveillance robot no longer requires a GPU or commercial robot platform; the full pipeline fits on two Raspberry Pi 4 boards.","Object detections from YOLOv3 can serve simultaneously as navigation cues and as surveillance event awareness, letting one perception stream do double duty.","The Kinect alone can provide the depth information needed for obstacle cues in the tested indoor setting, removing the need for extra range sensors at the base level.","Because speech recognition is paired with multilingual translation, the same robot can accept commands and respond in several languages, which is directly useful for remote monitoring across language barriers.","A design built entirely from off-the-shelf hardware and open software gives other teams a reproducible starting point for adding sensors, faster models, or autonomous behaviors."],"supporting_citations":[],"fun_headline_variants":["Low-cost robot takes voice commands and streams video","Voice-steered surveillance bot built from two Pis and a Kinect","Speak to steer: budget robot uses CPU-only YOLOv3","Interactive surveillance: voice commands on a Raspberry Pi duo","Two Pis, a Kinect, and YOLOv3 give a robot ears and eyes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a pretrained YOLOv3 running on a Raspberry Pi CPU plus a single Kinect depth stream is sufficient for reliable indoor navigation and event awareness; if lighting, clutter, obstacle types, or network conditions defeat the depth data or the detector, the claim that commands are translated to actions without manual control no longer holds.","fun_headline_variants_meta":{"raw":{"variants":["Low-cost robot takes voice commands and streams video","Voice-steered surveillance bot built from two Pis and a Kinect","Speak to steer: budget robot uses CPU-only YOLOv3","Interactive surveillance: voice commands on a Raspberry Pi duo","Two Pis, a Kinect, and YOLOv3 give a robot ears and eyes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1434,"prompt_tokens":821,"completion_tokens":613,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":523}},"tokens_in":437,"tokens_out":613,"duration_ms":6331,"temperature":1.0,"reasoning_tokens":523,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:13:27.798896+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same robot in a room with strong window light and a dark floor, or with low-reflectivity obstacles such as a glass coffee table, and count how often the Kinect depth map misses the obstacle or YOLOv3 fails to detect a person; frequent collisions or misdetections in these unreported conditions would falsify the claim of reliable no-manual-control operation.","supporting_citations":[],"review_version":2}