{"id":"b03569d6-d9f7-4aec-b682-ad3bf91a6f5f","arxiv_id":"2501.15727","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Gensors lets everyday users define personalized visual sensors by decomposing their sensing goal into testable criteria, and a user study shows improved perceived control and understanding over prompt-only authoring.","lead":"This paper introduces Gensors, a system that lets people create custom camera sensors by describing a task in plain language and refining simple yes/no criteria. A 12-person study found that users felt more control and understanding with this approach than with editing a single prompt.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The baseline condition's model is unspecified: Section 4.3 details Gensors' two-stage Flash+Pro pipeline, but Section 5 never states which MLLM and temperature evaluated the baseline single prompt, so the claimed benefits could stem from model capability rather than criteria-based authoring.","rationale":"The reader identified sample generalizability as the weakest assumption, but I consider a more load-bearing threat to be internal validity: the baseline condition's model configuration is never specified, creating a potential confound between the interface (criteria vs. prompt) and the underlying MLLM pipeline. If the baseline used a weaker model than Gensors' Flash+Pro combination, the observed differences in perceived control, understanding, and communication could be attributed to model capability rather than to the criteria-based authoring approach. This concern is concrete and testable, and it is not covered by the paper's stated limitations. The paper's counterbalancing and within-subjects design mitigate order effects, and the non-significant 'Test' metric suggests participants did not rate everything higher, but neither addresses the model mismatch. I still agree with the overall CONDITIONAL verdict: the paper should be accepted only after the authors clarify the baseline implementation or release code/prompts. Since the verdict does not change, I mark it UNCHANGED, but the condition is more specific than the reader's generalizability concern.","tokens_in":26522,"tokens_out":4743,"duration_ms":45856,"concrete_test":"Inspect the study implementation or supplementary material to identify the exact Gemini model, temperature, and prompt template used for the baseline condition. Then run a matched-model ablation: for each participant's final baseline prompt, evaluate it through the same two-stage Flash+Pro pipeline used by Gensors (or, conversely, run Gensors' criteria prompts through the baseline's single-model configuration), and re-measure the four questionnaire ratings. If the significant differences in control, understanding, and communication disappear or shrink substantially under matched models, the central claim must be qualified as an effect of the model pipeline rather than the criteria interface alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Gensors improves perceived control, understanding, and communication rests on the comparison between Gensors and a prompt-editor baseline. Section 4.3 specifies that Gensors runs each criterion with Gemini 1.5 Flash at temperature 0 and the final verdict with Gemini 1.5 Pro at temperature 0.8. Section 5 describes the baseline only as 'a prompt-editor version, similar to the one used in the formative study,' without specifying which Gemini model, temperature, or inference configuration processed the baseline prompt. If the baseline prompt was evaluated by Gemini 1.5 Flash alone (or by a less capable model), then participants' lower ratings in the baseline could reflect the model's poorer adherence to complex, multi-clause prompts rather than the lack of criteria decomposition. This would confound the interface manipulation with model capability, undermining the internal validity of the significant differences reported in Section 6.1. The paper's own limitation section (Section 8) addresses sample generalizability but does not mention this potential model mismatch. The non-significant 'Test' metric does not rule out a model confound, since control, understanding, and communication ratings may be more sensitive to the quality of the model's explanations than to the interface alone. Without clarifying the baseline implementation or making code/prompts available, the central attribution of the effect to Gensors' criteria-based authoring is not fully established.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Gensors, a system that lets end-users create personalized AI-powered visual sensors by decomposing high-level sensing tasks into individual criteria, with support for automatically generated criteria, example-based criteria discovery, test-case suggestions, and configurable verdict logic. The authors report a formative study (n=6) that motivates four design goals, then present a within-subjects user study (n=12) comparing Gensors to a prompt-editor baseline. They find significantly higher self-reported ratings for control, understanding of model capabilities, and ability to communicate requirements, while the testing metric did not reach significance. Qualitative analysis describes how criteria-level debugging, auto-generated criteria, and example-diff features helped users handle blind spots and MLLM idiosyncrasies.","tokens_in":26795,"tokens_out":4496,"duration_ms":42952,"significance":"If valid, the central comparative result is a useful contribution to end-user authoring of intelligent sensors: it provides evidence that criteria-based decomposition, rather than single-prompt iteration, can improve perceived control and understanding for non-expert users. The paper is transparent in reporting that the Test metric was not significant, applies full Bonferroni correction, and grounds its claims in both quantitative and qualitative data. The main threat to the central claim is that the baseline condition's underlying model configuration is not specified, leaving open a potential confound between the interface manipulation and model capability. This issue is load-bearing and needs to be addressed before the central claim is fully established.","major_comments":[{"comment":"The baseline condition's MLLM configuration is unspecified. Section 4.3 states that Gensors runs individual criteria with Gemini 1.5 Flash at temperature 0 and the final verdict with Gemini 1.5 Pro at temperature 0.8, while Section 5 describes the baseline only as 'a prompt-editor version, similar to the one used in the formative study' and Figure 3 does not state which model or temperature evaluated the baseline prompt. If the baseline prompt was evaluated by Gemini 1.5 Flash alone, or by a different model than the one used for Gensors' final verdict, the significant differences in Control, Understand Capabilities, and Communicate Requirements reported in Section 6.1 could be explained by the stronger reasoning and explanation quality of Gemini 1.5 Pro rather than by the criteria-based authoring workflow. The paper must specify the exact model, temperature, and inference pipeline used in the baseline, and ideally demonstrate that both conditions used comparable model capability. Section 8's limitation list does not mention this potential confound.","section":"Sections 4.3, 5.3, and Figure 3"},{"comment":"The self-reported Test metric did not reach statistical significance (Section 6.1 reports no p-value, just 'not statistically significant'), yet Section 6.3 makes strong qualitative claims that Gensors enabled 'robustly test and debug' and that users 'could robustly test and debug with Gensors.' Given that the quantitative evidence for testing is weak, the qualitative claims should be tempered or accompanied by effect sizes and exact p-values for all four metrics so that readers can assess the strength of the evidence. The non-significant result also does not rule out the model-capability confound raised above, since perceived control and understanding may be more sensitive to the quality of the model's explanations than to the interface alone.","section":"Sections 6.1 and 6.3"}],"minor_comments":[{"comment":"There is a duplicated word: 'feeling feeling that her efforts were counterproductive' should read 'feeling that her efforts were counterproductive.'","section":"Section 6.2"},{"comment":"There is a duplicated phrase: 'changed this “and” to “or”, which then caused the sensor to briefly classify the desk as cluttered, though not consistently so' reads fine, but elsewhere the text has 'described as 'chaotic and disorganized ” to be triggered...' with a stray space before the quotation mark.","section":"Section 6.3.1"},{"comment":"The phrase 'assessed the the sensor's overall effectiveness' contains a duplicated 'the.'","section":"Section 6.3.2"},{"comment":"The figure labels are dense and small; consider enlarging the UI screenshots and using higher-contrast labels for the numeric callouts, as some callouts (e.g., k1–k3) are difficult to distinguish in the printed version.","section":"Figure 2"},{"comment":"The study procedure says participants spent 50 minutes creating two sensors with 25 minutes per condition, but the earlier description says a 90-minute total commitment; consider clarifying how the 30-minute tutorial, 50-minute building, questionnaire, and interview fit into 90 minutes.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: Gensors is a well-made HCI systems paper with a genuinely new mechanism, and the user study is decent. The main risk is that the baseline condition is under-specified, so you can't fully rule out model capability as the cause of the effect. It deserves a serious referee, but it needs a revision that tightens the comparison and the claims.\n\nWhat's actually new: making sensor criteria first-class entities—auto-generate criteria from the task and camera view, isolate and test each one in parallel, add image examples per criterion, generate criteria from positive/negative examples, and get suggested test cases. That combination is not in ConstitutionMaker, EvalLM, or Selenite. The design goals from the formative study are reasonable and the system genuinely follows them. The user study is competent: 12 participants, counterbalanced, within-subjects, Wilcoxon with Bonferroni, and three significant effects on control, understanding, and communication. The qualitative section gives plausible mechanisms for those effects, like isolating a criterion to see the model's reasoning. They also honestly report that the 'Test' metric did not reach significance.\n\nSoft spots, in order of seriousness. First, the baseline model is unspecified. Section 4.3 says Gensors runs criteria on Gemini 1.5 Flash (temp 0) and the final verdict on Gemini 1.5 Pro (temp 0.8). But the prompt-editor baseline in Section 5 is only described as 'similar to the one used in the formative study'; nowhere does the paper say which Gemini model and temperature evaluated the single prompt. If the baseline used Flash alone, the lower ratings could come from model capability, not the interface. That's a genuine internal-validity gap. It may be harmless—Gensors also mostly uses Flash for criteria—but the paper can't rule it out without stating the baseline configuration or running a same-model control. This is fixable in revision. Second, the conclusion says Gensors leads to 'more robust' sensors, but the Test metric was non-significant, so 'robust' overstates the data. The results section is transparent; the conclusion just goes a bit further than the evidence. Third, the usual small-tech-savvy-sample and self-report limitations, which the paper itself acknowledges. No code or data is released, which makes the baseline issue harder to check.\n\nBottom line: this deserves a serious referee. The core idea is useful, the study is mostly clean, and the baseline confound is reportable and addressable. I'd accept it with a request to specify the baseline model/settings and soften the robustness claim.","headline":"Worth a serious referee: Gensors' criteria-decomposition workflow is a real step forward for end-user MLLM sensors, but the baseline model is under-specified and the conclusion overclaims robustness.","tokens_in":27328,"tokens_out":3678,"would_cite":false,"duration_ms":32191,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users feel more control, understanding, and clarity defining camera-based AI sensors when they work from testable criteria than when they iterate on a single prompt.","keywords":["human-AI interaction","multimodal large language models","intelligent sensing","criteria-based authoring","end-user programming","prompt debugging","visual sensors","personalization"],"falsifier":"Run a pre-registered study with a larger, less technically oriented sample in which users must make a sensor meet five fixed behavioral specifications using either Gensors or a single-prompt baseline, then measure the fraction of specification checks passed; the central claim would not hold if the criteria interface fails to outperform the baseline on objective correctness or if its control and clarity advantages disappear once age, technical background, and prior experience are controlled.","tokens_in":1410,"feed_emoji":"📹","tokens_out":1971,"duration_ms":77146,"temperature":0.7,"pith_summary":"This paper argues that everyday users can define personalized visual sensors more effectively when the sensing task is broken into individual, testable criteria rather than refined as one large prompt. It presents Gensors, a system that lets users generate, edit, isolate, and test these criteria against live camera frames, and reports a 12-person study in which participants rated Gensors significantly higher than a single-prompt baseline on sense of control, understanding of the model, and ability to communicate their requirements. The paper also shows that MLLM quirks, such as hallucinations and response flickering, shape the sensor-creation process and can even serve as debugging signals. If the claim is right, the path to broader customization of AI sensing lies in requirement scaffolding more than in better prompt phrasing.","feed_headline":"Criteria beat prompts for building AI camera sensors","feed_subtitle":"Criteria break a sensing prompt into testable checks; users report more control, understanding, and clarity.","key_machinery":"The central object is the criterion, an atomic unit of reasoning: a minimal natural-language question targeting one aspect of the sensing task, evaluated by an MLLM that returns a valence and a short explanation. Gensors runs all active criteria in parallel on recent camera frames, then has a second MLLM, or a user-chosen boolean rule, combine the individual results into a final verdict. This divergent-then-convergent pipeline, together with the Examples-Diff feature that converts user-labeled positive and negative image frames into new criteria, carries the argument by making each requirement independently inspectable and by surfacing blind spots.","core_discovery":"The central claim is that decomposing a high-level sensing goal into atomic criteria, each evaluated separately by a multimodal large language model and shown as a color-coded result, gives end users more control, understanding, and communicative clarity than refining a single free-text prompt. In a within-subjects study with 12 participants, Gensors significantly outperformed a prompt-editor baseline on the self-reported measures of control, understanding of underlying model capabilities, and ability to think through and communicate requirements, while the difference on testing was directionally positive but not statistically significant. Qualitatively, participants could localize failures to one criterion, disable or revise it without touching others, and use automatically generated criteria or image-based difference examples to uncover conditions they had not considered. The paper presents this as evidence that MLLM-powered sensing can be made customizable by non-experts when requirements are treated as first-class objects.","pith_inferences":["The criteria-based authoring pattern likely transfers beyond visual sensing to any MLLM task where a long prompt is hard to debug, such as document triage or email classification, since atomic testable conditions provide the same isolation benefits.","A larger, less tech-savvy sample would clarify whether the interface helps or adds overhead: managing many criteria may become a new burden for users who are not comfortable with parallel, modular thinking.","One testable extension is to apply the Examples-Diff mechanism to non-binary tasks and non-visual modalities, such as audio or sensor streams, converting labeled examples into criteria in those domains.","An objective measure, such as the number of edits needed to reach a fixed behavioral specification or the accuracy of final verdicts on held-out scenes, would test whether the self-reported gains in control and understanding translate into better downstream sensor behavior."],"forward_implications":["Users can specify high-level, subjective sensing tasks by editing criteria instead of wrestling with one long prompt, lowering the barrier for non-technical people.","Debugging becomes systematic because a misbehaving sensor can be traced to a single criterion and fixed without risking unintended changes to other behavior.","Automatically generated criteria and image-difference reasoning reduce omissions, prompting users to consider edge cases and hazards they would not anticipate on their own.","MLLM stochasticity becomes visible as flickering outputs, and users can treat flickering as a signal that a criterion or prompt needs refinement rather than only as a reliability problem.","The design goals should remain relevant as models improve, because the bottleneck shifts from raw model capability to requirement articulation and testing."],"supporting_citations":[{"why":"Establishes that non-AI experts struggle to design LLM prompts, motivating the criteria-based interface in this paper.","marker":"[79]"},{"why":"Shows that complex, multi-faceted prompts cause models to overlook details, supporting per-criterion granularity.","marker":"[45]"},{"why":"Describes prior work that embeds principles into prompts, motivating the need for isolating and testing individual principles as criteria.","marker":"[61]"},{"why":"Presents Zensors as the prior general-purpose sensing approach that is constrained by CV models and hard to debug, which this work builds beyond.","marker":"[35]"},{"why":"Represents related prompt-evaluation work that extracts user requirements for evaluation but does not support requirement articulation during construction.","marker":"[28]"}],"fun_headline_variants":["Break AI sensors into testable criteria for better control","Atomic criteria give users control over AI camera sensing","Decompose prompts into criteria for better smart cameras","Criteria-based sensors: more control, less guesswork","Turn one big prompt into testable checks for AI vision"],"cache_read_input_tokens":29440,"weakest_assumption_plain":"The 12 tech-savvy, self-selected participants and their self-reported Likert ratings fairly represent the everyday users the system is meant to serve; the paper itself flags that this sample may limit generalizability.","fun_headline_variants_meta":{"raw":{"variants":["Break AI sensors into testable criteria for better control","Atomic criteria give users control over AI camera sensing","Decompose prompts into criteria for better smart cameras","Criteria-based sensors: more control, less guesswork","Turn one big prompt into testable checks for AI vision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2565,"prompt_tokens":1021,"completion_tokens":1544,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":637,"completion_tokens_details":{"reasoning_tokens":1468}},"tokens_in":637,"tokens_out":1544,"duration_ms":9132,"temperature":1.0,"reasoning_tokens":1468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:59:24.020176+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a pre-registered study with a larger, less technically oriented sample in which users must make a sensor meet five fixed behavioral specifications using either Gensors or a single-prompt baseline, then measure the fraction of specification checks passed; the central claim would not hold if the criteria interface fails to outperform the baseline on objective correctness or if its control and clarity advantages disappear once age, technical background, and prior experience are controlled.","supporting_citations":[],"review_version":1}