{"id":"7c7b4bb2-b71a-45e8-9918-7084b26dfd11","arxiv_id":"2606.30697","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LUMOS provides a semantic interaction layer converting accessibility metadata into machine-readable UI blueprints for grounded AI agent actions.","lead":"LUMOS introduces a semantic layer that converts operating system accessibility metadata into structured, machine-readable UI blueprints for AI agents to observe and act without screenshots. A smart generalist might read it to see how existing OS features could make AI computer-use agents faster, cheaper, and more reliable.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Accessibility metadata completeness across apps is the untested foundation","rationale":"The reader's weakest_assumption matches the load-bearing precondition exactly. Because the manuscript is framed as a proposal whose utility is conditional on metadata quality, and no supporting measurements appear in the abstract, the UNVERDICTED verdict with low confidence remains appropriate. No internal inconsistency or stronger technical flaw is visible from the given material.","tokens_in":1708,"tokens_out":307,"duration_ms":12269,"concrete_test":"Traverse accessibility trees via OS APIs (e.g., Windows UI Automation or macOS NSAccessibility) on 8 representative apps (Chrome, VS Code, Excel, Finder/Explorer, Slack, terminal, settings panel, PDF viewer); for each, compute the fraction of visible elements that expose all fields required by the blueprint (stable identifier, role, name, value, bounds, action list). Report per-app coverage and the rate of missing fields.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LUMOS can produce reliable semantic blueprints from native accessibility metadata and browser structures, enabling an observe-act loop with constrained primitives. This holds only if OS-provided trees are sufficiently complete, consistent, and stable across applications. The abstract states the conversion occurs but supplies no coverage metrics, failure modes, or cross-app evaluation; the paper explicitly conditions success on cases where semantic structure is already present, leaving the frequency and reliability of that condition unaddressed.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces LUMOS, a semantic interaction layer that converts native OS accessibility metadata and browser UI structures into machine-readable semantic blueprints containing stable identifiers, roles, names, values, bounds, and action affordances. It enables an LLM-driven observe-act loop using constrained visible-UI primitives and live semantic pointer grounding via OS automation APIs, with the goal of reducing token costs, latency, visual ambiguity, and coordinate uncertainty compared to screenshot/OCR-based agents. The approach is positioned as complementary to visual methods, applying only where semantic structure is already present, and is claimed to suggest a path toward AI-native operating systems.","tokens_in":1778,"tokens_out":559,"duration_ms":19931,"significance":"If the architectural claims hold and the accessibility metadata proves sufficiently complete and stable, LUMOS could meaningfully lower the cost and error rate of computer-use agents by replacing pixel-based interpretation with grounded semantic primitives. The proposal is a clean systems-level idea that directly addresses a practical pain point in current LLM agents, but its significance remains prospective because the manuscript contains no empirical validation, coverage statistics, or cross-application measurements.","major_comments":[{"comment":"Abstract: The statements that LUMOS 'reduces dependence on screenshots,' 'reduces token costs, latency, and ambiguity,' and enables reliable operation 'when operating systems already provide semantic structure' are presented as results, yet the manuscript supplies no quantitative evaluation, token-count comparisons, latency measurements, success-rate metrics, or coverage analysis across applications to support these claims.","section":"Abstract"},{"comment":"Abstract and introduction: The central premise that native accessibility trees are 'sufficiently complete, consistent, and stable' to support an observe-act loop without fallback is asserted but left unquantified; no data on failure modes, fraction of UI elements missing semantic metadata, or cross-app reliability is provided, rendering the practical scope of the proposal indeterminate.","section":"Abstract"},{"comment":"Abstract: The claim that the system 'supports live semantic pointer grounding by querying the UI element under or near the cursor' is described at a high level but lacks any specification of the OS APIs used, error-handling strategy when the query returns incomplete data, or how the resulting blueprint is serialized for the LLM.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract refers to 'these results' but the manuscript appears to be an architectural proposal without an evaluation section; clarifying whether the work is intended as a position paper or a systems contribution with forthcoming experiments would help readers.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive review. The comments correctly identify that the abstract presents design objectives in language that could be read as reporting completed results. We address each point below and will revise the manuscript accordingly to ensure claims accurately reflect the conceptual nature of the proposal.","responses":[{"response":"We agree that the abstract phrasing implies empirical outcomes. The manuscript is a systems proposal describing an architecture and does not contain evaluations. We will revise the abstract to use prospective language (e.g., 'is designed to reduce dependence on screenshots' and 'suggests a path toward AI-native operating systems') and remove any implication of measured results.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The statements that LUMOS 'reduces dependence on screenshots,' 'reduces token costs, latency, and ambiguity,' and enables reliable operation 'when operating systems already provide semantic structure' are presented as results, yet the manuscript supplies no quantitative evaluation, token-count comparisons, latency measurements, success-rate metrics, or coverage analysis across applications to support these claims."},{"response":"The manuscript already qualifies applicability to settings 'when operating systems already provide semantic structure' and states that LUMOS is complementary to visual methods. We acknowledge that the abstract does not sufficiently foreground variability in accessibility metadata. We will add an explicit limitations paragraph discussing known incompleteness of accessibility trees and the need for fallback mechanisms, drawing on existing platform documentation.","revision_made":"partial","referee_comment":"[Abstract] Abstract and introduction: The central premise that native accessibility trees are 'sufficiently complete, consistent, and stable' to support an observe-act loop without fallback is asserted but left unquantified; no data on failure modes, fraction of UI elements missing semantic metadata, or cross-app reliability is provided, rendering the practical scope of the proposal indeterminate."},{"response":"The abstract is intentionally high-level. The full manuscript describes the observe-act loop; we will expand the abstract with a brief reference to the relevant platform APIs (UI Automation, Accessibility frameworks) and add a short implementation subsection covering error handling for incomplete queries and the JSON serialization format used for blueprints.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that the system 'supports live semantic pointer grounding by querying the UI element under or near the cursor' is described at a high level but lacks any specification of the OS APIs used, error-handling strategy when the query returns incomplete data, or how the resulting blueprint is serialized for the LLM."}],"tokens_in":1456,"tokens_out":553,"duration_ms":37478,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper proposes turning native OS accessibility metadata and browser structures into semantic blueprints with stable IDs, roles, bounds, and actions, then letting an LLM drive an observe-act loop on those primitives. It also adds cursor-based live grounding. That framing is new in the agent literature and directly targets the token and ambiguity costs of screenshot-based agents.\n\nThe paper does a solid job stating the problem: current agents pay for visual parsing when the OS already exposes structured data for humans. The suggestion to use constrained visible-UI primitives rather than app-specific scripts is practical, and the explicit note that LUMOS only applies where semantic structure already exists avoids overclaiming.\n\nThe soft spot is exactly what the stress-test flagged. The whole approach rests on accessibility trees being sufficiently complete, consistent, and present across applications and browsers. The abstract states the conversion happens but supplies zero coverage numbers, failure cases, or cross-app measurements. Without that, the claimed reductions in cost and ambiguity remain assertions. The paper conditions success on the metadata already being there, yet never quantifies how often that condition holds.\n\nThis is for researchers building computer-use agents who want to explore OS-level interfaces. A reading group could usefully discuss the architecture and the open question of metadata reliability. It is not yet useful for anyone needing reproducible results or deployment guidance.\n\nI would not cite the work in its current form. It deserves peer review if the full manuscript contains an implementation plus even basic evaluation of tree coverage and failure modes; otherwise the idea stays too preliminary to justify referee time.","headline":"LUMOS is a clean architectural sketch for grounding agents in accessibility trees instead of pixels, but the paper gives no data on whether those trees are complete or stable enough to matter.","tokens_in":2224,"tokens_out":400,"would_cite":false,"duration_ms":16743,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LUMOS converts accessibility metadata into semantic blueprints so AI agents can observe and act on UI elements without screenshots.","keywords":["accessibility metadata","AI agents","semantic UI","operating systems","LLM interaction","UI automation","machine-readable semantics","observe-act loop"],"falsifier":"Implement LUMOS on a standard desktop and run agents on common applications such as web browsers and office software; count the fraction of tasks that fail due to missing, incomplete, or inconsistent accessibility data.","tokens_in":2609,"feed_emoji":"💻","tokens_out":674,"duration_ms":27848,"temperature":0.7,"pith_summary":"Current operating systems present interfaces through pixels and visuals that suit humans but create high costs and ambiguity for AI agents. LUMOS addresses this by turning native accessibility metadata and browser structures into compact, machine-readable semantic blueprints that list stable identifiers, roles, names, values, bounds, and action options. Language models then follow an observe-act loop that issues grounded commands on these visible UI primitives. The approach matters because it lowers token usage, removes coordinate uncertainty, and enables consistent operation across applications without writing per-app scripts. It frames existing accessibility support as the basis for more direct machine interaction with computers.","feed_headline":"Accessibility metadata powers AI agents without screenshot reliance","feed_subtitle":"LUMOS turns native OS data into stable semantic blueprints with identifiers and actions for direct control.","key_machinery":"LUMOS semantic interaction layer that converts accessibility metadata into machine-readable blueprints carrying stable identifiers and action affordances.","core_discovery":"LUMOS converts native accessibility metadata and browser UI structures into machine readable semantic blueprints with stable identifiers, roles, names, values, bounds, and action affordances. It also supports live semantic pointer grounding by querying the UI element under or near the cursor through operating-system automation APIs. An LLM then acts through an accessibility grounded observe act loop using constrained visible-UI primitives rather than application-specific scripts. LUMOS does not claim to replace visual agents; instead, it reduces dependence on screenshots when operating systems already provide semantic structure.","pith_inferences":["Similar semantic layers could be built for mobile platforms that already expose accessibility APIs.","A hybrid system could fall back to visual methods only when accessibility metadata is absent for a given element.","Wider use of the layer would create market pressure for developers to make accessibility metadata more complete.","The observe-act pattern could transfer to other domains such as robotic control where sensor data needs semantic grounding."],"forward_implications":["AI agents incur lower token costs because they process structured semantic data instead of image pixels or OCR output.","Action reliability increases through explicit element roles, bounds, and affordances rather than visual interpretation.","Agents achieve cross-application consistency by using OS-provided primitives instead of writing application-specific code.","The layer supplies a concrete route toward operating systems that expose machine-readable interaction as a native feature."],"fun_headline_variants":["LUMOS provides semantic blueprints from accessibility metadata","Accessibility data enables direct AI agent actions via LUMOS","LUMOS reduces screenshot use with machine readable UI semantics","Semantic OS layer for AI using native accessibility structures","LUMOS grounds LLM actions in stable accessibility identifiers"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Operating systems provide sufficient, complete, and reliable accessibility metadata across applications and browsers to support reliable AI agent operation without fallback to visual methods.","fun_headline_variants_meta":{"raw":{"variants":["LUMOS provides semantic blueprints from accessibility metadata","Accessibility data enables direct AI agent actions via LUMOS","LUMOS reduces screenshot use with machine readable UI semantics","Semantic OS layer for AI using native accessibility structures","LUMOS grounds LLM actions in stable accessibility identifiers"]},"model":"grok-4.3","cost_usd":0.004756,"raw_usage":{"total_tokens":2351,"prompt_tokens":682,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":47562000,"prompt_tokens_details":{"text_tokens":682,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1595,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":682,"tokens_out":74,"duration_ms":17609,"temperature":1.0,"reasoning_tokens":1595,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T01:52:35.177252+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Implement LUMOS on a standard desktop and run agents on common applications such as web browsers and office software; count the fraction of tasks that fail due to missing, incomplete, or inconsistent accessibility data.","supporting_citations":[],"review_version":1}