REVIEW 5 cited by
Human-In-the-Loop Software Development Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, Large Language Models (LLMs)-based multi-agent paradigms for software engineering are introduced to automatically resolve software development tasks (e.g., from a given issue to source code). However, existing work is evaluated based on historical benchmark datasets, rarely considers human feedback at each stage of the automated software development process, and has not been deployed in practice. In this paper, we introduce a Human-in-the-loop LLM-based Agents framework (HULA) for software development that allows software engineers to refine and guide LLMs when generating coding plans and source code for a given task. We design, implement, and deploy the HULA framework into Atlassian JIRA for internal uses. Through a multi-stage evaluation of the HULA framework, Atlassian software engineers perceive that HULA can minimize the overall development time and effort, especially in initiating a coding plan and writing code for straightforward tasks. On the other hand, challenges around code quality remain a concern in some cases. We draw lessons learned and discuss opportunities for future work, which will pave the way for the advancement of LLM-based agents in software development.
Forward citations
Cited by 5 Pith papers
-
DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds
Fine-tuning a 30B coding agent on 576 planning-aware trajectories collected under Claude Code improves its SWE-bench score on unseen harnesses (OpenCode +3.4, mini-swe-agent +7.0 with self-plans).
-
Where Is the Cost of Third-Party API Routers in Agentic Software Development?
Router-side response tampering yields 0% defense success on Claude Code, Codex, Cursor, and OpenCode; whitelist and LLM review only partially restore control.
-
Anticipating Bugs: Ticket-Level Bug Prediction and Temporal Proximity Effects
Bug-inducing tickets can be predicted better than random at ticket creation, and accuracy improves as the ticket approaches implementation, with no single feature family dominant at every stage.
-
Governed AI-Assisted Engineering: Graduated Human Oversight for Agentic Code Generation in Regulated Domains
GAIE introduces an Oversight Classification Model to route code generation tasks to human-in-the-loop, human-over-the-loop, or automated-with-monitoring tiers based on regulatory impact, customer proximity, reversibil...
-
CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows
CURATE is a human-in-the-loop, catalog-backed multi-agent system that composes and deploys scientific workflows; a prototype recreated SeBS-Flow workflows and automated a PyADM1 simulation pipeline.
Discussion (0). Continue with ORCID to comment.