Pith. sign in

REVIEW 5 cited by

Human-In-the-Loop Software Development Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12924 v2 pith:4TP7XO6O submitted 2024-11-19 cs.SE cs.AIcs.HCcs.LG

classification cs.SEcs.AIcs.HCcs.LG
keywords softwaredevelopmentcodehulaagentsframeworkatlassiancoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, Large Language Models (LLMs)-based multi-agent paradigms for software engineering are introduced to automatically resolve software development tasks (e.g., from a given issue to source code). However, existing work is evaluated based on historical benchmark datasets, rarely considers human feedback at each stage of the automated software development process, and has not been deployed in practice. In this paper, we introduce a Human-in-the-loop LLM-based Agents framework (HULA) for software development that allows software engineers to refine and guide LLMs when generating coding plans and source code for a given task. We design, implement, and deploy the HULA framework into Atlassian JIRA for internal uses. Through a multi-stage evaluation of the HULA framework, Atlassian software engineers perceive that HULA can minimize the overall development time and effort, especially in initiating a coding plan and writing code for straightforward tasks. On the other hand, challenges around code quality remain a concern in some cases. We draw lessons learned and discuss opportunities for future work, which will pave the way for the advancement of LLM-based agents in software development.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds

    cs.SE 2026-08 conditional novelty 7.0 of 10

    Fine-tuning a 30B coding agent on 576 planning-aware trajectories collected under Claude Code improves its SWE-bench score on unseen harnesses (OpenCode +3.4, mini-swe-agent +7.0 with self-plans).

  2. Where Is the Cost of Third-Party API Routers in Agentic Software Development?

    cs.SE 2026-07 conditional novelty 6.5 of 10

    Router-side response tampering yields 0% defense success on Claude Code, Codex, Cursor, and OpenCode; whitelist and LLM review only partially restore control.

  3. Anticipating Bugs: Ticket-Level Bug Prediction and Temporal Proximity Effects

    cs.SE 2025-06 conditional novelty 6.0 of 10

    Bug-inducing tickets can be predicted better than random at ticket creation, and accuracy improves as the ticket approaches implementation, with no single feature family dominant at every stage.

  4. Governed AI-Assisted Engineering: Graduated Human Oversight for Agentic Code Generation in Regulated Domains

    cs.HC 2026-06 unverdicted novelty 5.5 of 10

    GAIE introduces an Oversight Classification Model to route code generation tasks to human-in-the-loop, human-over-the-loop, or automated-with-monitoring tiers based on regulatory impact, customer proximity, reversibil...

  5. CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows

    cs.SE 2026-08 conditional novelty 5.0 of 10

    CURATE is a human-in-the-loop, catalog-backed multi-agent system that composes and deploys scientific workflows; a prototype recreated SeBS-Flow workflows and automated a PyADM1 simulation pipeline.

Pith tools