REVIEW 2 cited by
SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent work has highlighted the potential of modelling interactive behaviour analogously to natural language. We propose interactive behaviour summarisation as a novel computational task and demonstrate its usefulness for automatically uncovering latent user intentions while interacting with graphical user interfaces. To tackle this task, we introduce SummAct, a novel hierarchical method to summarise low-level input actions into high-level intentions. SummAct first identifies sub-goals from user actions using a large language model and in-context learning. High-level intentions are then obtained by fine-tuning the model using a novel UI element attention to preserve detailed context information embedded within UI elements during summarisation. Through a series of evaluations, we demonstrate that SummAct significantly outperforms baselines across desktop and mobile interfaces as well as interactive tasks by up to 21.9%. We further show three exciting interactive applications benefited from SummAct: interactive behaviour forecasting, automatic behaviour synonym identification, and language-based behaviour retrieval.
Forward citations
Cited by 2 Pith papers
-
MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users
A hierarchical auto-annotation pipeline plus episode-aware federated aggregation lets mobile GUI agents be trained on automatically labeled user trajectories at about 1% of human annotation cost.
-
Bi-Fact: A Bidirectional Factorization-based Evaluation of Intent Extraction from UI Trajectories
Bi-Fact, a bidirectional fact-level LLM-based metric, reports higher agreement with human judgments than existing metrics when scoring intent extraction from GUI trajectories.
Discussion (0). Continue with ORCID to comment.