Pith. sign in

REVIEW 4 cited by

Exploring the Limits of ChatGPT for Query or Aspect-based Text Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.08081 v1 pith:HUQQ6KTW submitted 2023-02-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords summarizationchatgptsummariestextperformancebeenchatgpt-generateddiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text summarization has been a crucial problem in natural language processing (NLP) for several decades. It aims to condense lengthy documents into shorter versions while retaining the most critical information. Various methods have been proposed for text summarization, including extractive and abstractive summarization. The emergence of large language models (LLMs) like GPT3 and ChatGPT has recently created significant interest in using these models for text summarization tasks. Recent studies \cite{goyal2022news, zhang2023benchmarking} have shown that LLMs-generated news summaries are already on par with humans. However, the performance of LLMs for more practical applications like aspect or query-based summaries is underexplored. To fill this gap, we conducted an evaluation of ChatGPT's performance on four widely used benchmark datasets, encompassing diverse summaries from Reddit posts, news articles, dialogue meetings, and stories. Our experiments reveal that ChatGPT's performance is comparable to traditional fine-tuning methods in terms of Rouge scores. Moreover, we highlight some unique differences between ChatGPT-generated summaries and human references, providing valuable insights into the superpower of ChatGPT for diverse text summarization tasks. Our findings call for new directions in this area, and we plan to conduct further research to systematically examine the characteristics of ChatGPT-generated summaries through extensive human evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond In-Context Learning: Aligning Long-form Generation of Large Language Models via Task-Inherent Attribute Guidelines

    cs.CL 2025-06 conditional novelty 7.0 of 10

    LongGuide automatically learns task-specific quality and length guidelines from small training sets, significantly improving LLM long-form generation.

  2. VA-Blueprint: Uncovering Building Blocks for Visual Analytics System Design

    cs.HC 2025-08 conditional novelty 6.0 of 10

    A semi-automated methodology and public knowledge base catalog the building blocks of 101 urban visual analytics systems, using GPT-4 for extraction with human-in-the-loop correction and expert validation.

  3. Adapting Online Customer Reviews for Blind Users: A Case Study of Restaurant Reviews

    cs.HC 2025-06 conditional novelty 6.0 of 10

    QuickCue, an LLM-powered browser extension that reorganizes restaurant reviews into aspect-sentiment summaries, significantly improved usability and reduced workload for blind screen reader users in a 10-user study.

  4. A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis

    cs.CL 2025-06 conditional novelty 4.0 of 10

    An LLM agent that reframes beam analysis as OpenSeesPy code generation reaches over 99 percent reliability on a small benchmark, but chiefly because the prompt contains a near-identical solved example.

Pith tools