Pith. sign in

REVIEW 5 cited by

Can ChatGPT Understand Too? A Comparative Study on ChatGPT and Fine-tuned BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.10198 v2 pith:JRX4KEOB submitted 2023-02-19 cs.CL

classification cs.CL
keywords chatgptabilitybertmodelstasksunderstandinganalysisattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, ChatGPT has attracted great attention, as it can generate fluent and high-quality responses to human inquiries. Several prior studies have shown that ChatGPT attains remarkable generation ability compared with existing models. However, the quantitative analysis of ChatGPT's understanding ability has been given little attention. In this report, we explore the understanding ability of ChatGPT by evaluating it on the most popular GLUE benchmark, and comparing it with 4 representative fine-tuned BERT-style models. We find that: 1) ChatGPT falls short in handling paraphrase and similarity tasks; 2) ChatGPT outperforms all BERT models on inference tasks by a large margin; 3) ChatGPT achieves comparable performance compared with BERT on sentiment analysis and question-answering tasks. Additionally, by combining some advanced prompting strategies, we show that the understanding ability of ChatGPT can be further improved.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 148 citations worldwide. Full citation record

  1. Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Recast predicts the turn distribution of future multi-turn LLM safety failures from dual-scale trajectory evidence, catching 88.3% of failures 2.41 turns early at 12.3% false alarms.

  2. Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Teaching an LLM to emit a fixed four-stage reasoning chain during fine-tuning makes single-pass multi-hop knowledge editing robust to distractor facts.

  3. Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A comprehensive survey of cross-lingual aspect-based sentiment analysis that catalogs tasks, datasets, modeling paradigms, and cross-lingual transfer techniques, and identifies research gaps.

  4. Distilled Large Language Model in Confidential Computing Environment for System-on-Chip Design

    cs.AI 2025-07 reject novelty 4.0 of 10

    Running lightweight distilled LLMs inside Intel TDX secure VMs reportedly gives higher tokens per second than plain CPU execution for sub-3B models, with Q4 quantization reaching about 3x FP16 throughput.

  5. Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

    cs.IR 2025-07 reject novelty 4.0 of 10

    LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.

Pith tools