Pith. sign in

REVIEW 2 cited by

Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.19248 v2 pith:Z2O6724X submitted 2024-02-29 cs.CL

classification cs.CL
keywords llmschinesedynamicanswerbenchmarkcdqalatestability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How to better evaluate the capabilities of Large Language Models (LLMs) is the focal point and hot topic in current LLMs research. Previous work has noted that due to the extremely high cost of iterative updates of LLMs, they are often unable to answer the latest dynamic questions well. To promote the improvement of Chinese LLMs' ability to answer dynamic questions, in this paper, we introduce CDQA, a Chinese Dynamic QA benchmark containing question-answer pairs related to the latest news on the Chinese Internet. We obtain high-quality data through a pipeline that combines humans and models, and carefully classify the samples according to the frequency of answer changes to facilitate a more fine-grained observation of LLMs' capabilities. We have also evaluated and analyzed mainstream and advanced Chinese LLMs on CDQA. Extensive experiments and valuable insights suggest that our proposed CDQA is challenging and worthy of more further study. We believe that the benchmark we provide will become one of the key data resources for improving LLMs' Chinese question-answering ability in the future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs

    cs.LG 2025-02 conditional novelty 6.0 of 10

    LLMs score poorly on CounterMATH, a new counterexample-based university math benchmark, and a 1,025-sample counterexample fine-tune yields small and partly inconsistent gains.

  2. Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Emotionally styled verbal prompts (encouraging, provocative, critical) modestly improve zero-shot LLM accuracy on many tasks, with the best style varying by model and task zone.

Pith tools