REVIEW 1 cited by
Towards Standard Criteria for human evaluation of Chatbots: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Human evaluation is becoming a necessity to test the performance of Chatbots. However, off-the-shelf settings suffer the severe reliability and replication issues partly because of the extremely high diversity of criteria. It is high time to come up with standard criteria and exact definitions. To this end, we conduct a through investigation of 105 papers involving human evaluation for Chatbots. Deriving from this, we propose five standard criteria along with precise definitions.
Forward citations
Cited by 1 Pith paper
-
Communication is All You Need: Persuasion Dataset Construction via Multi-LLM Communication
A six-role multi-LLM communication framework generates persuasive dialogue data that human judges find nearly indistinguishable from human-written rewrites.
Discussion (0). Continue with ORCID to comment.