ChatGPT Evaluation (3rd in a series of articles about ChatGPT & AI)
ChatGPT Evaluation (3rd in a series of articles about ChatGPT and AI)
Dear Mr./Ms. GB, regarding the intelligence demonstrated by ChatGPT in various fields that have been hot topics in WAG GB recently, it seems that a fairly comprehensive and quantitative evaluation of ChatGPT is needed, so that we can understand its strengths and weaknesses, and can use it more wisely and carefully, although this is also not easy. I believe that through the refinement of the AI algorithm (language model) and retraining with a more complete and up-to-date dataset, ChatGPT will become even smarter in the future. The following paper, released several months after ChatGPT was launched, is an attempt for the first time to conduct a comprehensive assessment of ChatGPT that touches on aspects of reasoning, hallucinations, and its interactive nature in dialogue. Yesterday I attended a webinar about this paper by one of its co-authors, an Indonesian student.
https://arxiv.org/pdf/2302.04023.pdf
The paper titled “A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity” evaluates ChatGPT comprehensively and quantitatively using 21 datasets covering several different natural language processing tasks, namely multitask, multilingual, and multimodal aspects. The multitask aspects of NLP include text summarization, language translation (from English to other languages and from other languages to English), sentiment analysis, question and answer, misinformation detection, task-oriented dialogue, etc., with performance varying from 4.1 to 93.3 on a scale of one hundred according to the respective metrics for this task. This paper also found that ChatGPT outperforms other large-scale language models based on zero-shot learning (the ability of a language model to perform tasks for which it is not explicitly trained). ChatGPT has advantages in understanding non-Latin languages (e.g., Korean, Japanese, Chinese) compared to its ability to generate sentences in those languages. ChatGPT is able to generate multimodal content (images) from user instructions through (intermediate) coding steps, for example with HTML canvas. On average, ChatGPT has an accuracy rate of 64.33% in ten (10) types of reasoning tasks, such as logical reasoning (deductive, inductive, abductive), non-textual reasoning (temporal, spatial, mathematical), common sense reasoning, logical reasoning, and causal reasoning. Thus, in terms of reasoning, ChatGPT is considered not yet good enough. In this context, ChatGPT has better deductive reasoning capabilities than inductive ones. Like other large-scale language models, ChatGPT has the weakness of sometimes (quite often?) hallucinating, meaning it provides new facts that are incorrect because it does not have access to an external knowledge base. Furthermore, this paper concludes that ChatGPT has the ability to collaborate with humans/users in turn-based dialogue interactions with an improvement of 8% for text summarization and 2% for language translation. Details of the quantitative evaluation and examples of the dialogues used can be seen in detail in this paper.
Bandung, 18 Februari, 2023
Bambang Riyanto