ChatGPT & AI (1st of a series of short articles about ChatGPT in particular and AI in general)
https://openai.com/blog/chatgpt/
Rupanya ChatGPT (https://openai.com/blog/chatgpt) It's quite popular and hotly discussed in various WhatsApp groups. Here's my exploration and in-depth analysis of how ChatGPT was created. Sorry for the length.
———————
ChatGPT is currently going viral everywhere, thanks to its ability to respond to questions/topics in human-friendly natural language and demonstrates the incredible AI advancements made by the OpenAI research team. It's perhaps worth understanding how ChatGPT was developed to understand its capabilities, strengths, and limitations. I've been following OpenAI's development for the past few years and have observed the AI findings and technologies they produce, especially in the realms of NLP (Natural Language Processing) and Computer Vision, which are incredibly impressive, like ChatGPT.
AI observers and enthusiasts may be more interested in how the AI technology behind ChatGPT was developed. ChatGPT is essentially a large-scale AI language model capable of generating text for various natural language processing needs, particularly interactive conversations/dialogues. The development of this language model is based on next token prediction and masked language modeling, key tasks in NLP. Compared to previously developed language models, ChatGPT has higher precision, detail, and coherence, meaning it is more human-friendly in generating dialogue. From an AI perspective, OpenAI uses the GPT3.5 deep learning model, an extension of GPT3, trained with supervised learning based on labeled data (specifically, prompt-response data) combined with Reinforcement Learning (RL). In ChatGPT's development, reinforcement learning is applied using human feedback to reduce incorrect and/or biased predictions. The specific reinforcement learning algorithm used in ChatGPT is PPO (Proximal Policy Optimization). PPO is a policy-based reinforcement learning algorithm, not a value-based one (Q learning). This use of reinforcement learning enabled the OpenAI team to produce a large-scale, high-performance language model on ChatGPT based on a pretrained model trained with supervised learning on a relatively small dataset. The word "relatively small" in this last sentence is intentionally put in quotation marks because the dataset is quite large, but not as large as it would be if fully trained with supervised learning, without using reinforcement learning. As a note, reinforcement learning can be simply analogized to when we teach a child to act/behave. If the action is good, we can reward it, but if the action is bad, we can give it punishment as needed. Through several exercises, over time, the knowledge accumulates in the child's brain to be more inclined to perform good actions.
In short, there are three steps taken by the OpenAI team in developing ChatGPT: 1) sampling input (prompt) from the language dataset and humans provide the desired answers and then these input/prompt and answer pairs are used to train the initial GPT3.5 model using the supervised learning algorithm, 2) sampling output/answers to an input/prompt based on the initial model developed in stage 1, and humans rank the answers from best to worst, then this data is used as a reward model from reinforcement learning, 3) Samples new input/prompts from the dataset, initializes PPO with a supervised policy, then generates output and the reward model calculates the reward value for this output. This reward value is then used to update the reinforcement learning policy using PPO.
By utilizing a combination of supervised learning and reinforcement learning with human feedback on a large dataset, ChatGPT can produce a fairly natural, precise, and coherent dialogue system. However, this is not without its drawbacks, one of which is that ChatGPT is based on a dataset that, although very large, is still limited and subject to bias, especially when compared to human answers/predictions. Bias is also very likely to be generated by human feedback. As the dataset and RL human feedback grow, ChatGPT can be further trained in iterative development to produce an increasingly intelligent dialogue system in the coming years. For specific topics that require in-depth answers, ChatGPT often falls short, either because the answer to a query is out of context or because the topic requires special expertise. In this regard, fine-tuning GPT3 can actually be done for very specific applications/topics with local/specific datasets using a technique known in machine learning as transfer learning.
For the record, GPT3 is architecturally a deep learning model. developed from Transformer—a Deep Learning architecture (Deep Neural Networks) that adopts a self-attention mechanism and has been widely researched/developed recently, especially in the NLP realm and recently also developed in the computer vision realm with Vision Transformer.
Bandung, 14 Januari, 2023
Bambang Riyanto
STEI ITB