OpenAI, an AI (artificial intelligence) research and deployment company, has introduced GPT-4o (“o” for “omni”), a new model designed to enhance human-computer interaction through multimodal capabilities.
This model from OpenAI accepts text, audio, and image inputs and generates outputs in the same formats. GPT-4o is designed to render quick response times, process audio inputs in as little as 232 milliseconds, and average 320 milliseconds, comparable to human conversational speed.
GPT-4o matches the performance of GPT-4 Turbo on text in English and coding tasks while offering significant improvements in non-English text processing. It also promises to excel in vision and audio understanding. The model operates much faster and at a 50% lower cost through the API than previous models.
Enhanced multimodal processing
Previous voice interaction models, such as Voice Mode in GPT-3.5 and GPT-4, experienced latencies of 2.8 and 5.4 seconds respectively. These models used a pipeline of separate models for transcribing audio to text and converting text back to audio, resulting in the loss of contextual information. GPT-4o integrates text, vision, and audio processing into a single neural network, preserving more contextual information and enabling more natural interactions.
On traditional benchmarks, GPT-4o achieves GPT-4 Turbo-level performance in text, reasoning, and coding. It has set new high scores on multilingual, audio, and vision capabilities. For instance, GPT-4o scored 88.7% on the 0-shot Chain of Thought (CoT) MMLU and 87.2% on the 5-shot no-CoT MMLU, surpassing previous models.
Safety is a critical focus for GPT-4o, incorporating safety measures across all modalities. This includes filtering training data and refining post-training model behavior. Evaluations according to OpenAI’s Preparedness Framework ensure that GPT-4o does not exceed Medium risk in cybersecurity, persuasion, or autonomy categories. Extensive external red teaming has also been conducted to identify and mitigate risks associated with the new modalities.
Controlled release and future plans
Currently, GPT-4o supports text and image inputs and text outputs, with plans to release audio outputs gradually. Initial audio outputs will be limited to preset voices adhering to existing safety policies. OpenAI will continue to refine and enhance the technical infrastructure, usability, and safety features of GPT-4o’s multimodal capabilities.
OpenAI said it plans to continue addressing the limitations and exploring the full potential of GPT-4o in the coming months.