Introducing GPT-4o: ChatGPT Elevates Your Chat Experience with Improved Speed and Object Description Capability
The new robot update promises more agile responses and will be available free of charge to all users. Additionally, the company revealed the launch of a computer application.
OpenAI launched ChatGPT-4, a new generation of its artificial intelligence model capable of "conversing", identifying objects and performing simultaneous translation - Photo Reproduction/Youtube
OpenAI announced on Monday (13) the launch of the latest iteration of the artificial intelligence (AI) model that powers ChatGPT, a conversational bot that has gained notoriety in recent months. This new version, GPT-4o, stands out for its improved ability to respond to real-time audio commands and improved ability to describe images.Since its inception, GPT language models have continually impressed and transformed the way we interact with AI systems. Now, with the launch of GPT-4o, the next milestone in the evolution of this technology, users can expect an even more detail-rich conversational experience.
The Evolution of ChatGPT
ChatGPT has been a reference for those seeking natural and meaningful conversational interactions with an AI system. From GPT-2 to GPT-3.5, remarkable advances have been made in natural language understanding and contextual responsiveness. The implementation ofGPT-4o will be gradual, reaching all users, including those using the free version of the service.
During a demonstration, the model was able to analyze the user's appearance and offer clothing suggestions for a job interview. In another test, the GPT-4o was used to compose a song. This is OpenAI's first model to integrate text, images and audio in real time autonomously. Previous models relied on other AI to process voice commands and images. The promise is that this evolution will make ChatGPT even more agile.
Improved Speed
One of the most notable features of GPT-4o is its improved speed. Thanks to improvements in the underlying hardware and refinements in the processing algorithms, users will be able to enjoy faster responses and more fluid interactions. Whether it's live chats, virtual assistance, or real-time text analysis, ChatGPT will deliver near-instant interactivity.According to OpenAI, GPT-4o responds to audio commands in an average of 320 milliseconds, with a minimum of 232 milliseconds. The company claims it is significantly faster than its predecessors, with GPT-3.5 taking an average of 2.8 seconds and GPT-4 taking 5.4 seconds.
Until then, ChatGPT went through multiple phases to process and respond to voice commands. Initially, it was necessary to employ a model to transcribe the audio into text. The content was then interpreted and responded to by GPT-3.5 or GPT-4. Lastly, another model converted the response back into audio. As disclosed by the company, GPT-4o will be developed with a single model that covers text, vision and audio, implying that all inputs and outputs will be processed by the same neural network. OpenAI CEO Sam Altman declared that this is the most advanced model ever designed by the company. “It is smart, fast and intrinsically multimodal,” he said.
Object Description Capability
Another advancement is ChatGPT's improved ability to describe objects. While previous versions were proficient in understanding concepts and answering questions about them, GPT-4o goes further, providing detailed and accurate descriptions of a wide variety of physical and abstract objects. Whether describing an object's appearance, its function, or even its history and cultural context, ChatGPT will now be able to provide comprehensive and insightful information that further enriches the user experience.
Potential Applications
According to the company, GPT-4o has greater ability to understand texts, images and audio compared to its predecessor, GPT-4, which was launched in March 2023. In addition, the company announced the launch of a ChatGPT application for computers, complementing the browser versions and applications for Android and iOS already available.
With these advances, the possibilities of ChatGPT are vast and diverse. Since more responsive and useful virtual assistants to more effective customer support systems, GPT-4o has the potential to improve a wide range of human interactions with AI. Additionally, in fields such as education, research, and entertainment, the ability to describe objects in detail can open up new opportunities for interactive learning, data exploration, and captivating content creation.
On social media, users have compared this new version to the virtual assistant from the film "Her" ("Her" in the original title), in which the protagonist becomes emotionally involved with an operating system. The reaction was so significant that it reached Sam Altman, who mentioned the name of the film on his profile on platform X (formerly Twitter).
When will GPT-4 be released?
OpenAI announced that it began making GPT-4o's text and image resources available this Monday. These features are also now available for developers to integrate into their own applications. Free version users will be able to use GPT-4, but with an unspecified message limit, while ChatGPT Plus subscribers will have access to a higher limit. Using GPT-4 with voice commands will be available in the coming weeks for ChatGPT Plus subscribers. The company has not yet announced when the video features will be available to all users, but has stated that they will initially be rolled out to a select group of developer partners.


























