Say a sentence to your computer, and it will create code threads, submit Pull Requests, and identify bug root causes—all in one go. OpenAI has launched voice control for the desktop version of ChatGPT, and this time, voice isn't just for chatting—it's for getting work done.
One Sentence, Three Tasks, All Done
On August 9, 2026, OpenAI added voice control to the desktop version of ChatGPT, turning voice from a "chat entry" into a "work entry."
In OpenAI's demo video, a developer gives a single command, and ChatGPT proceeds to create a code thread, submit a Pull Request(a method to merge your code into a project), and then identify the root cause of a bug—executing all three steps seamlessly.
Powering this capability is the new model ChatGPT-Live(an AI model designed for smooth voice conversations). Unlike the voice feature on the mobile version, the desktop version aims not just to answer questions but to execute tasks.
It can also operate the computer itself: access websites, open applications; on macOS, with the help of Appshots, it can even read screen content, including alternative text prepared for visually impaired users. Verbal commands translate into real actions on the screen.
The positioning becomes clearer when you think about it: while the mobile version's voice feature focuses on chat and Q&A, the desktop version is all about execution.
This is the next step in pushing voice towards productivity after the mobile voice feature.
What Layer of "Understanding" Is This?
Voice assistants have been around for a while, but this time, the difference lies in the "understanding" layer.
Traditional voice assistants understand "action + object": setting an alarm, checking the weather—one step and done. ChatGPT Voice needs to understand "tasks": a sentence containing multiple steps that it must first break down and then execute in order. The command in the demo was broken down into three actions: creating a thread, submitting code, and checking for bugs.
After breaking it down, it must also be able to act. ChatGPT-Live not only processes voice but also connects to computer operation capabilities—accessing websites, applications, and on macOS, even reading screen content. Understanding and acting are now integrated into the same model.
This screen-reading capability may seem insignificant, but it's actually crucial: commands can now directly refer to what's on the screen, and the AI can see what you're seeing; even images can be "read" through alternative text, so they don't become blind spots.
The Shortcoming of Voice Is Exactly the Breakthrough This Time
In the past, voice assistants couldn't handle complex tasks, and everyone thought it was because they "couldn't hear clearly." But the real bottleneck was elsewhere.
Traditional voice assistants can't handle multi-step tasks, and most people assume the problem is with recognition: noisy environments, heavy accents, or unclear speech. However, the command in the demo wasn't long or unclear. The real hurdle is in execution: understanding "submit a Pull Request" is different from actually submitting it into the project, which involves a whole set of computer operation capabilities. Setting an alarm or checking the weather—those are one-step tasks, representing the ceiling for old voice assistants; executing three steps in a row without error is the bar this time.
Recognition rate was never the issue. The real difficulty lies after understanding—having the hands to execute. This time, the model has been equipped with hands.
Looking back, it's clearer: voice recognition has improved significantly in recent years, yet voice assistants still can't handle complex tasks. If the bottleneck were truly in the ears, with better recognition, these tasks should have been doable by now. The fact that they haven't been, exactly indicates the bottleneck is elsewhere.
So, the essence of this upgrade is integrating "voice understanding" and "computer operation" into the same model. The difference for users is direct: the range of tasks voice commands can handle starts to approach real work scenarios, no longer limited to queries and settings.
Competitors Have Caught Up; It's About Who Goes Deeper
OpenAI has just brought voice control to the desktop, and Anthropic's Claude voice mode has also been updated. The race is on.
Claude's voice mode can call upon Opus, Sonnet, and Haiku model tiers(AI models of different capability levels from the same company), performing tasks in Gmail, Calendar, Slack, Notion, and Canva. However, its performance in multi-step tasks has not yet been disclosed by the official sources.
The next few months will be worth watching for two things: First, whether Anthropic will provide actual test data on multi-step tasks; second, whether the list of applications that can be accessed on both sides will continue to grow. The answers to these questions will largely determine who leads in this race.
The desktop is the next stop, but the phone hasn't been left behind: on iOS, you can use ChatGPT Voice through remote access in Codex—say something on the phone, and it executes on the computer.
Turning "speaking" into "doing" is the next form of desktop AI, and both companies see it. The current advantage of ChatGPT Voice lies in multi-step execution, but whether this advantage can be maintained depends on its performance in professional scenarios—which, for now, neither side has provided an answer to.
Install the Desktop Version and You Can Start Talking
The feature is already live, and those with the desktop version can try it directly. Start with a simple command and see how much it understands.
Update to the latest version of the ChatGPT desktop application, log in to your account, and enable ChatGPT Voice in the settings.
Start with a two-step command, such as "Open the calendar and add a meeting for tomorrow afternoon at three," and see if it confirms with you before proceeding.
Then try a work-related multi-step command, similar to the demo's "create a thread, submit code, and check for bugs," and see how many steps it can complete.
macOS users can try Appshots: have it read screen content and help you process what's in the current window.
One reminder: it will proceed with execution, so for important operations (such as submitting code, sending messages, or deleting things), remember to monitor the confirmation step. The faster the AI acts, the more important human oversight becomes.
This article is based on the original IT Home article (2026-08-09). The numbers released by the manufacturer (such as benchmark scores, reductions, etc.) are official figures and have not been independently verified by third parties unless otherwise noted.