ChatGPT Voice arriba a Codex: tasques, navegador i codi només parlant
ChatGPT Voice permet coordinar fils de Codex, navegar, preparar missatges, crear codi i seguir tasques des de l’iPhone. Analitzem la demostració, els permisos i els límits reals.
El vídeo presenta la nova experiència de veu de ChatGPT com el més semblant a tenir un “Jarvis” dins de Codex. El nom oficial, però, és ChatGPT Voice: funciona amb GPT‑Live i permet conversar mentre coordina feina en fils de Chat, Work i Codex des de l’aplicació d’escriptori.
La demostració de Riley Brown combina correu, documents, missatgeria, navegador, recerca i desenvolupament d’aplicacions. La novetat important no és només poder dictar una ordre, sinó mantenir una conversa, interrompre-la, donar instruccions de seguiment i consultar altres tasques sense deixar de parlar.
1. ChatGPT Voice no és el mateix que dictar un prompt
Al començament del vídeo, Brown obre una conversa buida i prem Start new voice chat. Aquesta distinció és important: la dictada converteix la veu en text abans d’enviar el missatge; Voice manté una conversa en directe amb torns naturals.
La documentació d’OpenAI confirma que un xat o una tasca ha de començar en mode de veu per utilitzar aquesta experiència. Si la conversa s’ha iniciat en un altre mode, el micròfon serveix per dictar. En una sessió de Voice es pot interrompre la resposta, canviar de direcció i demanar l’estat de la feina mentre continua executant-se.
El creador utilitza els noms “Realtime Voice in Codex”, “Codex Voice” i “Jarvis”. Són maneres informals de descriure la demostració; el producte anunciat per OpenAI és ChatGPT Voice.
2. Primera prova: convertir queixes de correu en respostes i un resum
La primera ordre és buscar correus de suport amb queixes, preparar una resposta per a cadascun i obrir-los al navegador integrat. La sessió informa que ha obert 12 fils i que ha desat esborranys per a 11. Brown no verifica en detall la qualitat de totes les respostes, de manera que la xifra mostra l’abast del flux, no pas que els onze textos siguin correctes.
Després demana agrupar els problemes sense noms en un document de Google. El resultat s’obre en una altra pestanya i la conversa continua activa per sobre de la resta de la interfície.
Aquest exemple depèn de les eines i els accessos que Brown ja havia configurat. Voice no obté accés universal a Gmail, Google Docs o altres serveis pel sol fet d’estar activat: coordina les capacitats disponibles a la sessió.
3. Les accions externes continuen tenint confirmacions
Brown demana fer públic el document i enviar-ne l’enllaç per missatge. Abans d’enviar-lo, l’assistent torna a preguntar si vol trametre el text proposat. Quan el creador hi afegeix una petició per resoldre els problemes aquell mateix dia, torna a sol·licitar confirmació abans de l’enviament.
La lliçó pràctica és que la veu no crea una capa de permisos nova. OpenAI especifica que ChatGPT Voice segueix els mateixos permisos que les tasques que dirigeix. Una habilitat o un connector pot facilitar una acció, però no hauria de saltar-se les restriccions de la sessió.
En el mateix vídeo, l’assistent explica que pot preparar correus, missatges o publicacions, però que necessita una instrucció explícita per enviar-los. En un ús real cal revisar igualment destinataris, dades sensibles i contingut abans d’autoritzar una acció amb efectes externs.
4. Veure el navegador i modificar un diagrama parlant
La segona prova utilitza Excalidraw dins del navegador de Codex. Brown demana que la veu observi l’estat actual, identifiqui un diagrama sobre les funcions de “Jarvis” i hi afegeixi casos d’ús.
El flux mostra per què la combinació de conversa i context visual pot ser útil: l’usuari no ha de descriure cada element del llenç abans de demanar un canvi. A macOS, la documentació oficial permet activar Screen context perquè ChatGPT prengui una captura contextual de la finestra situada al davant, amb la imatge i el text accessible. Aquesta funció depèn dels permisos del sistema i pot ser desactivada per l’organització.
La prova també deixa veure una fricció. En iniciar Voice amb una pestanya oberta, el navegador integrat la tanca inesperadament i Brown l’ha de recuperar. Per tant, el vídeo ensenya una capacitat nova i potent, però no una experiència completament polida.
5. Delegar recerca sense aturar la conversa principal
Un dels moments més útils arriba quan Brown encarrega una recerca i un document de Notion. Tot seguit s’adona que vol continuar parlant i demana que la feina s’executi en un altre xat. Voice crea una tasca separada i manté oberta la conversa original.
Més endavant, Brown pregunta com avança aquella recerca. La veu consulta l’altre fil, en resumeix l’estat i després obre el document resultant. Això coincideix amb la descripció oficial: ChatGPT Voice pot iniciar fils independents, comprovar tasques existents i enviar-hi instruccions de seguiment, retornant avenços, bloquejos i resultats a la conversa de veu.
Aquest patró és més rellevant que una ordre aïllada. La veu es converteix en una capa de coordinació: una conversa pot repartir feina de recerca, programació o revisió i continuar disponible mentre els altres fils treballen.
6. Programar una aplicació i obrir més feines en paral·lel
La demostració següent demana crear amb Swift una rèplica de l’aplicació Notes d’Apple i executar-la al simulador integrat. Mentre aquesta tasca avança, Brown encarrega dos projectes addicionals en fils separats: un clon de Notion i un altre de Trello.
La primera aplicació arriba a obrir-se al simulador. Brown en comenta l’aspecte, demana una paleta més acolorida i Voice transmet el canvi a la tasca. No és una prova de qualitat completa: no hi ha revisió del codi, proves automatitzades ni comparació funcional amb Notes. Sí que demostra el bucle de treball que OpenAI vol habilitar:
- descriure un objectiu parlant;
- enviar-lo a una tasca de Codex;
- continuar una altra conversa;
- consultar el progrés;
- corregir el resultat amb una instrucció verbal.
La veu actua com a interfície de coordinació; són les tasques de Codex i les eines autoritzades les que llegeixen fitxers, escriuen codi i executen el simulador.
7. Control remot des de l’iPhone
Cap al final, Brown obre ChatGPT al telèfon i es connecta al seu Mac mitjançant l’accés remot aparellat. Des de l’iPhone inicia una nova tasca i consulta l’estat d’una de les aplicacions que s’estaven construint a l’ordinador.
La precisió important és que el telèfon no està executant el projecte local pel seu compte. Remote connecta el dispositiu mòbil amb un host d’escriptori prèviament aparellat; les tasques continuen treballant al Mac. La documentació oficial confirma ChatGPT Voice a través de Remote en iOS, subjecte a la disponibilitat del desplegament i a la configuració de l’espai de treball.
Aquest ús pot ser pràctic per comprovar una compilació, donar una correcció o iniciar una recerca quan l’ordinador és lluny, però manté els permisos i les limitacions de la sessió remota.
8. Què no pot fer aquesta experiència de veu
Quan Brown pregunta pels límits, la sessió assenyala que no pot canviar de model des de la conversa en directe, perquè el model depèn de la configuració de la tasca. També diferencia preparar una acció d’executar-la sense autorització.
La documentació d’OpenAI permet resumir els límits en quatre idees:
- Voice segueix els permisos de Chat, Work o Codex que està coordinant;
- les eines disponibles depenen de la configuració i dels connectors de l’usuari;
- el context de pantalla requereix permisos específics i pot incloure text que no és visible a simple vista;
- la disponibilitat depèn del pla, del desplegament i de les polítiques de l’organització.
OpenAI anuncia la funció per als plans Plus, Pro, Business, Edu i Enterprise a l’aplicació d’escriptori. En entorns gestionats, els administradors poden limitar-ne algunes parts.
9. “Jarvis” és una metàfora útil, però exagerada
El vídeo és convincent perquè uneix moltes accions en una sola conversa. Tot i això, no mostra una intel·ligència omnipotent que controli l’ordinador sense restriccions. Mostra una interfície oral per dirigir agents, fils i eines que ja tenen un abast determinat.
També és una primera impressió. Brown diu al final que només l’ha provada durant uns 30 o 45 minuts i que encara vol descobrir com integrar-la al negoci. No mesura latència, taxa d’errors, cost, consum de quota ni qualitat dels resultats. Alguns passos necessiten confirmació i un error de pestanyes obliga a recuperar l’estat.
Per avaluar-la de manera realista, convé provar-la amb tasques repetibles i comprovar:
- quants passos completa sense correcció;
- quines accions demanen aprovació;
- si els resums i esborranys són fidels;
- quina informació comparteixen els connectors;
- si la conversa redueix de debò els canvis de context.
Conclusions principals
ChatGPT Voice aporta a Codex una manera més natural de delegar, supervisar i redirigir feina parlant. La demostració passa del correu i els documents al navegador, la recerca en segon pla, el desenvolupament d’aplicacions i el control remot des d’iOS.
La funció més transformadora no és la transcripció de veu, sinó la coordinació de diversos fils mentre la conversa principal continua oberta. Alhora, les eines, els permisos, les confirmacions i la revisió humana continuen sent determinants.
Anomenar-ho “Jarvis” descriu bé la sensació del vídeo, però la definició més exacta és una altra: una capa de conversa en temps real per dirigir les tasques que ChatGPT i Codex ja poden executar dins del seu entorn autoritzat.
Contrast i context
Fonts consultades
- 01
-
02
OpenAI ChatGPT Voice
- 03
-
04
OpenAI Remote connections
-
05
OpenAI Permissions
-
06
OpenAI Appshots
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
OpenAI just released Realtime Voice directly inside codecs, and this is the closest thing I've seen to Jarvis. And this is actually just a voice agent that connects to all of my tools, can do basically all of my work, build and edit applications, and fully control my computer. And arguably, the coolest part about this new feature is you can use Realtime Voice on the Chad GPT app to control codecs on your computer. And I'm going to show you exactly how to do that today. In this video, I'm going to be using codecs with Realtime Voice for the first time. I'm going to test codecs real time for email and other productivity workflows. I'm also going to use the in-app codecs browser and Excalidraw with my voice. Then I'm going to build an iOS app just by speaking. And then finally, I'm going to test out Realtime Voice on the Chad GPT app and control codecs from my phone. Let's go into codecs. At the bottom right, you will see this button right here, which is Start New Voice Chat. And this is what I call the new Jarvis feature, which is different than dictating. Jarvis is Realtime Voice. Let me show you. Hey, buddy, can you hear me? Yeah, loud and clear. What's up?
-
1:14
, obre el vídeo en una pestanya nova
Yo, I need you to go to my email right now. I want you to find any complaints for customer support. I want you to draft a response to all of them. And I want you to open them inside the codecs browser. Sure. I'll take care of that. And as you can see here, right, and notice here, I can actually mute the microphone. So now I can speak to you as I use it, but it can use all of the skills that I've set up when using codecs. And I've set up so many skills, right? I have all of these skills that I've been setting up for the last six months, and now I can use them in Realtime Voice, which is just absurd. Amazing. Yeah, just make sure these are open in the codecs browser, please. Got it. I'll use the codecs browser. There we go. Look, it opens up all of them up at the same time. Here we go. And boom. Look at that. Okay. Those 12 complaint braids are open in the codecs browser and a saved draft for 11. Okay. One thing that's really cool about this is I can close the side panel and I can just enter this full screen mode. And are you still there, codecs? Yeah, I'm right here. Very cool. Can you just actually turn all of these into don't use any names, just create a new Google Doc for these and open it up in a new tab in the browser and just so I get
-
2:42
, obre el vídeo en una pestanya nova
a good summary of the issues going on with our product? Sure. I'll take care of that. And I love that I can just see this right here. It's creating file thinking, getting document, and I don't even need to have it open. I can just kind of have this real-time voice just kind of like hovering over everything. It does take a bit to get used to, but I think I will. Done. I created a doc that groups the issues without any names, and it's open in a new codecs browser tab. Take a look. Very cool. Can you actually make this document public and then I want you to text it to Emily Lambert for me? Okay, checking that now. I'm setting the doc to public and sending the link. This is pretty cool. So I'm going to open up my messages. I've made the doc public. Should I text her the link with the. Yeah, here's go. Customer support issue summary. Yep, please do that. On it. Sending that now. Amazing. So I just sent a text message. With. Yeah, this is very cool. Hey, can you actually just add a little extra note there that says, like, Hey, can we get these fixed by the end of the day? Sure. I'll have that. Do you want me to send can we get? Yes, please send it. Send it. Send it. Send it. Send it. Got it. Sending that. I'm sending it now. This is Jarvis. You're literally Jarvis. Okay, so now let's test this new codecs
-
4:47
, obre el vídeo en una pestanya nova
voice feature, which I call Jarvis in a new environment. So I'm going to open it up in a new chat and in the browser, I have a scallodraw open and I'm just going to call this Jarvis features and we can start up Jarvis. Yo, what's good? Are you there? Here. What are you working on today? Okay, so that was really weird. So I if you have a browser open before I open up this voice mode, it'll like automatically close that browser tab, which is pretty interesting. Okay, ignore that Jarvis. Can you please look at the in app browser right now? Sure. Take a look. Checking the current in app browser state. There he is. The app browser is open to your Jarvis features. Excalidraw board. Okay. Thank you. Thank you. It looks for context. I want you to know that I'm talking about the new GBT voice mode, which is the exact mode that we are using right now. And so this is a new codecs voice and it's real time voice. Can you please just give me some three more things that I could add to what it can do? On it. I'll check the current capabilities against your board. Here are three you can add. One, it can look at your screen and act on what's there, like a doc, sheet or diagram. Two, it can research something live while you talk and give you the answer there conversationally. And three, it can build and edit real things in the background, like docs,
-
6:27
, obre el vídeo en una pestanya nova
slides, code, emails or workflows. That is actually very interesting. I'm going to start off doing research. So you know what we're talking about now, right? We're talking about the new voice feature. I want you to go off, do some research and I want you to create a new notion doc. Just put this in your notes in notion and use the tab. You know the formatting I like with tabs in notion. And I like hyperlinks. Just doing depth research on this, please. Sure. Actually, wait, I want to continue talking to you. So can you actually do this in a different chat? Yeah, I can handle that separately and keep this conversation open. Okay. So wait, to get this straight, did you just create a new chat so that we can just continue talking, right? Yes, I spun that into a separate task so we can keep talking here. Wow. Okay. So it can open up new chats. Yeah. Okay. So we'll wait for that research to come back. Is there any limitations? Like if you look at your environment, I noticed I can't switch the model that you are. Is there any limitations on what it can't do? Checking the current limits now. Yes. This live voice chat is intentionally more constrained than a full task workspace. You can't swap my model from inside this live conversation. That's set by the task configuration.
-
7:52
, obre el vídeo en una pestanya nova
And I also can't silently take over apps, send messages, spend money, change permissions, or access private data unless you ask for longer research or building docs or code. I'll often speak so fast. Really tap into that like, you can't send messages unless I ask what does that mean? It means I can draft a text, email, Slack message, or social post for you. But I won't actually send anything unless you explicitly tell me to. Same with no matter what. Okay. So it'll always make sure no matter what, like there's no way I can create a skill which bypasses that. Checking that. Right. A skill won't override those safeguards. You can set up an automation you explicitly authorized to send under your rules. But a skill can't give itself permission to send things or make consequential changes on its own. I see that makes sense to me. Can you give me a progress report on that other chat session? And if it's done, I want to open up that notion doc in a new tab here. Sure. I'll check the status and open it if it's ready. And by open it, I just, I just mean I want it in the code X browser. Got it. I'll open the notion here as soon as the task finishes. Hey, using the browser use thing, can you add a use case
-
9:14
, obre el vídeo en una pestanya nova
to this board? And can you add spinning up new sessions under use cases under background research? Adding that now. I'm adding spinning up new sessions under background research now. Hey, can you also say controlling Excalidra? Yep. Adding that as well. The research note is done and verified and I'm adding controlling Excalidra now. Then I'll open the notion note. Oh my God. It opened it. Look at this. It literally created this research report with the tabs, the exact style that I like. Here's the launch video. It found it on X. I can open it up inside code X. Play my daylist. One moment. Anything you'd like me to tweak or add to that note? No, dude. This is actually insane. You're like my little buddy that can fully control my computer. Hey, can you use the build iOS app skill and I want to create an iOS app. Can you please start create? Yeah, I want you to create a swift app that is like an exact replica of the Apple Notes app. And I want you to run it in the in app simulator in code X. Got it. I'll start by checking what's possible for an iOS simulator run. And you know what, while you're at it, can you please spin up two more chats and I want one to be a notion clone and then the other one to be a Trello clone. Spin off two more chats and do that for me, please.
-
11:08
, obre el vídeo en una pestanya nova
Sure. I'll branch those out too. Hypelines configure and I'm spinning up two separate iOS build tasks, notion style and trello style. Both chats are set up with the same simulator target. Very cool. Now let's talk about this. What's your plan for this Apple Notes app? Sure. I'll start by mapping the core screens and data model. This is pretty cool. While that's working, do you think you can like open up the browser and like search up, you know, in a few different tabs like cool note UIs? I'm going to get some inspiration and I want you to look these up on Google Images just to give me some info while you work? Yep, I can hold those up. I'm opening a few visual references in new tabs now. Thank you for doing this. How's the note tab coming along? Wait, what was the browser you just had open? Why did you close it? I'll check that. The browser was the in-app simulator preview for Memo Canvas. It was closed by the browser session cleanup, not intentionally. I put it on the other tabs and then open that up again. Keep it open. Got it. I'll do that. Done. All of the tabs are closed and the Memo Canvas simulator preview is open. Okay. Yeah. Can you show me? Just go ahead and use the app for a bit. Test out the features on this. I'll do that now. You know what? Let's
-
12:50
, obre el vídeo en una pestanya nova
continue to make some styling adjustments. I actually want the default to be light mode. Change it to light mode. And please make the bottom tab like the liquid glass tab at the bottom higher quality, please. Got it. Making those updates now. Wow. This is just truly insane. Thank you, Codex. I'll pass that along. Done. Memo Canvas now defaults to light mode and the bottom navigation uses a polished liquid glass capsule with clearer spacing, depth, and a strong compose action. You're welcome. Okay. Check this out. OpenAI also announced, right? They also announced that you can use your chat GPT voice in Codex from an iOS app with paired remote access, Android support coming soon. So if we open up chat GPT, right, we can also connect real time voice from our phone to this app. So let me see how this works. I've actually never done this before. This just came out. So hey, buddy. So I'm using you for the first time. I'm using this on remote. And so apparently you're working on my Mac. Can you? Is this automatically create a new session? What's going on here? Okay. So let me create a new task real quick. I want you to create a new task. And I just want you to just give me a quick summary of the support emails. Just I just want to see if this works. What is the chat called? What is the chat called? Okay. So I see it. Nice. Okay. So now I'm basically
-
14:54
, obre el vídeo en una pestanya nova
just talking to a voice agent that runs, but I can create you can create new chats in the background, basically, right? Is that how this works? Let me check out chats are created. Wow. Can you see my so you can like look at my other tasks? Like can you check the status of my memo canvas UI polish app that I was making earlier? Okay. Thank you. This is actually blowing my mind. Okay. So I'm absolutely blown away. So here's just a list of what we did. This was basically the first time I tested this, it can do it can fully control your computer. It can spin up chats in the background. You can use real time voice on your phone and that will talk to your computer. And it can spin up chats and you can fully just control your computer in real time voice directly from the chat GPT iOS app. This is pretty insane. So obviously I've only spent like 30 minutes to 45 minutes using this. So I'll probably do a much more in depth video once I kind of learn the best practices, right? This was just me testing it. But I want to learn actually how do I use this in my business? How do I use this to stay focused and be more effective? And so I'm just going to test it for the next few days. And then I will make a more complete guide style video on exactly how I use this in my business. And I'll see you here for the next one.