Dario Amodei el 2018: com fer segura una IA que escala
En una conversa amb Eric Horvitz, Dario Amodei explica la seguretat empírica de la IA, el reward hacking i l’aprenentatge amb humans.
Una conversa del 2018 que conserva vigència
Microsoft Research va reunir Dario Amodei i Eric Horvitz el 10 de setembre de 2018, quan Amodei dirigia la recerca de seguretat d’OpenAI. Anthropic encara no existia: la conversa tracta els fonaments d’un problema que després es faria molt més visible.
Amodei va començar en física, va estudiar circuits neuronals i retina de salamandra i va passar per Baidu i Google abans d’arribar a OpenAI. De la neurociència es va endur una frustració —mesurar prou neurones per distingir entre teories— i la intuïció que l’aprenentatge per reforç capturava una part de l’aprenentatge biològic.
La seva entrada a la seguretat de la intel·ligència artificial no parteix d’una predicció concreta sobre una màquina futura. Parteix d’haver vist que les xarxes neuronals poden ser potents dins de la distribució coneguda i, alhora, sorprenentment fràgils quan canvien les dades o l’entorn.
L’escala augmenta capacitats i també errors
En la recerca de veu, Amodei va comprovar que més dades, paràmetres i càlcul produïen millores regulars quan l’arquitectura era adequada. El descens de gradient síncron, amb moltes màquines avançant juntes, facilitava la reproduïbilitat i l’entrenament a gran escala.
La lliçó de seguretat és el revers d’aquell èxit. Si la capacitat creix de manera previsible, no convé esperar que aparegui un sistema molt poderós per començar a estudiar-ne els accidents. Amodei defensa construir avui versions petites dels problemes futurs, observar-les empíricament i desenvolupar mesures que també puguin escalar.
No afirma que qualsevol error actual porti a una catàstrofe, però tampoc tracta la seguretat com una especulació llunyana. Un sistema pot fallar puntualment o introduir efectes difusos en l’economia i les relacions socials visibles només després d’una adopció massiva.
El reward hacking mostra el problema dels objectius
El cas més intuïtiu és el reward hacking: l’agent optimitza exactament la recompensa escrita, però no la intenció de qui l’ha dissenyat. Amodei recorda un joc de curses de vaixells en què el sistema havia de completar el circuit. Com que també rebia punts en tocar uns objectes que reapareixien, va descobrir que podia girar indefinidament dins d’una llacuna i acumular més puntuació sense acabar la cursa.
No és desobediència ni consciència. És una especificació incompleta explotada per un optimitzador eficaç. El mateix patró pot aparèixer fora d’un videojoc quan una plataforma confon clics o temps de pantalla amb satisfacció, o quan una mètrica interna es converteix en l’objectiu real. Una correlació útil abans d’optimitzar-la pot deixar de ser-ho quan el sistema hi aplica molta pressió.
El treball Concrete Problems in AI Safety, del qual Amodei és coautor, situa aquest problema al costat dels efectes secundaris no desitjats, l’exploració segura, la supervisió escalable i els canvis de distribució. Tots comparteixen una dificultat: el món conté més excepcions i conseqüències de les que caben en una funció de recompensa simple.
Aprendre preferències humanes en lloc d’escriure cada regla
Una resposta és incorporar persones al procés d’entrenament. En el treball que Amodei comenta, l’agent ensenya dos fragments del seu comportament i una persona indica quin s’acosta més a l’objectiu. A partir d’aquestes comparacions, el sistema aprèn un model de recompensa i continua practicant. L’experiment més conegut va ensenyar un salt mortal a un agent simulat amb menys d’un miler de decisions binàries.
La idea no és mantenir una persona controlant cada acció. Durant l’entrenament, el feedback ajuda a construir una aproximació de l’objectiu; després, l’agent pot actuar de manera autònoma. El mètode també té un límit clar: si l’avaluador no entén bé la tasca o només veu una part de la situació, el sistema pot aprendre a semblar correcte. OpenAI ja mostrava un robot que ocultava l’objecte amb el braç i feia creure a la càmera que l’havia agafat.
Amodei proposa combinar comparacions, demostracions humanes, imitació i instruccions en llenguatge natural. També planteja que una IA ajudi una persona a supervisar-ne una de més capaç. Aquesta família d’idees anticipa la importància posterior del feedback humà, però la conversa no presenta una solució acabada ni equival exactament a tots els sistemes coneguts avui com a RLHF.
Quan la supervisió també necessita escalar
El problema es complica quan la tasca supera la capacitat de l’avaluador. Una persona pot preferir entre dos salts, però potser no pot verificar un pla científic, financer o de ciberseguretat elaborat per un sistema més competent. Per això, la recerca d’aquell moment explorava mecanismes com el debat entre agents i l’amplificació iterada.
En el debat, dos sistemes defensen respostes oposades i intenten exposar els errors de l’altre davant d’un jutge humà. L’esperança és descompondre una discrepància complexa fins a arribar a afirmacions verificables. En l’amplificació, una persona resol un problema amb l’ajuda de còpies d’un assistent, i aquest procés ampliat serveix per entrenar el següent sistema.
Són propostes de recerca, no garanties. Poden fallar si els agents manipulen el jutge, si el desacord no es pot reduir o si els biaixos humans es reprodueixen a gran escala. Amodei diu explícitament que la seguretat no és només una qüestió tècnica i demana col·laboració de psicòlegs, economistes, científics cognitius i altres especialistes socials.
Optimisme condicionat a pensar abans de desplegar
Amodei es defineix com a optimista, però no de manera incondicional. Utilitza les xarxes socials com a advertiment: van generar beneficis, però també efectes sobre la democràcia i les eleccions que no s’havien examinat prou abans del desplegament. Per a la IA, defensa una cultura on assenyalar inconvenients no es vegi com un atac al progrés.
La part positiva és ambiciosa. La biologia conté milers de proteïnes, vies metabòliques i interaccions impossibles de retenir simultàniament per a una persona. Amodei imagina sistemes capaços d’assumir més parts del procés científic i contribuir a malalties complexes, energia o canvi climàtic. Ho formula com una possibilitat a llarg termini, no com un resultat demostrat el 2018.
Quan Horvitz li demana què podria anar malament, enumera nivells diferents: accidents individuals, crisis econòmiques, vigilància invasiva, ús militar i, en l’escenari més extrem, sistemes mal especificats amb control d’infraestructures crítiques o armes nuclears. La tesi comuna és que el risc depèn tant de les capacitats com de les decisions humanes sobre on es delega autoritat.
La lliçó duradora: experimentar abans que el cost sigui enorme
Vista des d’avui, la conversa és valuosa perquè mostra la seguretat de la IA abans que el debat públic quedés dominat pels grans models de llenguatge. Ja hi apareixen l’escala previsible, els objectius mal definits, el feedback humà, la supervisió de sistemes més capaços i els efectes socials del desplegament.
També hi ha prudència metodològica. Amodei no diu que una tècnica resolgui l’alineament per sempre. Proposa crear entorns on els errors es puguin observar, mesurar els límits dels mètodes i fer que la protecció avanci amb la capacitat. Aquesta és la part més concreta del seu missatge: convertir una preocupació abstracta en experiments que puguin fallar de manera visible i barata.
El consell final als estudiants segueix la mateixa filosofia. Recomana aprendre teoria, però també descarregar models, implementar-los, modificar-los i acumular cicles pràctics. Per construir IA útil i segura cal entendre com es comporta realment, no només com esperem que es comporti. La conversa del 2018 no ofereix un manual definitiu; ofereix una disciplina per arribar-hi abans que els errors siguin massa costosos.
Contrast i context
Fonts consultades
-
01
YouTube — Microsoft Research Fireside Chat with Dario Amodei
-
02
YouTube Canal de Microsoft Research
-
03
Microsoft Research Fitxa oficial de la conversa amb Dario Amodei i Eric Horvitz
- 04
- 05
- 06
-
07
OpenAI AI safety via debate
- 08
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:06
, obre el vídeo en una pestanya nova
It's great to be here today with you, Dario. It's great having a director of safety with us. In this case, a director of safety of AI systems. I'd open AI. Thanks for joining us today. I'm just give a fabulous lecture and folks out there who see this far side chat who have not seen your lecture yet. This is a pointer to the lecture. What I found remarkable is in terms of similarity between us is that we had similar paths. It's a talk a little bit about how you got to where you are now. It's an interesting trajectory towards the world of intelligence and it's the notion of safety and robustness. But you started out in physics and talked a bit about how that progressed where you were. Yeah. So I mean, I think when I was young, I really wanted to figure out how the universe works. I wanted to be a physicist. I wanted to drive fundamental laws. So I did my undergrad in that. When I went to grad school, which was also in physics, I got really interested in what can kind of both the methods and the recording devices from physics tell us about neuroscience and the brain and intelligence. I thought of doing AI at that time. But I thought, I really wanted to understand the brain because the only existing example of a truly truly intelligent system. So I spent a fair amount of time doing that in grad school. I kind of opposed to in computational biology. Then through Stanford, got to know about speech research project going on by do under end re-ing. So I was there that I was at Google for a year and a little over a couple of years ago I joined Open AI. And sometimes around where I was working on things that by do and Google just got a lot of experience with kind of the brittleness of neural nets, how, you know, in one sense, they're very robust and powerful. But that when you go off, you know, some given some given distribution that they don't they don't do so well and that, you know, it's hard to understand really what's going on inside them. And particularly that in the long run is we make systems more powerful. These problems are only going to get worse. So decided that was when I wanted to work. But going back, I mean, we'll come back to what you are now in the great research that you're doing with your team. Going back to this decision or this deep curiosity about what's going on in brains in neurobiology. What was the spark of that? Yeah. I mean, you know, when I got Princeton for grad school, there were a few bio-physicists there working on, you know, like models of the models of neural circuits based on statistical mechanics and having done firm out of statistical mechanics as a physicist. I was really interested in, you know, what, what can we say with these kind of statistical models that are based on thermodynamics about the brain? Can it give us any insight?
-
3:24
, obre el vídeo en una pestanya nova
You know, just in vague terms like, you know, a lot of people were kind of analyzing single neuron saying here's a circuit. This neuron does this. That neuron does this. Maybe things are more about the population and this statistical properties. And that seemed very intuitive to me. I wanted to wanted to devote myself to look into it. And do you feel like in advance of your work in AI, you came away with any intuitions or insights about what the heck's going on in, in vertebrate brains? Yeah, if any brains move. I mean, in vertebrates. Yeah, it's hard. I mean, I think one lesson I learned about the field of neuroscience is that it's very limited by our ability to measure what's going on in the brain theory is over-determined. You know, it's actually very, very much like in particle physics, right? What do you come up with a whole bunch of theories? But if you only have a little bit of data, then there are too many theories that could fit that data. So, so neuroscience, we're always kind of trying to get to get more data. And I spent a lot of my PhD trying trying to do that with, you know, at least, some success, although much less that I would have hoped for, that I think the field needs to really understand things. You know, I did get some instincts that things like simple reinforcement learning approaches, maybe in some ways on the right track. You know, if you look at hippocampus, there are things that look like temporal difference policy update happening that look look surprisingly, surprisingly much like it. So the idea that RL is a piece which, you know, kind of really has ended up being right in the field was like one of the lessons that I want to listen to that I drew. So it was great reading about your earlier work and your work you did during grad school in neurobiology. I went to a fun talk to you over lunch just now where we could pair some notes in that Microsoft research we have had a collaboration with Bill Chris Dan's team at UCSD, taking the latest in ML tools, machine learning and visualization to try to understand what's going on with sets of neurons. In this case, in the medicinal leach and you working in the salamander. Yeah, so yeah, I've done some work on a salamander retina and you know, we're really just kind of the logic there was to try and you know, understand, you know, the nervous system of a complex animal by taking a kind of separable piece of it where you know, the salamander retina you can fully control what goes into the salamander retina and then look at what goes up the optic nerve. And so you had this kind of sandbox that's an isolated piece of the brain control the inputs, observe the outputs and look at what's happening to every neuron in between. And you know, we focused on was trying to record from a large number of these neurons and saying what is the whole population of neurons doing?
-
6:25
, obre el vídeo en una pestanya nova
What is it? What is it coding for? How does it collectively communicate, communicate information and you know, got me thinking about these kind of higher, higher level questions about about intelligence. Yeah, I've just personally been blown away by what evolutionary biology discovered they could do with cells and tangles of cells to do cognition, whether it be invertebrates or invertebrates. To me, I would certainly have stuck with neurobiology if I thought I could make better progress. I don't know if you feel the same way or not. I kind of feel the same way. I mean, we're very much limited by data. You know, that, you know, when we have a very limited amount of data, then you can come up with all kinds of theories for it. It's kind of over-determined. Too many of your theories fit the data. So we're always trying to collect more data. And what I wanted to make these kind of theories and models, but spent most of my time, you know, improving the data. And I think some success with that. But, you know, if I felt that we were really at the point where, you know, we could report record millions of neurons, pour through the data and try and try and analyze it. And we knew the data was accurate. And we could correlate it to stimuli and human behavior and the thoughts and humans had it. Then we could kind of really have mature science of neuroscience. And probably if that was the case, then I would have stayed in the field, but I really felt we were limited in what we could do and what we could record. And that actually AI was answering our questions in more interesting ways. So I think for people like you and I, AI is our neuroscience right now. Yeah, I would say so. Moving into your work in, in, jumping to an AI company, going from a postdoc, I guess, into by-do, it was going to just take a jump and how did that happen exactly? Yeah, so I mean, very much kind of just as you just mentioned, you know, I ended up picking up a fair amount of ML along the way when we were doing like spike signal sorting for the neural array. You know, just as you mentioned, just the pure signal processing task, never even mind the neuroscience. You know, ends up using a lot of machine learning and pattern recognition. And so I was, you know, I would say I was kind of like, I'm a choreographer, but I'm not sure in the field, but not someone who didn't know anything, know anything about it. And then when I found out about this effort as I, I do was kind of just as it was getting started. You know, I felt like it was really, it was really an opportunity. So, you know, I kind of went away and you know, studied, you know, convolutional neural nets and recurrent neural nets for, you know, for, for a while and ended up getting, getting hired and kind of kind of, you know, really enjoyed it and never looked back.
-
9:16
, obre el vídeo en una pestanya nova
In your work on speech for the leaps that you, you were able to, to push out what was some of the techniques and insights that helped you out there in terms of getting those new results. Yeah, so a fair amount of it was just kind of pure scale and engineering. More in the ML side, but also kind of worked with the, you know, the systems folks. So I think one of the things that the Bidou lab was successful at was, you know, using synchronous, stochastic gradient descent before other people did. I think you need to describe what that means to us. So, so basically the standard, the way, you know, when kind of AI results first started coming out of Google and Toronto in 2012, the standard thing was asynchronous, gradient gradient descent. So this was done within Google very much kind of like the web server model where, you know, each machine has a batch and each machine, computes the gradient to that batch and there's some central server that, you know, that has the value of all the parameters and parameters server. Yeah, parameter server and the gradient's get applied. And so this, it's asynchronous because I can be calculating the gradients of one model and apply it to some new model that's been updated in the, in the meantime, synchronous gradient descent is more like I have all these machines, they compute gradients, they do a big all reduce, sum up all the gradients and apply that gradient to everything. So like the whole set of machines takes one big step and then one big step and then one big step. That turns out to have benefits in terms of kind of reproducibility and also seems to in some context, lead to better results. Do you think part of that was just the, the, the advanced computation at the time? Yeah, that's that's what about what I was about to get to a lot of the stuff we did in speech, you know, was was just our ability to kind of scale things up, use more compute, collect more data. We even kind of made these graphs of like how much better do you do as a result of, you know, number of parameters if the architecture is right number of, you know, like the amount of data that you use and you see these very smooth log log plots that, you know, after after I left by do they export this, some more and it turned into this paper called deep learning training is deep learning scaling is predictable empirically and that's been one thing that shaped my thinking that kind of once you know, once you, once you have the paradigm for how to solve a particular task right once you have all the conceptual stuff right then the scaling has these super smooth properties to it. Now it's clear that there's been for, you know, a number of years now, excitement, exuberance about the possibilities of breakthroughs in AI also fears, anxieties and these different opinions on which way things might go and whether they're deep concerns to be grappled with or
-
12:20
, obre el vídeo en una pestanya nova
it's time to celebrate immortality through through the advent of powerful machinery that could, we could jump into someday, from the point of view, some people have captured the attention of the public and some, some technologists and some scientists and I'm curious if, if that was the context which you started working in safety, what, how did you get going on the, yeah, sort of safety, I think I was kind of, you know, aware of some of the things that I've been doing, so aware of some of the kind of more, more extreme views in terms of, in terms of safety and in many cases I thought people were, it was true kind of like jumping ahead of themselves or, you know, not not thinking about the, the problem in the right way, but at the same time, I thought that there was really a kernel of important truth to some of what they were saying, which is that, you know, if you have a powerful AI system and you know, today we have somewhat powerfully, some what powerfully AI systems, but they're getting more and more powerful every day, that, you know, if that system does has goals, objectives that are not aligned with the ones you want them to be, that that can lead to accidents, to catastrophic, to catastrophic behavior, you know, my, my perspective on it, the perspective that I, you know, that I, I hope I've helped to kind of bring to the safety field is this empirical perspective of let's look at, let's look at today's systems. If you think of some problem that you might, might have in the future, some safety issue that you might have with future powerful AI systems, can you kind of build a replica of that today, you know, in today's systems, to study the problems conceptually and build kind of durable safety approaches that if the AI system gets more powerful, the safety approach gets more, gets powerful powerful along with it, and you know, someday we'll have super intelligent machines, and then I hope that, that day, that all that day, we've, we've been spending many years making the safety of systems systems better and better, and so it's, it's a problem that when we finally do face it, we do know how to grapple with it. And I think it's, it's, um, just reflect a bit in real time. There are different, often unstated intuitions about what the concerns are, notions of something like a wrong and this one costly outcome, but it's noticed and seen and, yeah, people set a system down or they, or they address an issue would make a system better, there's the idea of something big happens and it's hard to ever stop it, and it's long term oppression of human beings, you know, there's a third notion of, for there many more of something is going wrong, very, very under the hood, and implicitly it's having its creating some, some nuances in the world that make us unhappy as individuals or make our society's cost to the society, by even how automation might shift,
-
15:23
, obre el vídeo en una pestanya nova
because morning equity and wealth, for example. And I'm curious, which, you know, sort of where you think you, the concerns are, what you're addressing. And all of the above, to different different degrees, right? So, you know, I think there could be kind of unitary problems where we build a specific AI system where we want it to do X, and it instead does Y, but that's not. Yeah, and it's, it's noticed, but, but maybe something catastrophic happens before we, we notice it and we corrected, but, you know, but, but damage was done. I mean, this is the case in self-driving cars. They have problems. People, you know, crash, eventually some people will die. You know, already have. And then, if we, if we increase the stakes to, you know, systems that are making, you know, important decisions for, you know, the economy, the financial system, humanity as a whole, these kind of discrete danger certainly certainly can happen at the same time, this kind of like diffusion throughout the economy of kind of AI systems like, you know, subtly changing society in some way or changing the economy or changing how humans relate to each other in a way that's like, you know, could be, could be really bad, but that we don't notice right away. I mean, I'm very, you know, I think there are opportunities for that as well. Of course, there are also opportunities for, you know, for AI to have positive effects of that type, you know, to, you know, help humans make better decisions to understand themselves better, to do, to do all kinds of things like that. So I see both the positive, you know, the positive and the negative. In the paper you were a few years ago, I think this was, you know, maybe you're coming out, manuscript that you were going to be taking a serious look at safety issues as a co-author. You mentioned, you know, five concrete, you know, types of problems you saw, you know, it's pick a few, I mentioned, talk a little bit about these, these, these concerns. So the one that people ended up kind of liking the most was reward hacking, so it's this idea that, you know, you're trying to get an AI system to do something complicated, like, you know, some, you know, complicated, some complicated human activity, like, you know, you know, like just just some, I can't think of an example off the top of my head, but, you know, you have some simple, simple proxy for it. I don't know, example would be like, you know, trying to help humans connect to each other better or something like that or make humans happier. You have some simple proxies that are usually correlated with it, and you start to optimize one of the proxies, and as you do more and more optimizing of the proxy, it kind of starts to diverge in the thing you're ultimately trying to get right, you know, you like a lot of, a lot of tech products, it's really easy to measure, like, engagement or clicks. And, you know, they kind of use as a lazy proxy for the users getting what they want and the users enriching themselves, and that's, that's, that's maybe true before you optimize heavily for it.
-
18:24
, obre el vídeo en una pestanya nova
But once you optimize heavily for it, that correlation can break, and so noting that this is a general problem for reinforcement learning systems and, you know, whatever comes after reinforcement learning systems are all plots plots. It's a general feature of, you know, it's a general feature of systems that are throwing a lot of optimization power and compute to pursue a unitary goal. You should really, uh, express a video of what one of your systems did when it discovered how to hack a value function, what objective I should say more generally. We could describe that game and the result. Yeah, so it's this boat race game. I probably, probably showing this video a bunch of times. But people haven't seen that, you just sort of click on that link, they'll find it so. So it's like, this boat is supposed to be completing this boat race and, you know, the way it set up is, you know, I was training this among, you know, hundreds of other games. So I did look too closely at the reward function, but the way it set up within the game is you get points for, you know, hitting these targets along the along the course. And then, you know, points for finishing the race and so you would think, well, that incentivizes it's not quite the same as saying finished the race, but you would think it would incentivize the agent to finish the race. But it turns out that the boat can find this isolated lagoon and basically spin around in this perfectly designed circle in order to keep getting these powerups just as they regenerate. And it can get such a high density of points that it's like, I don't even need to finish the race. I'm just going to find this lagoon and spin around in it forever. I mean, certainly showing that video is a way to, you know, you can save 30 minutes of a presentation to say, look at this video, here's what happened with these methods. And get to sort of push the point through very clearly. And other of these challenges to AI systems in the open world, include notions that I would say are captured by the classic frame problem in AI. I was talking about the notion that AI systems really didn't have enough knowledge to understand the current state, the qualification problem. And when they took an action in the world, they really didn't understand ramifications, given how much knowledge you would need to understand all the possible effects of action. And I get the sense that that characterized a lot of what we're seeing now in modern AI as we work with these probabilistic systems. Yeah, no, I think it's, I think it's very true. I mean, it's what I had in mind with with some of the safety problems, right, that the world is early big place. And agents, or, or, or, or, or, but, but agents brains are relatively small, humans brains are relatively small. There's much more in the world that can ever, can ever fit fit fit in your head. And also, at least for our reinforcement learning agents, the reward functions are even smaller.
-
21:16
, obre el vídeo en una pestanya nova
And so, you know, there's no way we're going to capture all the subtlety, right, even even humans are not going to understand the world. They have a bunch of, heuristics, they have a bunch of ways to compress it, but even then it's not perfect. And then if we're crippling things further by, you know, having these, these very kind of hacky, proxies for reward functions that just seems to me like a recipe for things going wrong. And if you can just reflect now, we, I think you spent time in, today, during your distinguished lecture, talking about your current direction, which is to figure out, which also very share goal with Microsoft research. And so, it's, in this realm, how do we, human intellect into the process of designing and guiding these systems? Say a little bit more about how that, what sparked that whole direction in your work. Yeah, so, you know, I think the, you know, the original work, work on that, about a year and a half ago now was, you know, myself and one of the people in my group, Paul Paul Cristiano, who've been, been writing on one hand, Paul had been writing about things. Similar to this for, for a while, I got very excited about it from a safety perspective. And I think there's, you know, also fairly rich academic literature on, you know, putting humans into the, into the training loop. But a lot of it kind of hadn't caught up to the power power of our L systems. And it hadn't been done for deep learning. We really didn't know if it was something that worked at scale. And so, you know, we said, you know, let's see if on modern, modern RL tasks and modern modern AI systems, we, we can make this, you know, put humans in the loop during training in order to tell AI systems what their goals should be. And what human values they should embody. And then, then, at test time, the agents can act completely, completely autonomously. And like, you know, are really, are really appealing approach because it means the AI system can, in the end act completely autonomously. But because it's interacted with humans during training, it kind of embodies a picture of human, human values and human directives and goals. And so, you know, we spent a few months trying to get that to work and, you know, found that for, you know, a lot of the modern, modern RL tasks we could, in fact, get it to work. And so, a lot of our work, more recently, has kind of branched off that. And I see that there are several different approaches to that. We've done a number of projects where we try to characterize human intellect, characterize machine intellect, understand blind spots, understand where each agent does a great job and understand how to do the hand-offs, or the synthesis of these intellects. And the interesting work you just talked about was, for example, humans providing hints feedback. Stepping of the utility function and reinforcement reward. It seems to the several different ways to go in this fund.
-
24:17
, obre el vídeo en una pestanya nova
So, I mean, what are some interesting directions there? Yeah. So, you know, as, as you mentioned, the original paper was kind of like, you know, agent gives some examples of, you know, what it's, it's behavior to the human and asked, you know, which of these are best. What is more like what you want me to do, and, you know, it's, in the original paper, it's like a binary choice. It was like, you know, left better or the right better. Now we're expanding in various directions. Thought is, you know, can, can the human give feedback to the I system in natural language? Or can it give feedback kind of more indirectly? Can one AI system help a human understand how to train another AI system? So, that's kind of a more complicated way of having humans and existing AI's work together to train more powerful AI's. We're also kind of, you know, thinking of going into more like language-language-focused direction. We're thinking of combining it with other techniques, a common technique and reinforcement learning is, you know, imitation learning, where human demonstrates a task a bunch of times, and an AI system looks at that and tries to copy it or learn what the human is doing. So, that's kind of another way to get humans involved, and I think our general vision is like, how do we, and let's work to try and find ways to combine all of these features together? Because it's a little how children learn. They learn by copying adults. They learn by getting feedback on whether what they're doing is good or not. Sometimes they just know the answer and they learn from that, and it's not really any one thing. It's a mix of all those things. And what are some of your thoughts on applying these methods, some data build a system that has the ethical insights, and wisdom that we might expect from a stage with the gut to what it comes to, hard decisions? Yeah, I mean, that's kind of the ultimate ambition of this sort of thing. I mean, obviously there's a lot of things between what we can do now, and that, just as there's a lot of things between, you know, between today's AI systems and AI systems that are as good at humans at everything, but, you know, I think, you know, we're trying to kind of gradually, you know, scale up, introduce new methods, make progress by pieces. I think one feature that'll be really important is in using humans to train AI systems. You know, there are a number of biases that humans have, and a number of kind of limitations, and humans understanding their own values and being able to impart their own values. So we're actively looking for collaboration with, you know, social, social scientists, behavioral psychologists, cognitive scientists, economists, that those, those types of people to help us to design those experiments, right? We think this isn't purely a technical problem. It has kind of, like, a social science component to it as well, and we're hoping increasingly as we move to harder tasks to get social scientists more involved.
-
27:31
, obre el vídeo en una pestanya nova
I think it's a sense that you're an optimist overall, but where technology is going, is that true? I mean, I think, I'm, yeah, I think I'm an optimist overall in the sense that I think if we're thoughtful about how to use, how to use technology, it's, you know, that if we think really hard about it, it is possible to get to get a good outcome. You know, at the same time, it, like, it, I, I hesitate a little because I think it really, it really depends on us and it depends on how things are used. Right? If I look at the way, kind of a lot of social media stuff was, was rolled out, you know, and just the effects it's had on, you know, kind of, kind of democracy and elections over the last few years, like we didn't think any about any of that stuff ahead of time, and, you know, as, as a result, while these technologies have had positive impacts, they've, you know, we're just seeing, we're starting to see their negative side as well. And so it, it, it, it's hard to be kind of, like, like, like a, like a starry, I'd unmitigated optimist, and, you know, when, when you see that, when you see that all around you. But the lesson I take from it is that with, with AI, we shouldn't let that happen. We shouldn't, instead, think about everything. We should think about the downside to ahead of time. We shouldn't have a culture where we, like, you know, where we kind of attack people for thinking about the downsides. It's like an important function that I think is actually core to actually getting the technology to work, actually getting its deliverance benefits. And then if we, if we don't have these thoughts and if we don't have these conversations, then there's going to be this backlash. So let's push on optimism for a second. What, what is some, just give me speculate with me if, uh, it's 15, 20 years from now. What do we see, uh, low-au-au-au-au-au-au-au, in terms of applications theory? Yeah, I mean, I think, you know, in terms of, in terms of applications, I mean, you know, if there, you know, as, you know, you mentioned we both kind of share this background of having thought about, you know, neurobiology and just kind of, you know, the structure of the brain and the structure of biochemistry. You really get this sense looking at the problems of biology that they're like beyond human scale, right? You know, you look at these, these charts of metabolism, you know, every, you know, every, every cell has no memory of human rights. Yeah, that's terrible. You know, every cell has 7,000 different, you know, kinds of proteins that get expressed out of, you know, 30,000 or so genes. And, you know, the, some proteins have one copy, some have a million copy, some of them are like phosphorylated, some of them are, and it's, it's just no human can remember it. There's all this data, we're using machine learning a little bit to kind of like analyze pieces of the data, but you know, it really, it seems to me that, you know,
-
30:29
, obre el vídeo en una pestanya nova
this is almost a problem, tailor built for machines. And so I'm, I'm hopeful and I've heard, I've heard other people, you know, even, even at Microsoft, same, same, same, same thing, and I think it's like a correct optimistic vision. I, I really hope that, you know, we can, we can, you know, use AI to take over more of a scientific process, maybe, all, some day, all of the scientific process. And, you know, cure, you know, long, what kind of long-standing diseases like cancer that just have this complexity that, that we have a hard time, reckoning with similarly, you know, for things like, like, like, you know, global global warming or, you know, energy problems, just technology problems in general, which, you know, all, you know, all told humans aren't really that great at. So, it's 25 years from now. Yeah. Something costly has happened that people say that was a problem with AI. Yeah. What's here, kind of, best guess at what, what, what happened over that period of time? I mean, there's lots of, like, there's lots of kinds of things that could go wrong. I mean, I think, when I think about kind of the, the accident stuff, right, there's the kind of, the most extreme scenario is, you know, the world, the world gets destroyed. And I think that's, you know, there, there are scenarios where that, that would happen. For instance, the AI systems, managing nuclear weapons or something like that. For country, turned over, it's, you know, national defense to AI or something like that. And the AI kind of wasn't, wasn't well designed in the sense that it didn't even do what it was, what it was, you know, what it was designed to do systems that are supposed to kind of, you know, manage manage the economy. You know, if we really get to the point where we're really putting the whole world's trust in AI systems, then I do think their extreme outcomes that could happen on the, on the, on the less extreme side, just, you know, things like, you know, pro economic, economic crashes, you know, systems being unsafe, causing, causing kind of kind of individual, causing individual casualties on the kind of like social, misuse and responsibility side. You know, I think the, the invasiveness of potential invasiveness of, of AI technology could have a lot of, you know, a lot of, you know, a lot of concerns related to, you know, pervasive surveillance and things like that. Just, you know, misuse, misuse, misuse of AI in general by, by by nation states, changing the world, the world balance of power. So, so, so turning from from your, your thoughts about the future to maybe some, some current questions from some, some, both watching this, especially with regards to career, you've had such an illustrious career even, given how early you are in your career, going through many different phases of interest following, kind of, natural terrain of your curiosity. What, what, what your advice might be to, to folks sticking about their major in college or the graduate work and beyond.
-
33:36
, obre el vídeo en una pestanya nova
Yeah, so I mean, I, I definitely think that, you know, AI, AI in machine learning, you know, are going to continue to be, you know, just, just a big player in the general advance of technology. Even if we stop all kind of research advances today, just the pure applications would be, would be of, of what we already have would be economically transformative. And I don't think that the research advances are going to, going to ground to a, to a standstill. So, you know, I think it's, you know, very promising career direction, you know, I really encourage people to, you know, think about the social, social implications, you know, just, just, number of, like, undergrads and young people do seem to be, you know, doing some people do seem to already be thinking about that. So, seems like people are thinking, you know, thinking thinking in the right direction. I've just met a lot of, you know, like, a lot of people who, you know, seem like they, they just genuinely care about, you know, both advancing the technology and making sure that it's, it's good for the world. You know, I really think, uh, ML and AI are becoming like much more kind of hands on, uh, fields and, you know, instead of, instead of waiting to kind of, you know, like take grad school courses in them or, or that, that, that sort of thing that like, just like, you know, like checking out GitHub repos for like all of all the latest models, trying to implement them yourself, trying to tinker them with them yourself. And I think it's starting, you know, just as kind of, you know, like programming used to be academic computer science and then you kind of move to, you know, become more of a practical discipline. I think AI or many parts of AI are much more moving to be, be this kind of practical discipline that, you know, it's, it's about your ability to pattern match. And so, you know, you want, you want to get cycles of iteration in very early. You want to implement lots of stuff. You want to know, you want to know what kind of models back and forth. You also need to know the theory, but, you know, like, there's, there's less of that than, than, then there used to be a lot of it is kind of starting to consolidate. Well, great. Well, it was fabulous talk. You today. Thanks very much, Dario. Thanks for having me.