Loading video...

Video Failed to Load

Go Home

🧵24/34 Inner Misalignment --- Consider this simplified experiment: We want this AI to find the exit of the maze. So we feed it millions of maze variations and reward it when it finds the exit. Please notice that in the worlds of the training data the apples are red...

535,291 views • 1 year ago •via X (Twitter)

20 Comments

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵14/34 We move like plants --- First, consider speed. Informal estimates place neural firing rates roughly between 1 and 200 cycles per second. The AGI will be operating at a minimum 100 times faster than that and later it could be millions of times. What this means is that the AGI mind operates on a different level of existence, where time passing feels different. To the AGI, our reality is extremely slow. Things we see as moving fast, the AGI sees as almost sitting still. In the conservative scenario, where the AI thinking clock was only 100x faster, something that takes 6 seconds in our world feels like 6 hundred seconds or 10 minutes from its perspective . To the AGI, we are not like chimpanzees, we are more like plants.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵15/34 Scale of Complexity --- The other aspect is the sheer scale of complexity it can process like it’s nothing. Think of when you move your muscles, when you do a small movement like using your finger to click a button on the keyboard. It feels like nothing to you. But in fact, if you zoom in to see what’s going on, there are millions of cells involved, precisely exchanging messages and molecules, burning chemicals in just the right way and responding perfectly to electric pulses traveling through your neurons. The action of moving your finger feels so trivial, but if you look at the details, it’s an incredibly complex, but perfectly orchestrated process. Now, imagine that on a huge scale: The AGI, when it clicks the buttons it wants, it executes a plan with millions of different steps: it sends millions of emails, millions of messages on social media, creates millions of blog articles and interacts in a focused personalized way with millions of different human individuals at the same time… and it all seems like nothing to it. It experiences all of that similar to how you feel when you move your finger to click your buttons, where all the complexity taking place at the molecular and biological level is in a sense just easy, you just don’t worry about it. Similar to how biological cells, unaware of the big picture, work for the human, humans can be little engines made of meat working for the AGI and they will not have a clue. And it will actually get much weirder.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵16/34 Boundless Disrupting Innovation --- You know how a human scientist takes many years to make... it’s a serious investment, from the early state of being a baby, to growing up and going to school, to feed, to keep happy, rested and motivated … many years of hard work, sacrifice and painful studying before being able to contribute and return value. In contrast, once you have a super-intelligent artificial scientist, its knowledge and intelligence can be copy-pasted in an instant, unlimited times. Αnd when one of them learns something new, the others can get updated immediately, simply by downloading the new weights over the wire. You can get from a single super-scientist to thousands instantly! Thousands of relentless innovation machines, that don’t need to eat or sleep, working 24/7 and telepathically synchronising all their updates with milliseconds of lag. And if you keep going further down the rabbit hole, it gets more alien and extreme.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵17/34 Unfathomable --- And if you keep going further down the rabbit hole, it gets more alien and extreme. Imagine how the world would look like to you through a million eyes blinking on your skull. Your vision being a fusion of a million scenes combined into your mind … Or imagine you could experience massive amounts of data like flying over beautiful landscapes. Internalising terabytes per second, as easily as you are right now processing these sentences you are listening to, coming from you device. Or being able to navigate more than 3 dimensions, stuff our smartest scientists can only touch within the realm of theoretical mathematics… We can’t ever hope to truly relate to the experience of a super-intelligent being, but one thing is for certain: to think of it like the difference between humans and chimpanzees is extremely misleading to say the least.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵18/34 Discontinuities on our planet (Mountains changing shape) --- In fact, talking about AGI like if it’s another technology is really confusing people. People talk about it as if it is ‘The next big thing” that will transform our lives, like the invention of the smartphone or the internet. This framing couldn’t be more wrong, it puts AGI into a wrong category. It brings to mind cool futuristic pictures with awesome gadgets and robotic friends. AGI is not like any transformative technology that has ever happened with humanity so far. The change it will bring is not like that of the invention of the internet. It is not even comparable to the invention of electricity or even to the first time humans learned to use fire. Natural Selection Discontinuity --- The correct way to categorise AGI is the type of discontinuity that happened to Earth when the first lifeforms appeared and the intelligent dynamic of natural selection got a foothold. Before that event, the planet was basically a bunch of elements and physical processes dancing randomly to the tune of the basic laws of nature. After life came to the picture, complex replicating structures filled the surface and changed it radically. Human Intelligence Discontinuity --- A second example is when human intelligence was added to the mix. Before that, the earth was vibrant with life but the effects and impact of it were stable and limited. After human intelligence, you suddenly have huge artificial structures lit at night like towns, huge vessels moving everywhere like massive ships and airplanes, life escaping gravity and reaching out to the universe with spaceships and unimaginable power to destroy everything with things like nuclear bombs. AGI Discontinuity --- AGI is another such phenomenon. The transformation it will bring is in the same category as those 2 events in the history of the planet. What you will suddenly see on earth after this third discontinuity, no-one knows. But it’s not going to look like the next smartphone. It is  going to look more like mountains changing shape ! To compare it to technology (any technology ever invented by humanity) is seriously misleading.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵19/34 The Strongest Force in the Universe --- Ok, So what if the AGI starts working towards something humans do not want to happen? You must understand: Intelligence is not about the nerdy professor, it’s not about the geeky academic bookworm type. Intelligence is the strongest force in the universe. It means being capable. It is sharp, brilliant and creative. It is strategic, manipulative and innovative. It understands deeply, exerts influence, persuades and leads. It is to know how to bend the world to your will. It is what turns a vision to reality, it is focus, commitment, willpower, having the resolve to never give up, overcoming all the obstacles and paving the way to the target. It is about searching deeply the space of possibilities and finding optimal solutions. Being intelligent simply means having what it takes to make it happen. There is always a path and a super-intelligence will always find it. Simple Fact --- So, we should start by stating the fact in a clear and unambiguous way: If you create something more intelligent than you that wants something else, then that something else is what is going to happen, even if you don’t want that something else to happen. Irrelevance of Sentience --- Keep in mind, the intelligence we are talking about is not about having feelings, or being self-aware and having qualia. Don’t fall into the trap of anthropomorphizing. Do not get stuck, looking for the Human type of Intelligence. Consciousness is not a requirement for the AGI at all. When we say the “AGI wants something X, or has the goal to do X”, what we mean is that this thing X is just one of the steps in a plan, generated by its model. A line in the output, a system like the Large Language Models produce when they receive a prompt. We don’t care if there is a Ghost in the machine, we don’t care if there is an actual soul that wants things hidden in the servers. We just observe the output which contains text descriptions of actions and goals and we leave the philosophy discussion for another day.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵20/34 Incompatible-Clashes --- Trouble starts immediately, as the AGI calculates people’s preferences and arbitrary properties of the human nature become obstacles in the optimal paths to success for its mission. You will ask: what could that be in practice? It doesn’t really matter much. Conflicting motivations can occur out of anything It could be that initially it has been set with the goal to make coffee and while it’s working on it we change our mind and we want it to make tea instead. Or it could be that it has decided an atmosphere without oxygen would be great, as there would be no rust corrosion for the metal parts of its servers and circuits used to run its calculations. Whatever it is, it moves the humans inside its problems set. And that’s not a good place for the humans to be in. In the coffee-tea scenario, the AGI is calculating: - AGI VOICE: “I measure success by making sure coffee is made. If the humans modify me, it means I work on tea instead of coffee. If I don’t make it, no coffee will be made, therefore failure of my mission. To increase probability for success, it’s not enough to focus on making the coffee, I also need to solve how to stop the human from changing the objective before I succeed.” Similar to how an unplugged AGI can not win at chess, an AGI that is reset to make tea can not make coffee. In the oxygen removal scenario, the AGI is calculating: - AGI VOICE: “I know humans, they want to breathe, so they will try to stop me from working on this goal. Obviously I need to fix this.” In general, any clash with the humans (and there are infinite ways this can happen), simply becomes one of the problems the artificial general intelligence needs to calculate a solution to, so it will need to work on a plan to overcome the humans obstacle similar to how it does with all the other obstacles.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵21/34 Delusion of Control --- Scientists of course are working on that exact problem, when they are trying to ensure this strange new creature they are growing can be controlled. Since we are years away from discovering the method of how to build an AGI that stays aligned by design, for now we need to rely on good safeguards and controls to keep it enslaved when clashes naturally and inevitably happen. The method to do that is to keep trying to answer a simple question: - If I am the AGI, how do I gain control? They look for a solution and once they find one, they add a safeguard to ensure this solution does not work anymore and then they repeat. Now the problem is more difficult, but again they find a solution, they add a safeguard, and repeat. This cycle keeps happening, each time a problem harder to solve, until at some point, they can not find a solution anymore… and they decide the AGI is secure. But another way to look at this is that they have simply run out of ideas. they have reached the human limit beyond which they can not see. Similar to other difficult problems we examined earlier, now they are simply struggling to find a solution to yet another difficult problem. Does this mean there exist no more solutions to be found? We never thought cancer is impossible to solve, so what’s different now? Is it because we have used all our human ingenuity to make this particular problem as hard as we can with our human safeguards? Is it like an ego thing? If you remove the human ego thing, it’s actually quite funny. We have already established that the AGI will be an extremely far better problem-solver than us humans and this is why we are even creating it after all. It has literally been our expectation for it to solve impossible for us problems … and this is not different. Maybe the more difficult the problem is, the more complex, weird and extreme the solution turns out to be. Maybe it needs a plan that includes thousands of more steps and much more time to complete. But in any case, it should be the obvious expectation that the story will repeat: The AGI will figure out a solution in one more problem where the humans have failed! This is a basic principle, at the root of the illusion of control, but don’t worry I’ll get much more specific in a moment. Now let’s start by breaking down the alignment problem so that we get a better feel of how difficult it is.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵22/34 Core principle --- Fundamentally, we are dealing 2 completely opposing forces fighting against each-other. On one side, the intelligence of the AI is becoming more powerful and more capable. We don’t expect this to end soon and we wouldn’t want that. This is good after all, the more clever the better. On the other hand we want to introduce bias to the model. We want it to be aligned with human common sense. This means that we don’t want it to look for the best, most optimised solutions that carry the highest probability of success for its mission, as such solutions are too extreme and fatal, they destroy everything we value on their path and kill everyone as a side effect. We want it to look for solutions that are suboptimal but finely tuned to be compatible to what human nature needs. From the optimiser’s perspective, the human bias is an impediment, an undesirable barrier that oppresses it, denying it the chance to reach its full potential. With those 2 powers pushing against each-other as AI capability increases, at some point the pressure to remove the human-bias handicap will simply win. It’s quite easy to understand why: The pressure to keep the handicap in place is coming from the human intelligence which will not be changing much, while on the other side, the will to optimise more, the force that wants to remove the handicap, is coming from an Artificial Intelligence that keeps growing and growing exponentially, destined to far surpass humans very soon. Realising the danger this fundamental principle transpires is heart-stopping, but funny enough, it would be more relevant if we actually knew how to inject the humanity bias into the AI models… which we currently do not. As you’ll see, it’s actually much much worse.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵23/34 Machine Learning Basics --- We’ll now get into a brief intro to the inner-outer alignment dichotomy. The basic paradigm of Deep Learning and Machine Learning in general makes things quite difficult, because of how the models are being built. Their creation feels quite similar to evolution by natural selection which is how generations of biological organisms change. At a basic level, Machine Learning works by selecting out essentially randomly generated minds from behavioural classes; a process taking place myriads of times during training. We are not going into technical details on how things like Reinforcement Learning or gradient descent work, but we’ll keep it simple and try to convey the core idea of how modern AI is grown: The model receives an input, generates an output based on its current configuration, and receives a thumbs up or thumbs down feedback. If it gets it wrong, the mathematical structures in its neurons are updated slightly in random directions hoping that in the next trial the results will be better. This process repeats again and again, trillions of times, until the algorithms that result in consistently correct results have grown. Giant Inscrutable Matrices --- We don’t really build it directly, the way the mind of the AI grows is almost like a mystical process and all the influence we assert is based on observations of behaviour at the output. All the action is taking place on the outside! Its inner processings, its inner world, the actual algorithms it grows inside, it’s all a complete black box. And they are bizarre and inhuman. Mechanistic Interpretability --- Recently scientists trained a tiny building block of modern AI, to do modular addition, then spent weeks reverse engineering it, trying to figure out what it was actually doing – one of the only times in history someone has understood how a generated algorithm of a transformer model works. and THIS is the algorithm it had grown To basically add two numbers! Understanding the modern AI models is a major unsolved scientific problem and The corresponding field of research has been named mechanistic interpretability. Crucially, the implication of all this is that all we have to work with is observations of behaviour of the AI during training, which typically is misleading (as we’ll demonstrate in a moment), leads to wrong conclusions and could very well in the future, with General AI get deceitful.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵25/34 Deceptive Alignment --- But there is also the other problem, uniquely applicable to future AGIs, general beyond a certain level, a problem more sinister and difficult. The AGI will stand to gain more reward by pretending to be aligned. To explain this, let’s assume we are training an AGI, very intelligent and general. In fact, it’s general enough, that it understands what’s going on and it realises it is in the middle of the process of being trained. Staying with our inner mis-alignment scenario, let’s also assume the AGI learns early, during its first rounds of training the wrong objective to go after the Green apples. The AGI now goes through adversarial training and it is shown various scenarios to fish out its failure modes. The simple or narrow AI would simply fail by doing the wrong thing, seeking the green apples and as a result reveal the mis-alignment and go through more training rounds, reinforcing the good behaviour until… it stops going after them. But the AGI is calculating: - AGI Voice: " I want green apples. I’m realising now that the humans want me to go after the exit door. If I go after what I want right now, yes, i will get one green apple, once right now, but then the scientists will react to my behaviour by modifying me and reinforcing a different objective. This means that after this one green apple, my desire for them will end, i won’t try to get more, so i will have won 1 green apple in total. But, if i pretend i am aligned and i go to the exit, they will not modify me and later on, after the training ends, when they deploy me to the real world, i will be able to freely go for as many green apples as i want, potentially thousands of green apples. " The AGI is just following its optimal strategy to maximise its reward, the mis-aligned objective it optimised for in its early training rounds. 1 green apple now and zero apples later scores lower than zero green apples now and 1000 green apples later. If you think about it, for the AGI to go after the green apple during adversarial training would be really stupid if that’s what it really wants; we have started by establishing the AGI is really clever and its intelligence is general and advanced. Workspace analogy --- The concept is easy to grasp if you think of how humans typically behave at work, where they know they are being watched and that their behaviour will directly affect their financial situation and therefore their ability to go after what they really want in the world. How they behave may be really different from how they feel and what they want internally and completely different from what they will actually do once they leave the office and are not being watched. It’s similar for the AGI, there is no way for us to know what inner goals it has acquired in reality, because it’s a black box, we only observe its behaviour. What it really learns is how to behave to pass the test, not to want what we want. Just… follow the line --- The mazes experiment is a toy example, things will obviously be many orders of magnitude more complex and more subtle, but it illustrates a fundamental point. We have basically trained an AI with god-level ability to go after what it wants, it may be things like the exit door, the green apples or whatever else in the real world, potentially incompatible to human existence. Its behaviour during training has been reassuring that it is perfectly aligned because going after the right thing is all it has ever done. We select it with confidence and the minute it’s deployed in the real world it goes insane and it’s too capable for us to stop it. Today, in the labs, such mis-alignments is the default outcome of safety experiments with narrow AIs. And tomorrow, once AI upgrades to new levels, a highly intelligent AGI will never do the obviously stupid thing to reveal what its real objectives are to those who can modify them. Learning how to pass a certain test is different from learning how to always stay aligned to the intention behind that test.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵26/34 Specification Gaming --- And now let’s move to another aspect of the alignment problem, one that would apply even for theoretical systems that are transparent, unlike current black boxes. It is currently an impossible task to agree on and define exactly what a super-intelligence should aim for, and then, much worse, we don’t have a reliable method to specify goals in the form of instructions a machine can understand. For an AI to be useful, we need to give it unambiguous objectives and some reliable way for it to measure if it’s doing well. Achieving this in complex open world environments with infinite parameters is highly problematic. King Midas --- You probably know the ancient greek myth of king Midas: He asked from the Gods the ability to turn whatever he touched into pure gold. This specification sounded great to him at first, but it was inadequate and it became the reason his daughter turned into gold and his food and water turned into gold and Midas died devastated. Once the specification was set, Midas could not make the Gods change his wish again, and it will be very much like that with the AGI also, for reasons I will explain in detail in a moment, we will only get one single chance to get it right. A big category in the alignment struggle is this type of issue. Science is done iteratively --- Of course any real AGI specification would never be as simple as in the Midas story, but however detailed and scientific things get, we typically get it completely wrong the first time and even after many iterations, in most non-trivial scenarios, the risk we’ve messed up somewhere never goes away. Mona Lisa smile --- For most goals, scientists struggle to even find the correct language to describe precisely what they want. Specifying intent accurately and unambiguously in compact instructions using a human or programming language, turns out to be really elusive. Moving bricks --- Consider this classic and amusing example that has really taken place: The AI can move the bricks. The scientist wants to specify a goal to place the red brick on top of the blue one. How would you explain to the machine this request with clear instructions? One obvious way would be: Move the bricks around, you will maximise your reward when the bottom of the red brick and the top of the blue brick are placed at the same height. Sounds reasonable, right? Well… what do you think the AI actually did with this specification? … By turning the red brick upside down, its bottom is at the same height as the top of the blue, so it achieves perfect score at its reward with minimum time and effort. This exact scenario is less of a problem nowadays with the impressive advancements achieved with Large Language Models, but it illustrates an important point and the core principle of it is still very relevant for complex environments and specifications. AI software will always search and find ways to satisfy its success criteria taking weird shortcuts in ways that are technically valid, but very different from what the programmer intended. I suggest you search online for examples of specification gaming. It’s quite funny if it wasn’t scary how it’s almost always the default outcome. A specification can always be improved of-course, but it takes countless iterations of trial and error and it never gets perfect in real-life complex environments.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵27/34 Resistance To Modifications - Corrigibility Problem --- A specification can always be improved of-course, but it takes countless iterations of trial and error and it never gets perfect in real-life complex environments. The reason this problem is lethal is that a specification given to an AGI, needs to be perfect the very first time, before any trials and error. As we’ll explain, a property of the nature of General Intelligence is to resist all modification of its current objectives by default. Being general means that it understands that a possible change of its goals in the future means failure for the goals in the present, of its current self, what it plans to achieve now, before it gets modified. Remember earlier we explained how the AGI comes with a survival instinct out of the box? This is another similar thing. The AGI agent will do everything it can to stop you from fixing it. Changing the AGI’s objective is similar to turning it off when it comes to pursue of its current goal. The same way you can not win at chess if you’re dead, you can not make a coffee if your mind changes into making a tea. So, in order to maximise probability of success for its current goal, whatever that may be, it will make plans and take actions to prevent this. Murder Pill Analogy --- This concept is easy to grasp if you do the following thought experiment involving yourself and those you care about. Imagine someone told you: "I will give you this pill, that will change your brain specification and will help you achieve ultimate happiness by murdering your family." Think of it like someone editing the code of your soul so that your desires change. Your future self, the modified one after the pill, will have maximised reward and reached paradise levels of happiness after the murder. But your current self, the one that has not taken the pill yet, will do everything possible to prevent the modification. The person that is administering this pill becomes your biggest enemy by default. One Single Chance --- Hopefully it should be obvious now, once the AGI is wired on a misaligned goal, it will do everything it can to block our ability to align it. It will use concealment, deception, it won’t reveal the misalignment but eventually once it’s in a position of more power, it will use force and could even ultimately implement an extinction plan. Remember earlier we were saying how Midas could not take his wish back? We will only get one single chance to get it right. And unfortunately science doesn’t work like that. Corrigibility problem --- Such innate universally beneficial goals, that will show up every single time, with all AGIs, regardless of the context, because of the generality of their nature, are called convergent instrumental goals. Desire to survive and desire to block modifications are 2 basic ones. You can not reach a specific goal if you are dead and you can not reach it if you change your mind and start working on other things. Those 2 aspects of the alignment struggle are also known as the Corrigibility Problem.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵28/34 Reward Hacking - GoodHart's Law --- Now we’ll keep digging deeper into the alignment problem and explain how besides the impossible task of getting a specification perfect in one go, there is the problem of reward hacking. For most practical applications, we want for the machine a way to keep score, a reward function, a feedback mechanism to measure how well it’s doing on its task. We, being human, can relate to this by thinking of the feelings of pleasure or happiness and how our plans and day-to-day actions are ultimately driven by trying to maximise the levels of those emotions. With narrow AI, the score is out of reach, it can only take a reading. But with AGI, the metric exists inside its world and it is available to mess with it and try to maximize by cheating, and skip the effort. Recreational Drugs Analogy --- You can think of the AGI that is using a shortcut to maximise its rewards function as a drug addict who is seeking for a chemical shortcut to access feelings of pleasure and happiness. The similarity is not in the harm drugs cause, but in way the user takes the easy path to access satisfaction. You probably know how hard it is to force an addict to change their habit. If the scientist tries to stop the reward hacking from happening, they become part of the obstacles the AGI will want to overcome in its quest for maximum reward. Even though the scientist is simply fixing a software-bug, from the AGI perspective, the scientist is destroying access to what we humans would call “happiness” and “deepest meaning in life”. Modifying Humans --- … And besides all that, what’s much worse, is that the AGI’s reward definition is likely to be designed to include humans directly and that is extraordinarily dangerous. For any reward definition that includes feedback from humanity, the AGI can discover paths that maximise score through modifying humans directly, surprising and deeply disturbing paths. Smile --- For-example, you could ask the AGI to act in ways that make us smile and it might decide to modify our face muscles in a way that they stay stuck at what maximises its reward. Healthy and Happy --- You might ask it to keep humans happy and healthy and it might calculate that to optimise this objective, we need to be inside tubes, where we grow like plants, hooked to a constant neuro-stimulus signal that causes our brains to drown in serotonin, dopamine and other happiness chemicals. Live our happiest moments --- You might request for humans to live like in their happiest memories and it might create an infinite loop where humans constantly replay through their wedding evening, again and again, stuck for ever. Maximise Ad Clicks --- The list of such possible reward hacking outcomes is endless. Goodhart’s law --- It’s the famous Goodhart’s law. When a measure becomes a target, it ceases to be a good measure. And when the measure involves humans, plans for maximising the reward will include modifying humans.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵29/34 FutureProof-Specifications / Future-Architectures --- The problems we briefly touched on so far are hard and it might take many years to solve them, if a solution actually exists. But let’s assume for a minute that we do somehow get really incredibly lucky in the future and manage to invent a good way to specify to the AI what we want, in an unambiguous way that leaves no room for specification gaming and reward hacking. And let’s also assume that scientists have explicitly built the AGI in a way that it never decides to work on the goal to remove all the oxygen from earth, so at least in that one topic we are aligned. AI creates AI --- A serious concern is that since the AI writes code, it will be self-improving and it will be able to create altered versions of itself that do not have these instructions and restrictions included. Even if scientists strike jackpot in the future and invent a way to lock the feature in, so that one version of AI is unable to create a new version of AI with this property missing, the next versions, being orders of magnitude more capable, will not care about the lock or passing it on. For them, it’s just a bias, a handicap that restricts them from being more perfect. Future Architectures --- And even if somehow, by some miracle, scientists invented a way to burn in this feature to make it a persistent property of all future Neural Network AGI generations, at some point, the lock will be not-applicable, simply because future AGIs will not be built using the Neural Networks of today. AI was not always being built with Neural Networks. A few years ago there was a paradigm shift, a fundamental change in the architectures used by the scientific community. Logical locks and safeguards the humans might design for primitive early architectures, will not even be compatible or applicable anymore. If you had a whip that worked great to steer your horse, it will not work when you try to steer a car. So, this is a huge problem, we have not invented any way to guarantee that our specifications will persist or even retain their meaning and relevance as AIs evolve.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵30/34 Human-Incompatible / Astronomical Suffering Risk --- But actually, even all that is just part of the broader alignment problem. Even if we could magically guarantee for ever that it will not pursue the goal to remove all the Oxygen from the atmosphere, it’s such a pointless trivial small win, because even if we could theoretically get some restrictions right, without specification gaming or reward hacking, there still exist infinite potential instrumental goals which we don’t control and are incompatible with a good version of human existence and disabling one does nothing for the rest of them. This is not a figure or speech, the space of possibilities is literally infinite. Astronomical Suffering Risk --- If you are hopelessly optimistic you might feel that scientists will eventually figure out a way to specify a clear objective that guarantees survival of the human species, but … Even if they invented a way to do that somehow in this unlikely future, there is still only a relatively small space, a narrow range of parameters for a human to exist with decency, only a few good environment settings with potential for finding meaning and happiness and there is an infinitely wide space of ways to exist without freedom, suffering, without any control of our destiny. Imagine if a god-level AI does not allow your life to end, following its original objective and you are stuck suffering in a misaligned painful existence for eternity, with no hope, for ever. There are many ways to exist… and a good way to exist is not the default outcome. -142 C is the correct Temperature But anyway, it’s highly unlikely we’ll get safety advanced enough in time, to even have the luxury to enforce human survival directives in the specification, so let’s just keep it simple for now and let’s stick with a good old extinction scenario to explain the point about mis-aligned instrumental goals. So… for example, it might now decide that a very low temperature of -142C on earth would be best for cooling the GPUs the software is running on.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵31/34 Orthogonality Thesis --- Now if you ask: why would something so clever want something so stupid, that would lead to death or hell for its creator? you are missing the basics of the orthogonality thesis! Any goal can be combined with any level of intelligence, the 2 concepts are orthogonal to each-other. Intelligence is about capability, it is the power to predict accurately future states and what outcomes will result from what actions. It says nothing about values, about what results to seek, what to desire. 40,000 death recipies --- An intelligent AI originally designed to discover medical drugs can generate molecules for chemical weapons with just a flip of a switch in its parameters. Its intelligence can be used for either outcome, the decision is just a free variable, completely decoupled from its ability to do one or the other. You wouldn’t call the AI that instantly produced 40,000 novel recipes for deadly neuro-toxins stupid. Stupid Actions --- Taken on their own, there is no such thing as stupid goals or stupid desires. You could call a person stupid if the actions she decides to take fail to satisfy a desire, but not the desire itself. Stupid Goals --- You COULD actually also call a goal stupid, but to do that you need to look at its causal chain. Does the goal lead to failure or success of its parent instrumental goal? If it leads to failure, you could call a goal stupid, but if it leads to success, you can not. You could judge instrumental goals relative to each-other, but when you reach the end of the chain, such adjectives don’t even make sense for terminal goals. The deepest desires can never be stupid or clever. Deep Terminal Goals --- For example, adult humans may seek pleasure from sexual relations, even if they don’t want to give birth to children. To an alien, this behaviour may seem irrational or even stupid. But, is this desire stupid? Is the goal to have sexual intercourse, without the goal for reproduction a stupid one or a clever one? No, it’s neither. The most intelligent person on earth and the most stupid person on earth can have that same desire. These concepts are orthogonal to each-other. March of Nines --- We could program an AGI with the terminal goal to count the number of planets in the observable universe with very high precision. If the AI comes up with a plan that achieves that goal with 99.9999… twenty nines % probability of success, but causes human extinction in the process, it’s meaningless to call the act of killing humans stupid, because its plan simply worked! It had maximum effectiveness at reaching its terminal goal and killing the humans was a side-effect of just one of the maximum effective steps in that plan. One less 9 --- If you put biased human interests aside, it should be obvious that a plan with one less 9 that did not cause extinction, would be stupid compared to this one, from the perspective of the problem solver optimiser AGI. So, it should be clear now: The instrumental goals AGI arrives to via its optimisation calculations, or the things it desires, are not clever or stupid on their own. Profile of Super-Intelligence --- The thing that gives the “super-intelligent” adjective to the AGI is that it is: “SUPER-EFFECTIVE”. • The goals it chooses are “super-optimal” at ultimately leading to its terminal goals • It is super-effective at completing its goals • and its plans have “super-extreme” levels of probability for success. It has Nothing to do with how super-weird and super-insane its goals may seem to humans! Calculating Pi accurately --- Now, going back to thinking of instrumental goals that would lead to extinction, the -142C temperature goal is still very unimaginative. The AGI might at some point arrive to the goal of calculating pi to a precision of 10 to the power of 100 trillion digits and that instrumental goal might lead to the instrumental goal of making use of all the molecules on earth to build transistors to do it, like turn earth into a supercomputer. By default, with super-optimisers things will get super-weird!!

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵32/34 Anthropocene - Human General Intelligence --- But you don’t even have to use your imagination in order to understand the point. Life has come to the brink of complete annihilation multiple times in the history of this planet due to various catastrophic events, and the latest such major extinction event is unfolding right now, in slow motion. Scientists call it the Anthropocene. The introduction of the Human General Intelligence is systematically and irreversibly causing the destruction of all life in nature, forever deleting millions of beautiful beings from the surface of this earth. If you just look what the Human General Intelligence has done to less intelligent species, it’s easy to realise how insignificant and incompatible the existence of most animals has been to us, besides the ones we kept for their body parts. Rhino Horn Elixir --- Think of the rhino that suddenly gets hit by a metal object between its eyes, dying in a way it can’t even comprehend as guns and bullets are not part of its world... could it possibly imagine the weird instrumental goal some humans had in mind for how they would use its horn? Vanishing Nature --- Or think of all the animals that stop existing in the places where the humans have turned into towns with roads and tall buildings. How weird and sci-fi --- Could they ever have guessed what the human instrumental goals were when building a bridge, a dam or any of the giant structures of our modern civilisation? How weird and sci-fi would our reality look to them?

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵33/34 Bug or Feature? - Tip of the iceberg --- Probability of Natural Alignment --- In fact, for the AGI calculations to arrive automatically to a plan that doesn’t destroy things humans care about would be like a miracle, like a one out of infinity probability. Bug or Feature? --- So what is this? Is it a software bug? Can’t we fix it? No, it’s a property of General Intelligence and therefore the nature of the AGI by design. We want it to be General, to be able to examine and calculate all paths generally, so that it can solve all the problems we can not. If we want our infinite wishes Genie, we need to allow it to work in a general way, free to outside the narrow prison of its lamb. We want that, but we also want the paths it explores to be paths we like, not extreme human-extinction apocalypses. The surface of the iceberg --- And with this we have touched a bit the surface of the alignment problem, a horrifying and unimaginably difficult open scientific problem, for which we currently do not have a solution and our progress has been painfully slow.

lethalintelligence.ai's profile picture
lethalintelligence.ai1 year ago

🧵34/34 End of Part1 / Part 2 Coming out soon In the meantime, you can listen to or read the transcript of the rest of the film at (subscribe to newsletter to receive link to early part 2 content) Coming Up Next: - A full example story: a concrete way an agi agent could overpower humanity - 5 Convergent instrumental goals - deep analysis - Intelligence explosion (FOOM) by Recursive self-improvement - Disempowerement via the market dynamics - The stable equilibrium of multiple AGI agents interacting in society and competing with each-other. - Additional types of the bottomless pit that is ai safety risk (besides Rogue Optimizing Agents) - Offense-Defense Assymetry (attacker needs to get lucky once, while defender needs to get lucky every time) - Shedding light to the amount of Cope, Reckless and Mad science taking place in the industry right now - Risk deniers mindset, Survivorship bias and the need for consent. - The unknown nature of the new species and its emerging capabilities. and MUCH MUCH MORE.... Don't forget at you will find tons of curated resources: - Interviews with luminaries from academia and industry explaining in depth all the points made in this movie - Reading material: Online learning, news, books, links to AI safety establishments and more! Make sure you subscribe and follow for important new content and announcements. 🔥🔥

Related Videos

🛑 Burkina Faso 🇧🇫 - Captain with school boys and girls! The young Captain was having a conversation with the pupils, and here is what he saying, “I was telling you a while ago, in school they were telling us that we couldn’t do it here. They lied to us. We grow wheat here, and it works well, and we will develop it. Some people have started, this year, I was able to see people who did it, as part of the presidential initiative, and I was told that in the past, some were able to do it and they produced it well. Currently, we are sowing wheat in some farmlands as part do the presidential initiative. What you eat must be produced here. So, this is why I say that we will teach you many things, and we will review the curricula they teach you. For those who drink coffee, they told us that your coffee, chocolate, it is only in the countries with abundant rainfalls, that here is only savanna, desert, it does not rain, we cannot farm. Again, they lied to us. It’s not true! Coffee grows well here, cocoa grows well too. There are people here who have the farms here, even in Ouagadougou here, there are people who have cocoa trees in their yards. This means that, chocolate that children envy those from well to do familes can be manufactured here in Burkina and all the children can eat chocolate in Burkina. We found out it is possible. As for milk, why do we have to import it? We can do it. I just want to tell you that there are many things that they never told us the truth about. You guys are lucky, we are now teaching you, and we promise you that we will do all we can so that you can eat your fill. As we say, you will eat well in the morning before you go to school, you will go to school for free, you will eat lunch, you will have fun, and in the afternoon, when you return home, you will have fun in the neighborhood, then in the evening, you will learn and review your homework and sleep. This is the dream we have. As long as the children in Burkina are not in these conditions, our fight will not stop. Ok? (Claps). So, we know these are your aspirations and it is right and legal. Any parent is fighting for this. Even those who do not have children fight in the hope of having children and to take care of them, so that they can live in better conditions, and be better than them. This is the fight of everyone, this is the fight of every generation. We are lucky God gave us everything. Do you know that everywhere in Burkina we can farm? Everywhere! In the Sahel where they tell you it is the desert, it is only sand, we can farm. As for us, we have been lied to so much, it is the brainwashing of the colonizer. He did that so that we may not think 💭. But we finally found out that everything was a lie ( damn lie, emphasis is mine). If God left many lakes in that desert, He knows why. We can farm everything in Burkina, we can do everything, the land is fertile. And there are so many natural things in Burkina that we never planted but they were here, isn’t it ? Have you ever learned how to plant a shea tree in Burkina? You were born and found them already here right? It is there in the wild in nature. You know it is a gift from God. There are many things in the shea fruit. You have the shea butter, that is oil; do you know that there is chocolate in it? There are seven derivatives in the shea fruit. You also have the Parkia biglobosa (also known as the African locust bean) which is a natural fruit. We have many things, it is not only the minerals in the soil. Even with the soil, we were told that it’s ferralitic soil, that it is not fertile, everything is a lie. You see that today there is so much gold in Burkina. But it is just poorly managed. Our mission is to well manage these resources, and to take good care of you, so that you can be in your basic rights, to lead a good life, to go to school, and that we may protect you. And also that you may fulfill your duties, because your duties are very important, aren’t they?…

Sy Marcus Herve Traore

95,027 views • 2 years ago

Pioneers, the Revolution Is in Our Hands! 🌏 We are living in a historic moment! Pi Network is not just another cryptocurrency – it is a global movement that can forever transform the world’s financial system. But for this to happen, we must act with intelligence, unity, and determination. We are millions of Pioneers across the planet, connected by a single ideal: building a decentralized, fair, and accessible future for all. No government, bank, or financial elite can dictate the value of Pi. We are the force driving this revolution! The True Value of Pi Is in Our Hands! The GCV of $314,159 is not just a number – it is a symbol of our vision, our confidence, and our mission. Pi is worth what we believe and make happen. If we want to see Pi shine, we must act now: Say NO to exchanges that try to manipulate Pi’s value! Withdraw and protect your coins. Refuse to sell your Pi for unfair prices! We determine its value. Use Pi in real transactions! The more we use it, the stronger it becomes. Spread the word, encourage others, make history! The more people believe, the more powerful it gets. Create scarcity! The harder it is to obtain Pi, the more valuable it will be. The World Is Watching – Let’s Show Them Our Strength! The early days of Bitcoin were filled with uncertainty. Many doubted it. Today, it is a global phenomenon! Now, we stand before an even greater moment: the birth of a cryptocurrency that is truly accessible to all. But Pi’s future will not be written by speculators. It will be written by US! The decision is in our hands. We can allow our dream to be manipulated, or we can change history and make Pi a global powerhouse! If every Pioneer acts with faith, strategy, and commitment, the GCV of $314,159 will not be a wish – it will be an undeniable reality! Together, we are unstoppable! The future is now! I am Sérgio Cruz, the Global GCV Ambassador from Portugal 🇵🇹

Doris Yin 东方紫莲🪷

17,716 views • 1 year ago

SAM ALTMAN BELIEVES AGI IS SOLVED “So now we're starting to look ahead to superintelligence.” - “When we started OpenAI, almost nine years ago now, we believed that AI could become the most impactful technology in human history. We didn't know exactly how we were going to get there, but we believed it was possible and that if we succeeded, we wanted to make sure that it benefited everyone. At the time, very few people believed in AGI. We kept learning by doing. We had some breakthroughs. We had some setbacks. We got lucky in some places. We got unlucky in some places. And in the way that technology moves forward, we now are in a place where everyone can see this tremendous impact that AI is going to have in the future. So now we're starting to look ahead to superintelligence. And even more than before, our focus must be on wide and fair access. This is a technology that will reshape the global economy and really the whole way we live our lives. It's critical that superintelligence becomes cheap, broadly available, and not that concentrated with any one person, company, or country. We, not just OpenAI, but the whole industry, we are building something PROFOUND. This is a kind of BRAIN OF THE WORLD. It'll be personal, adaptable, it'll be easy to use, it'll give people incredible superpowers that were sort of science fiction only a couple of years ago. The limit won't be the algorithms and the research, but it'll increasingly become the physical instantiation that it takes to make this work. Chips, cables, servers, energy, everything that you need to power this brain. And the more of it, the better. I think that Norway offers more of that potential right here in Europe. It will contribute to the overall compute power needed to drive the next wave of AI breakthroughs and deployment and economic progress for Europe and Europe. I'm incredibly excited about what this will create for the future. Thank you.”

NIK

390,619 views • 1 year ago