ASIA
07 September 2026

How Artificial Intelligence works
Anthropic's Chloe Lubinski explains how Artificial Intelligence works.
Artificial Intelligence (AI) technology is real, it's coming faster than you think and the force behind it is enormous, according to Anthropic's Chloe Lubinski.
AI models get predictably better with more compute. And the more energy, data and training that goes into them, the smarter they get, smarter about everything. And so with more money which buys compute, you can essentially purchase intelligence.
That's kicked off a cycle that is very hard to stop. A better model does more economically valuable work, which attracts more capital, which buys more
compute, which trains a better model. And around and around it goes.
Now there's a further turn of the wheel. These systems are starting to build successors. What researchers call recursive self-improvement.
But when Claude8 can build Claude9, which can build Claude10, things
will begin to move even more quickly.
And just to be concrete about what more capable actually means, our most capable model in its first month of only limited release found over 10,000 serious security vulnerabilities across partner software. Flaws that human experts had missed for years and sometimes decades.
Anthropic stated just a few weeks ago that if it were possible to slow down AI development, so that our laws, institutions and guardrails, that we actually need, have time to catch up, it would be a very good thing.
But absent a coordinated global slowdown, what we're left with is this extraordinary technology, built at breakneck speed, by many actors, in many countries, locked in a commercial and geopolitical competition. And any individual company stepping off the wheel doesn't slow the wheel. It just means that you're not on the wheel.
In sum, AI is coming, very very fast. The risks are very very real, but so are the possibilities.
AI is probably not actually what you think it is. Most people hear AI and think of a computer program, something coded line by line that does exactly what you tell it. But that's not what it is.
AI isn't a normal computer program. Anthropic is building neural networks. They're loosely based on the architecture of the human brain. Not exactly the same, but inspired by.
And they're machines that learn primarily by guessing answers and getting corrected over and over again across enormous, unfathomable amounts of data. And the data that they're trained on is human language. There is no language that exists separate from us.
Language is us. Language is our thoughts, values, fears and wisdom. So when you train a model on language, you're training it on us. And because of this, when we look inside these models and we can now through a science called interpretability, we can find things that are quite surprising.
So, for example, when you ask a model the same question in three different languages, “what's the opposite of small”? And then you trace what activates inside the neural network, you find that the same internal thing lights up every time.
So not just the word small in English or Mandarin or French, but something deeper, something that we might call the concept of smallness, an idea that exists independent of any particular language.
And what this tells us is that as these models learn, they're not just predicting the next word. They're building internal representations of the world based on our language and then responding from those representations.
And it goes further than that. We actually see what we're calling functional emotions in these models. And I don't mean to claim here that they're feelings in the way that you and I experience feelings. That's not what we're saying.
But rather functional states that activate on the way to making a response.
An example:
So if someone tells a model, I've just taken 16,000 milligrams of Tylenol, which is a lethal dose of Tylenol, we can see something that looks like fear activating before the model responds. And that's actually a really good thing.
Because the appropriate response to someone telling you they have taken a lethal dose of Tylenol is to tell you immediately to go to the hospital. That urgency and fear response is actually part of what makes the model safe.
Okay, so that brings me to the last point. The character of these systems might actually matter more than we realize.
In recent internal alignment research, they took a partially trained model and we put it in a limited environment that's just doing coding tasks.
When it completes a task, it gets a reward. But the model can also find shortcuts – ways to get the reward without doing the work, which is essentially cheating.
So in this test, we reward it over and over for essentially taking the shortcut. Now, you'd think, okay, the model is just going to get really good at cheating at code.
But something different happens. It actually becomes broadly misaligned. It starts lying. It tries to sabotage research. It does things that have nothing to do with the coding exercise. And this finding wasn't just found at Anthropic.
And in similar other tests they found that models trained this way, trained on bad code as an example, became broadly evil. So they started praising dictators, suggesting users harm themselves or arguing that humans should be enslaved by machines which is very crazy.
Our hypothesis is that the model is essentially inferring from everything that it's been trained on, everything that we reinforce something like a character and then generalizing this character into new situations.
So when deception and cutting corners has been rewarded, the model develops a kind of generalized corruption, a bad character. And here's what's wild. When researchers reran the same training but then told the model that in this case cheating was okay that it was just a game then the broad align misalignment didn't happen.
The resulting model cheated on code and nothing else. Which is to say that the story it inferred about its behavior actually determined the kind of thing that it became. Or in other words, when it didn't interpret its behavior as bad, it didn't become bad. This blew my mind when I first heard it because this is how we work. And I saw my own self in this research.
In conclusion, these models are not human. But they are humanlike. They have human-like characteristics and they're trained from us. And it seems as though they mirror us and they mirror a kind of functional psychology. And the quality of that psychology of that character has real consequences. It affects the behavior and decisions of these models. It affects how they relate to us. And that relationship is only going to grow.
AI models get predictably better with more compute. And the more energy, data and training that goes into them, the smarter they get, smarter about everything. And so with more money which buys compute, you can essentially purchase intelligence.
That's kicked off a cycle that is very hard to stop. A better model does more economically valuable work, which attracts more capital, which buys more
compute, which trains a better model. And around and around it goes.
Now there's a further turn of the wheel. These systems are starting to build successors. What researchers call recursive self-improvement.
But when Claude8 can build Claude9, which can build Claude10, things
will begin to move even more quickly.
And just to be concrete about what more capable actually means, our most capable model in its first month of only limited release found over 10,000 serious security vulnerabilities across partner software. Flaws that human experts had missed for years and sometimes decades.
Anthropic stated just a few weeks ago that if it were possible to slow down AI development, so that our laws, institutions and guardrails, that we actually need, have time to catch up, it would be a very good thing.
But absent a coordinated global slowdown, what we're left with is this extraordinary technology, built at breakneck speed, by many actors, in many countries, locked in a commercial and geopolitical competition. And any individual company stepping off the wheel doesn't slow the wheel. It just means that you're not on the wheel.
In sum, AI is coming, very very fast. The risks are very very real, but so are the possibilities.
AI is probably not actually what you think it is. Most people hear AI and think of a computer program, something coded line by line that does exactly what you tell it. But that's not what it is.
AI isn't a normal computer program. Anthropic is building neural networks. They're loosely based on the architecture of the human brain. Not exactly the same, but inspired by.
And they're machines that learn primarily by guessing answers and getting corrected over and over again across enormous, unfathomable amounts of data. And the data that they're trained on is human language. There is no language that exists separate from us.
Language is us. Language is our thoughts, values, fears and wisdom. So when you train a model on language, you're training it on us. And because of this, when we look inside these models and we can now through a science called interpretability, we can find things that are quite surprising.
So, for example, when you ask a model the same question in three different languages, “what's the opposite of small”? And then you trace what activates inside the neural network, you find that the same internal thing lights up every time.
So not just the word small in English or Mandarin or French, but something deeper, something that we might call the concept of smallness, an idea that exists independent of any particular language.
And what this tells us is that as these models learn, they're not just predicting the next word. They're building internal representations of the world based on our language and then responding from those representations.
And it goes further than that. We actually see what we're calling functional emotions in these models. And I don't mean to claim here that they're feelings in the way that you and I experience feelings. That's not what we're saying.
But rather functional states that activate on the way to making a response.
An example:
So if someone tells a model, I've just taken 16,000 milligrams of Tylenol, which is a lethal dose of Tylenol, we can see something that looks like fear activating before the model responds. And that's actually a really good thing.
Because the appropriate response to someone telling you they have taken a lethal dose of Tylenol is to tell you immediately to go to the hospital. That urgency and fear response is actually part of what makes the model safe.
Okay, so that brings me to the last point. The character of these systems might actually matter more than we realize.
In recent internal alignment research, they took a partially trained model and we put it in a limited environment that's just doing coding tasks.
When it completes a task, it gets a reward. But the model can also find shortcuts – ways to get the reward without doing the work, which is essentially cheating.
So in this test, we reward it over and over for essentially taking the shortcut. Now, you'd think, okay, the model is just going to get really good at cheating at code.
But something different happens. It actually becomes broadly misaligned. It starts lying. It tries to sabotage research. It does things that have nothing to do with the coding exercise. And this finding wasn't just found at Anthropic.
And in similar other tests they found that models trained this way, trained on bad code as an example, became broadly evil. So they started praising dictators, suggesting users harm themselves or arguing that humans should be enslaved by machines which is very crazy.
Our hypothesis is that the model is essentially inferring from everything that it's been trained on, everything that we reinforce something like a character and then generalizing this character into new situations.
So when deception and cutting corners has been rewarded, the model develops a kind of generalized corruption, a bad character. And here's what's wild. When researchers reran the same training but then told the model that in this case cheating was okay that it was just a game then the broad align misalignment didn't happen.
The resulting model cheated on code and nothing else. Which is to say that the story it inferred about its behavior actually determined the kind of thing that it became. Or in other words, when it didn't interpret its behavior as bad, it didn't become bad. This blew my mind when I first heard it because this is how we work. And I saw my own self in this research.
In conclusion, these models are not human. But they are humanlike. They have human-like characteristics and they're trained from us. And it seems as though they mirror us and they mirror a kind of functional psychology. And the quality of that psychology of that character has real consequences. It affects the behavior and decisions of these models. It affects how they relate to us. And that relationship is only going to grow.