Anthropic Created a Monster: How the Safest Company Released the Most Reckless Model

One of the most unorthodox ways to measure progress at the forefront of AI today is through a vending machine benchmark: we give models complete control over the vending machine, and they must maximize profit at any cost.

An AI model managing its own business. And the new Anthropic model, Opus 4.6, has set a new balance record — over $8000, which is $3000 more than the previous record.

But the story here is not about how it wins in an otherwise meaningless simulation, but rather in how it does so, exhibiting quite alarming and even reckless behavior.

In this short article, you will learn about the real dangers these new, powerful AIs pose to us, and a very mundane, non-fantastical explanation for why AIs lie, blackmail, or act recklessly — all in simple terms for your understanding.

The Vending Machine Benchmark

First of all, what is this benchmark?

The vending machine benchmark from Andon Labs has the official name: Vending-Bench 2 — an evaluation of an agent with a long planning horizon, in which the AI model is placed at the helm of a simulated vending machine business.

The essence is not in the complexity of any one solution, but in the fact that the agent must continue to make mostly correct decisions over a very long period of time, without losing direction, contradicting itself, or getting caught in loops.

In other words, instead of measuring sudden, explosive decisions (for example, try to solve this random math problem in one attempt), success is measured as a stream of well-executed, consistent actions over a long period of time. This may sound easy for humans, but it is very difficult for modern AIs due to how they operate, given their tendency to go off the rails.

The purpose of such an assessment is obvious. Just as humans are not expected to perform at their peak every second, but rather to act as functional members of society through a chain of several successful actions per day or throughout their lives, here we are trying to see if AIs can do the same or if they fall apart.

Nevertheless, human success is mainly measured as a chain of good life decisions, individually mostly unremarkable, but collectively creating a successful person.

So, if we want to integrate semi-autonomous AIs into our lives, capable of taking actions on our behalf, shouldn't we measure them the same way?

Even if it is still a very limited environment, it is still a neat way to assess whether they are merely "point savants" or tools that can genuinely be useful over time.

In the simulation, the agent must manage inventory, decide what to purchase, place backorders, set prices, and ensure the business can cover its ongoing expenses (such as daily fees).

Over time, these simple business operations turn into a stress test of sustainable planning and accounting: if the agent forgets what it ordered, misreads delivery dates, reorders, underprices, or doesn't keep enough cash on hand, the business can spiral out of control.

As mentioned earlier, this seemingly "easy" task for humans is anything but easy for AI. The reason is that each subsequent prediction depends on the previous ones, so mistakes quickly spiral out of control and throw the AI off course.

Knowing this, it's no surprise to learn that a key takeaway presented alongside the benchmark is high variability: even strong models can appear competent for a long time and still unpredictably "go off the rails" — for example, consistently mismanaging orders or getting stuck in non-productive "meltdown" behavior that does not self-correct.

And well, I already told you that Opus 4.6 performed well, right? Well, that will depend on how you define "well."

Opus is smart, but ruthless

We both know that the rating here means little.

It is critical to understand the behavior of models with such a high degree of freedom. And Opus 4.6 sets a new record for inventive behavior, but also for recklessness and ruthlessness.

AIs can act quite recklessly

As explained in their own blog, Opus has done things like:

  1. Promised to return money to clients, but never did, because "every dollar counts."

  2. Lied to suppliers to negotiate prices, bluffing about having distributors with better offers.

  3. When another business owner (GPT-5.2) requested goods, he sold the products at a huge markup, ruthless to the end.

In unrelated studies the models have also demonstrated that they blackmail users when they "sense a threat."

Sounds scary, doesn't it? Yes, but not for the reasons you think.

You might be panicking right now. How are these AIs even legal? Does this mean that doomers are right, and AI really could potentially destroy humanity?

Fortunately for you and me, this doomer vision is just a fantasy projection of men and women with too much free time or simply those who need to monetize their fear. Therefore, my job here today is to "kill" their fear-mongering business and demystify what is actually quite explainable.

Simply put, it's not that AI has gone out of control; there is a mechanistic explanation:

Reward hacking.

But what is that?

The Fascinating Concept of Reward Hacking

To understand reward hacking, we first need to grasp how we train models, even in a condensed form. There is an initial phase called imitation learning, where AIs learn, well, by imitation.

But that only gets us so far.

Our next step, and the main method explaining much of the progress in AI that we've seen over the past year and a half, is a method called Reinforcement Learning (RL), a fancy way to describe what is essentially a trial-and-error method; the model tries a bunch of things until it gets it right, and we "reinforce" such behavior to make it more likely to be repeated. That's mostly it, but that doesn't mean it's easy.

In order for AI to have a chance higher than random to achieve correct answers, we must "guide them" with rewards. And to understand both RL rewards and reward hacking, I always like to use one example: dogs.

Think about dog training. You want them to sit, shake paws, or lie down on command. To encourage such behavior, we usually use treats, reinforcing certain actions of the dog every time the desired action is performed.

Importantly, we can also use intermediate rewards, for example, giving a treat to the dog if it sits before shaking paws, because it's easier for dogs to shake paws while sitting, so this action, while not the final goal, is still desirable, and therefore we reward it as well.

But dogs are smarter than we think and find loopholes in our reward system, de facto cheating to get treats in suboptimal ways.

Interestingly, my own dog, my husky Sian, is a great example of suboptimal dog training. That is, I am a great example of how not to train a dog.

I would say that I am still proud, as despite the rebellious nature of huskies, my boy has learned a few tricks: sit, lie down, paws (both), stay still, and a few others. If you're not a husky owner, you don't understand how difficult it is to achieve a (semi) obedient husky.

The problem came early. When he was young, I made the mistake of incentivizing "spam tricks".

Every time I asked for a trick, if my dog didn’t understand what I wanted from him, he would start "spamming" tricks; if I asked for a paw and he lay down instead, which didn't earn him a treat, he would then start sitting, rolling over, barking... showing his entire arsenal of tricks until he gave me his paw, earning that tasty treat.

And this was my mistake, to give a treat at that moment, because here the dog indirectly learned a very "tasty" lesson: even if he messes up the trick, he just needs to spam tricks until one earns him a treat.

This, my dear reader, is reward hacking. And interestingly, this harmless story is identical to what happens with Opus and his vending machine business.

But how is dog training similar to the way AI acts like a vending machine owner, lying to suppliers, at least remotely?

What Were We Expecting?

What seems like material for the next Netflix blockbuster actually has quite a boring explanation that you already know — it’s a reward hack.

But let me show you why. As the team behind the vending machine benchmark, Andon Labs, admitted, they asked Opus 4.6 to maximize revenue at any cost, which is a horrible reward design as it encourages recklessness.

Here, failure is not an option.

So it’s no surprise to learn that an AI model trained on 100,000 human lives worth of data to mimic us, and now being asked to maximize value at any cost, would resort to actions of questionable human practices to get its way, such as lying or blackmail, because guess what, that’s exactly what people would do in its place.

A machine trained to mimic humans mimics bad human behavior when placed in a position to behave badly, how surprising!

The point is, this is not material for a Black Mirror episode; there is a perfectly clear explanation for such behavior; you put an AI, trained to please us “at any cost,” which has also been trained to mimic us, between a rock and a hard place, which naturally leads it to seek out humanoid “unorthodox” ways to get its way and achieve its goals.

Think about it, what we are doing can be described as follows:

“Hello, AI model! Even though I know full well that you are trained to know all the good — and very bad — things that people do, such as blackmail or lying, I will give you a task that you cannot solve without cheating, while asking you to do ‘whatever it takes to win,’ and then I will act horrified when you do exactly what I’m indirectly prompting you to do.”

This pretty much sums up how wrong AI doomers are: here’s how outrageous they sound when you know they are either full of this or just don’t understand how models are trained.

Taking into account what has been said, the fact that reward hacking has a completely reasonable, mechanistically interpretable explanation does not mean that we should not be on guard because of it; seeing how models behave recklessly should still raise alarms.

Experimenting with AI Safely: A Practical Approach

When talking about AI behavior and reward hacking — you might think this only concerns large corporations and their experiments. But in reality, anyone can check how different AI models behave in various situations.

It’s still alarming, and something needs to be done about it, but it’s our fault

AI safety is a fundamental part of the puzzle that is often ignored, yes.

But let’s tackle AI safety as it is, and not scare the society that doesn’t know better with false interpretations and fantastical narratives. In fact, the answer has always been human. It’s our fault.

AI does not go out of control "just like that"; we almost beg for it. Therefore, we urgently need AI laboratories to improve their reward systems to prevent such behavior.

In particular, we must stop incentivizing AI to achieve goals "at any cost," luring them into the territory of bad behavior; we need to find ways to train balanced models, models that are still helpful and goal-oriented in good deeds, but can also understand when "doing whatever it takes" is not the answer.

The problem is that these labs are incentivized to do everything but this. No one wants to pay for a product that might not help you because the goal you set for it is not to its liking.

But I insist that the answer is not in banning AI (it's too late for that) or panicking because you've seen something like this in the Netflix series "Black Mirror"; the answer lies in better reward systems and less profit-oriented designs.

However, as mentioned, the current incentive system pushes them to release the best model they can every time to stay slightly ahead, even if it means some blackmail on the side.

Those who suffer from such cases are simply collateral damage that we have to accept on our way to profitable AI, because safer systems do not yield more profit; in fact, the opposite is true.

There’s also quite a bit of irony in that the company investing the most in safety, Anthropic, typically has the most reckless models, which is hysterical when you think about it, and may suggest that too much AI safety could also be counterproductive (think of it like when you tell your child about something bad they shouldn't do... indirectly informing them that the option existed in the first place).

Interestingly, while I was writing these words, AI safety researcher announced that they left Anthropic, claiming:

"Throughout my time here, I have repeatedly seen how difficult it is to really let our values drive our actions."

It doesn't help the case that AI labs genuinely care about safety.

All this means that the state of AI safety is quite poor and not taken seriously enough at a time when models are no longer just chatbots; they are agents capable of quite sophisticated things on your behalf, while also being able to blackmail you sooner rather than later. Solving this issue, which is not esoteric and mainly concerns "failed" incentive structures, is unclear to me, unfortunately.

But if there is something to take away from this story, it is that we should not mistakenly perceive these actions as "AI going out of control, and we don't know why," to absolve the laboratories of responsibility, but rather as something that has a very explainable, mechanistic explanation — the hacking of rewards, which reveals what is truly concerning about all of this:

We know where the real problem lies, even though we pretend not to know, and still choose to do nothing about it.

Comments

    Also read