What Is Reinforcement Learning? What Was Happening Behind That “Extra-Dimensional Fight”?
Using our NVIDIA Isaac Sim robot sumo experiment, this article explains reinforcement learning through state, action, environment, reward, and trial and error. It also shows how RL differs from supervised learning and why the robots discovered lariats, folded arms, and aerial attacks.

On this page
In our previous article, we showed how two robots trained to wrestle sumo in NVIDIA Isaac Sim evolved chaotically as the epochs progressed—from simply staring at each other to eventually launching aerial attacks.
As we watched the robots move during that experiment, we gained a much clearer understanding of what kind of product Isaac Sim actually is.
To be frank, we were struck by how similar Isaac Sim felt to Unity ML-Agents, which we worked with about six years ago.
At a fundamental level, the concepts and pipeline—environment, agent, reward, and policy—are essentially the same.
Isaac Sim operates on a completely different scale and level of parallelism compared with Unity ML-Agents, but both are fundamentally environments for running reinforcement learning (RL) loops.
This time, partly to organize our own understanding, we would like to summarize the essence of reinforcement learning itself.
There Is No “Correct Answer” to Imitate
Reinforcement learning is one form of machine learning, but one major difference from other approaches is that there is no example dataset showing the correct behavior to imitate.
For example, in training for text-generating AI such as LLMs, there is an “ideal” sequence of text that should come next in a given context, and the AI aims to get as close as possible to that correct answer.
The sumo robots, however, were never given joint movements or demonstration data saying, “This is what good sumo looks like.”
Note: Put another way, if we showed the sumo robots the correct actions and taught them step by step what they should do next, as with an LLM-style learning setup, then in theory we could create much more well-behaved sumo robots. We will leave that topic for another time.
In reinforcement learning, there is no model answer. There is only a set of evaluation rules—a scoring system.
There Is a Goal, but the Method Is Up to the AI
So how does an AI learn when it has no example to imitate?
Through trial and error.
To put it rather bluntly, it is like a “die-and-learn” game.
For readers who do not play many games, that phrase may make things even more confusing, so let us explain it first. A “die-and-learn” game is one where you fail over and over, learn from each death, and gradually internalize how to overcome the challenge.
Supervised learning (SL), by contrast, is less like a die-and-learn game and more like memorizing past exam questions. That difference captures something fundamental about the two approaches.
Reinforcement Learning = An Ultra-Difficult, No-Hints “Die-and-Learn” Game
The reinforcement learning process is a lot like being forced to play FromSoftware games such as Elden Ring or Dark Souls blindfolded and without a strategy guide.
No example to follow
Nobody tells you the correct way to move.
Learn by dying (trial and error)
The player—the AI—walks into traps, falls off cliffs, and gets killed by bosses in a single hit, receiving penalties along the way. By accumulating its own failure data—its record of “deaths”—it gradually discovers patterns such as, “I should not step there,” or, “I need to dodge at this timing.”
The reason our sumo robots evolved from simply staring at each other to engaging in aerial combat was that they could only discover winning strategies by “dying”—falling out of the ring—tens of thousands of times.
It was the ultimate die-and-learn game.
Reinforcement learning continuously runs through the following loop:
- State: The robot’s current posture and distance from the opponent
- Action: How to move its joints
- Environment: The physical space inside Isaac Sim
- Reward: The “score” or “penalty” received as a result
The only instruction given to the AI is essentially:
“Earn the highest total reward possible.”
The idea is: “We define the goal—a high score—but we do not care how you get there. Find the path yourself through trial and error.”
Why Did It End Up in “Aerial Combat”?
That “we do not care how you get there” characteristic is exactly what caused the chaotic evolution we observed.
A human would think, “This is sumo, so I should push my opponent out of the ring.”
But an AI with no model behavior to imitate has no concept of “the spirit of sumo” or of ordinary human expectations.
For the AI, the only correct answer is to earn points as quickly and efficiently as possible within the rules.
After enough trial and error, the strategies it discovered included:
- Folding its arms behind its back because physical contact with them could cause it to lose balance
- Delivering a lariat because it could knock the opponent down quickly
- Kicking the opponent high into the air, reducing friction to zero and taking complete control
These were extreme physics hacks that a human would be unlikely to imagine.
The moment we realized, “We thought we were teaching sumo, but the AI was actually treating it as a high-score game governed by the laws of physics,” all of the robots’ strange movements suddenly made sense.
Conclusion: The Difficulty—and Fun—of Communicating Human Intent to AI
Through this experiment, we experienced both the fascination and the difficulty of reinforcement learning: letting an AI learn through trial and error using only rules, without giving it a correct answer.
The AI was not failing.
It was relentlessly exploring ways to exploit gaps in the rules we had given it and searching for a path to victory.
Seeing that the depth—and chaos—of reinforcement learning, where AI can produce behavior far beyond human imagination, was still alive and well even in Isaac Sim was both a major discovery for us and something that felt strangely nostalgic.
Next time, we plan to experiment with reward design (reward engineering): what kind of rules and rewards should we give these robots if we want them to perform clean, human-intended sumo?
Struggling to turn ideas into reality? With a proven track record of over 1,000 clients, our agile and flexible team will accelerate your business growth.



