We Taught Robots "Sumo"—and They Mastered Lariats, Cossack Dancing, and Aerial Combat
We trained two robots to fight sumo in NVIDIA Isaac Sim and tracked their evolution from Epoch 1 to 100. The experiment explores reinforcement learning, Sim-to-Real, domain randomization, and the unexpected strategies the AI discovered—from lariats and kicks to low stances and aerial combat.

On this page
- Part 1: The World of Isaac Sim and Sim-to-Real
- Why NVIDIA Isaac Sim?
- The Challenge of Sim-to-Real
- Record of the Battle
- Epoch 1–6: The Staring Contest Era
- Epoch 7: Who Needs Arms?
- Epoch 8–10: A Brief Moment of Actual Sumo
- Epoch 11–14: The Lariat Is Unleashed
- Epoch 15–20: The Fusion of Leg Techniques
- Epoch 25: The Reign of Kicks
- Epoch 40: Chaos and Breaking the Limits
- Epoch 50: The Low-Center-of-Gravity Era
- Epoch 100: Into Aerial Combat
- Conclusion
In this reinforcement learning experiment, we trained two robots to wrestle sumo in NVIDIA Isaac Sim. Following their evolution from Epoch 1 to 100, we explore why we chose Isaac Sim, the concepts of Sim-to-Real and Domain Randomization, and how the AI developed unexpected strategies ranging from lariats and leg techniques to aerial combat.
As complete beginners in robotics development, we first asked ourselves:
What can we actually do right now just to step onto the dohyo?
There was only one answer.
Make robots wrestle sumo in a virtual environment.
With that decision, we launched a reinforcement learning project using NVIDIA's physics simulation platform, Isaac Sim, with two robots competing against each other.
As the limits of physics collided with the AI's relentless obsession with winning, what unfolded inside the sumo ring went far beyond conventional sumo and became an ultimate evolution of combat.
In the first half of this article, we look at why we chose Isaac Sim and the possibilities of Sim-to-Real. In the second half, we follow the dramatic evolution of the robots from Epoch 1 all the way to Epoch 100.
First, take a look at the video.
Part 1: The World of Isaac Sim and Sim-to-Real
Why NVIDIA Isaac Sim?
In robot reinforcement learning, training through repeated failures in the real world is extremely expensive and time-consuming.
A robot can fall and break parts. Servo motors can burn out. Repeating thousands or millions of trials with physical hardware is simply not practical.
This is where NVIDIA's robotics simulation environment, NVIDIA Isaac Sim, comes in.
- High-fidelity physics simulation: Reproduces physical phenomena such as gravity, friction, and collisions.
- Photorealistic rendering: Uses RTX technology to create highly realistic visual environments.
- High-speed, large-scale parallel training: Simulation time can be accelerated far beyond real-world speed, allowing many training trials to run efficiently.
As a result, trial-and-error processes that might otherwise take months can potentially be compressed into hours or days.
The Challenge of Sim-to-Real
The process of transferring an AI model trained in simulation to a physical robot is known as Sim-to-Real.
It sounds simple at first, but the real world contains differences that simulations cannot perfectly reproduce, such as numerical errors, unexpected friction, and sensor noise.
This difference is known as the Sim-to-Real Gap.
One technique used to reduce this gap is Domain Randomization, where physical parameters in the simulation—such as friction coefficients or robot weight—are randomly varied during training.
In this sumo experiment as well, our goal was to build a robust decision-making algorithm capable of handling harsh and unpredictable physical conditions.
However...
The AI's idea of the optimal strategy for winning began evolving in a direction far beyond anything we had imagined.
Record of the Battle
From here, we follow the robots' sumo training process epoch by epoch.
Epoch 1–6: The Staring Contest Era

- Battle: The robots slowly approach each other, then stop in the center of the ring. The time limit expires with almost nothing happening.
- Commentary: These are the robots in their earliest stage. They are still figuring out how to move at all. We saw an extreme form of AI risk aversion: "Standing still and waiting is safer than moving, falling over, and losing."
Epoch 7: Who Needs Arms?
- Battle: The robots finally begin stepping forward, but their charges are easily avoided. Looking closely, however, something strange becomes obvious: both arms are folded behind their backs.
- Commentary: The robots have become slightly more aggressive. What is particularly interesting is that the AI appears to have concluded that arms are obstacles that can make the robot lose its balance through physical contact. So it seals away its arms and charges using only its torso, producing an extremely surreal fighting style.
Epoch 8–10: A Brief Moment of Actual Sumo
- Battle: Their charges finally connect perfectly, producing a genuine head-on pushing contest.
- Commentary: The two robots collide directly, push against one another, and eventually force the opponent outside the ring. This was the only period in all 100 epochs that could truly be called "sumo." As developers, we were watching happily and thinking, "Great, everything is going according to plan."
...At least, for the moment.
Epoch 11–14: The Lariat Is Unleashed

- Battle: The arms that had been folded away are suddenly released and swung violently. The age of the one-hit lariat begins.
- Commentary: The AI seems to have realized something: "It's faster to hook the opponent with an arm and knock them down than to push them." During this period, whichever robot landed the first sharp lariat achieved a 100% win rate, plunging the ring into a dark, post-apocalyptic era.
Epoch 15–20: The Fusion of Leg Techniques
- Battle: The robots aim for lariats with their upper bodies while simultaneously using their legs to hook the opponent's legs.
- Commentary: The strategy evolves from simple strikes into coordinated upper- and lower-body attacks. Advanced hooking techniques naturally emerge as the robots attempt to break their opponent's balance.
Epoch 25: The Reign of Kicks
- Battle: Repeated smaller lariats drive the opponent toward the edge of the ring, followed by a merciless kick that sends them flying outside.
- Commentary: The AI has fully understood the rule of forcing the opponent out of the ring. It drives the opponent into the dangerous edge zone and then selects a kick as the finishing counterattack, creating a brutally effective combination.
Epoch 40: Chaos and Breaking the Limits
- Battle: Movement becomes incredibly fast. Even the frame rate struggles to keep up, making the action almost impossible to follow.
- Commentary: The robots enter an era of bizarre, ultra-high-speed movement that seems to challenge the limits of the physics simulation itself. It becomes almost impossible to explain their motion in words.
Epoch 50: The Low-Center-of-Gravity Era

- Battle: The robots suddenly begin lowering their posture as far as possible. While maintaining a stance resembling a Cossack dance, they unleash sharp Taekwondo-like kicks combined with precise footwork reminiscent of fencing.
- Commentary: To avoid attacks to the upper body, the AI arrives at a simple solution: lower the hips. The high-speed leg techniques launched from this low center of gravity have moved far beyond sumo and into the territory of mixed martial arts.
Epoch 100: Into Aerial Combat

- Battle: As the robots begin raising their posture again, powerful attacks that kick the opponent high into the air become the dominant strategy.
- Commentary: The AI's final strategy was a hack of the laws of physics:
"If you lift your opponent into the air, friction becomes zero, and you can take 100% control."
The experiment ends with spectacular launching attacks and a kind of transdimensional battle that could only exist inside a simulation.
Conclusion
We tried to teach the robots sumo.
What we ended up with was an extra-dimensional combat AI that had mastered aerial warfare.
At first glance, this experiment might look like a failure.
But this is exactly what makes reinforcement learning so fascinating.
AI continuously searches through the rules humans provide, exploiting every gap it can find in pursuit of the physical behavior that maximizes its probability of winning.
Because Isaac Sim provided a high-fidelity physics simulation environment, we were able to witness strategies evolve in ways that humans would never have imagined.
To see the actual battles for yourself, be sure to watch the YouTube video at the beginning of this article.
Struggling to turn ideas into reality? With a proven track record of over 1,000 clients, our agile and flexible team will accelerate your business growth.



