01 logo

Figure AI Found a Robot Scaling Law. 56% of the Time.

Figure AI Found a Robot Scaling Law. 56% of the Time.

By JinPublished about 11 hours ago • 8 min read

In robotics, 56% is an awkward number.

It is high enough for Figure AI to say it found a “robotics version of the scaling law.” It is low enough for rivals to say the system cannot generalize. Helix 2.5 sits on that number: 420 tests, 237 successes, a 56% full-task success rate. Supporters see a jump from 9% to 56%. Critics see 183 failures.

Figure AI did not stage this one in a model home. It sent robots into 30 unfamiliar houses in the San Francisco Bay Area and had them tidy living rooms, make beds, and fold towels. The company says it did not train or adapt the robots for those homes or the objects inside them. Founder Brett Adcock said in the launch video: “This is the most important project since Figure AI was founded.”

In the video, a robot is making a bed. One corner of the sheet is not tucked in properly.

From model homes to unfamiliar houses

Previous Helix demos left memorable images. Two robots pulled a blanket flat together. One used its rear end to push a dishwasher shut when its hands were full. But the demos always happened in well-lit model homes.

Helix 02 could run long tasks like dishwashing and logistics sorting. Its training data came from the places where the robots would work. It learned to wash dishes in one kitchen and move items on one production line. In a new place, the team often had to collect data again.

Helix 2.5 tries to cross that line.

It starts from random initialization. It first pretrains on Figure AI’s human behavior dataset, Index. Then it uses task data to adapt to three behaviors: tidying toys, folding towels, and making beds. After training, the robots went into 30 homes. The team had not collected robot data in those homes. It did not fine-tune for the beds, pillows, towels, or toys inside them. Each behavior ran on the same fixed version. Nothing was retrained for a specific house.

Figure AI calls this zero-shot. The term has limits. Helix 2.5 did not hear a natural-language instruction and invent bed-making on its own. Figure AI still gave it training data for the three tasks. What it had not seen were the test homes, the room layouts, and the specific objects in them.

The limit does not make the test meaningless. Homes are hard to standardize the way factories are. Bed height, towel material, where toys land, and where the robot can stand all change from one house to the next. If a home robot has to collect data and retrain at every new address, it will not become a product people can buy.

Helix 2.5 proves one thing: a robot can carry experience into an unfamiliar place and finish some tasks.

How 56% was calculated

Figure AI ran 140 tests per task.

For the living room, the robot had to put all 13 to 15 toys into a basket. For towels, it had to fold four towels and put all of them into a basket. For beds, it had to meet rules for pillow position, blanket direction, and bed coverage. A human safety intervention or a timeout counted as failure.

The results:

Bed-making: 94 successes out of 140, about 67%.
Towel-folding: 87 out of 140, about 62%.
Toy-tidying: 56 out of 140, 40%.
Total: 237 successes out of 420, 56%.

During testing, Helix 2.5 showed whole-body corrections. When its position was wrong, it stepped back and reset its stance. On beds, it moved to the other side to keep going. It was not replaying a fixed motion trajectory. It was searching for a position in a new layout.

A control model without Index pretraining succeeded about 9% of the time. Helix 2.5 reached 56%. That comparison convinced Figure AI that pretraining on human behavior gave the robot transfer across homes and objects.

But 56% means that in nearly half of cases, the robot did not finish the full task.

Figure AI’s scaling law

Large language models fascinate tech companies because they got smarter and because their progress was predictable for a long time. Add data, parameters, and compute, and training loss falls along a fairly stable curve. A company can estimate what the next round of investment will buy before it runs the largest training job.

Robotics has not had that luck. Robot data is expensive, small, and tied to different hardware, tasks, and environments. A company can push one task to a high success rate and still not know whether ten times more data will help other tasks.

Figure AI says it now sees a similar pattern.

It trained models on four sizes of the Index dataset: 1x, 2x, 4x, and 8x. Model size, downstream task data, and evaluation stayed fixed. Only the human behavior pretraining data grew. Each time the data doubled, the model’s error in predicting the robot’s next action fell. Figure AI used the first three runs to predict the largest run. The final error was 0.54% of the full range.

Figure AI calls it the first “human-to-robot transfer scaling law” measured on a humanoid robot.

The comparison is simple: with everything else equal, Index pretraining raised the full-task success rate from about 9% to 56%.

That result explains the spending. Figure AI’s deal with Nscale includes an initial $3.5 billion compute commitment. Deployment is planned for the second half of 2027. The long-term target is up to 100,000 Nvidia Vera Rubin GPUs. Index collects human experience at about 35 minutes per second.

If Figure AI is right, Helix 2.5’s failures are a scale problem: not enough data, not a large enough model, not enough compute. Keep scaling all three, and 56% climbs the way early language models climbed.

Peers did not applaud

Tony Zhao attacked the reliability behind 56%. He is a core author of ACT and ALOHA and co-founder and CEO of Sunday Robotics. After Figure AI released the data, he wrote: “Useful work = generalization + reliability.”

Figure AI reads 237 successes as a robot that can work in unfamiliar homes. Tony reads the other 183 failures. In a home, a robot may touch fragile objects. It may work around children and pets. Whether it can occasionally finish a task and whether a user can trust it to work alone are different questions.

Sunday Robotics reported 778 successes out of 785 autonomous attempts in a clothing-folding test, a 99.1% success rate. The numbers are not directly comparable. One Sunday test is one garment. One Figure AI towel test required four towels in a row and all four in a basket. Tony still chose reliability as his attack. Sunday’s product Memo also targets home robots. The two companies have clashed publicly before.

Nikolai Ensslen’s criticism is a route fight. He co-founded the robot motion control company Synapticon and served as its CEO. He is now developing Physical AI and humanoid robot systems. He first acknowledged that Figure AI disclosed its success rates. Then he said the numbers show imitation learning cannot generalize.

In his definition, generalization means the system works in a different place and works every time. Helix 2.5 succeeded in 30 homes, but 44% of tasks failed. So it “did not do it at all.”

A few years ago, that judgment would have been harsh. A humanoid robot completing long, whole-body tasks in 30 unfamiliar homes is progress. But GPT-6 has changed what some practitioners expect from “generalization.”

GPT-6’s learning style is what shook the industry. It uses context to understand a goal it was not trained on. It tries, gets feedback, reflects, and corrects. Many practitioners see its grasp of 3D space, task meaning, and the results of its own actions as closer to emergent intelligence.

Traditional vision-language-action models learn a mapping between vision and action trajectories. They can move fast on tasks from training. When a situation is not covered, they may repeat the same trajectory instead of understanding why they failed.

That context makes Figure AI’s 56% delicate. Supporters can say Helix 2.5 moved large-scale human experience onto a full-size humanoid robot without collecting data in every home. Critics can say the model learned to imitate inside a larger data distribution. When a scene leaves that distribution, it still lacks the ability to understand the task and escape failure.

Scale problem or route problem?

The split is simple. Is 56% a starting point that will rise, or a stop on a route that leads nowhere?

Figure AI says scale. Not enough data, not a large enough model, not enough compute. Expand Index, expand the model, expand compute, and the success rate rises along the scaling law. The company has bet a $3.5 billion initial compute commitment and a growing Index dataset on Helix. Its latest round raised more than $1 billion at a $39 billion post-money valuation. It says it has produced more than 350 Figure 03 units and cut production to one unit per hour.

Critics say route. Imitation learning can make a robot look better across a larger distribution. It may not make the robot understand the task. To work in a different place every time, it needs understanding, reflection, and error correction, not just more data.

Generalist was the earlier robotics startup that made scaling laws its main story. Its 2025 GEN-0 release showed relationships between pretraining data, model size, training compute, and downstream performance. It claimed robotics had entered a pretraining era like large language models. GEN-1 and GEN-1.5 kept scaling data and model size. Figure AI has joined that story.

More embodied AI companies are moving from “show a demo” to “tell a pretraining and scaling law story.” The capital logic changed. If robot capability can be predicted like a language model, R&D becomes an engineering effort that can absorb more investment.

Robots are not language models. Their data comes from the physical world. It involves hardware, tasks, and environments. It is expensive and unevenly distributed. Whether human video or behavior data transfers reliably to robot actions is the central question. Helix 2.5’s 56% is evidence. It is not a verdict.

Next, 56% must prove itself

Figure AI has taken the story to the scaling law. The next data has to answer four questions.

First, does the success rate rise steadily as Index grows? If 56% becomes 80%, 90%, or higher, Helix 2.5 becomes a starting point for robot pretraining. If it stays near half, Tony Zhao and Nikolai Ensslen’s criticism becomes a challenge to the whole imitation-learning route.

Second, what do the failures look like? Are the 44% minor misses or systematic breakdowns? Can the robot recover on its own, or does it need a human? In a home, the safety intervention rate may matter more than the full-task success rate.

Third, does cross-task generalization hold? Helix 2.5 tested three tasks: toys, towels, and beds. Can Index pretraining help on new tasks without dedicated data? If not, it is still a more complex imitation learning system.

Fourth, can cost become product? A $3.5 billion initial compute commitment, a long-term target of 100,000 GPUs, and human experience collected at 35 minutes per second have to become robots that ship, get maintained, and scale. They cannot stay as curves in a technical blog.

Figure AI used 30 unfamiliar homes, 420 tests, and a 56% success rate to tell a capital story about a robot scaling law. Peers see 44% failure and the generalization ceiling of imitation learning.

56% does not need to be high. It needs to rise along the curve. Robotics needs predictable progress, not another demo. Figure AI has placed its bet. The next data will answer.

startupgadgetsthought leadersproduct reviewappsmobiletech newsfact or fictionfuture

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin