Benchmarks: Evaluating Embodied AI
Humans win the sprint; machines win the marathon. Humanoids shift the focus from peak speed to continuous operational reliability, and the model layer is now what moves that bar.
The hurdle for humanoid deployment is not outperforming human dexterity, but achieving continuous, non-fatiguing execution inside human-built environments. The measurement hurdle is separate and harder: most benchmarks hold the policy fixed and test the machine, when the policy is the part improving fastest and the part a demo video hides.
Man vs. Machine: the 10-hour sorting sprint
In a landmark May 2026 challenge, a human intern and a Figure AI F.03 humanoid robot competed in a 10-hour parcel-sorting task.
The human processed 12,924 packages (2.79 seconds/package), while the F.03 processed 12,732 packages (2.83 seconds/package). The human won the sprint by a mere 0.04-second margin, highlighting the narrow performance gap of modern embodied systems.
The physiological barrier
While the human intern won, they finished the shift with severe blisters, arm and back pain, stating they could not have sustained another 30 minutes.
Conversely, the F.03 units operated continuously. By hot-swapping units into charging docks, the robotic fleet maintained an uninterrupted flow, proving that consistency and non-fatigability outvalue raw peak speed in multi-shift operations.
The 200-hour autonomous endurance test
To validate long-term reliability, Figure followed the challenge with a 200-hour continuous autonomous run where a fleet of F.03 robots successfully sorted 250,000 packages without a hardware failure.
This shifts the industrial benchmark from experimental demonstrations to robust, predictable Mean Time Between Failures (MTBF) in physical logistics.
The human infrastructure moat
Humanoids are commercial drop-in capex because they are designed to fit human spaces: aisles, shelves, staircases, and tool handles.
Traditional warehouse automation requires massive, fixed conveyor rebuilds that lock up capital. Humanoids allow operators to automate logistics dynamically without altering the real estate blueprint.
The same benchmark, run against the model instead of the robot
Every figure above holds the policy fixed and asks what the machine can endure. The more interesting version holds the machine fixed and asks what a better policy is worth, because that is the bar that moves fastest and the one a buyer cannot see in a demo video.
Physical Intelligence scores its models exactly this way, on throughput rather than success rate, on the grounds that success rate alone rewards a policy for being slow. On a real cardboard-box workflow taken from a chocolate factory a few blocks from their office, moving from the SFT-like stage to reinforcement-learning post-training roughly doubled throughput. Espresso making, which demands forceful portafilter insertion and accurate timing, cleared 90% success and then ran 13 hours straight.
The hardware did not change across those numbers. That is the point: a warehouse operator comparing two identical arms is increasingly comparing two policies, and the benchmark has to be able to see the difference.
Source: Chelsea Finn, Physical Intelligence, on Root Access (Y Combinator), 12 August 2026; transcript in `docs/articles/transcripts/Chelsea Finn - Next Decade in Robotics - Root Access.md`
A benchmark that only measures the trained task measures nothing
Parcel sorting is a repetitive, single-task benchmark. It is the right test for a fixed logistics contract and the wrong test for a general-purpose claim, because a specialist trained on exactly that workflow will always win it. The harder question is what happens off the training distribution.
Three results define that test today. A generalist model interacting with an air fryer, an appliance appearing in three episodes of a dataset it was not deliberately collected for, opening it, loading it and closing it. Garment folding transferring to a large industrial Byarm UR5E with no folding data collected on that platform, approaching human teleoperation performance. And an emergent one: during pinwheel assembly, after a mistake left the pin in the wrong hand, the robot inserted it with the left gripper despite every demonstration, in both post-training and pre-training, using the right.
None of those is a throughput number, and all three predict deployment cost better than one. A model that generalises across objects and embodiments amortises its data across a fleet. A model that wins one benchmark has to be paid for again at every new site.
Source: Chelsea Finn, Physical Intelligence, on Root Access (Y Combinator), 12 August 2026
The disclosure that any robot benchmark needs
The single largest source of noise in robotics benchmarking is that impressive footage rarely states its autonomy level. Teleoperated, teleoperation-assisted, autonomous-with-intervention and fully autonomous are four different products, and a laundry-folding clip looks identical across all four.
So the honest way to read any number on this page, including the ones above, is to ask three things of it. Was a human in the loop, and how often. How long did it run unbroken, rather than what was the best single attempt. And was it the same policy across every trial, or one tuned per site.
Deployment claims should be read the same way. Physical Intelligence reports that the YC companies Ultra and Weave run post-trained π models in production for warehouse packaging and laundry folding, alongside applications spanning surgical robots, drones and agricultural tractors. That is a first-party claim of production use; the autonomy level within it is the number that decides what it is worth.
Source: Chelsea Finn, Physical Intelligence, on Root Access (Y Combinator), 12 August 2026