About Humanoid Index
Why it exists · what it refuses to do · what makes it hard to copyHumanoid Index is an independent benchmark that scores robots 0–100 from published, cited specifications — 175 machines from 93 companies across 20 countries, built on 493 source citations. It is not funded by, affiliated with, or reviewed by any robotics manufacturer. Nothing on it is a matter of opinion that isn't shown as one.
Why we built it
The humanoid-robot field is drowning in demo videos and valuation headlines. Anyone trying to answer a simple question — which of these machines actually works, and how would I know? — runs into four failures that no amount of reading press coverage fixes.
1 · A demo doesn't show autonomy
A robot folding laundry on stage may be running its own policy or may be a puppet worn by a person in a motion-capture suit in the next room. The footage looks identical. The distinction is the entire question — it separates a product from a prototype — and it is almost never in the caption.
So autonomy is the heaviest pillar here at 25%, scored from a published ladder rather than an impression, and teleoperation is capped near the floor of it: research 20, teleoperated 35, supervised 70, task-autonomous 92. No profile for any robot category is permitted to redefine those rungs.
2 · Spec sheets are marketing documents
Payloads are quoted at best case, runtimes without load, and the numbers that would embarrass a machine are simply absent. A ranking that fills the gaps with estimates quietly rewards whoever discloses least.
So every figure here is cross-checked against at least two independent
sources on different registrable domains, carries a per-robot confidence grade
and an as_of date, and — the rule that does the real work —
a missing spec scores zero for that sub-metric and is never imputed.
Absence is a finding, not a blank to fill in charitably.
3 · One score can't span every robot
The obvious move — one grand index over every robot — produces nonsense the moment you leave humanoids. Score a quadruped against a human hand's 22 degrees of freedom and manipulation returns 0.0 for lacking thumbs. Score a surgical arm on walking speed and mobility returns 0.0 for being bolted to the floor. Both are arithmetically correct and semantically worthless.
So each category carries its own index — its own weights, its own reference scales — and composites are ranked within a category, never across. The 10 categories →
4 · Rankings go stale faster than they're rebuilt
This field ships something material every month. A hand-built ranking is substantially wrong within a quarter, and the ones you find are usually a blog post with a date on it and no way to tell what changed since.
So this one is continuously re-researched, every change is checked against its sources before it lands, and every change that lands is published on a public changelog and RSS feed with those sources attached. You can see what moved, when, and on whose evidence.
What that adds up to — a worked example
The most-covered humanoid programme in the world is Tesla Optimus Gen 2. On this index it ranks #14 of 156 humanoids with an XHI of 72.2 — not because we dislike it, but because its publicly demonstrated autonomy mode is supervised-autonomy and the ladder caps that at 35 regardless of how good the hardware underneath is. Every step of that arithmetic is on its page, with the sources it came from. Disagree with the weighting and you can recompute the whole index yourself — the formula is published.
Kept current, not rebuilt
Most rankings are a snapshot: assembled once, accurate on publication day, quietly wrong six months later. This one is re-researched continuously, and every change to it is published with the sources it was checked against.
Figures are cross-checked against at least two independent sources before they reach the leaderboard, unverifiable specs are left blank rather than guessed, and anything that fails those checks is held back rather than published. The full record of what changed and when is on the changelog, with citations, and in the RSS feed.
The second structural bet: one index per category
Why the taxonomy is load-bearing
Most benchmarks that expand do it by adding rows. Adding a drone to a humanoid leaderboard produces a number that is arithmetically valid and means nothing, so this platform expands by adding indices instead — 10 live scoring profiles across 10 declared categories, each with its own weights and its own reference scales, anchored to that category's frontier rather than a human body.
A profile that omits a pillar is saying not applicable to this category, not "scores zero" — the remaining weights are renormalised, so a quadruped is not docked a fifth of its index for lacking thumbs. And a category with no registered profile is refused at every admission gate: a record filed there would silently fall back to being scored as a humanoid, which is acceptable for rendering an old row and unacceptable for admitting a new one.
What survives across categories
Exactly three pillars, and the reason is worth stating precisely: they are scored from ladders rather than physical measurements. A teleoperated submarine and a teleoperated humanoid are equally teleoperated; shipping means shipping; money is money.
- Autonomy — the same four rungs everywhere, teleop capped in all of them
- Commercial readiness — production maturity plus named, cited deployments
- Ecosystem & viability — funding, valuation, corporate backing
That makes a cross-category autonomy comparison honest where a cross-category composite never could be. The published indices today:
How to check us
Every claim on this site is traceable
- The methodology renders each scoring profile verbatim from the code that computes it — weights, ladders, reference scales and what each denominator is anchored to.
- Each robot page lists its source URLs, its confidence grade, and the date its figures were last verified.
- The full dataset is a JSON endpoint. Recompute the index yourself and tell us where we're wrong.
- The changelog records every applied change; entries applied by an agent rather than a human reviewer are badged as such, never presented as human-reviewed.
- All 156 published humanoid scores are locked by a snapshot test, so a change to the scoring code that silently moves a robot fails the build instead of shipping.
What this is not
- Not a buyer's guide. A high index score means a machine is well documented and capable on published evidence — not that it suits your process, your safety case or your budget.
- Not investment advice. The ecosystem pillar reads funding and valuation as signals of survival odds, nothing more.
- Not vendor-funded, and not affiliated with any manufacturer. No company pays to appear, to rank, or to be re-checked, and no listing is removable by request — corrections are, with sources. See the trademark notice.
- Not a measure of what a robot could do. It scores what has been demonstrated and published. Unreleased capability is real and is not evidence.
Corrections
If a figure here is wrong, we want the correction more than we want to have been right — including from the manufacturer. Send the claim and at least one primary source to governance@autogovern.io. It enters the same review queue as everything else, gets the same validity checks, and if it's applied it appears on the changelog like any other change. Operated by autogovern.io · legal.