Skip to content
CMA-ES Explainer
1/9
Live physics · zero gradients

Teach a whole-body G1 model to walk

Twenty-nine source actuators, 480 Hz articulated dynamics, and 5,040 learned locomotion weights — optimized live in your browser by CMA-ES with no gradient ever computed.

Physical plant

29 source joints

All leg, waist, and arm bodies carry the pinned mode-11 inertias, joint axes, and hard limits.

Learned controller

15 × 42 × 8

Fifteen locomotion rows multiply 42 physical signals by eight periodic basis terms: 5,040 weights.

Disclosed reflex

14 arm joints

The arms add real mass and reaction forces while a deterministic swing-and-balance reflex drives them.

Interactive Story Mode · Guided Walkthrough
Choose an experiment, then inspect its measured outcome

No completed owner rollout is available yet. Chapter descriptions explain experiments and goals.

walking curriculum mean
Cyan / violet rings are contact booleans. 1928 Sears Craftsman Living Room · Drag to orbit · pinch to zoom.Rose arrow: owner lateral pulse during playback; manual vector preview is display-only · arm joints are kernel-posed with real mass (head/hands: display-only)
0.0005
0.0002 (refine)0.001 (calibrated)0.01 (aggressive)

Calibrated launch radius (default): the owner improved on every declared seed in the 16-generation flat-walking sweep, and over equal wall time this band learned faster than any wider one.

1928 Sears Craftsman Bungalow

Honor Bilt Kit Architecture

Parametric architectural reconstruction with 70+ authentic period furnishings & multi-room navigation

Total Corridor Distance: 18.2 m

Full estate traversal from front veranda through parlor, dining, kitchen, hallway, and bedroom suite.

Living Room (Parlor) Physical Environment Profile

Room Area: 23.5 m² · Ceiling Height: 2.85 m

Central gathering room with quartersawn white oak flooring, 1.35m board-and-batten wainscoting, inglenook brick fireplace, and exposed coffered ceiling box beams.

✨ Clinker brick fireplace✨ Coffered ceiling box beams✨ Stickley Morris armchair✨ Dirk van Erp mica lamp

Frankensim G1 flagship

Optimize a 5,040-D walking policy

Fifteen learned actuator rows each read 42 physical signals through eight gait-phase basis terms:15 × 42 × 8 = 5,040 learned weights

A disclosed full-CMA curriculum learned 105 meaningful owner coordinates: standing bias, periodic foot unloading, then pelvis feedback. Live search expands that curriculum to all 5,040 weights. Every candidate is scored on the same 1.5-second, 720-step task and challenge you watch.

15-Dstand
+90-Dtransfer weight
5,040-Drefine live
Physical objective

Advance with the full owner-defined walking objective. Awaiting owner admission.

Full CMA is implemented on the 128-D arm below, but its O(n²) covariance would contain 25,401,600 entries here; the browser boundary honestly refuses it above 256-D.

The wave field and half-sine lateral shove are deterministic owner inputs. Survival is lexicographically primary: one extra integrated physics step beats every possible shaping-score difference. “Recovery” is horizon-censored when the robot never returns to the disclosed upright band.
Challenge
Continuous learning16 physical rollouts / generation

LM-CMA keeps its mean, search radius, and direction history hot until you press Stop. The best policy is replayed on stage every 32 generations.

Loading the owner-composed G1 experiment…
Keep this gaitno policy yet

Both are exact. The file is the archival form; the link carries the policy inside the URL fragment, so it is never uploaded anywhere. The link is long — around 50 kB — because anything smaller stopped reproducing the gait that was trained: a rollout this long amplifies rounding, and a lossy link came back walking 0.57 m instead of 0.66 m.

Receipt Objective Equalizer
Reweight the same 11 reported channels; owner optimization is unchanged

These presets recompute the visible receipt analysis only. They do not alter the owner's fixed task objective, rerun CMA-ES, or claim that the rendered gait changed.

Scalable variants, one physical budget

A live 5,040-D walking, flat race

Same curriculum mean, Philox seed, population of 16, physical evaluator, and evaluation budget. Full CMA is absent only because the owner correctly refuses dense covariance above 256 dimensions; all four families race on the 128-D arm.

Full CMA-ES

Every covariance interaction

O(n²) storage · O(n³) update

Separable CMA-ES

One variance per coordinate

O(n) storage · O(n) update

LM-CMA

A bounded history of directions

O(mn) storage · O(mn) update

LM-MA

A bounded moving transform

O(mn) storage · O(mn) update

What the kernel actually does (and doesn't)

tap to expand

Modeled

  • · 15 actuated DoFs (legs + waist)
  • · Free-floating base, SE(3) poses
  • · Semi-implicit Euler, fixed dt = 1/480 s
  • · Penalty-based normal contact + Coulomb friction (μ ≈ 0.6)
  • · Static Hertz preload at simulation start
  • · Five terminal-guard detectors (horizon, height, tilt, contact, joint-limit)

Simplified

  • · Four compliant foot patches, not full soles
  • · No torso arms, hands, or upper shell
  • · No motor torque curves or thermal limits
  • · No joint belt-elasticity or backlash
  • · Terrain is a 1-D heightfield, not a mesh
  • · Policy is a periodic basis, not a neural net

Not modeled

  • · No rolling or sliding friction asymmetry
  • · No slip detection or recovery reflex
  • · No inertial measurement, encoder, or actuator lag
  • · No environment wind, vibration, or camera noise
  • · No sim-to-real transfer or hardware validation
  • · No learned controller beyond the periodic basis

A walker that survives the kernel can still fall on real hardware. The page deliberately stops at a deterministic explainer experiment; treating it as a Unitree validation would be a category error.

Two architectures on the same contract — phase prior vs transformer

The flagship above uses a 5,040-D linear residual policy on a hand-designed phase basis — a strong, sample-efficient prior. Beside it runs a 2.9M-parameter causal transformer on the same action-causal contract, where moving forward costs real actuator work. Its committed PPO+Muon run never learned — reward flat across all 60 iterations, checkpoint from iteration 0, and a policy head exported entirely zero, so it emitted no action at all. Rather than display that as a result, the shipped artifact keeps that trunk frozen, repairs its observation normalisation, and trains only the 29×256 output layer — cloned from the CMA-ES gait, then searched against this environment's own reward — until it walks 7.09 m, past the phase prior's 7.05 m at the same budget. The contract caps speed at 0.65 m/s, so 7.80 m is the most anything can travel here; give the 105-parameter prior three times the search and it reaches 7.29 m and leads again. Which architecture wins is a question about budget, not about architecture, and that is the point worth taking away. Artifacts remain under public/robots/g1/transformer/.

Measuring both policies in a background worker (live CMA-ES search + 720-step transformer rollout) — the page stays interactive.

The same transformer, measured on the real robot

Everything above is measured in the stand-in, and a stand-in whose forward speed is a formula can only ever settle an argument about the formula. So the same architecture was searched against the owner the flagship actually runs — articulated-body dynamics, real contact, and the identical objective CMA-ES minimises in the demo at the top of this page. The transformer contributes a residual on top of the tuned controller and its output layer starts at zero, so the search begins at that controller's exact behaviour and has to earn every step from there.

The search itself is the thing this site is about: LM-CMA from fs-dfo, the same optimizer family the flagship offers, with IPOP restarts because a converged run spends its remaining budget standing still. A whole generation is one parallel batch across every core. That combination replaced a hand-rolled evolution strategy which, on the arithmetic of the day, reached a 45.8% improvement over the tuned controller on flat ground in eighty times the wall-clock time. The search reaches 165.4% there now. The two figures are not strictly comparable — the transcendentals underneath them changed in between, for the reason given below — but the wall-clock difference and the change in kind are real, and the current number is the one the panel below will reproduce on your own machine.

Transformer residual versus the tuned controller on real G1 physics, averaged over flat and terrain-with-push
PolicyObjectiveDistanceSteps
Tuned controller (5,040-D phase residual)-63.460.3188 m720
Transformer residual on top of it-98.180.5036 m720

Lower objective is better. Averaged over both challenges the residual improves the tuned controller by 54.7% on the owner's own verdict and walks 58% further without falling, from 44,952 evaluations in 759 s on ten CPU cores — no GPU and no backprop, which you cannot run through a contact solver anyway. The optimiser is this project's own LM-CMA with IPOP restarts (fs-dfo), restarted 7 times. On flat ground alone the same search reaches 165% and 1.06 m; that is the easier condition, so the cross-challenge figure is the one quoted above.

Only the 960-parameter output layer moves in the policy above, and how much of the network is worth searching turns out to depend on what it is being asked to do. Across both challenges at a matched budget the output layer reached 54.7%, the final transformer block — 38,016 parameters — reached 71.8%, and turning all 77,696 loose reached 43.2%. On flat ground alone the order reverses: the output layer reaches 165.4% and walks 1.06 m, while the final block manages 155.2% and only 0.62 m. Extra capacity earns its keep on the harder pair and costs distance on the easy one, so there is no single answer here to how big a policy should be — only a measured one per task. Re-running the output layer and the block at 45,000 evaluations, more than three times the budget, returned the same figures to the digit: these are ceilings, not snapshots.

The output layer is what ships, because it wins the condition the in-browser trainer starts on and is the only scope that trainer can resume: a wider policy carries its own trunk, so its head alone means nothing without it. An earlier version of this card claimed a wider search simply does worse, from a sweep run before the transformer's arithmetic was made portable across targets. Those numbers were measured against host maths the browser did not share and did not survive re-measurement. Weights and receipts ship under public/robots/g1/transformer/.

Now run that search yourself

Everything above reports a search that already finished. This runs it here: the same LM-CMA over the same 960 parameters, against the same articulated-body physics and the same objective, in a worker on your machine. The kernel ships with SIMD enabled and the model is 64 units wide and two layers deep, which is the whole reason a laptop suffices. It starts on flat ground, where the first improvement typically lands within about ten seconds and roughly forty rollouts; the harder cross-challenge run the figures above are measured on is one dropdown away. There is no fixed budget, so it keeps going until you stop it — on one laptop it passed 120% better than the tuned controller inside two minutes — and the policy it finds downloads as an FSGT weights file. It is a 64-wide, two-layer model, so it is deliberately not interchangeable with the 256-wide artifact the comparison above loads; that loader pins its audited architecture and refuses anything else, which is the point of the audit. A run also survives a reload, and the policy fits in a link: only the 960 trained parameters travel, because the rest of the network is fixed by the kernel's own seed. They ride in the URL fragment, so the policy never reaches a server, and whoever opens it has the owner re-run it on their own machine before believing the number.

A policy trained on the workstation scores the same here, to the digit. That took fixing. The physics owner was already identical on both targets, but the transformer called the host's exp, tanh, sin, cos and powf; macOS and the WebAssembly build ship different implementations of those, and 720 steps of contact-rich dynamics turns a last-bit disagreement into a different trajectory — a policy that walked on one machine fell over on the other. Routing those five through the same Rust implementation everywhere closed it: the shipped policy reproduces here at −157.68085560636536 and 1.0565011518224396 m, the digits the workstation recorded. The owner still re-measures every policy it is handed, so every figure in the panel is what your machine measured — it now simply agrees.

The curve appears once the search has closed a generation.
Training results: flat ground
PolicyObjectiveDistance
Tuned controller (this policy before the search)——
Best transformer residual so far——

Nothing is running. Training searches the 960-parameter output layer with LM-CMA against the owner's own objective; each rollout is 1.5 s of simulated walking. Leave it running and the curve keeps falling — the run has no fixed budget and stops when you stop it.

What this simulation actually does

Every candidate policy is a vector of 5,040 learned weights: 15 lower-body and waist actuators each read 42 physical signals through 8 gait-phase basis terms (15 × 42 × 8 = 5,040). The policy outputs bounded residual efforts; an articulated multibody kernel with SE(3) integration, contact, and friction integrates all 29 source joints at a fixed timestep — the same 1.5-second, 720-step experiment for every candidate and for the winner you watch.

CMA-ES never sees derivatives. It samples a population from a Gaussian search distribution, scores each walk (upright distance, foot contact schedule adherence, energy, and hard guards for falls and joint limits), then reshapes its covariance toward the successful candidates. Full CMA-ES is refused above 256 dimensions because a dense 5,040² covariance would need 25,401,600 entries; the live flagship therefore uses the separable and limited-memory variants you can compare directly.

The source boundary is precise. Frankensim transcribes Unitree's current 29-DoF mode-11 description; the Three.js scene projects the 30 emitted world-from-link poses and never recomputes robot kinematics. The fixed head and hand shells are visual geometry. This remains a deterministic explainer, not a validated hardware controller or sim-to-real result. See the official model guide and its mode-11 URDF.