In _update_gradient (solver.py, around line 3220), the elliptic-cone Hessian term _update_gradient_JTCJ_dense is launched with dim=(dim_block, ndof_tri). On CUDA, dim_block comes from the SM count and each thread strides over contact slots. On every other device the fallback ("fall back for CPU") is dim_block = d.naconmax: one thread per contact slot per lower-triangle Hessian entry, launched once per Newton iteration. The launch therefore scales with the contact capacity rather than with the contacts that exist, and every active thread accumulates into ctx_h with +=, which Warp lowers to atomics.
For the Go2 from MuJoCo Menagerie with elliptic cones at 4096 worlds (naconmax 196,608, 171 Hessian entries) that is 33.6 M threads per launch, of which more than 99 % read nacon and return (about 5 contacts per world exist).
Measured on the Warp CPU device with MuJoCo Warp main at cc97eea (warp-lang 1.15.0, MuJoCo 3.14.0): Go2 scene with its own cone="elliptic" impratio="100", position actuators, 4 ms step, Newton 10 iterations / 20 line-search iterations, 32 worlds, 100 timed steps, three runs each: 10.2 to 10.4 ms/step on main, 5.7 to 5.8 ms/step with a world-major kernel (below). On a non-CUDA GPU backend (a Metal port of Warp, M4 Max, MuJoCo Warp v3.14.0 with the same change) the cone term took 72 % of the step before the change, and throughput went from 240,281 to 553,510 steps/s at 4096 worlds.
The fix I have is a separate kernel, _update_gradient_JTCJ_dense_world, launched with dim=(nworld, ndof_tri) off CUDA. Each thread owns one (world, Hessian entry), scans its world's constraint rows (nefc[world]) for elliptic contacts in the CONE state, and writes its entry once. It uses the same _elliptic_hessian_entry_from_projections contraction as the existing kernel, needs no atomics, and visits contacts in constraint-row order, so the sum is deterministic. The CUDA path is unchanged. The commit is 6d8e29a on branch elliptic-jtcj-offcuda of my fork, one commit on top of cc97eea. The full CPU test suite passes on it except io_test::test_put_data_nefc_zero_dense, which fails identically on main. Go2 (dense) and G1 (sparse) with elliptic cones against mj_step on the CPU device, 2 worlds, 40 steps: max |dq| 1.8e-7 to 2.4e-6 rad, the same envelope as main (1.8e-7 to 2.3e-6). I plan to open it as a pull request once my CLA is in place.
In
_update_gradient(solver.py, around line 3220), the elliptic-cone Hessian term_update_gradient_JTCJ_denseis launched withdim=(dim_block, ndof_tri). On CUDA,dim_blockcomes from the SM count and each thread strides over contact slots. On every other device the fallback ("fall back for CPU") isdim_block = d.naconmax: one thread per contact slot per lower-triangle Hessian entry, launched once per Newton iteration. The launch therefore scales with the contact capacity rather than with the contacts that exist, and every active thread accumulates intoctx_hwith+=, which Warp lowers to atomics.For the Go2 from MuJoCo Menagerie with elliptic cones at 4096 worlds (
naconmax196,608, 171 Hessian entries) that is 33.6 M threads per launch, of which more than 99 % readnaconand return (about 5 contacts per world exist).Measured on the Warp CPU device with MuJoCo Warp
mainat cc97eea (warp-lang 1.15.0, MuJoCo 3.14.0): Go2 scene with its owncone="elliptic" impratio="100", position actuators, 4 ms step, Newton 10 iterations / 20 line-search iterations, 32 worlds, 100 timed steps, three runs each: 10.2 to 10.4 ms/step onmain, 5.7 to 5.8 ms/step with a world-major kernel (below). On a non-CUDA GPU backend (a Metal port of Warp, M4 Max, MuJoCo Warp v3.14.0 with the same change) the cone term took 72 % of the step before the change, and throughput went from 240,281 to 553,510 steps/s at 4096 worlds.The fix I have is a separate kernel,
_update_gradient_JTCJ_dense_world, launched withdim=(nworld, ndof_tri)off CUDA. Each thread owns one (world, Hessian entry), scans its world's constraint rows (nefc[world]) for elliptic contacts in the CONE state, and writes its entry once. It uses the same_elliptic_hessian_entry_from_projectionscontraction as the existing kernel, needs no atomics, and visits contacts in constraint-row order, so the sum is deterministic. The CUDA path is unchanged. The commit is 6d8e29a on branchelliptic-jtcj-offcudaof my fork, one commit on top of cc97eea. The full CPU test suite passes on it exceptio_test::test_put_data_nefc_zero_dense, which fails identically onmain. Go2 (dense) and G1 (sparse) with elliptic cones againstmj_stepon the CPU device, 2 worlds, 40 steps: max |dq| 1.8e-7 to 2.4e-6 rad, the same envelope asmain(1.8e-7 to 2.3e-6). I plan to open it as a pull request once my CLA is in place.