Skip to content

Elliptic cones off CUDA: the cone Hessian term is launched with one thread per contact slot (dim_block = naconmax) #1704

Description

@pulipakaa24

In _update_gradient (solver.py, around line 3220), the elliptic-cone Hessian term _update_gradient_JTCJ_dense is launched with dim=(dim_block, ndof_tri). On CUDA, dim_block comes from the SM count and each thread strides over contact slots. On every other device the fallback ("fall back for CPU") is dim_block = d.naconmax: one thread per contact slot per lower-triangle Hessian entry, launched once per Newton iteration. The launch therefore scales with the contact capacity rather than with the contacts that exist, and every active thread accumulates into ctx_h with +=, which Warp lowers to atomics.

For the Go2 from MuJoCo Menagerie with elliptic cones at 4096 worlds (naconmax 196,608, 171 Hessian entries) that is 33.6 M threads per launch, of which more than 99 % read nacon and return (about 5 contacts per world exist).

Measured on the Warp CPU device with MuJoCo Warp main at cc97eea (warp-lang 1.15.0, MuJoCo 3.14.0): Go2 scene with its own cone="elliptic" impratio="100", position actuators, 4 ms step, Newton 10 iterations / 20 line-search iterations, 32 worlds, 100 timed steps, three runs each: 10.2 to 10.4 ms/step on main, 5.7 to 5.8 ms/step with a world-major kernel (below). On a non-CUDA GPU backend (a Metal port of Warp, M4 Max, MuJoCo Warp v3.14.0 with the same change) the cone term took 72 % of the step before the change, and throughput went from 240,281 to 553,510 steps/s at 4096 worlds.

The fix I have is a separate kernel, _update_gradient_JTCJ_dense_world, launched with dim=(nworld, ndof_tri) off CUDA. Each thread owns one (world, Hessian entry), scans its world's constraint rows (nefc[world]) for elliptic contacts in the CONE state, and writes its entry once. It uses the same _elliptic_hessian_entry_from_projections contraction as the existing kernel, needs no atomics, and visits contacts in constraint-row order, so the sum is deterministic. The CUDA path is unchanged. The commit is 6d8e29a on branch elliptic-jtcj-offcuda of my fork, one commit on top of cc97eea. The full CPU test suite passes on it except io_test::test_put_data_nefc_zero_dense, which fails identically on main. Go2 (dense) and G1 (sparse) with elliptic cones against mj_step on the CPU device, 2 worlds, 40 steps: max |dq| 1.8e-7 to 2.4e-6 rad, the same envelope as main (1.8e-7 to 2.3e-6). I plan to open it as a pull request once my CLA is in place.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions