Skip to content

Upstream issue draft: CUDA illegal memory access (kernel data race) under concurrent multi-process load on the same GPU #1711

Description

@starimpact

Upstream issue draft: CUDA illegal memory access (kernel data race) under concurrent multi-process load on the same GPU

提交地址: https://github.com/google-deepmind/mujoco_warp/issues/new
署名邮箱: starimpact@126.com (用户指定)
状态: 待用户粘贴提交 (本机无 gh/GitHub 登录通道)
日期: 2026-09-29


Title

CUDA illegal memory access in one process while another mujoco_warp process runs/exits on the same GPU (data race, immune to CUDA_LAUNCH_BLOCKING=0 only)

Environment

  • GPU: NVIDIA RTX PRO 6000 Blackwell Server Edition, 96 GB, sm_120 (2×, both affected)
  • CUDA Toolkit 12.9, Driver 13.0 (as reported by warp init)
  • warp-lang 1.17.0
  • mujoco 3.13.0, mujoco-warp 3.12.0 (also reproduced on 3.11.0)
  • mjlab 1.6.0 (manager layer; put_data(nworld=4096/500, nconmax=128, njmax=256))
  • PyTorch 2.8.0+cu128, DDP 2 ranks (NCCL), CUDA graphs disabled
  • Ubuntu 22.04, kernel 5.15

Summary

Two independent processes each run a mujoco_warp simulation on the same
GPU
(no MPS; separate CUDA contexts). While process B (a short-lived
eval job, nworld=500) is running — and sometimes right after it exits —
process A (a long-running training job, nworld=4096 per rank, plus NCCL
allreduce between 2 ranks) hits a CUDA illegal memory access. The IMA
is asynchronously reported, so the reported site varies (NCCL watchdog,
clip_grad_norm_ sync, _build_sequence indexing). In some cases the IMA
leaves rank 0's CUDA context wedged instead of aborting, producing a
permanent hang.

Frequency: roughly 1 in 3 co-location runs triggers within ~10 minutes;
in pure single-process training the same crash family appears spontaneously
at ~1 per 600 iterations.

Reproduction sketch

Process A (persistent): make_data(nworld=4096), continuous
forward/step loop (~50 physics steps/s) + a PyTorch DDP allreduce every
~70 s.

Process B (churn): same script shape, make_data(nworld=500), runs ~3-10
minutes then exits; repeat B in a loop.

Crash appears in A while B is running or within ~2 min after B's exit.

Signatures observed in A:

  1. Abort form: NCCL watchdog thread catches
    CUDA error: an illegal memory access was encountered
    (ProcessGroupNCCL.cpp:2068), rank SIGABRTs.
  2. Hang form: rank 0 blocks forever inside a trivial CUDA sync
    (torch.nn.utils.clip_grad_norm_ → total-norm kernel / .item());
    rank 1 (which received the IMA exception first) blocks in
    destroy_process_group, its NCCL kernel spinning at GPU 100%. Stack
    dump via faulthandler confirms both call sites; nothing in our own
    kernels is on-stack — the corruption happened earlier, asynchronously.

What we ruled out

  • PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True — no effect.
  • TORCH_NCCL_ASYNC_ERROR_HANDLING=1 — converts some hangs to aborts but
    does not prevent the IMA.
  • Not the 3.11 nefc overflow: we verified 3.11's
    _update_gradient_JTDAJ_dense_tiled* lacked the wp.min(njmax, nefc)
    clamp (fixed in 3.12); upgrading removed that bug but the co-location
    IMA above persists on 3.12.
  • Out-of-bounds in our own contact-scan kernels — hardened with explicit
    bounds; crashes persist.

Key observation pointing to a kernel data race

With CUDA_LAUNCH_BLOCKING=1 in process A, 6 consecutive co-location
injection rounds all pass
(vs ~1/3 crash rate without it), at 2.1×
slowdown. I.e. serializing kernel launches in A immunizes it against load
from B on the same GPU. This suggests a timing-dependent race inside a
kernel executed by A (mujoco_warp/warp-compiled kernels are the dominant
GPU work besides NCCL), aggravated by SM contention from B.

Suspicion

A data race in one of the mujoco_warp (or warp 1.17-compiled) kernels that
only manifests under concurrent SM pressure from a second process on the
same device, e.g. shared-memory reuse across kernel instances or
wp.launch scratch buffers. We cannot pin the exact kernel because the
error surfaces asynchronously; compute-sanitizer racecheck under
co-location would be the next step for us, but filing first in case the
team recognizes the pattern (the 3.12 nefc-clamp fix suggests related
hardening history).

Happy to provide full logs, stack dumps, and a standalone reproducer
script if useful.

Contact: starimpact@126.com

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions