Upstream issue draft: CUDA illegal memory access (kernel data race) under concurrent multi-process load on the same GPU
提交地址: https://github.com/google-deepmind/mujoco_warp/issues/new
署名邮箱: starimpact@126.com (用户指定)
状态: 待用户粘贴提交 (本机无 gh/GitHub 登录通道)
日期: 2026-09-29
Title
CUDA illegal memory access in one process while another mujoco_warp process runs/exits on the same GPU (data race, immune to CUDA_LAUNCH_BLOCKING=0 only)
Environment
- GPU: NVIDIA RTX PRO 6000 Blackwell Server Edition, 96 GB, sm_120 (2×, both affected)
- CUDA Toolkit 12.9, Driver 13.0 (as reported by warp init)
- warp-lang 1.17.0
- mujoco 3.13.0, mujoco-warp 3.12.0 (also reproduced on 3.11.0)
- mjlab 1.6.0 (manager layer;
put_data(nworld=4096/500, nconmax=128, njmax=256))
- PyTorch 2.8.0+cu128, DDP 2 ranks (NCCL), CUDA graphs disabled
- Ubuntu 22.04, kernel 5.15
Summary
Two independent processes each run a mujoco_warp simulation on the same
GPU (no MPS; separate CUDA contexts). While process B (a short-lived
eval job, nworld=500) is running — and sometimes right after it exits —
process A (a long-running training job, nworld=4096 per rank, plus NCCL
allreduce between 2 ranks) hits a CUDA illegal memory access. The IMA
is asynchronously reported, so the reported site varies (NCCL watchdog,
clip_grad_norm_ sync, _build_sequence indexing). In some cases the IMA
leaves rank 0's CUDA context wedged instead of aborting, producing a
permanent hang.
Frequency: roughly 1 in 3 co-location runs triggers within ~10 minutes;
in pure single-process training the same crash family appears spontaneously
at ~1 per 600 iterations.
Reproduction sketch
Process A (persistent): make_data(nworld=4096), continuous
forward/step loop (~50 physics steps/s) + a PyTorch DDP allreduce every
~70 s.
Process B (churn): same script shape, make_data(nworld=500), runs ~3-10
minutes then exits; repeat B in a loop.
Crash appears in A while B is running or within ~2 min after B's exit.
Signatures observed in A:
- Abort form: NCCL watchdog thread catches
CUDA error: an illegal memory access was encountered
(ProcessGroupNCCL.cpp:2068), rank SIGABRTs.
- Hang form: rank 0 blocks forever inside a trivial CUDA sync
(torch.nn.utils.clip_grad_norm_ → total-norm kernel / .item());
rank 1 (which received the IMA exception first) blocks in
destroy_process_group, its NCCL kernel spinning at GPU 100%. Stack
dump via faulthandler confirms both call sites; nothing in our own
kernels is on-stack — the corruption happened earlier, asynchronously.
What we ruled out
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True — no effect.
TORCH_NCCL_ASYNC_ERROR_HANDLING=1 — converts some hangs to aborts but
does not prevent the IMA.
- Not the 3.11
nefc overflow: we verified 3.11's
_update_gradient_JTDAJ_dense_tiled* lacked the wp.min(njmax, nefc)
clamp (fixed in 3.12); upgrading removed that bug but the co-location
IMA above persists on 3.12.
- Out-of-bounds in our own contact-scan kernels — hardened with explicit
bounds; crashes persist.
Key observation pointing to a kernel data race
With CUDA_LAUNCH_BLOCKING=1 in process A, 6 consecutive co-location
injection rounds all pass (vs ~1/3 crash rate without it), at 2.1×
slowdown. I.e. serializing kernel launches in A immunizes it against load
from B on the same GPU. This suggests a timing-dependent race inside a
kernel executed by A (mujoco_warp/warp-compiled kernels are the dominant
GPU work besides NCCL), aggravated by SM contention from B.
Suspicion
A data race in one of the mujoco_warp (or warp 1.17-compiled) kernels that
only manifests under concurrent SM pressure from a second process on the
same device, e.g. shared-memory reuse across kernel instances or
wp.launch scratch buffers. We cannot pin the exact kernel because the
error surfaces asynchronously; compute-sanitizer racecheck under
co-location would be the next step for us, but filing first in case the
team recognizes the pattern (the 3.12 nefc-clamp fix suggests related
hardening history).
Happy to provide full logs, stack dumps, and a standalone reproducer
script if useful.
Contact: starimpact@126.com
Upstream issue draft: CUDA illegal memory access (kernel data race) under concurrent multi-process load on the same GPU
Title
CUDA illegal memory access in one process while another mujoco_warp process runs/exits on the same GPU (data race, immune to CUDA_LAUNCH_BLOCKING=0 only)
Environment
put_data(nworld=4096/500, nconmax=128, njmax=256))Summary
Two independent processes each run a mujoco_warp simulation on the same
GPU (no MPS; separate CUDA contexts). While process B (a short-lived
eval job, nworld=500) is running — and sometimes right after it exits —
process A (a long-running training job, nworld=4096 per rank, plus NCCL
allreduce between 2 ranks) hits a CUDA illegal memory access. The IMA
is asynchronously reported, so the reported site varies (NCCL watchdog,
clip_grad_norm_sync,_build_sequenceindexing). In some cases the IMAleaves rank 0's CUDA context wedged instead of aborting, producing a
permanent hang.
Frequency: roughly 1 in 3 co-location runs triggers within ~10 minutes;
in pure single-process training the same crash family appears spontaneously
at ~1 per 600 iterations.
Reproduction sketch
Process A (persistent):
make_data(nworld=4096), continuousforward/steploop (~50 physics steps/s) + a PyTorch DDP allreduce every~70 s.
Process B (churn): same script shape,
make_data(nworld=500), runs ~3-10minutes then exits; repeat B in a loop.
Crash appears in A while B is running or within ~2 min after B's exit.
Signatures observed in A:
CUDA error: an illegal memory access was encountered(
ProcessGroupNCCL.cpp:2068), rank SIGABRTs.(
torch.nn.utils.clip_grad_norm_→ total-norm kernel /.item());rank 1 (which received the IMA exception first) blocks in
destroy_process_group, its NCCL kernel spinning at GPU 100%. Stackdump via
faulthandlerconfirms both call sites; nothing in our ownkernels is on-stack — the corruption happened earlier, asynchronously.
What we ruled out
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True— no effect.TORCH_NCCL_ASYNC_ERROR_HANDLING=1— converts some hangs to aborts butdoes not prevent the IMA.
nefcoverflow: we verified 3.11's_update_gradient_JTDAJ_dense_tiled*lacked thewp.min(njmax, nefc)clamp (fixed in 3.12); upgrading removed that bug but the co-location
IMA above persists on 3.12.
bounds; crashes persist.
Key observation pointing to a kernel data race
With
CUDA_LAUNCH_BLOCKING=1in process A, 6 consecutive co-locationinjection rounds all pass (vs ~1/3 crash rate without it), at 2.1×
slowdown. I.e. serializing kernel launches in A immunizes it against load
from B on the same GPU. This suggests a timing-dependent race inside a
kernel executed by A (mujoco_warp/warp-compiled kernels are the dominant
GPU work besides NCCL), aggravated by SM contention from B.
Suspicion
A data race in one of the mujoco_warp (or warp 1.17-compiled) kernels that
only manifests under concurrent SM pressure from a second process on the
same device, e.g. shared-memory reuse across kernel instances or
wp.launchscratch buffers. We cannot pin the exact kernel because theerror surfaces asynchronously;
compute-sanitizer racecheckunderco-location would be the next step for us, but filing first in case the
team recognizes the pattern (the 3.12 nefc-clamp fix suggests related
hardening history).
Happy to provide full logs, stack dumps, and a standalone reproducer
script if useful.
Contact: starimpact@126.com