This repository bootstraps and smoke-validates practical ROCm cloud VMs for ML and GPGPU development. Provider setup scripts keep cloud-specific orchestration explicit, shared helpers own only genuinely common primitives, and the main validator checks provider-neutral workload capabilities.
The design favors reproducible minimal environments, conservative system mutation, small real workload checks, and maintenance costs that remain realistic for a small project. Optional and experimental paths stay separate from baseline guarantees.
External review and in-scope issue reports are welcome. The project currently accepts issues but not pull requests because the maintainer cannot commit to reviewing arbitrary external diffs. Review invitations are read-only unless the maintainer explicitly establishes a different boundary. See CONTRIBUTING.md for the reporting policy, review lenses, and a reusable handoff brief for human and AI-assisted reviewers. Report sensitive vulnerabilities through the private route documented in SECURITY.md, not through a public issue.
setup/amd_devcloud/,setup/azure/, andsetup/hot_aisle/contain the explicit provider setup paths. The experimental DevCloud path includes its canonical manualdoctlprovisioning guide.setup/common/contains shared setup stages and requirements used by more than one provider.lib/contains small shared shell primitives.validate/contains the provider-neutral validator and representative workload smokes.tests/contains the local test driver, repository-wide ShellCheck entry point, and the human-facing testing guide.security/contains narrow deterministic checks for selected direct upstream provenance receipts; scheduled automation reports drift without accepting new trust metadata.README.md,CONTRIBUTING.md, andSECURITY.mddescribe support, public reporting, and confidential vulnerability reporting respectively.
Installing packages, importing a library, or detecting a GPU does not by itself establish that a ROCm environment is usable. Driver, runtime, compiler, ABI, package, and provider-image interactions can still break real work. This project therefore probes the live environment, runs small representative workloads, and records what actually worked instead of treating installation success as sufficient evidence.
The project is cloud-first because maintaining edge-case or unsupported local ROCm hardware can consume substantial effort for limited practical value. CUDA remains the maintainer's local environment for daily development, while bounded ROCm cloud environments provide portability validation, comparative understanding, and experience with another GPU ecosystem. Different implementations can expose hidden assumptions and provide an independent correctness and portability check, though agreement across them is not proof of correctness.
Cloud images, kernels, drivers, ROCm releases, and package combinations change. The repository therefore treats measured versions and validation results as version-stamped evidence for particular environment generations, not permanent promises about every image a provider has offered or will offer. It validates the VM and host GPU stack intentionally: containers may help selected workloads, but they do not remove the kernel, driver, device-access, permissions, and runtime boundary on which ROCm execution depends.
The goal is to automate recurring setup and debugging pain without becoming production infrastructure as code or a universal cloud abstraction. By capturing experience from systems and cloud environments that many users may not have the hardware access, budget, or time to investigate together, the repository offers reusable, empirically validated guidance for similar use cases while keeping its support and maintenance surface deliberately bounded.
- Target cloud environments:
- Hot Aisle MI300X, single-GPU VM instances
- (optional) Azure
Standard_NV24ads_V710_v5instances- Assumed base image: NVV5 V710 ROCm Linux Image, Gen2 variant as of mid-2026; independently verified to be built on Ubuntu 24.04 (product page on Microsoft Marketplace)
- (experimental) AMD DevCloud Single GPU Plan MI300X Droplet
instances provisioned with the Ubuntu 24.04 Bare OS image
- The root-to-user handoff and in-VM bootstrap through the common minimum baseline and packaged AMD RAPIDS gate are accepted for the tested environment generation. First-class provider support remains pending.
- Provision the provider resource through the canonical
manual
doctlguide; keep that provider-side workflow separate from guest setup.
- Exact kernel, ROCm, and
amdgpukernel module versions are intentionally reported by probe logic in scripts rather than hard-coded here. - Provider support applies to validated environment generations, not every image or software version a provider serves. Supporting a replacement generation does not imply continued support for the previous one. If provisioning cannot reliably select a supported generation, provider support may temporarily be marked transitional or suspended.
- Reproducible, minimal environment setup with Python dependencies constrained selectively when demonstrated compatibility, ABI, or reproducibility reasons justify it
- Safe, resumable bootstrap scripts with reboot handling
- Minimal validation to detect obviously broken environments (e.g., ROCm availability, basic workload execution)
- External inputs that the bootstrap owns use explicit source identities and verification appropriate to their privilege and reproducibility role. Mutable network content is not piped directly into a privileged shell, and selected unexpected origin, redirect, digest, or package-metadata changes fail closed for manual review.
- Project-specific ROCm library and runtime selection remains process-scoped:
- Setup does not add project ROCm paths to system dynamic-loader configuration or run
ldconfigmerely to make them globally visible. - Setup does not persist
LD_LIBRARY_PATH,ROCM_HOME,ROCM_PATH,HIP_PATH, or similar runtime and build-selection variables in shell startup files. - A deliberately dedicated environment with a pinned system stack may add that stack's exact versioned
bindirectory toPATH. This narrow executable-selection exception does not permit persistent library or runtime-selection variables. The conventional/opt/rocmpath is used for a CuPy-family command only after the repository proves that it resolves to the complete ROCm root selected throughPATH; this does not authorize unverified reliance on a mutable alternative. - AMD DevCloud setup maintains that project-owned
/opt/rocm-7.2.3/binprofile block immediately after the ROCm userland stage and prints a temporary current-shellPATHcommand, the pinned root, command-scopedLD_LIBRARY_PATH=/opt/rocm-7.2.3/lib, and CuPy-familyROCM_HOME=/opt/rocmguidance. DevCloud accepts that CuPy path only while the alternatives link resolves exactly to/opt/rocm-7.2.3. A rerun prints the guidance again while completed stages remain skipped; no separate path-report log is required. Setup does not change those runtime-selection variables in the invoking shell. After setup writes the profile block, end the existing SSH login and establish a new one before relying on or accepting the persistentPATH; the printed export is only the immediate-shell alternative. - The common main validator asks the
hipconfigselected throughPATHfor its complete ROCm root, resolves and reports the canonical directory, and supplies that root plus itslibdirectory only within the validator process and its children. Activated upstream CuPy,amd-cupy, and hipCIM checks instead receive command-scopedROCM_HOME=/opt/rocmonly after it resolves to that same canonical root; their loader path remains the explicit versioned library directory. - The validator entry script owns that environment selection and passes selected paths explicitly into helpers that need them; helpers do not implicitly choose or modify the caller's ROCm environment.
- Setup does not add project ROCm paths to system dynamic-loader configuration or run
- Genuinely production ready environment setup
- Bundling the complete application projects or workload implementations listed in Target Workloads, whether the workloads are public OR private; small repository-owned validation smokes remain in scope
- Redistribution of proprietary drivers, runtimes, binaries, and source code
- Academic and course provided program implementations that have not been explicitly approved for public release (please see Private ONLY section under Target Workloads)
- General-purpose firewall reconciliation, source-IP allowlisting, dynamic DNS, or VPN infrastructure
- A universal supply-chain security framework, package manager, attestation service, or guarantee that canonical upstreams, release systems, runners, or published artifacts cannot be compromised. HTTPS and repository-recorded hashes provide bounded evidence, not absolute provenance.
- Expansion to entirely new cloud providers, GPU-platform families, or workload categories is currently frozen while the existing baseline is completed and stabilized. The already-planned experimental AMD DevCloud path remains in scope.
- Acting as a general unofficial ROCm platform-enablement layer:
- Narrow unsupported-platform experiments may be used for concrete diagnostics, upstream bug reproduction, or explicitly approved compatibility investigations.
- An isolated successful build, JIT kernel, example, or smoke test does not create a baseline support promise, maintenance obligation, or precedent for additional unsupported platform combinations.
- The existing Azure
gfx1101hipCollections diagnostic patch and optional CuPygfx1100fallback with a manual HSA override are narrow exceptions, not precedents for broader unsupported-platform support. - Azure Pro V710 support does not currently include hipCIM, hipDF, or general RAPIDS-on-RDNA support. Such support would require sufficiently explicit and mature upstream targeting followed by deliberate adoption within this repository's scope and maintenance budget.
The projects below motivate the environment capabilities checked here. Each repository-owned smoke exercises only a small, explicit library or runtime contract; it does not emulate or exhaustively cover the corresponding application. Listing a workload also does not imply that every provider or setup tier supports every optional dependency it may require.
In this documentation, an execution motif is one deliberately isolated GPU operation and data-access pattern taken from a larger applied workload. It is useful for comparing how hardware and software stacks execute that pattern, but it is not the complete application, an application correctness test, or a bootstrap support guarantee. The longer-term MI300X-versus-H200 motif study is a separate project; this repository records only the environment context that overlaps its existing bootstrap scope.
- Custom Canny edge detection operator written with CuPy from first principles:
- This bootstrap repository does not bundle or run that complete application. Its current bundled NumPy/CuPy main smoke test checks boolean-mask indexing and paired multi-axis integer advanced-index gathering, both with normally allocated results, alongside the documented custom-kernel and numerical checks.
- That bundled smoke test does not currently call NumPy or CuPy
takeexplicitly or validate a take-style gather into preallocated output storage. The separate Canny operator is already transitioning to explicit NumPy/CuPy take usage. The packaged hipCIM smoke is an independent end-to-end comparison of upstream cuCIM Canny with scikit-image Canny, not execution of the custom application.
- ResNet-50 vs ViT-B/16 classification and resource usage performance project
- ComfyUI 2D image generation:
- FLUX.x [dev] (where "x" is 1 or greater)
- Ollama-powered LLM inferencing:
- Use local
open-webuiinstall for GUI frontend, for streamlined experience and software maintenance; please see official Open WebUI documentation for more relevant configuration details - Target model families:
- Gemma (3 and above)
- gpt-oss (2025 version and newer)
- Mistral Small (3 and above)
- (optionally) Llama (3.3 and above)
- Please see Additional Notes sub-section for why Ollama runtime setup and validation are intentionally excluded from bootstrap scripts.
- Use local
- Applied GMV aggregation:
- Filter completed order items, derive
order_datefromcreated_at, group by(order_date, category), and calculateSUM(quantity * unit_price),SUM(quantity), andCOUNT(*). - This workload requires hipDF. Only the experimental AMD DevCloud path is eligible to support it, after packaged hipDF is deliberately adopted and validated; it is not a current Hot Aisle, Azure, or baseline guarantee.
- This full workload is distinct from the isolated GPU execution motif used to study one low-arithmetic-intensity, irregular-memory-access portion of hash-based GROUP BY.
- Filter completed order items, derive
- HIP micro-benches:
- Elias Konstantinidis's mixbench
- Hash-based GROUP BY group discovery and probing with composite keys,
represented by cuco/hipCollections
static_set-style insertion and probing. Integer-encoded composite keys may be prepared outside the isolated primitive's timed region. - The existing hipCollections
static_mapdirect-aggregation and host-bulkinsert_or_applyimplementation remains a separate experiment, not the current representative execution motif. - (optional) gather-GEMM via Triton with scrambled row maps
- MLP in CuPy from first principles trained with vanilla mini-batch SGD and MSE loss, with training, validation, and testing on MNIST_784 dataset provided by Scikit-Learn OpenML
- Comparison of classification performance and loss curves of various custom Voice Activity Detection models implemented in PyTorch
- CUDA/HIP custom co-rank based iterative mergesort from first principles
- Ollama is intentionally not installed by the bootstrap scripts:
- Ollama installation and model pulls are manual because, as of mid-2026, models pulled from the default Ollama registry cannot be pinned to exact immutable versions.
- There has also been documented partial coupling between Ollama runtime versions and model versions hosted on the default registry.
- On ROCm systems, fragile driver/runtime/hardware/etc. combinations may also cause GPU hangs or other inference-time failures. Therefore, this repository also does not treat Ollama inference as an automated bootstrap validation step.
- High-capacity solid state storage (100GB+) recommended for Flux-class and LLM model weights when configuring cloud environment(s).
- Version-stamped CuPy custom-kernel acceptance evidence:
- On 2026-09-05, the tiny semantic and million-attempt scale
atomicCAS/template cases passed forint32,int64, and high-rangeuint64on a Hot Aisle MI300X VF with ROCm 7.2.4, source-built CuPy 14.1.1, and NumPy 2.5.2. - On 2026-09-09, the same cases passed on an Azure Radeon Pro V710 MxGPU with ROCm 7.2.0, source-built CuPy 14.1.1, and NumPy 2.5.3. The direct Numba 0.67.0 smoke also confirmed that Numba selected the intended TBB threading layer.
- The direct CuPy and Numba smokes and the complete strict main validator passed on Azure. The complete validator took approximately 2 minutes 35 seconds. These results record the tested environments rather than promising compatibility with every future image or package generation.
- On 2026-09-05, the tiny semantic and million-attempt scale
- The optional source-built CuPy environment is requested during setup with
--source-built-cupy-env-setup.- Requesting this environment also installs the selected oneAPI TBB libraries during the earlier APT phase, before any Python virtual environment is created.
- When present, validation runs both the CuPy custom-kernel smoke and the Numba smoke as one complete environment check.
- If the environment is absent, validation skips that entire check unless
--fail-on-no-source-built-cupy-envrequires it to be present. - The Numba smoke requires the selected threading layer to be TBB. Canny requires CuPy and Numba-relevant behavior, while the explicit TBB selection is retained for the private MLP workload's parallel CPU activation functions.
- Cross-platform execution of a representative smoke can provide an independent portability and correctness check because different compiler, runtime, and hardware implementations may expose different hidden assumptions.
- A CUDA-side development check does not make CUDA a supported bootstrap environment and does not substitute for ROCm acceptance.
- Keep bootstrap coverage at the level of durable platform capabilities rather than bundling private workload regressions or reenacting historical upstream bugs.
- Version-stamped experimental AMD DevCloud handoff evidence:
- On 2026-09-09, the root-to-user handoff was accepted on a disposable Ubuntu 24.04 VM. Testing covered fresh account creation, the exact
*password marker, a separate key-only SSH login, full passwordless sudo, idempotent reruns, and conservative refusal and preservation of a conflicting existing account. - On 2026-09-10, the root-only sudoers-rejection test passed both locally and on the disposable VM. The real ordinary-user preflight then passed from separate SSH sessions for two independently created valid users, including their home, repository access, effective required groups, and full noninteractive-sudo contract.
- The retained root and ordinary-user clones were checked after credential-prompted HTTPS cloning. None had a configured credential helper, a standard Git credential-store file, or credentials embedded in the origin URL.
- These results establish the initial account handoff and preflight behavior only; they do not establish ROCm, GPU, or packaged-workload readiness on AMD DevCloud.
- On 2026-09-09, the root-to-user handoff was accepted on a disposable Ubuntu 24.04 VM. Testing covered fresh account creation, the exact
- Version-stamped AMD SMI architecture-query evidence:
- On 2026-09-10, Azure ROCm 7.2.0 with AMD SMI 26.2.1 reported an AMD Radeon Pro V710 MxGPU, native
gfx1101, and driver 6.16.13 throughamd-smi static --gpu 0 --asic --driver --json. - On the same date, Hot Aisle ROCm 7.2.4 with AMD SMI 26.2.2 reported an AMD Instinct MI300X VF, native
gfx942, and driver 6.16.13 through the same JSON fields. - The common architecture query now uses AMD SMI rather than deprecated ROCm SMI. The matching field shape across those two observations supports one narrow shared parser; it does not make either provider's package names or driver version a cross-provider requirement.
- On 2026-09-10, Azure ROCm 7.2.0 with AMD SMI 26.2.1 reported an AMD Radeon Pro V710 MxGPU, native
- Experimental packaged hipCIM recipe evidence:
- As of 2026-09-10, the maintainer reports that hipCIM Canny from the ROCm 7.2.0 AMD Python index passed on the intended ROCm 7.2.3 system-stack combination.
- This supports developing one pinned DevCloud recipe. It is not yet repository acceptance of the automated setup, the future packaged RAPIDS gate, other hipCIM functions, hipDF, or a general cross-version package matrix.
- Version-stamped pristine AMD DevCloud evidence:
- On 2026-09-12, an actual single-GPU Bare OS instance reported Ubuntu 24.04.4 under KVM, kernel 6.8.0-124, and AMD PCI device
1002:74b5before AMDGPU installation. - The AMD device appeared as a PCI processing accelerator, with no loaded
amdgpumodule or/dev/kfd. A Virtio display already provided/dev/dri, so generic DRI-directory presence is not evidence that the AMD GPU stack is installed or ready. - No AMDGPU/ROCm packages, AMD repositories, matching
/optpaths, DKMS command, or ROCm tools were detected. The DigitalOcean Ubuntu mirror and droplet-agent repository were normal provider state rather than AMD stack state. - UFW matched the exact shared FRESH fingerprint and the SSH session used server port 22. These observations define the initial admission fixture without pinning incidental hostname, CPU, storage, kernel patch, Ubuntu point release, mirror URL, or pending-upgrade details.
- On 2026-09-12, an actual single-GPU Bare OS instance reported Ubuntu 24.04.4 under KVM, kernel 6.8.0-124, and AMD PCI device
- Version-stamped experimental AMD DevCloud driver-checkpoint evidence:
- On 2026-09-13, the complete repository-local shell suite, its separate root-only scenario, and the applicable root handoff, separate SSH login, UFW transition, and reboot-acknowledgement checks passed on a disposable Ubuntu Server 24.04 VMware VM. That VM supplied host-side workflow evidence only, not AMDGPU or GPU evidence.
- An actual fresh AMD DevCloud Bare OS instance then completed the root-to-user handoff, ordinary-user preflight, UFW FRESH-to-BASELINE transition, existing-package upgrade, tmux installation, acknowledged system-upgrade reboot, pinned AMD repository bootstrap, pinned AMDGPU DKMS and versioned AMD SMI installation, and acknowledged driver reboot.
- After reboot, the running
6.8.0-124-generickernel reported installed DKMS moduleamdgpu/6.16.13-2327507.24.04, a loadedamdgpumodule,/dev/kfdowned by therendergroup, and one AMD Instinct MI300X VF with nativegfx942and AMD SMI driver version6.16.13. - The transitive
rocm-core7.2.3dependency maintained/opt/rocmthrough Debian alternatives with canonical target/opt/rocm-7.2.3. Setup recognizes that exact link as a driver-stage artifact but continues to use the versioned root directly. A completed-state rerun skipped every prior stage and repeated the post-driver verifier successfully. - This acceptance establishes the experimental setup through the driver checkpoint. It does not yet establish the ROCm userland, common validator baseline, packaged RAPIDS environment, or workload readiness.
- Version-stamped experimental AMD DevCloud common-baseline evidence:
- On 2026-09-17, a fresh single-GPU Bare OS instance at reviewed commit
49df627completed the root-to-user handoff, exact UFW baseline, both acknowledged reboot boundaries, pinned repository and driver stages, completerocm7.2.3userland, versioned ROCm profile block, common system prerequisites, and baseline Python environment. The system upgrade changed the reported point release from Ubuntu 24.04.4 to 24.04.5 while retaining kernel6.8.0-124-generic. - The post-setup environment report observed AMDGPU DKMS
6.16.13-2327507.24.04, AMD SMI26.2.2.70203-90~24.04, one AMD Instinct MI300X VF with nativegfx942,rocm7.2.37.2.3.70203-90~24.04, HIP7.2.53211-c2d9476115, Python 3.12.3, and PyTorch 2.11.0 with ROCm 7.2. - The common default validator passed its strict UFW and ROCm checks, HIP and hipCollections smokes, PyTorch audio/codec CPU ABI smoke, ResNet-50 and ViT-B/16 GPU forward-and-backward checks, and Triton fp16 matmul smoke. The absent optional source-built CuPy and ComfyUI environments were skipped as designed.
- The accepted post-setup workflow is
./validate/bin/validate_main.shfollowed by the read-only./setup/amd_devcloud/bin/amd_devcloud_acceptance_probe.sh. Its standard output may be retained as acceptance evidence when useful; setup does not own a persistent report artifact and does not depend on creating one. - This acceptance establishes the experimental automated ROCm userland and common minimum baseline. It does not yet establish the packaged AMD RAPIDS environment, hipCIM correctness gate, optional hipDF components, or first-class DevCloud support.
- On 2026-09-17, a fresh single-GPU Bare OS instance at reviewed commit
- Version-stamped experimental AMD DevCloud packaged-environment evidence:
- On 2026-09-18, a fresh single-GPU Bare OS instance that began from reviewed commit
146edaccompleted the optional packaged AMD RAPIDS setup. After the requirements correction, commitc6850c5rebuilt the packaged stage withamd-cupy13.5.1,amd-hipcim25.10.0, scikit-image 0.25.2, scikit-learn 1.8.0, NumPy 2.5.3, and Numba 0.67.0;pip checkreported no broken requirements. - The common validator with
--fail-on-no-packaged-amd-rapids-envpassed the complete baseline plus the packaged CuPy custom-kernel, Numba/TBB, and hipCIM Canny checks. The Canny comparison disagreed on zero of 20,480 pixels, confirming the adopted zero-percent tolerance on the tested MI300X VF after the same fixture had also measured zero disagreement during CUDA-side development. - The read-only environment report reconfirmed Ubuntu 24.04.5, kernel
6.8.0-124-generic, AMDGPU DKMS6.16.13-2327507.24.04, AMD SMI driver 6.16.13, nativegfx942,rocm7.2.37.2.3.70203-90~24.04, HIP7.2.53211-c2d9476115, and Python 3.12.3. A completed-state setup rerun recognized the UFW baseline, repeated the driver and path checks, skipped every setup and reboot stage including the packaged environment, and exited successfully. - This acceptance establishes the automated packaged hipCIM/CuPy environment and its strict validation gate for this experimental generation. It does not establish optional hipDF components, arbitrary ROCm/Python package combinations, or first-class DevCloud support.
- On 2026-09-18, a fresh single-GPU Bare OS instance that began from reviewed commit
- Version-stamped experimental AMD DevCloud current-head evidence:
- On 2026-10-04, a fresh single-GPU Ubuntu 24.04 Bare OS instance at reviewed commit
a6b6dedcompleted the root-to-user handoff, a separate key-only ordinary-user SSH login, the exact UFW baseline, both acknowledged reboot boundaries, ROCm userland, the common baseline, and the packaged AMD RAPIDS environment. A completed-state rerun recognized the managed state, repeated the driver and path checks, skipped every completed setup and reboot stage, and exited successfully. - The acceptance probe observed Ubuntu 24.04.5, kernel
6.8.0-142-generic, AMDGPU DKMS6.16.13-2327507.24.04, AMD SMI26.2.2.70203-90~24.04, one AMD Instinct MI300X VF with nativegfx942,rocm7.2.37.2.3.70203-90~24.04, HIP7.2.53211-c2d9476115, Python 3.12.3, and PyTorch 2.11.0 with ROCm 7.2. The packaged environment containedamd-cupy13.5.1,amd-hipcim25.10.0, NumPy 2.5.3, and Numba 0.68.0;pip checkreported no broken requirements. - The strict common validator passed the complete baseline and packaged AMD RAPIDS gate after initial setup and again after the completed-state rerun. The packaged CuPy custom-kernel cases, Numba/TBB checks, and hipCIM Canny comparison passed; the Canny comparison disagreed on zero of 20,480 pixels.
- Setup confirmed
/opt/rocm-7.2.3as the pinned root, accepted command-scopedROCM_HOME=/opt/rocmonly while that conventional path resolved to the pinned root, and retained/opt/rocm-7.2.3/libas the command-scoped loader path. The validator independently reported/opt/rocm-7.2.3as both its PATH-selected and canonical ROCm root. A separate SSH login after setup prepended/opt/rocm-7.2.3/bintoPATH, selected the versionedhipconfig, and reported/opt/rocm-7.2.3fromhipconfig --path; the preexisting setup shell correctly remained unchanged. - A raw tmux-backed terminal transcript was retained outside the repository and its SHA-256 matched after retrieval. During the reboot workflow, termination of the SSH/tmux session left the maintainer's local XFCE Terminal 1.1.3 under Xfce 4.18 needing the local
resetcommand. This is a client-terminal recovery observation, not a VM or bootstrap failure; behavior in other terminals was not tested. - This acceptance closes the current-head revalidation requirement for the experimental DevCloud generation. It does not establish optional hipDF components, arbitrary ROCm/Python package combinations, a download-throughput baseline, or first-class DevCloud support.
- On 2026-10-04, a fresh single-GPU Ubuntu 24.04 Bare OS instance at reviewed commit
- Version-stamped experimental AMD DevCloud provisioning evidence:
- On 2026-10-03,
doctl1.177.0 authenticated an AMD SSO-backed DevCloud API token againsthttps://api.devcloud.amd.com. Read-only discovery exposed the account-visiblegpu-mi300x1-192gb-devcloudsize inatl1at the observed price of $1.99 per hour and theubuntu-24-04-x64image. These identifiers, availability, and price are an account- and date-specific receipt rather than a provider promise. - An explicitly confirmed
doctlcreate used the discovered size, region, image, and an existing account SSH key. The separately managed existing cloud firewall was attached before SSH access. The resulting Droplet appeared in the AMD DevCloud portal at the first manual refresh approximately five to eight minutes after creation; this proves portal visibility, not the exact propagation time. - After the complete guest acceptance and evidence retrieval, an explicitly confirmed
doctldelete succeeded. A subsequent API lookup of the exact Droplet returned404, and the Droplet was absent at the first manual portal refresh less than four minutes later. The repository did not automate either lifecycle decision. - The tested path made no project-management API calls and relied on the account's existing default-project placement. The canonical guide should preserve that boundary rather than request project-mutation scopes or automate project creation, selection, or reassignment. A user who needs non-default project organization must review and perform that separate provider-side operation deliberately.
- On 2026-10-03,
- On a clean Ubuntu 24.04 x86-64 control environment, cross-check the
current upstream
doctlinstallation documentation and selected release, then execute the exact pinned standalone-install example insetup/amd_devcloud/README.mdend to end.- This is a local documentation check; it requires neither an API token nor a paid DevCloud resource.
- Keep the example explicitly non-authoritative even after it passes. Do
not generalize one result to other operating systems, architectures, or
doctlreleases. - Do not schedule another paid lifecycle round solely to repeat the provisioning, firewall attachment, SSH, and deletion commands already accepted on 2026-10-03 and 2026-10-04.
- Replace the current Kitware and Intel signing-key streams with controlled unprivileged downloads, exact expected origins and redirect behavior, and reviewed key fingerprints or artifact digests before installing system trust state. Preserve the existing package-source and provider boundaries.
- Give the optional fastfetch release path an explicit reviewed version, canonical repository identity, redirect policy, and artifact verification, or skip it when those checks cannot be maintained. Do not make it a required workload dependency.
- Review direct AMD Python-package ownership across the AMD and default
Python indexes. Prevent same-name source substitution without using
--no-deps, mirroring transitive metadata, or changing the accepted packaged recipe before resolver behavior is understood and tested. - Audit source clones and recursive submodules for exact host, owner, repository, and immutable-ref contracts where reproducibility warrants them. Keep this bounded to current dependencies; do not adopt vendoring or a general dependency framework by default.
- Perform one bounded public-release audit of the current tree and reachable Git history:
- No committed credentials, tokens, private keys, private URLs, unintended personal data, unredacted logs, environment artifacts, private application or coursework code, or unexplained binary files were found.
- Included patches and derived workflow material have documented upstream provenance and licensing treatment. Public support claims and the issue-only external review policy remain consistent with the repository contract.
- Existing Git author metadata includes the maintainer's institutional email address; the maintainer reviewed and accepted that disclosure, so no history rewrite is required.
- This was a concrete release checkpoint, not a speculative forensic review or precedent for routine history rewriting. A future history rewrite still requires an actual finding.
- Add instructions for using this repository on Hot Aisle MI300X:
- Document the Hot Aisle setup command.
- Explain
--show-plan-onlyand the optional setup flags. - Explain expected reboot and rerun behavior.
- Explain that setup installs tmux before the intentional reboot, then document starting or resuming a named tmux session after reconnecting for long post-reboot stages.
- Document the standard validation command and optional validation flags.
- Document the manual Ollama setup and validation path.
- Document the final manual ComfyUI setup and validation steps.
-
Document recovery from a failed setup stage:
- Inspect and resolve the original error.
- Start with inexpensive, read-only inspection and the narrowest reversible recovery that can address the failed stage. Distinguish required recovery from optional diagnostics, and do not assume spare hardware or a fresh paid-cloud deployment when preserved state can answer the question.
- Identify the virtual environment, repository clone, or other artifacts owned by the failed stage.
- Delete stage-owned artifacts only when rebuilding them is necessary.
- Delete only the relevant completed or pending marker when the documented recovery procedure specifically requires it.
- Rerun the setup script so completed stages remain skipped.
- Document the exact marker and artifact locations instead of recommending broad directory deletion.
-
Document approximate setup and validation times:
- Minimal setup and validation on Hot Aisle MI300X.
- Full setup and validation on Hot Aisle MI300X.
- Minimal setup and validation on Azure Pro V710.
- Full setup and validation on Azure Pro V710.
- Present these as rough observations rather than guarantees.
- Revalidate and document the
filelockshutdown warning observed with ComfyUI version 0.19.0 on Azure Pro V710:- An
Exception ignored ImportErrormessage may appear when stopping ComfyUI withCTRL+C. - Previous testing found no resulting corruption of model weights or workflows.
- Users should still save open workflows before exiting the Web UI or stopping the server.
- Confirm that the linked upstream issue comment still supports describing this as harmless Python dependency noise.
- An
- Document that some Hot Aisle VM images may require user confirmation during
apt-get upgradebecause of preinstalled kernel upgrades.- Revalidate this behavior against the current noninteractive upgrade logic before publishing the note.
- Document the experimental attention warning emitted by PyTorch 2.11.x on Azure Pro V710.
- Explain that the repository intentionally leaves the experimental attention implementation disabled by default.
- Treat enabling it as an optional per-workload experiment.
- Audit whether the ordinary Azure bootstrap user can inspect system logs needed for routine ROCm and cloud debugging without routinely using sudo:
- Test the required
journalctlaccess on the current Azure image. - Treat non-root log access as the behavioral requirement rather than assuming a particular group is sufficient.
- Inspect the current group and ACL behavior; do not assume membership in
admalone solves it. - Keep this as a bounded permissions audit rather than a general Linux authorization framework.
- Test the required
-
Document why hipDF source builds are excluded from the supported bootstrap:
- Local testing found compatibility failures across newer ROCm ecosystem combinations.
- Installing an older ROCm stack over provider-installed ROCm libraries risks package conflicts or package stomping.
- Forward-porting hipDF and maintaining downstream compatibility patches is outside the repository's scope.
- Keep packaged hipDF validation as a possible later extension of the explicitly aligned experimental AMD DevCloud path, separate from the supported Hot Aisle and Azure baseline.
-
Adopt packaged hipDF later within the experimental DevCloud environment when a concrete workload needs it:
- Keep its direct requirements and ROCm 7.2.3 AMD Python index separate from the ROCm 7.2.0 hipCIM/CuPy requirements input.
- Resolve the more patch-sensitive hipDF group first and the hipCIM/CuPy group second, using ordinary dependency metadata for both, followed by
pip checkand complete workload validation. This tested ordering is not a claim of general index interoperability or wheel provenance. If the later operation breaks hipDF, fail and revisit demonstrated constraints rather than using--no-depsor mirroring the transitive closure. - Packaged hipDF may join the broader RAPIDS gate when that keeps orchestration simpler. Do not source-build or forward-port hipDF, create a compatibility solver, or require feature parity with Hot Aisle and Azure.
-
Revalidate conventional CuPy ROCm-root selection on the next applicable Hot Aisle and Azure provider acceptance rounds:
- Confirm the supported Hot Aisle and Azure source builds use command-scoped
ROCM_HOME=/opt/rocmonly after it resolves to thehipconfig-selected canonical root. - This verifies alignment with upstream CuPy's conventional-path recommendation and AMD's corresponding ROCm CuPy-fork build guidance; it does not make
/opt/rocma general unverified runtime selector.
- Confirm the supported Hot Aisle and Azure source builds use command-scoped
-
Document why the supported Hot Aisle and Azure paths build CuPy from source:
- Current ROCm versions and less commonly validated targets such as Azure Pro V710
gfx1101may not be adequately supported by prebuilt wheels. - The source build uses the native detected HIP architecture by default.
- Current ROCm versions and less commonly validated targets such as Azure Pro V710
-
Document
CUPY_BUILD_GFX11_FALLBACK=1:- When the native architecture is
gfx1101, this option also builds CuPy forgfx1100. - This makes an optional
HSA_OVERRIDE_GFX_VERSION=11.0.0experiment possible because the requiredgfx1100code objects are present. - Without compatible code objects, using the override may cause a segmentation fault or another runtime failure.
- Never enable the HSA override by default.
- When the native architecture is
-
Document the motivation and limits of the
gfx1100fallback:- Some tested
gfx1101problem shapes selected substantially slower native rocBLAS or hipBLAS kernels. - Treat override-assisted results as optional compatibility or performance experiments, not baseline validation results.
- Do not promise that the override improves every workload.
- Some tested
-
Document the automated validation boundary:
- Default validation does not set
HSA_OVERRIDE_GFX_VERSION=11.0.0automatically. - The main validator refuses to run when it inherits a non-empty
HSA_OVERRIDE_GFX_VERSION, so its results remain native baseline validation. - Validation of the dual-target build is left to the user because there is currently no known documented and reliable way to independently determine, from the resulting CuPy build itself, which HIP architectures were compiled into it.
- Document how users can manually run
cupy_numpy_smoke.pywith the override as a runtime sanity check. - Make clear that this manual check confirms only whether the tested workload runs successfully. It does not independently enumerate or verify every architecture compiled into CuPy.
- Warn that running with an incompatible override may cause a segmentation fault or another runtime failure.
- Default validation does not set
-
Document the exact architecture scope of the dual-target build:
CUPY_BUILD_GFX11_FALLBACK=1addsgfx1100only when the detected native architecture isgfx1101.- Other RDNA 3 targets, including discrete GPU targets such as
gfx1102and integrated GPU targets such asgfx1103andgfx1104, do not receive the same dual-target build behavior. - There are no equivalent fallback environment variables for those targets or for other architecture families.
- Broader support for these architectures in newer ROCm releases does not imply support by this repository's CuPy build logic.
- This repository currently limits the fallback behavior to the validated Azure Pro V710
gfx1101use case. - Supporting additional targets remains subject to the scope freeze and requires a validated repository use case rather than being added solely for architecture-family symmetry.
- Document that unauthenticated Hugging Face Hub warnings are expected:
- Environment setup does not assume that the user will provide an API token.
- Users may configure a token manually after considering the security implications of storing credentials on a cloud VM.
- Document the fastfetch validation boundary:
- Azure users may optionally install fastfetch from an official GitHub release during setup.
- fastfetch is not required by any supported workload, so the main validation script does not validate it.
- Users can run
fastfetchmanually to confirm the optional installation. - After the DevCloud setup path and Hot Aisle quick-start documentation are complete, decide whether DevCloud should offer the same optional GitHub-release convenience. If adopted, share the identical release pin and package filename rather than duplicating them across provider variables.
- Add a note that links to custom public projects will be added when those projects are ready to be showcased with this repository.
-
After the Hot Aisle common-use path is established, add equivalent Azure Pro V710 setup and validation instructions.
-
Add structured validation receipt output, including the shared UFW classification, after a common summary and reporting design is justified.
-
If the maintainer explicitly decides to permit additional maintainers, document a lightweight maintainer role and selection policy:
- Treat demonstrated collaboration, access to enough cloud resources for meaningful validation, and relevant ROCm, GPU, Linux, cloud, or workload experience as candidate considerations rather than current acceptance criteria.
- Keep important project knowledge and decisions reconstructible from the repository rather than private conversations or one maintainer's memory.
- Do not create governance machinery before the role and actual need are approved.
-
If recurring in-scope remote HIP/C++ development demonstrates a need, evaluate optional provider-aware clangd setup guidance:
- Determine how clangd should use the provider or system ROCm installation before deciding whether bootstrap automation is justified.
- Treat it as an optional developer convenience rather than a baseline capability or validation gate unless a later supported workflow proves otherwise.
- Do not promise it for every provider. The separate MI300X-versus-H200 execution-motif study does not currently create a tooling requirement for this repository.
-
Add a linked table of contents if the README becomes too long to navigate comfortably.
- All code and documentation in this repository were drafted with assistance from ChatGPT and Gemini models publicly available circa 2026. Architecture, support boundaries, semantics, interfaces, invariants, acceptance or rejection decisions, and final integration approval remain human-owned.
- Human review is contract-first and risk-weighted. It includes deeper inspection of high-risk behavior and selective implementation review supported by tests and behavioral summaries; it does not claim exhaustive line-by-line inspection of every agent-assisted change.
- This disclosure provides transparency about the development process; it is not by itself proof of originality, complete provenance, or license compliance.
- Reports of suspected similarity to third-party material or licensing concerns are welcome so that affected code or documentation can be reviewed and, when appropriate, replaced.