Describe the bug
Summary
On the Hygon DCU mixed FlagGems route, aten::argmin.default over dimension 1
returns an invalid index for a (32, 20) input. Valid column indices are
0..19, but FlagGems returns 20. DAS boxing returns the correct index 19
for the identical input on the same DCU card.
There is no exception or device fault. This is a silent incorrect result. It
matches the symptom observed in Foldseek 3Di encoding, where a 20-column
distance matrix produced state indices outside its valid range.
Environment
- Hardware: Hygon DCU / BW1000, device
flagos:0
- Torch-FL backend:
backends_dcu_flaggems.conf
- FlagGems vendor:
hygon
- PyTorch: 2.10.0
Minimal reproducer
Save the following complete program as repro_argmin_width20.py. The route is
selected before import torch_fl, so each invocation starts a fresh process
with the intended backend configuration.
import argparse
import os
parser = argparse.ArgumentParser()
parser.add_argument("--route", choices=("boxing", "flaggems"), required=True)
args = parser.parse_args()
if args.route == "flaggems":
os.environ["FLAGOS_USE_FLAGGEMS"] = "1"
else:
os.environ.pop("FLAGOS_USE_FLAGGEMS", None)
import torch
import torch_fl
rows, columns, rounded_width = 32, 20, 32
# Logical shape is still (32, 20). The extra physical tail makes the bad
# 32-wide load deterministic without relying on a final-row OOB access.
storage = torch.full((rows * rounded_width,), 1000.0, dtype=torch.float32)
values = storage.as_strided((rows, columns), (columns, 1))
values[:, 19] = 1.0
values[1:, :12] = -1000.0
actual = torch.argmin(values.to("flagos:0"), dim=1)
torch.flagos.synchronize()
reference = torch.argmin(values, dim=1)
actual_cpu = actual.cpu()
different = actual_cpu != reference
print("route:", args.route)
print("actual:", actual_cpu.tolist())
print("reference:", reference.tolist())
print("invalid:", (actual_cpu >= columns).sum().item())
print("mismatched elements:", different.sum().item(), "/", different.numel())
if different.any():
position = different.nonzero()[0].item()
print("first mismatch:", position, actual_cpu[position].item(), reference[position].item())
torch.testing.assert_close(actual_cpu, reference)
Run the same program twice:
export FLAGOS_LOG_DISPATCH=1
python repro_argmin_width20.py --route boxing
python repro_argmin_width20.py --route flaggems
DAS boxing versus mixed FlagGems
Both commands used the same node, flagos:0, logical tensor shape, dtype,
physical storage, values, and program. The only difference was the selected
backend before importing Torch-FL.
| Route |
Dispatch log |
Comparison with CPU reference |
| DAS boxing |
argmin -> cuda |
0 / 32 mismatched elements; largest returned index is 19 |
| mixed FlagGems |
argmin -> flagos_python |
1 / 32 mismatched elements; element 0 is invalid index 20 |
DAS boxing output
[flagos dispatch] argmin -> cuda
actual: [19, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
reference: [19, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
invalid: 0
mismatched elements: 0 / 32
Mixed FlagGems output
[flagos] loading backend config from backends_dcu_flaggems.conf
[flagos dispatch] argmin -> flagos_python
actual: [20, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
reference: [19, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
invalid: 1
mismatched elements: 1 / 32
first mismatch: 0 20 19
AssertionError: Tensor-likes are not equal!
Root cause
File: src/flag_gems/ops/argmin.py, argmin_kernel_opt_k1.
For N=20, the kernel selects a reduction block width of 32:
def heur_block_n(args):
return min(4096, triton.next_power_of_2(args["N"]))
However, it loads all 32 lanes with an unconditional mask:
n_offset = start_n + tl.arange(0, BLOCK_N)
offset = m_offset[:, None] * N + n_offset[None, :]
inp_vals = tl.load(inp + offset, mask=True)
Lanes 20 through 31 must not participate in a 20-column reduction. Instead,
for row 0 they read the first 12 values of row 1. In this reproducer those
values are -1000.0, less than row 0's valid minimum 1.0 at column 19, so
the reduction returns invalid index 20.
Expected behavior and suggested fix
For every reduction width, torch.argmin(values, dim=1) must return an index
in [0, values.size(1) - 1] and match DAS boxing/PyTorch.
Mask both row and column bounds when loading, and use a positive-infinity or
dtype-maximum other value for invalid lanes. For example:
valid = (m_offset[:, None] < M) & (n_offset[None, :] < N)
inp_vals = tl.load(inp + offset, mask=valid, other=max_value)
Please add this (32, 20) case, plus other non-power-of-two widths, as
regression tests.
Describe the bug
Summary
On the Hygon DCU mixed FlagGems route,
aten::argmin.defaultover dimension 1returns an invalid index for a
(32, 20)input. Valid column indices are0..19, but FlagGems returns20. DAS boxing returns the correct index19for the identical input on the same DCU card.
There is no exception or device fault. This is a silent incorrect result. It
matches the symptom observed in Foldseek 3Di encoding, where a 20-column
distance matrix produced state indices outside its valid range.
Environment
flagos:0backends_dcu_flaggems.confhygonMinimal reproducer
Save the following complete program as
repro_argmin_width20.py. The route isselected before
import torch_fl, so each invocation starts a fresh processwith the intended backend configuration.
Run the same program twice:
export FLAGOS_LOG_DISPATCH=1 python repro_argmin_width20.py --route boxing python repro_argmin_width20.py --route flaggemsDAS boxing versus mixed FlagGems
Both commands used the same node,
flagos:0, logical tensor shape, dtype,physical storage, values, and program. The only difference was the selected
backend before importing Torch-FL.
argmin -> cuda0 / 32mismatched elements; largest returned index is19argmin -> flagos_python1 / 32mismatched elements; element 0 is invalid index20DAS boxing output
Mixed FlagGems output
Root cause
File:
src/flag_gems/ops/argmin.py,argmin_kernel_opt_k1.For
N=20, the kernel selects a reduction block width of 32:However, it loads all 32 lanes with an unconditional mask:
Lanes 20 through 31 must not participate in a 20-column reduction. Instead,
for row 0 they read the first 12 values of row 1. In this reproducer those
values are
-1000.0, less than row 0's valid minimum1.0at column 19, sothe reduction returns invalid index 20.
Expected behavior and suggested fix
For every reduction width,
torch.argmin(values, dim=1)must return an indexin
[0, values.size(1) - 1]and match DAS boxing/PyTorch.Mask both row and column bounds when loading, and use a positive-infinity or
dtype-maximum
othervalue for invalid lanes. For example:Please add this
(32, 20)case, plus other non-power-of-two widths, asregression tests.