Skip to content

# [Silent incorrect result] Hygon FlagGems aten::argmin.default does not mask the tail of a non-power-of-two reduction #6447

Description

@wzx226226

Describe the bug

Summary

On the Hygon DCU mixed FlagGems route, aten::argmin.default over dimension 1
returns an invalid index for a (32, 20) input. Valid column indices are
0..19, but FlagGems returns 20. DAS boxing returns the correct index 19
for the identical input on the same DCU card.

There is no exception or device fault. This is a silent incorrect result. It
matches the symptom observed in Foldseek 3Di encoding, where a 20-column
distance matrix produced state indices outside its valid range.

Environment

  • Hardware: Hygon DCU / BW1000, device flagos:0
  • Torch-FL backend: backends_dcu_flaggems.conf
  • FlagGems vendor: hygon
  • PyTorch: 2.10.0

Minimal reproducer

Save the following complete program as repro_argmin_width20.py. The route is
selected before import torch_fl, so each invocation starts a fresh process
with the intended backend configuration.

import argparse
import os

parser = argparse.ArgumentParser()
parser.add_argument("--route", choices=("boxing", "flaggems"), required=True)
args = parser.parse_args()
if args.route == "flaggems":
    os.environ["FLAGOS_USE_FLAGGEMS"] = "1"
else:
    os.environ.pop("FLAGOS_USE_FLAGGEMS", None)

import torch
import torch_fl

rows, columns, rounded_width = 32, 20, 32
# Logical shape is still (32, 20). The extra physical tail makes the bad
# 32-wide load deterministic without relying on a final-row OOB access.
storage = torch.full((rows * rounded_width,), 1000.0, dtype=torch.float32)
values = storage.as_strided((rows, columns), (columns, 1))
values[:, 19] = 1.0
values[1:, :12] = -1000.0

actual = torch.argmin(values.to("flagos:0"), dim=1)
torch.flagos.synchronize()
reference = torch.argmin(values, dim=1)
actual_cpu = actual.cpu()
different = actual_cpu != reference
print("route:", args.route)
print("actual:", actual_cpu.tolist())
print("reference:", reference.tolist())
print("invalid:", (actual_cpu >= columns).sum().item())
print("mismatched elements:", different.sum().item(), "/", different.numel())
if different.any():
    position = different.nonzero()[0].item()
    print("first mismatch:", position, actual_cpu[position].item(), reference[position].item())
torch.testing.assert_close(actual_cpu, reference)

Run the same program twice:

export FLAGOS_LOG_DISPATCH=1
python repro_argmin_width20.py --route boxing
python repro_argmin_width20.py --route flaggems

DAS boxing versus mixed FlagGems

Both commands used the same node, flagos:0, logical tensor shape, dtype,
physical storage, values, and program. The only difference was the selected
backend before importing Torch-FL.

Route Dispatch log Comparison with CPU reference
DAS boxing argmin -> cuda 0 / 32 mismatched elements; largest returned index is 19
mixed FlagGems argmin -> flagos_python 1 / 32 mismatched elements; element 0 is invalid index 20

DAS boxing output

[flagos dispatch] argmin -> cuda
actual:    [19, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
reference: [19, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
invalid: 0
mismatched elements: 0 / 32

Mixed FlagGems output

[flagos] loading backend config from backends_dcu_flaggems.conf
[flagos dispatch] argmin -> flagos_python
actual:    [20, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
reference: [19, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
invalid: 1
mismatched elements: 1 / 32
first mismatch: 0 20 19
AssertionError: Tensor-likes are not equal!

Root cause

File: src/flag_gems/ops/argmin.py, argmin_kernel_opt_k1.

For N=20, the kernel selects a reduction block width of 32:

def heur_block_n(args):
    return min(4096, triton.next_power_of_2(args["N"]))

However, it loads all 32 lanes with an unconditional mask:

n_offset = start_n + tl.arange(0, BLOCK_N)
offset = m_offset[:, None] * N + n_offset[None, :]
inp_vals = tl.load(inp + offset, mask=True)

Lanes 20 through 31 must not participate in a 20-column reduction. Instead,
for row 0 they read the first 12 values of row 1. In this reproducer those
values are -1000.0, less than row 0's valid minimum 1.0 at column 19, so
the reduction returns invalid index 20.

Expected behavior and suggested fix

For every reduction width, torch.argmin(values, dim=1) must return an index
in [0, values.size(1) - 1] and match DAS boxing/PyTorch.

Mask both row and column bounds when loading, and use a positive-infinity or
dtype-maximum other value for invalid lanes. For example:

valid = (m_offset[:, None] < M) & (n_offset[None, :] < N)
inp_vals = tl.load(inp + offset, mask=valid, other=max_value)

Please add this (32, 20) case, plus other non-power-of-two widths, as
regression tests.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions