Skip to content

arm/neon: add LSX implementations for vpaddlq_{s8,s16,u8,u16} - #1456

Merged
mr-c merged 1 commit into
simd-everywhere:masterfrom
jinboson:add-neon-2-lsx-for-vpaddlq
Oct 10, 2026
Merged

mr-c merged 1 commit into
simd-everywhere:masterfrom
jinboson:add-neon-2-lsx-for-vpaddlq

Conversation

@jinboson

@jinboson jinboson commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

These had no LSX path, so the 128-bit pairwise-add-long fell back to

lo = vshrq_n_u16(vshlq_n_u16(vreinterpretq_u16_u8(a), 8), 8);
hi = vshrq_n_u16(vreinterpretq_u16_u8(a), 8);
return vaddq_u16(lo, hi);

four vector operations for what is a single instruction on LSX: vhaddw__ adds adjacent elements pairwise across all sixteen bytes of both operands, so passing the same register twice yields exactly the eight pairwise sums.

This matters for libjpeg-turbo's h2v1/h2v2 downsampling, which is built on vpadalq_u8 (= vaddq_u16(a, vpaddlq_u8(b))). Measured on Loongson 3A5000HV with libjpeg-turbo built -DWITH_SIMDE=1 -DWITH_PROFILE=1 -O2, per-pass throughput from tjbench:

-subsamp 422 downsampling 3850 -> 4320 Msamples/sec (+12.2%)
-subsamp 420 downsampling 4963 -> 6822 Msamples/sec (+37.5%)

and the kernel itself shrinks from 60 to 52 (h2v1) and 76 to 64 (h2v2) instructions.

These had no LSX path, so the 128-bit pairwise-add-long fell back to

  lo = vshrq_n_u16(vshlq_n_u16(vreinterpretq_u16_u8(a), 8), 8);
  hi = vshrq_n_u16(vreinterpretq_u16_u8(a), 8);
  return vaddq_u16(lo, hi);

four vector operations for what is a single instruction on LSX:
vhaddw_<wide>_<narrow> adds adjacent elements pairwise across all sixteen
bytes of both operands, so passing the same register twice yields exactly
the eight pairwise sums.

This matters for libjpeg-turbo's h2v1/h2v2 downsampling, which is built on
vpadalq_u8 (= vaddq_u16(a, vpaddlq_u8(b))).  Measured on Loongson 3A5000HV
with libjpeg-turbo built -DWITH_SIMDE=1 -DWITH_PROFILE=1 -O2, per-pass
throughput from tjbench:

  -subsamp 422  downsampling   3850 -> 4320 Msamples/sec  (+12.2%)
  -subsamp 420  downsampling   4963 -> 6822 Msamples/sec  (+37.5%)

and the kernel itself shrinks from 60 to 52 (h2v1) and 76 to 64 (h2v2)
instructions.
@mr-c
mr-c merged commit c6bc9b8 into simd-everywhere:master Oct 10, 2026
145 of 146 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants