Conversation

would you trade having half the SIMD performance (with 128-bit vectors) for a 30-40% increase in perf on scalar code? 🫠

3
0
0

@never_released

Now I am curious which soc this is about.
Considering simd is used everywhere even for integer code and the vectorizer of both gcc and llvm are getting to the point where they can vectorize even some Lexing loops.

1
0
0

@pinskia Apple chips have that

Apple chips are still at 4x128b SIMD today

whereas the higher end x86 chips are at 2x512b (Xeon for intel and Zen 5 for AMD)

1
0
0
@never_released @pinskia IMO SVE is really way too big of a confounding variable there to treat that as a proper tradeoff...
1
0
1

@palmer @pinskia

given that they implemented SSVE for the SME engine with a 512-bit vlen and that the predicates manipulation is handled in the CPU core instead of the SVE engine, think they're quite a bit of the way there already

1
0
0

@never_released i’d take that trade off any day of the week personally

0
0
0
@never_released @ @hachyderm.io I'm not really an Arm-ologist so I might be wrong here, but IIUC the streaming SVE stuff doesn't allow executing the load/store SVE instructions and thus is much farther decoupled from the integer side of things than full SVE would be.
0
0
0