Conversation

Jarkko Sakkinen

Significant perf updates for GPT-OSS-20B “home-baked” inference in Goosedump:

engine: low-hanging fruit for GPT-OSS-20B

- Add ARM64 NEON and AVX-512 SIMD support.
- Reuse pre-quantized Q8Activation across matrix projections.
- Parallelize active MoE expert computations with Rayon.
- Parallelize GQA causal attention using Rayon.

Signed-off-by: Jarkko Sakkinen <jarkko.sakkinen@iki.fi>

Should be out soon :-)

0
0
0