Conversation

Jarkko Sakkinen

Edited 16 hours ago
There's like three ways to stream an "expert":

1. mmap
2. pre-warmed mmap
3. direct mapping of NVMe without roundtrip to page cache

I'm improving my benchmark in Goosedump to see which would be the best. It's already quite fast given a lot of work on getting SIMD and cache alignment play well.

An expert in Mixture-of-Experts actually mean syntactic elements such as commas. It's more like mixture of nitpickers :-) LLM is a text predictor; not like intelligence and that is only expertise it has at its core.
0
0
0