Bonsai 2 27B with 256k context running on RTX 4090 24 GiB, and hosted by Puu OS.
Puu OS also has preconfigured Qwen 3.8 27B, which fits easily to DGX Spark.
LLM models, when you actually host them, require their own (manual) shoehorning in order to fly nicely in the environment, and there is no exact template for this. Thus, I focus on one all-arounder model with two different quantizations in the base installation of Puu OS. For those that don't know this Bonsai 2 27B is a ternary quantization of Qwen 3.8 27B.
This is a personal achievement for me despite not being a huge achievement overall as I could never imagine I could have a something with a decent context size with the hardware that I have at hand.
I purposely want to do a small AI assisted contribution with this to kernel when I have a chance :-)