Conversation

Jarkko Sakkinen

Edited yesterday
Bonsai 2 27B with 256k context running on RTX 4090 24 GiB, and hosted by Puu OS.

Puu OS also has preconfigured Qwen 3.8 27B, which fits easily to DGX Spark.

LLM models, when you actually host them, require their own (manual) shoehorning in order to fly nicely in the environment, and there is no exact template for this. Thus, I focus on one all-arounder model with two different quantizations in the base installation of Puu OS. For those that don't know this Bonsai 2 27B is a ternary quantization of Qwen 3.8 27B.

This is a personal achievement for me despite not being a huge achievement overall as I could never imagine I could have a something with a decent context size with the hardware that I have at hand.

I purposely want to do a small AI assisted contribution with this to kernel when I have a chance :-)
2
0
1

@jarkko have you used laguna s2.1? in my experience it beats qwen 3.8 on speed and qualty.

1
0
0
@zygoon My target is a bit different than picking the best model.

Qwen 3.8 is a great reference model for testing a hosting platform, as it is fairly popular and when you combine Qwen 3.8 27B and Bonsai 2 27B you have wide set of quantization of the same base model.

Hosting any model is a small subproject by itself which depends on many variables such as hardware platform and hosting stack. And bunch of random peeking and poking as things develop rapidly and documentation lags behind.

That said, I'm aware of the Laguna model and Poolside, which is one of more relevant and interesting AI startups from US ;-)
1
0
1

@jarkko Have you compared it to one of the Gemma 4 26B-A4B MoE quants or Ornith-1.5-35B ? They're both pretty small and have the room for the context.

1
0
0
@penguin42 nope :-) the key factor here is good reference point not the AI model itself. and getting comprative figures. i just picked one. now i can focus e.g. litellm issues ;-) and see how concurrent sessions work. many things to polish.
1
0
1
@penguin42 https://gitlab.com/puu-os/bootc this is what i'm trying to get in condition for the context
1
0
0
@penguin42 [It might be also first Buildroot based operating system ever providing Wayland GNOME desktop]
0
0
0
@zygoon @penguin42 I started to work on this last March when I got my hands on DGX Spark. I booted up NVIDIA DGX OS and was shocked to see that it is not really a hosting platform. It's just a low quality Ubuntu fork with Ollama and shit. Shame on you NVIDIA ;-) They should really should be IMHO.

I looked up at options and OS projects for hosting AI locally are mostly done by AI lunatics with Claude and package random stuff like LM Studio for instance. They have nothing special or any type of architecture to make them particularly feasible hosting LLM workloads.

So I setup a goal that I want an OS that is like my NAS or OpenWRT in my router that focuses providing one service really well with a dedicated machine and that's I've tryin to achieve since :-)
2
0
0
@penguin42 @zygoon 8900 lines in seven'ish months. a shameful number by today's standards ;-)
0
0
0
@penguin42 @zygoon I've at the same time seen appearance of completely proprietary boxes with similar idea appear on the market with weird subscription deals etc. So before THAT becomes a thing I want to develop a superior solution.
0
0
0