Conversation

Jarkko Sakkinen

The screen cast below demonstrates ReadSeek 0.8.0 with hand-optimized inference engine for Qwen3-VL-2B. Given how shitty the CPU is (AMD Ryzen 5 PRO 4650) in this laptop, the performance is not all that bad, and it will of course only improve over time :-)

ReadSeek will retain itself as being based only on CPU inference as a core design decision; ReadSeek *is* smart but *does not* get in the way. That's why hand-optimized inference is even more so important aspect of this project.

The reasoning roots allowing to retain CUDA, Metal, unified memory and VRAM under user's full control. You can run your local models and ReadSeek will always do its job as an ubiquitous partner :-)

#pi #opencode #coding #agent
1
0
0

Jarkko Sakkinen

Edited 21 days ago
This work is highly inspired by the gist of Dwarf Star 4 and Colibri. I was thinking what could practical application for those ideas? Then it came to me that optimizing inference for distilled models opens a whole world of possibilities for command-line tools.
1
0
0
I also think there's a lot of future innovation to be done in model streaming from e.g. NVMe media.
1
0
0

@jarkko I think that's why there's 'high bandwidth flash' being talked about - the same stacked die arrangement they currently do for HBM ram.

1
0
1
@penguin42 yeah also flash prices are going high up to the skies :-)
1
0
1
@penguin42 It's funny. 128 GiB of RAM in a PC used to be a lot but not unheard. E.g., this 10 year old laptop that I'm typing right now has 64 GiB of RAM. Now 128 GiB of RAM is like something that only people with white coats and ties can access: -) Soon we probably have to stop selling cars thanks to big build.
2
0
0

@jarkko Yeh it's crazy; I'd heard there were getting that way on some types of PCB and even passive components.

0
0
1

@jarkko @penguin42 Sir, I present to you Asus PX13 - 128GB of ram in a very handy 13" laptop

1
0
1
@zygoon @penguin42@mastodon.org all my external displays are ProArt :-)
0
0
0