{"by":"carloslfu","descendants":115,"id":49524447,"kids":[49525120,49527160,49525218,49525192,49530561,49525114,49527422,49525564,49545169,49545904,49533837,49525005,49536475,49548870,49525061,49529004,49527207,49528769,49530355,49536413,49527514,49525577,49524757,49533380,49541096,49545519,49533469,49527204,49525240,49525174,49525191],"score":234,"text":"I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory&#x2F;RAM, thanks to expert-offloading&#x2F;ssd-streaming. Easy to install&#x2F;update, and mac-native using MLX and Swift.<p>It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I&#x27;ll be implementing and porting the MTP module for speculative decoding next","time":1788280966,"title":"Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s","type":"story","url":"https://github.com/carloslfu/slotstream"}