Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I am running it on a single RTX 3090 (24GB VRAM).

Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...

It uses an order of magnitude less VRAM at longer contexts which is a huge advantage over Qwen 3.6 27B



Seems like that's the tradeoff with this model. Close to 27b intelligence while using less vram.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: