PeeOnYou [he/him]

  • 0 posts
  • 5 comments
Joined 4 years ago
Cake day: April 16th, 2022
  • running a model locally on a macbook is going to be somewhat of a poor experience imo… the bandwidth for that memory is quite slow compared to a dedicated GPU, but it’s not unusable. GPUs are required if you want any sort of decent speed. I have a 5090 which is 1 card under the best you can get and it still struggles a bit with larger quants. RTX Pro 6000 is the king card and it will let you run some much larger quants at usable speeds but it also costs about $16k now.

    People who are REALLY serious about running local tend to have multiple graphics cards connected to a single machine. 3090s are popular because they have decent bandwidth and 24gb of vram each. But even those are selling for $2k+ used.

    Qwen3.8 27b or Qwen 35B A3B quantized are pretty good for local. With 48GB you could run the larger UD-Q6_K versions but if you go with a smaller one like Q5_K_M you’d have more room for a larger context window and it might run a little faster. https://huggingface.co/unsloth/Qwen3.8-27B-GGUF