#Ollama
The press says Muse Glimmer 30B requires a 5090. r/LocalLLaMA is posting screenshots of it running on a used 3090 from 2020. Both are right, and the explanation is in the VRAM budget: 17 GB of weights, 1.7 GB of KV cache, and an attention architecture designed to fit. Here's the math line by line, the tokens-per-second estimate on a 3090 with the work shown, and the verdict on when 24 GB is enough and when it isn't.
A step-by-step guide to running an LLM locally with Ollama: install it, run qwen3:8b in 2 commands, and plug it into your code through the OpenAI-compatible endpoint. It runs on 8 GB of RAM with no GPU required. As a bonus, the VRAM math by model size and when local beats the API.