Will this LLM fit your GPU or Mac? npx fitllm — accurate memory math for MLA/sliding-window/hybrid/MoE architectures that naive VRAM calculators get wrong by up to 18x. Single file, zero deps, conformance-vector tested. MIT.
cli amd inference nvidia moe quantization mlx mla vram kv-cache apple-silicon llm llama-cpp local-llm ollama gguf localllama memory-calculator vram-calculator will-it-run
-
Updated
Jul 16, 2026 - JavaScript