llama-swap alternative in Rust — OpenAI-compatible GGUF runtime with VRAM-pressure eviction, context auto-fallback, and usage tracking.
-
Updated
Jul 22, 2026 - Rust
llama-swap alternative in Rust — OpenAI-compatible GGUF runtime with VRAM-pressure eviction, context auto-fallback, and usage tracking.
Heuristic signals on whether an OpenAI-/Anthropic-compatible endpoint serves the model it claims — catch model-swapping, quantization & silent context truncation. Zero-dependency Python CLI. Signals, not proof.
Add a description, image, and links to the model-swapping topic page so that developers can more easily learn about it.
To associate your repository with the model-swapping topic, visit your repo's landing page and select "manage topics."