A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling
-
Updated
Aug 3, 2026 - Python
A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling
AI-powered analytics platform built with a futuristic Solaris Night Glassmorphism UI. Upload CSV datasets, run real-time RAG-based analysis, execute Pandas-powered AI queries, and visualize live LLM inference streams with physics-driven Framer Motion animations.
Add a description, image, and links to the streaming-inference-engine topic page so that developers can more easily learn about it.
To associate your repository with the streaming-inference-engine topic, visit your repo's landing page and select "manage topics."