h200
Here are 11 public repositories matching this topic...
Thermal-aware batch controller for vLLM/TensorRT-LLM. Prevents HBM thermal throttling from killing p99 latency on H100/H200. Monitors nvidia-smi, auto-cuts batch size at 85°C, migrates cold KV to DRAM. Prometheus + Grafana included. 4.2s -> 2.1s p99 at 128K context.
-
Updated
Apr 13, 2026 - Python
Monitor low-utilization time, idle-state episodes, and workload starvation signals on NVIDIA datacenter GPUs.
-
Updated
Apr 3, 2026 - Python
Daily-verified cloud GPU rental prices across 18+ providers (H100, H200, B200, A100, RTX 4090...). Open dataset, updated daily, CC BY 4.0. Live at gpurentalprices.com
-
Updated
Jul 30, 2026
GLM-5.2 744B at 4-bit on Modal 4x H200 via vLLM, plus a static streaming chat UI.
-
Updated
Jul 25, 2026 - HTML
dd-ready fully packed FreeDOS disk image for crossfalshing SAS2008 like Dell H200 to HBA
-
Updated
May 9, 2023
AMD MI300X vs NVIDIA H200: a controlled single-GPU benchmark of inference and training. AMD's 1.32x paper FLOPs advantage inverts on achieved efficiency, but its 1.36x memory becomes a 1.84x KV-cache advantage that wins large-model serving outright.
-
Updated
Jul 28, 2026 - Python
⚡ TIMTEH Model Forge — Uncensored, abliterated & reasoning-distilled GGUFs. Forged on 8×H200 SXM5 | 1.1TB VRAM
-
Updated
Mar 27, 2026 - Shell
SGPU — Simple GPU monitor for the SGVR H200 lab (MLXP / NAVER Cloud NKS). Zero-install kubectl dashboards, per-pod GPU process attribution, in-pod TUI, and per-user usage accounting.
-
Updated
Jul 14, 2026 - Python
Improve this page
Add a description, image, and links to the h200 topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the h200 topic, visit your repo's landing page and select "manage topics."