Skip to content

Pull requests: JustVugg/colibri

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Vulkan MoE GEMV backend for integrated/AMD GPUs (draft, complements #418) enhancement New feature or request vulkan Backend Vulkan/AMD
#729 opened Jul 31, 2026 by minne100 Draft
feat(win): fix silent CPU fallback, launcher suite, DirectStorage expert loads enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore
#670 opened Jul 28, 2026 by khalilswdp Contributor Loading…
8 tasks done
feat(core): model-architecture seam, chat templates, text antiprompt enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore
#667 opened Jul 28, 2026 by khalilswdp Contributor Loading…
5 tasks done
QLoRA training path: fine-tune GLM-5.2 (744B) in 64 GB RAM enhancement New feature or request feature Nuova funzionalità
#626 opened Jul 26, 2026 by pavolbauer Loading…
5 tasks done
Metal fmt=4 grouped-int4 decode: attention + routed experts (#585) metal Backend Metal/Apple needs-rebase Confligge, serve rebase dell'autore
#587 opened Jul 24, 2026 by RDouglasSharp Contributor Loading…
Preserve model-declared EOS tokens in serve mode enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore
#584 opened Jul 24, 2026 by saskw2010 Draft
Fix stateful KV tail at the NGEN limit bug Difetto verificato nel codice
#567 opened Jul 23, 2026 by winklemad Contributor Loading…
3 of 5 tasks
add persistence controls for lower SSD writes enhancement New feature or request
#555 opened Jul 23, 2026 by Skater1808 Loading…
5 tasks
CPU: KV cache quantization — KV8 (fp8 e4m3) + KV_TQ (rotated-int4 / PolarQuant) enhancement New feature or request performance Velocità / tok-s / ottimizzazioni
#553 opened Jul 23, 2026 by NeuralNotwerk Contributor Loading…
feat: add WebGPU expert workers enhancement New feature or request
#552 opened Jul 23, 2026 by gauravsaini Loading…
feat: add distributed expert workers enhancement New feature or request
#551 opened Jul 23, 2026 by gauravsaini Loading…
feat: add dense MLP activation sharding enhancement New feature or request performance Velocità / tok-s / ottimizzazioni
#550 opened Jul 23, 2026 by gauravsaini Loading…
feat: add Qwen3-30B-A3B engine model-support Supporto a nuovi modelli needs-rebase Confligge, serve rebase dell'autore
#544 opened Jul 23, 2026 by opxyc Contributor Loading…
4 of 5 tasks
pilot: multi-worker PILOT_REAL prefetch (PILOT_WORKERS) — byte-identical base for the #441 hardware A/B performance Velocità / tok-s / ottimizzazioni
#480 opened Jul 21, 2026 by cdhdt Contributor Draft
4 of 5 tasks
nix: CUDA support enhancement New feature or request
#416 opened Jul 19, 2026 by attilaolah Contributor Draft
5 tasks
KV cache quantization: fp8 (KV8) + 4-bit TurboQuant (KV_TQ) on CPU, CUDA, and Metal enhancement New feature or request performance Velocità / tok-s / ottimizzazioni
#399 opened Jul 18, 2026 by NeuralNotwerk Contributor Loading…
feat: Add NUMA-aware RAM-disk streaming enhancement New feature or request needs-rebase Confligge, serve rebase dell'autore performance Velocità / tok-s / ottimizzazioni
#377 opened Jul 17, 2026 by BColsey Loading…
Nearly double the hit-rate needs-rebase Confligge, serve rebase dell'autore performance Velocità / tok-s / ottimizzazioni
#223 opened Jul 14, 2026 by withinboredom Contributor Loading…
4 of 5 tasks
Add DeepSeek V4 Flash CPU inference with NVMe expert streaming model-support Supporto a nuovi modelli
#165 opened Jul 14, 2026 by DrewZt Loading…
9 tasks done
ProTip! Find all pull requests that aren't related to any open issues with -linked:issue.