Popular repositories Loading
-
-
llm-regression-gate
llm-regression-gate PublicA CI/CD-style quality gate for LLM features: runs a prompt+model over a golden dataset, catches quality regressions with paired stats (McNemar + noise band) and critical-slice hard-blocks, and fail…
Python
-
llm-cost-autopilot
llm-cost-autopilot PublicOpenAI-compatible gateway that routes each request to the cheapest capable model, serves a scope-gated semantic cache, and validates its own routing (labels offline, LLM-as-judge online). ~40% save…
Python
-
ai-pipeline-forensics
ai-pipeline-forensics PublicObservability for multi-step AI pipelines: traces each step as OpenTelemetry-shaped spans, pinpoints the first failing step (root cause, propagation-aware), and grows an eval dataset from flagged f…
Python
-
self-healing-docs
self-healing-docs PublicGitHub Action that detects docs a PR's code change made stale and flags or auto-fixes them — embeddings + retrieval + LLM-as-judge, in CI.
Python
-
llm-output-arbitration
llm-output-arbitration PublicMost systems generate answers — this one judges them. A jury of competing Claude critics scores an LLM output, then deterministic synthesis returns one confidence-scored verdict.
Python
If the problem persists, check the GitHub status page or contact support.