ai-mlops

Installation
SKILL.md

MLOps & LLMOps - Production Operations Hub

Operating posture: version every changeable artifact, gate every release with a regression-eval suite in CI, instrument the whole path with OpenTelemetry (pin the GenAI semantic-convention schema version you emit), treat tool/RAG context as untrusted input, and ship rollback plus incident playbooks before launch.

When To Use This Skill

Activate this skill when the user asks for:

  • Deploying an ML, LLM, RAG, or agent-backed system to production
  • Designing serving, batch, hybrid, or multi-region runtime architecture
  • Adding observability, drift detection, alerting, retraining, or release gates
  • Writing incident runbooks, rollback plans, or go/no-go checklists
  • Migrating an LLM system to a new model, model version, provider, or API surface
  • Hardening an AI system against prompt injection, RAG poisoning, tool abuse, or data leakage
  • Building governance artifacts for privacy, auditability, or regulated rollout
  • Choosing how to operate prompts, model artifacts, feature definitions, or agent graphs safely
  • Diagnosing why changes to an ML system keep rippling: entanglement, correction cascades, undeclared consumers, pipeline jungles, or config sprawl
  • Operating fairness, privacy-budget, human-oversight, appeal, watermark/provenance, copyright/memorization, or environmental controls
  • Deploying multimodal image, document, audio, video, vision-language, or diffusion systems with bounded media ingestion, safety, latency, and cost
Installs
199
GitHub Stars
90
First Seen
Jan 22, 2026