ai-mlops
Installation
SKILL.md
MLOps & LLMOps - Production Operations Hub
Operating posture: version every changeable artifact, gate every release with a regression-eval suite in CI, instrument the whole path with OpenTelemetry (pin the GenAI semantic-convention schema version you emit), treat tool/RAG context as untrusted input, and ship rollback plus incident playbooks before launch.
When To Use This Skill
Activate this skill when the user asks for:
- Deploying an ML, LLM, RAG, or agent-backed system to production
- Designing serving, batch, hybrid, or multi-region runtime architecture
- Adding observability, drift detection, alerting, retraining, or release gates
- Writing incident runbooks, rollback plans, or go/no-go checklists
- Migrating an LLM system to a new model, model version, provider, or API surface
- Hardening an AI system against prompt injection, RAG poisoning, tool abuse, or data leakage
- Building governance artifacts for privacy, auditability, or regulated rollout
- Choosing how to operate prompts, model artifacts, feature definitions, or agent graphs safely
- Diagnosing why changes to an ML system keep rippling: entanglement, correction cascades, undeclared consumers, pipeline jungles, or config sprawl
- Operating fairness, privacy-budget, human-oversight, appeal, watermark/provenance, copyright/memorization, or environmental controls
- Deploying multimodal image, document, audio, video, vision-language, or diffusion systems with bounded media ingestion, safety, latency, and cost