gke-ai-troubleshooting-handle-disruption-gpu-tpu

Installation
SKILL.md

Handle Disruption on GPUs and TPUs Troubleshooting

🔍 Diagnostic Workflow

Step 0: Context Acquisition

  • Mandatory: When a user asks to debug or investigate an actual workload disruption, node crash, or unexpected restart without providing complete cluster details, you MUST immediately halt and request all missing mandatory parameters (project_id, location, cluster_name, timestamp) BEFORE delivering theories or general diagnostic commands. Only skip context acquisition if the user explicitly requests a generic reusable runbook or provides a complete static telemetry/log dump for offline analysis.
  • Optional: node_name, workload_name, workload_namespace, nodepool_name.

Step 1: [Low Risk] Check for Upcoming Scheduled Maintenance

Installs
1.6K
Repository
google/skills
GitHub Stars
19.1K
First Seen
Jul 24, 2026
gke-ai-troubleshooting-handle-disruption-gpu-tpu — google/skills