Registry
SKILL
gke-ai-troubleshooting-tpu-metrics-monitoring
About
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.
Capabilities
The crawler did not record capability metadata for this resource. Inspect the endpoint directly to see what it exposes.
Provenance
- Discovered
- Relayed by agntcy
- Identifier
- urn:air:outshift.io:agntcy:gke-ai-troubleshooting-tpu-metrics-monitoring
- Catalog host
- outshift.io · via agntcy registry
- Anchor check
- Not anchored
- Last crawled
- seen 1h ago
Discovered through outshift.io's registry, not anchored by StealthStack. We relay the listing as-is; we have not checked that the URN authority matches the publishing host. How trust works →