Azure (Microsoft) logo

Skill

aks-cluster-setup

design and configure Azure Kubernetes Service clusters

Covers Azure Architecture Deployment Kubernetes

Description

Make the AKS-specific design decisions for a new production Azure Kubernetes Service (AKS) cluster — SKU (Automatic vs Standard), pod IP model (Azure CNI Overlay vs kubenet), API-server access, egress, identity, upgrades, node pools, and reliability — then delegate the actual provisioning to the Azure Skills deployment engine. WHEN: create AKS cluster, provision AKS environment, design AKS networking, choose AKS SKU, Day-0 AKS checklist, plan a production AKS cluster. DO NOT USE FOR: debugging a running cluster (use aks-troubleshooting); assessing an existing cluster for AKS Automatic (use aks-automatic-readiness); GPU / model-serving setup (see azure-skills airunway-aks-setup); generic non-AKS Azure resource deployment (use azure-skills azure-deploy directly).

SKILL.md

AKS Cluster Setup

This skill owns the AKS-specific design decisions for standing up a production cluster and then hands the provisioning execution to the Azure Skills deployment engine. It distinguishes Day-0 decisions (networking, API-server access — hard to change later) from Day-1 features (can be enabled post-creation). See references/cli-reference.md for the commands.

How this works with Azure Skills

This is a thin facade over provisioning — it owns the AKS design and delegates the generic provisioning execution to Azure Skills; it does not run azd or az deployment itself. It enriches the request with AKS-specific design decisions, then delegates the generic provisioning workflow to Azure Skills.

User: "Create a production AKS cluster"
  │
  ▼
aks-cluster-setup (this skill)
  │  Decides the AKS-specific design:
  │   - AKS Automatic vs Standard
  │   - Pod IP model (Azure CNI Overlay vs VNet-routable vs kubenet)
  │   - API-server access (public / private / authorized ranges)
  │   - Egress model, identity, node pools, upgrade strategy, zones
  │
  ▼
Hands off to Azure Skills (install alongside):
  ├─ azure-prepare   → analyzes requirements, generates the infra plan
  ├─ azure-validate  → validates the plan, provisions a preview
  └─ azure-deploy    → runs azd up / az deployment to create the cluster

Install both: aks-cluster-setup for the AKS design brain, Azure Skills as the deployment engine underneath. If Azure Skills is not present, this skill still produces the full design and the exact az aks create command in references/cli-reference.md so the user can provision manually. This skill never provisions non-AKS resources — for those, use Azure Skills directly.

Rules

  1. Start from the user's requirements for compute, networking, security, and scale.
  2. Inspect the host's available tools for an Azure MCP capability that advertises AKS operations. Use the matching capability under whatever name the host assigned, inspect its advertised schema or discovery surface, and choose the smallest operation that fits. Do not gate on or construct a mapping for a literal tool name. Fall back to az aks only when the host exposes no matching capability or its AKS surface lacks the needed operation.
  3. Default to AKS Automatic unless the user needs control not supported by Node Auto Provisioning. Standard is for full configurability at higher operational overhead.
  4. Record the rationale for every Day-0 decision (networking, API-server access) — these are expensive or impossible to change after creation.

Required inputs (ask only what is needed)

If the user is unsure, use the safe defaults below.

  • Environment type: dev/test or production
  • Region(s), availability zones, preferred node VM sizes
  • Expected scale (node/cluster count, workload size)
  • Networking requirements (API-server access, pod IP model, ingress/egress control)
  • Security and identity requirements, including image registry
  • Upgrade and observability preferences
  • Cost constraints

Design workflow

1. Cluster type

  • AKS Automatic (default): a curated experience with pre-configured best practices for security, reliability, and performance. Use unless you need custom networking, autoscaling, or node-pool configuration not supported by Node Auto Provisioning (NAP).
  • AKS Standard: full control over configuration, at additional setup and management overhead.

2. Networking (pod IP, egress, ingress, dataplane)

Pod IP model (key Day-0 decision):

  • Azure CNI Overlay (recommended): pod IPs from a private overlay range, not VNet-routable; scales to large environments and suits most workloads.
  • Azure CNI (VNet-routable): pod IPs directly from the VNet (pod subnet or node subnet); use when pods must be addressable from the VNet or on-premises. Docs: https://learn.microsoft.com/azure/aks/azure-cni-overlay

Dataplane & network policy:

  • Azure CNI powered by Cilium (recommended): eBPF-based, for high-performance packet processing, network policy, and observability.

Egress:

  • Static Egress Gateway for stable, predictable outbound IPs.
  • For restricted egress: UDR + Azure Firewall or an NVA.

Ingress:

  • App Routing add-on with Gateway API — recommended default for HTTP/HTTPS workloads.
  • Istio service mesh with Gateway API — advanced traffic management, mTLS, canary releases.
  • Application Gateway for Containers — L7 load balancing with WAF integration.

DNS:

  • Enable LocalDNS on all node pools for reliable, performant resolution.

3. Security

  • Use Microsoft Entra ID everywhere (control plane, Workload Identity for pods, node access). Avoid static credentials.
  • Azure Key Vault via the Secrets Store CSI Driver for secrets.
  • Enable Azure Policy + Deployment Safeguards.
  • Enable encryption at rest for etcd/API server and in-transit for node-to-node.
  • Allow only signed, policy-approved images (Azure Policy + Ratify); prefer Azure Container Registry.
  • Isolation: namespaces, network policies, scoped logging.

4. Observability

  • Use Managed Prometheus and Container Insights with Grafana for logs and metrics.
  • Enable Diagnostic Settings to collect control-plane and audit logs into a Log Analytics workspace.
  • Complement with Application Insights, Resource Health, and AppLens detectors for troubleshooting.

5. Upgrades & patching

  • Configure Maintenance Windows for controlled upgrade timing.
  • Enable auto-upgrade for the control plane and node OS to stay current on security patches and Kubernetes versions.
  • Consider LTS versions (2-year support, Premium tier) for enterprise stability.
  • Fleet upgrades: use AKS Fleet Manager for staged rollout from test to production.

6. Performance

  • Use Ephemeral OS disks (--node-osdisk-type Ephemeral) for faster node startup.
  • Select Azure Linux as the node OS (smaller footprint, faster boot).
  • Enable KEDA for event-driven autoscaling beyond HPA.

7. Node pools & compute

  • Dedicated system node pool: at least 2 nodes, tainted CriticalAddonsOnly for system workloads.
  • Enable Node Auto Provisioning (NAP) for cost savings and responsive scaling.
  • Use latest-generation SKUs (v5/v6) for host-level optimizations.
  • Avoid B-series (burstable) VMs — they cause performance and reliability issues.
  • Use SKUs with at least 4 vCPUs for production.
  • Set topology spread constraints to distribute pods across hosts/zones per SLO.

8. Reliability

  • Deploy across 3 availability zones (--zones 1 2 3).
  • Use the Standard tier for a zone-redundant control plane and the 99.95% API-server SLA.
  • Enable Microsoft Defender for Containers for runtime protection.
  • Configure PodDisruptionBudgets for all production workloads.

9. Cost controls

  • Use Spot node pools for batch/interruptible workloads (up to 90% savings). For ongoing spend reduction on a running cluster, use aks-cost-optimization.
  • Stop/Start dev/test clusters: az aks stop / az aks start.
  • Consider Reserved Instances or Savings Plans for steady-state workloads.

Guardrails

  • Do not request or output secrets (tokens, keys).
  • Do not ask the user to paste subscription IDs. Resolve subscription and resource scope via MCP tools or az account show / az account list.
  • For ambiguous Day-0 decisions, ask clarifying questions. For Day-1 features, propose 2–3 safe options with trade-offs and choose a conservative default.
  • Do not promise zero downtime; advise workload safeguards (PDBs, probes, replicas) and staged upgrades.

Error handling

Error / symptomLikely causeRemediation
MCP tool call fails or times outInvalid credentials, subscription, or AKS contextVerify az login, confirm the active subscription with az account show, check the target resource group without echoing identifiers back
Quota exceededRegional vCPU or resource limitsRequest a quota increase or select a different region / VM SKU
Networking conflict (IP exhaustion)Pod subnet too small for the CNI/overlay modelRe-plan IP ranges; may require cluster recreation (Day-0)
Workload Identity not workingMissing OIDC issuer or federated credentialEnable --enable-oidc-issuer --enable-workload-identity, configure the federated identity

More from Azure (Microsoft)

View publisher

© 2026 YourAI.tools. Every skill from an identity-verified publisher.

Independent catalog. Not affiliated with, endorsed by, or sponsored by Anthropic or any listed publisher. All trademarks belong to their respective owners.