Back to feed
Dev.to
Dev.to
7/12/2026
The original title is: "AWS puts gray zone failures into the EKS control loop"

The original title is: "AWS puts gray zone failures into the EKS control loop"

Original: AWS puts gray zone failures into the EKS control loop

Short summary

AWS now automates EKS zonal shift to redirect traffic away from impaired availability zones on suspicion rather than operator confirmation. Gray failures — zones that are slow but not down — cause the longest partial outages because health checks keep passing. The article argues CI/CD teams must treat slow-but-not-down as its own chaos drill scenario and ensure deploy controllers share the same traffic view as AWS's automatic shifts.

  • AWS automates EKS zonal shift on suspicion, removing one step of operator judgement
  • Gray failures (slow but not down) cause longer outages than hard failures
  • CI/CD pipelines need to account for mid-rollout zonal shifts when evaluating canary metrics

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more