Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm guessing your Chaos Gorilla helped to harden your architecture against this threat.

Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?



EDIT: I WAS WRONG. Chaos Monkey and Chaos Gorilla both exist and simulate different forms of chaos.


"Create More Failures

Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run this service because we want engineering teams to be used to a constant level of failure in the cloud. Services should automatically recover without any manual intervention. We don't however, simulate what happens when an entire AZ goes down and therefore we haven't engineered our systems to automatically deal with those sorts of failures. Internally we are having discussions about doing that and people are already starting to call this service "Chaos Gorilla"."

http://techblog.netflix.com/2011_04_01_archive.html


My apologies! I was wrong.


Chaos Monkey takes down instances and such. Chaos Gorilla takes down entire AZs. :)


Monkey takes down single hosts; Gorilla simulates the loss of an entire AZ.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: