Hook

Their other posts in the index, biggest breakout first.
So let's talk about some of the failure modes that you are going to run into as you scale your app from 500 users to 50,000 users to 500,000 users and beyond. I'm a software engineer, I work at a MAG7 company and these scaling challenges are things that all of the big tech companies are thinking about all the time and they are problems that you are absolutely going to run into as you grow and scale your app. It's a good problem to have because it means that what you've built is resonating with people and people want to access it, but a part of retaining users and sustaining your product growth is making sure that it's a high quality and reliable experience for the people that are coming to your service. So to that end, we're going to talk through failure modes and the common things that you might run into. So one of the most important concepts that you need to keep in mind is that when you are hosting your service, and by service I mean your web app that will be stored on a machine, I have another video where I talk about the concept of how all web apps at the end of the day are stored on a machine somewhere. This is what we refer to as the cloud or a cloud hosted service. Um, I'm happy to do a longer deep dive on that, but assuming that you have that foundational understanding, your web app is here and you are always going to have requests flying over the network to hit your service. When I'm saying request, it's users A, B, C, D and then all the way down to user 500,000 and they're all accessing your app at the same time. So in addition to users coming across the network and hitting your service, your service is also likely hitting some APIs on the back end here. So let's say you are building a music app and you're doing something using the Spotify API. You would have users come hit your service and then your service then goes and hits the Spotify API. Let's say you're building a Shopify app, user hits your service, you are then querying the Shopify API. Um, so on and so forth. But the entire point is that your ability to deliver value to these users is predicated on your ability to then hit these services reliably. And because all of these requests are traveling over the network, there are a number of problems that can happen in the wild because networks are not reliable 100% of the time. So because networks are not reliable 100% of the time, there's a concept in software engineering and software design called fallback mechanisms. And the entire concept of fallback mechanisms is what do you do when you are making a request across network to a downstream service and that network request drops and it fails. How do you ensure that you can still deliver value to your users over here? So the simplest technique that you can implement here is something called retry logic. And what retry logic means is that instead of just making a single request and then failing and then arrowing out to the user if this network request fails, you retry maybe one to three times. So let's say user A comes here, they say, hey, service, I want you to do your thing. You query Spotify API. For whatever reason, Spotify is having problems that day, maybe they're having an outage of their own, it goes down, your your service returns with a 500. Instead of just erroring and returning a failure state to the user, what you're going to do is say, okay, I'm going to wait 30 seconds and see if the Spotify servers have restored themselves. They are running on cloud services themselves, so they have a fleet of VMs that are handling these failure cases and restarting their machines in case of failure, auto scaling, cloud compute, it's a really interesting topic and I'll probably make some other videos about that as well. But the entire point is that, okay, let's wait 30 seconds and we'll try again before we deliver a failure. So you wait your 30 seconds, you try again. Okay, let's say in this case the Spotify servers are still starting back up, they're not 100% ready to deliver their content yet, you come back. Instead of failing this time, you say, okay, I'm going to wait now for a minute instead of 30 seconds and I'm going to try again. And then finally at this time point in time, you've waited a total of 90 seconds, the servers have have stood back up and they're ready to go, you get your content back and you can send it to the user. So typically, um, in software engineering, you'll retry three times is usually the typical benchmark. You can go up to five times. Um, beyond five times, it's not necessarily the best option because it's more likely that something is just corrupted with the the resource that you're trying to to hit. Um, but yeah, basically you want to do two concepts, trying more than one time, usually three to five times and then also having some sort of a delay in between your retry. So first time 30 seconds, another time 60 seconds, another time maybe 90 seconds. And the like amount of time that you wait before you try again is very variable on the API that you're trying to hit. So in some cases you might do a retry at, I don't know, five minutes instead of at 90 seconds. And so with all these concepts in software engineering, software design, system architecture, you see, understand the fundamental principle and then you apply it in a specific way to the use case that makes more sense to your web app. So I hope this helps. Um, there is a large number of other failure modes that can happen because the internet and networks, they are very dynamic systems. Um, I will continue to share more about this, but I hope this helps. And good luck, we believe in you, we want to see what you're building and you just need to scale it so that we can see it.