Check the status page
As I've covered somewhat extensively in prior writing (see A month of OpEx quick wins and Self-hosting our GitHub Action runners), this year has been marked by a careful and deliberate shrinking of our stack, button down, with the goals of, in order, one, reduce long-term maintenance burden. Two, reduce operational expenses. Almost without us really meaning to, this has materialized indirectly as adopting a lot of Cloudflare stuff, which makes sense given that Cloudflare's primary positioning for almost all of their non-DNS infrastructure is a more cost-effective alternative to entrenched legacy alternatives like AWS. These savings have been real, and minus some hiccups here or there, the migrations have been extremely painless. However, we do find ourselves with the new problem of having accidentally introduced what at times feels like an SPOF. When Cloudflare's DNS infrastructure goes down, it impacts basically the entire world, and therefore I don't mind that much that it impacts us. But now that we're using their more esoteric offerings, we are no longer shielded by the aegis of universality. All of this is prelude to the fact that we're discovering Cloudflare has more operational issues than I think their reputation suggests, or at least our exposure to those issues feels precarious: last week (recency bias, but also fresh wounds) hit us thrice: once because of general dashboard issues, once because of R2 (their S3 competitor), and once because all POST requests were silently/myseriously failing.
Complaints aside, even knowing what I now know, I think all of our decisions were the right one. Part of this is just getting used to a slightly different status quo. We were able to remediate the two customer-facing incidents of these three very well (thanks entirely to Matias and Steph!), and in doing so, it revealed that we're just missing a lot of obvious kill switches and documentation for failure modes that aren't the ones we've already been scarred by. So a lot of August will be enumerating and addressing those.
But as obvious as it sounds in retrospect, next time we evaluate a new vendor or a new SKU on an existing vendor, the first item on my list is going to be running a script on their status page to see what the actual stability of that SKU is and not just what the reputation suggests.