An Update on Aeries Service Reliability

and What We’re Doing About It

To the Aeries Community,

 

Over the past several weeks, you have experienced service issues that fell short of the standard you should expect from us. In some cases, we found out from you before we found out ourselves. That’s not acceptable, and I’m not going to dress it up as anything other than what it is: we fell short. We let you down, and we know how much disruption that causes when it happens at a critical time during the school year.

 

This letter is an acknowledgment of what went wrong, how we’re fixing it, and our commitment to what you can expect moving forward. We’re launching a program, to be completed by the end of the year, with five concrete commitments. I’d rather tell you exactly what we’re doing than offer vague reassurance.

 

Catching problems before you do. Too often lately, our alerts have gone off after you already noticed something was wrong. We’re starting focused monitoring on key workflows that matter to you. This proactivity means we can catch the next issue before it reaches your district. I can’t promise perfect visibility overnight, but improvements are already in progress.

 

Communicating better when something does go wrong. This is the one you’ve told us about most directly. We’re setting clear rules for when an issue needs a public status update versus a message to only those affected, committing to regular updates while an incident is active, and double-checking an issue is resolved before we say so.

 

Building in more room to grow. Some of our recent issues came from running out of headroom at peak demand, especially at school year peaks. Our infrastructure scales automatically with demand, but our scaling policies were tuned too aggressively for peak conditions, pulling capacity back faster than we should have. We’re adjusting those policies to hold more headroom in reserve for high-demand periods, and load-testing against realistic peak conditions, not just average-day traffic.

 

Making changes safer. We’re building the ability to roll out updates to a new set of servers and switch traffic over, rather than updating everything in place. Once that’s in place, if a change ever causes a problem, we’ll be able to roll it back quickly instead of slow roll backs.

 

Strengthening how we test. One recent issue got through because a specific scenario wasn’t covered during testing. We’ve increased our automated regression coverage and we’re doubling down specifically on edge-case and negative-path scenarios, not just the paths we expect to work, so gaps like that one are caught before they reach you.

 

Almost certainly an issue will, in some form, reach you. We run complex software at a scale that touches millions of students, and we won’t ever hit perfect uptime. What you should expect from us is that we catch it faster, fix it faster, and are transparent while we’re doing it.

 

We know trust gets rebuilt by what we do, not what we write in a letter. Let us know when we fall short. We’re committed to improving, and we mean it.

 

Steve Evans
Chief Product & Technology Officer