On September 3rd, 2026, at 11:52 AM EDT, Cloudli Communications identified an issue with the WorkerNode Auto-scaler incorrectly shutting down multiple instances hosting micro-services, including those supporting the VS API and Cloudli Quoting Tool, resulting in temporary service disruption. The affected services were restored at 12:11 PM EDT.
The WorkerNode auto-scaler determines the compute resources required to run micro-services in our hosting facility and scales instances accordingly. The process relies on resource-usage metrics. During this incident, the WorkerNode auto-scaler did not receive all required metric information and scaled the environment down, draining some micro-services nodes. Because the configured minimum node count was set too low, the environment was allowed to scale below the capacity required to re-allocate all affected micro-services to alternate instances.
Post-incident analysis identified incomplete metrics as the condition that caused the WorkerNode auto-scaler to reduce capacity beyond a safe operating threshold. The minimum auto-scaler node setting has been increased to a safe level to prevent the environment from scaling below the capacity required to maintain service availability.
Engineering has also identified the metrics-reading issue and is working on permanent correction. These safeguards are intended to prevent a similar metrics issue from causing an unsafe reduction in WorkerNode micro-services capacity in the future.
While the incident was short in duration, we take any interruption of service very seriously and are continuously evaluating new processes and mitigation measures that can be proactively implemented to ensure service continuity.
We thank you for your continued support. Please feel free to reach out if you would like to discuss the particulars of this incident report further.