Resolved
We've been below capacity for A100 hardware for the past 30 minutes and all queues have been drained. Thanks for your patience
Monitoring
We are seeing pods start up again now and they are working through the backlog.
Investigating
We're seeing most A100 hardware automatically cordoned due to possible hardware failure.