The post has been translated automatically. Original language: Russian
SLA (Service Level Agreement) is often read diagonally, paying attention only to the availability figure — 99.9% or 99.99%. But between these figures, the difference in downtime per year is a multiple, and this is far from the only thing worth checking.
What is important to look at in the contract: — how is it considered simple: does it include scheduled maintenance work, which is warned in advance; — compensation in case of violation of the SLA — a specific formula, rather than a vague formulation "by agreement"; — reaction time and time to eliminate the incident separately, rather than one general "response time"; — escalation procedure — who to contact if the first-line engineer does not solve the problem; — conditions of access to the site for own equipment, if it is a colocation; — the procedure for notification of planned works and their maximum frequency.
A good SLA is one where you can calculate exactly how much theoretical downtime a provider allows and what happens if it exceeds these limits. If the wording is vague, this is a reason for additional questions before signing, and not after the first incident.
SLA (Service Level Agreement) часто читают по диагонали, обращая внимание только на цифру доступности — 99,9% или 99,99%. Но между этими цифрами разница в часах простоя в год кратная, и это далеко не единственное, что стоит проверять.
Что важно смотреть в договоре: — как считается простой: включаются ли туда плановые технические работы, о которых предупреждают заранее; — компенсация при нарушении SLA — конкретная формула, а не расплывчатая формулировка «по согласованию»; — время реакции и время устранения инцидента отдельно, а не одно общее «время реагирования»; — порядок эскалации — к кому обращаться, если инженер первой линии не решает проблему; — условия доступа к площадке для собственного оборудования, если это colocation; — порядок уведомления о плановых работах и их максимальная частота.
Хороший SLA — тот, где можно точно посчитать, сколько теоретического простоя допускает провайдер и что произойдёт, если он превысит эти рамки. Если формулировки размытые — это повод для дополнительных вопросов до подписания, а не после первого инцидента.