Senior DevOps Engineer
Main requirements
• DevOps/SRE experience of at least 4 years and confident practical work with Linux, Docker and Kubernetes.
• Practical experience in Kubernetes cluster administration and application delivery via Helm or Kustomize.
• Experience with Yandex Cloud or other public cloud: Kubernetes, container registry, networks, IAM, object storage, monitoring and access management.
• Confident knowledge of CI/CD tools: GitLab CI, Jenkins, Argo CD or analogues; understanding the GitOps approach.
• Experience Terraform, Ansible, or other IaC tools.
• Understanding networks, TLS, reverse proxy, ingress, DNS, load balancing, and zero trust principles.
• Hands-on experience with Prometheus, Grafana, Loki/ELK, OpenTelemetry or similar observability tools.
• Understanding the principles of safe operation: RBAC, secrets management, image scanning, vulnerability management, backup and disaster recovery.
It will be an advantage:
• Experience in operating GPU loads, model serving, vLLM/TGI/Triton or similar inference services.
• Experience with Kafka/RabbitMQ, PostgreSQL, Redis, S3-compatible storage and service mesh.
• Knowledge of OPA, Vault, Keycloak, Istio or similar policy enforcement and access control tools.
• Work experience in an industrial company, in isolated segments and environments with high security requirements.
What you will do
• Deploy and maintain GenAI platform services in Kubernetes on-prem and in Yandex Cloud: API services, Agent Runtime, RAG, queues, databases, and observability components.
• Develop infrastructure as code: Helm, Kustomize, Terraform or similar tools; ensure reproducibility of dev, test, stage and production environments.
• Configure and maintain CI/CD: build images, quality and security checks, publish to registry, rollout/rollback, configuration version control.
• Ensure network connectivity and secure boundaries between on-prem and Yandex Cloud: DNS, ingress, TLS, routing, VPN/dedicated channels, segmentation, and secret management.
• Configure monitoring, logging, and tracing: availability and performance metrics, SLI/SLO, alerting, incident analysis, and capacity planning.
• Support GPU and CPU workloads, including resource planning, autoscaling, queues, and isolation of inference, sandbox, and batch tasks.
• Work with development, data/ML and information security teams: determine operational requirements, conduct technical reviews, participate in emergency work and postmortem.
What we offer
- Work in a stable company (resident Astana Hub);
- Assistance and support in professional and career growth (individual development plan, mentoring, internal training system);
- A friendly team atmosphere and work with excellent specialists;
- Professional Certificate Compensation Program;
- A good social package (free corporate English courses, sports subscription compensation, loyalty program, etc.).
| Name | Work Experience | Contacts | Response date | Notes | Response status |
|---|
There are no responses yet
When someone responds to a vacancy, a list of candidates will appear here