Vacancies
Astana
Full-time
Full-time
Senior/middle ML engineer
Archive
The vacancy has been archived and is no longer active
Main requirements
What a candidate should know:
- ViT & Multimodality: A deep understanding of Vision Transformers (ViT, Swin, DETR) and video Transformers (TimeSformer, ViViT). Experience working with VLM (LLaVA, Qwen-VL) at the interface of text and video streams, fine-tuning (LoRA/QLoRA).
- Video Analytics & DeepStream: Designing highload pipelines for real-time video analytics. Practical experience with NVIDIA DeepStream SDK (GStreamer, nvinfer, multi-streaming, tracking), processing RTSP streams with minimal latency.
- GPU Optimization & TensorRT: Profiling and acceleration of inference on GPU. Confident work with TensorRT (layer fusion, INT8/FP16 calibration, dynamic shapes, building engines via trtexec), understanding CUDA specifics and memory bottlenecks.
- OpenVINO & Quantization: Optimization of models for CPU/Edge deployment. Experience working with OpenVINO and NNCF for Post-Training Quantization (INT8/INT4), accuracy-aware tuning and minimizing quality degradation during compression.
- Dynamo & Compilation of graphs: Using PyTorch 2.0 Dynamo (torch.compile) for JIT compilation and graph merging (AOTAutograd, Triton kernels), accelerated learning and out-of-the-box inference.
- LLM & Custom Agent Systems: Designing agent-based architectures on top of LLM without strict reference to the "magic" of frameworks. Implementation of state graphs (LangGraph, finite automata), custom Tool Calling/Function Calling, ReAct/Plan-and-Execute patterns, agent context and memory management, multi-agent interaction orchestration. LLM Serving (vLLM, Continuous Batching).
- RAG & Retrieval: Pipeline design from chunking to generation. Understanding vector databases (HNSW), embedding types (Dense, Sparse, ColBERT), and rerun architecture.
- System Design & MLOps: End-to-end AI architecture design (FastAPI, K8s, Kafka), Cost/Performance trade-off calculation, LLMOps (Evaluation: RAGAS, LLM-as-a-Judge), CI/CD for ML models.
What you will do
What should a candidate be able to do:
- Build highload pipelines for video analytics: To design and launch real-time video processing systems (dozens/hundreds of RTSP streams) based on NVIDIA DeepStream. Be able to write custom GStreamer plugins, link tracking and detection, and minimize end-to-end latency (e2e latency).
- Getting the most out of hardware (GPU/CPU Optimization): Take the PyTorch model and speed it up 3-10 times. Independently convert models to ONNX/TensorRT (adjust dynamic dimensions, INT8/FP16 calibration) or OpenVINO (use NNCF for quantization while maintaining accuracy). Apply torch.compile (Dynamo) to speed up training and inference.
- Design custom agent systems: Create autonomous AI agents from scratch (or based on LangGraph), abandoning the "magic" of heavy frameworks where control is needed. Be able to link LLM with external APIs (Function Calling), build state graphs, manage the agent's context/memory, and handle its errors/hallucinations.
- Implement VLM and ViT in business processes: Fine-tune Vision Transformers and multimodal models (LLaVA, Qwen-VL) for specific domain data (specific frames, medical images, satellite photos). Glue CV pipelines (YOLO/DeepStream) with LLM to generate text reports on video.
- Build production systems based on RAG: Inject tons of unstructured data, select optimal chunking and embedding strategies, build hybrid search (BM25 + Dense) and re-routing. Be able to evaluate the quality of the RAG pipeline (RAGAS) and eliminate hallucinations.
- Output AI to Production (End-to-End): Package models in microservices (FastAPI/gRPC),orchestrate them in Docker/Kubernetes. Configure CI/CD for ML, monitor inference (TTFB, throughput, GPU utilization, drift metrics) and build processes for retraining/updating models without downtime.
- Make architectural decisions: Evaluate the Cost/Performance trade-off: choose between calling the provider's API and deploying the Open Source model on their GPUs; decide when to write a custom agent in pure Python and when to use a ready-made framework.
- Take technical leadership: Design the architecture of the AI components of the project, decompose tasks for middles and juniors, conduct code reviews of ML code and set quality standards (logging, testing, reproduction) in the team.
What we offer
Conditions:
- Work on a large-scale state/national project
- Modern AI stack (LLM, multimodal, CV)
- The ability to influence the architecture of solutions
- Competitive wages
ТОО "Yurt Tech"
Мы верим, что будущее создают люди. Наша миссия — объединять талантливых специалистов и создавать технологии, которые приносят пользу обществу и развивают нашу страну.
Participant of Astana Hub
Job Type
Full-time
Type employment
Full-time
Required education level
Bachelor's degree
Direction
Information Technology
Work experience
At least 3 years
Salary
-
All statuses
New
Viewed
Under review
Invitation
Confirm
Rejected
| Name | Work Experience | Contacts | Response date | Notes | Response status |
|---|
There are no responses yet
When someone responds to a vacancy, a list of candidates will appear here
123 Views