Astanahub Logo
Vacancies

Senior/middle ML engineer

Archive
Astana Full-time Full-time

The vacancy has been archived and is no longer active

Main requirements

What a candidate should know:

  • ViT & Multimodality: A deep understanding of Vision Transformers (ViT, Swin, DETR) and video Transformers (TimeSformer, ViViT). Experience working with VLM (LLaVA, Qwen-VL) at the interface of text and video streams, fine-tuning (LoRA/QLoRA).
  • Video Analytics & DeepStream: Designing highload pipelines for real-time video analytics. Practical experience with NVIDIA DeepStream SDK (GStreamer, nvinfer, multi-streaming, tracking), processing RTSP streams with minimal latency.
  • GPU Optimization & TensorRT: Profiling and acceleration of inference on GPU. Confident work with TensorRT (layer fusion, INT8/FP16 calibration, dynamic shapes, building engines via trtexec), understanding CUDA specifics and memory bottlenecks.
  • OpenVINO & Quantization: Optimization of models for CPU/Edge deployment. Experience working with OpenVINO and NNCF for Post-Training Quantization (INT8/INT4), accuracy-aware tuning and minimizing quality degradation during compression.
  • Dynamo & Compilation of graphs: Using PyTorch 2.0 Dynamo (torch.compile) for JIT compilation and graph merging (AOTAutograd, Triton kernels), accelerated learning and out-of-the-box inference.
  • LLM & Custom Agent Systems: Designing agent-based architectures on top of LLM without strict reference to the "magic" of frameworks. Implementation of state graphs (LangGraph, finite automata), custom Tool Calling/Function Calling, ReAct/Plan-and-Execute patterns, agent context and memory management, multi-agent interaction orchestration. LLM Serving (vLLM, Continuous Batching).
  • RAG & Retrieval: Pipeline design from chunking to generation. Understanding vector databases (HNSW), embedding types (Dense, Sparse, ColBERT), and rerun architecture.
  • System Design & MLOps: End-to-end AI architecture design (FastAPI, K8s, Kafka), Cost/Performance trade-off calculation, LLMOps (Evaluation: RAGAS, LLM-as-a-Judge), CI/CD for ML models.

What you will do

What should a candidate be able to do:

  • Build highload pipelines for video analytics: To design and launch real-time video processing systems (dozens/hundreds of RTSP streams) based on NVIDIA DeepStream. Be able to write custom GStreamer plugins, link tracking and detection, and minimize end-to-end latency (e2e latency).
  • Getting the most out of hardware (GPU/CPU Optimization): Take the PyTorch model and speed it up 3-10 times. Independently convert models to ONNX/TensorRT (adjust dynamic dimensions, INT8/FP16 calibration) or OpenVINO (use NNCF for quantization while maintaining accuracy). Apply torch.compile (Dynamo) to speed up training and inference.
  • Design custom agent systems: Create autonomous AI agents from scratch (or based on LangGraph), abandoning the "magic" of heavy frameworks where control is needed. Be able to link LLM with external APIs (Function Calling), build state graphs, manage the agent's context/memory, and handle its errors/hallucinations.
  • Implement VLM and ViT in business processes: Fine-tune Vision Transformers and multimodal models (LLaVA, Qwen-VL) for specific domain data (specific frames, medical images, satellite photos). Glue CV pipelines (YOLO/DeepStream) with LLM to generate text reports on video.
  • Build production systems based on RAG: Inject tons of unstructured data, select optimal chunking and embedding strategies, build hybrid search (BM25 + Dense) and re-routing. Be able to evaluate the quality of the RAG pipeline (RAGAS) and eliminate hallucinations.
  • Output AI to Production (End-to-End): Package models in microservices (FastAPI/gRPC),orchestrate them in Docker/Kubernetes. Configure CI/CD for ML, monitor inference (TTFB, throughput, GPU utilization, drift metrics) and build processes for retraining/updating models without downtime.
  • Make architectural decisions: Evaluate the Cost/Performance trade-off: choose between calling the provider's API and deploying the Open Source model on their GPUs; decide when to write a custom agent in pure Python and when to use a ready-made framework.
  • Take technical leadership: Design the architecture of the AI components of the project, decompose tasks for middles and juniors, conduct code reviews of ML code and set quality standards (logging, testing, reproduction) in the team.

What we offer

Conditions:

  • Work on a large-scale state/national project
  • Modern AI stack (LLM, multimodal, CV)
  • The ability to influence the architecture of solutions
  • Competitive wages
search
All statuses
New
Viewed
Under review
Invitation
Confirm
Rejected
Name Work Experience Contacts Response date Notes Response status
reach

There are no responses yet

When someone responds to a vacancy, a list of candidates will appear here

123 Views
We use cookies to ensure the proper functioning of our website and for analytics. By continuing to use the website, you confirm your acceptance of the use of cookies. More
Share