The post has been translated automatically. Original language: Russian
I ask myself the same question every time another "center of gravity" appears in architecture. Service-oriented architecture, then microservices, then Kubernetes. And I keep asking myself is this a real shift for a decade or a HYPE that will quietly blow away in two years?
The question is not an idle one - the answer depends on where to invest your time and how to design the systems that will have to be maintained for the next five years.
And when Jonathan Bryce, executive director of CNCF (Cloud Native Computing Foundation), says at the anniversary event that AI-inference will be the next such shift, it deserves our attention.
Not because he's necessarily right, but because CNCF is 240+ projects and an ecosystem that defines what cloud-native infrastructure looks like in enterprise.
If this ecosystem turns towards AI inference, the consequences will affect everyone who works with the Kubernetes platform, which essentially controls where and how your containers are launched.
The question I asked myself was, if AI agents really become the main way to consume backend services, what does this mean for the way we design systems today?
Context: CNCF celebrates 10 years
Bryce spoke at an event in New York on the occasion of the 10th anniversary of the founding of the consortium. CNCF currently oversees more than 240 projects, from Kubernetes and Prometheus to Argo and dozens of lesser-known tools. It is the largest cloud-native ecosystem, with over 700 corporate members, thousands of developers, and an infrastructure that de facto standardizes how enterprise builds and runs applications.

This is an important point. This is not an independent study with a sample of respondents and a verifiable methodology. This is the position of the person who manages the ecosystem and is interested in ensuring that Kubernetes remains relevant. When the head of CNCF says “Kubernetes will become the center of AI inference,” he simultaneously describes the trend and shapes it. You need to read adjusted for this interest. But the presence of interest does not make the arguments incorrect - it just requires double-checking the facts.
Kubernetes inference: not a hypothesis, but a direction
Bryce's central thesis is that as open source fundamental AI models improve, organizations will create smaller, specialized models based on them, trained for a narrower range of tasks.
Most of these models will be deployed on inference engines running on Kubernetes clusters. Bryce frames this as a “huge opportunity” for CNCF to become the center of gravity for inference.
I can imagine how teams that a year ago didn't even consider Kubernetes for ML workloads are now deploying vLLM and Triton Inference Server there. The reasons are clear: the same orchestrator, the same observability stack (Prometheus + Grafana), the same scaling mechanisms, the same expertise in the team. Deploying a separate infrastructure for inference when you already have mature Kubernetes is an additional operational burden that no one wants to bear.
But it's not just about technology. If the inference engine runs in the same cluster as other services, it is supported by the same platform engineering team, it is monitored by the same stack, and it obeys the same security policies. This reduces the cognitive burden for the entire organization. The alternative is separate ML platforms with separate teams, separate monitoring, and separate duties - this is the fragmentation that Kubernetes was originally designed to eliminate.
It's already happening.
The question is, how deep will it go beyond the top tech companies?
Source: CNCF Chief: AI Inference Will Drive Increased Cloud-Native Software Consumption (report from the CNCF 10th Anniversary Event, New York, February 2026, based on a speech by Executive Director Jonathan Bryce)
Я каждый раз задаю себе один и тот же вопрос, когда в архитектуре появляется очередной «центр гравитации». Сервис-ориентированная архитектура, потом микросервисы, потом Kubernetes. И постоянно задаюсь вопросом это реальный сдвиг на десятилетие или хайп, который через два года тихо сдуется?
Вопрос не праздный - от ответа зависит, куда вкладывать свое время и как проектировать системы, которые придется поддерживать следующие пять лет.
И вот когда Jonathan Bryce, исполнительный директор CNCF (Cloud Native Computing Foundation), на юбилейном мероприятии говорит, что AI-инференс станет следующим таким сдвигом, это заслуживает нашего внимания.
Не потому что он обязательно прав, а потому что CNCF - это 240+ проектов и экосистема, которая определяет, как выглядит cloud-native инфраструктура в enterprise.
Если эта экосистема разворачивается в сторону AI-инференса, последствия затронут каждого, кто работает с платформой Kubernetes, которая по сути рулит тем, где и как запускаются ваши контейнеры.
Вопрос, который я задал себе если AI-агенты действительно станут основным способом потребления backend-сервисов, что это значит для того, как мы сегодня проектируем системы?
Контекст: CNCF отмечает 10 лет
Bryce выступал на мероприятии в Нью-Йорке по случаю 10-летия основания консорциума. CNCF сегодня курирует более 240 проектов - от Kubernetes и Prometheus до Argo и десятков менее известных инструментов. Это крупнейшая экосистема cloud-native, с более чем 700 корпоративными участниками, тысячами разработчиков и инфраструктурой, которая де-факто стандартизирует то, как enterprise строит и запускает приложения.

Тут важный момент. Это не независимое исследование с выборкой респондентов и проверяемой методологией. Это позиция человека, который управляет экосистемой и заинтересован в том, чтобы Kubernetes оставался релевантным. Когда глава CNCF говорит “Kubernetes станет центром AI-инференса” - он одновременно описывает тренд и формирует его. Читать нужно с поправкой на эту заинтересованность. Но наличие интереса не делает аргументы неверными - просто требует перепроверки фактами.
Инференс на Kubernetes: не гипотеза, а направление
Центральный тезис Bryce звучит так: по мере того как open source фундаментальные AI-модели улучшаются, организации будут создавать на их основе меньшие специализированные модели, натренированные на более узкий диапазон задач.
Большинство этих моделей будут развернуты на inference engines, работающих на Kubernetes-кластерах. Bryce формулирует это как “огромную возможность” для CNCF стать центром гравитации для инференса.
Я представляю, как команды, которые год назад даже не рассматривали Kubernetes для ML-нагрузок, сейчас разворачивают vLLM и Triton Inference Server именно там. Причины понятны: тот же оркестратор, тот же стек наблюдаемости (Prometheus + Grafana), те же механизмы масштабирования, та же экспертиза в команде. Разворачивать отдельную инфраструктуру для инференса, когда у тебя уже есть зрелый Kubernetes - это дополнительная операционная нагрузка, которую никто не хочет нести.
Но дело не только в технике. Если inference engine работает в том же кластере, что и остальные сервисы, его поддерживает та же команда platform engineering, он мониторится тем же стеком, он подчиняется тем же политикам безопасности. Это снижает когнитивную нагрузку для всей организации. Альтернатива - отдельные ML-платформы с отдельными командами, отдельным мониторингом и отдельными дежурствами - это та фрагментация, которую Kubernetes изначально был призван устранить.
Это уже происходит.
Вопрос - насколько глубоко зайдет за пределы топовых tech-компаний?
Источник: CNCF Chief: AI Inference Will Drive Increased Cloud-Native Software Consumption (репортаж с мероприятия в честь 10-летия CNCF, Нью-Йорк, февраль 2026, на основе выступления исполнительного директора Jonathan Bryce)