The post has been translated automatically. Original language: Russian
Autonomous AI Agent hacked Huggingface
For those who are not aware, Huggingface is a great resource for AI, it is such an open platform where developers from all over the world share their AI models, code for their work, training data and any other related to this industry... I will try to describe what happened very simply, so that it would be clear to most. Go.
The other day, in Huggingface (I'll write HF later), the infrastructure was broken not by mom's hatskers, but by an autonomous AI agent. He climbed through a leaky (as it turned out) dataset loader and over the weekend he himself stole access to internal servers — thousands of actions in a row, without interruptions, without a single rule that would slow him down. The public models and websites were not touched — the company checked this separately. But some of the internal data and accesses have leaked. Paramampammmmmm....The most interesting thing is not the attack itself, but what happened afterwards.
HF security guards began to analyze 17 thousand recorded actions of the attacker and first tried the usual commercial models — the same ones that we all use (zhpt chats, anthropics, etc.). And these models simply refused to help. Because they don't know how to distinguish a security guard who parses malicious code from the attacker himself, and anyone who sends hacking commands and files is suspicious to AI by default. Whether you save a company or break it, the models don't matter, they block the same way. As a result, they had to raise the open model on their own servers, without other people's restrictions. Only then did the analysis go quickly — hours instead of days, and nothing else leaked outside the company.
This (case) differed from everything we had dealt with before in one important way: it was controlled, from start to finish, by an autonomous system of AI agents - and we discovered and analyzed it mainly with the help of our own AI.
It turns out to be a funny (and not very) picture: the attacker has no rules at all, and at this moment the defender is almost forced to prove to his own AI that he is not Dr. Evil. The rules of the game don't interfere with who they owe, and this is not a one-time bug, but rather a big systemic hole in the way security around AI is currently organized.
The hole in HF was closed, the servers were rebuilt, access was updated, and their specialists, law enforcement agencies, and possibly those whose names cannot be named, were connected... As a useful conclusion for you, if you have tokens on HuggingFace, update them immediately, without delay, because what is there and how it leaked. Autonomous, offensive tools based on artificial intelligence are no longer theoretical. This reduces the cost of running a broad, patient, multi-stage campaign and works at the speed of a machine. Protecting an online platform now means treating the data surface and models as a first-class attack surface and using AI in defense to keep up. We will continue to invest there and share what we learn.It's a pity that if we get hacked, we don't have our own AI that will agree to understand the attack, and won't block us as suspicious...an article from HF themselves about this incident.
Автономный ИИ-агент взломал Huggingface
Для тех, кто не в курсе Huggingface — большой ресурс для ИИ, это такая открытая платформа, где разработчики со всего мира делятся своими модели ИИ, кодом для их работы, обучающими данными и всякое другое связанное с этой отраслью... Постараюсь очень просто описать что произошло, чтобы большинству было понятно. Поехали.
На днях в Huggingface (далее буду писать HF) сломали инфраструктуру не мамины хацкеры, а автономный ИИ-агент. Залез через дырявый (как оказалось) загрузчик датасетов и за выходные сам растащил доступы по внутренним серверам — тысячи действий подряд, без перерывов, без единого правила, которое бы его тормозило. Публичные модели и сайты не тронул — в комипании это проверили отдельно. А вот часть внутренних данных и доступов утекла. Парампампаммммм....Самое интересное — не сама атака, а то, что случилось потом.
Безопасники HF начали разбирать 17 тысяч записанных действий атакующего и сначала пробовали обычные коммерческие модели — те же, которыми пользуемся мы все (чаты жпт, антропики и т.д.). И эти модели просто отказались помогать. Потому что не умеют отличить безопасника, разбирающего вредоносный код, от самого злоумышленника и любой, кто присылает хакерские команды и файлы, для ИИ подозрителен по умолчанию. Спасаешь компанию или ломаешь её — модели без разницы, она блокирует одинаково. В итоге им пришлось поднимать открытую модель на своих же серверах, без чужих ограничений. Только тогда разбор пошёл быстро — часы вместо дней, и ничего чужого не утекло за пределы компании.
Этот (случай) отличался от всего, с чем мы имели дело раньше, одним важным образом: он управлялся, от начала до конца, автономной системой агентов ИИ - и мы обнаружили и проанализиромвали его в основном с помощью нашего собственного ИИ.
Получается забавная (и не очень) картина: у атакующего никаких правил вообще, а защитник в этот момент вынужден чуть ли не доказывать своему собственному ИИ, что он не доктор Зло. Правила игры мешают не тому, кому должны и это не разовый баг, а походу большая системная дыра в том, как сейчас устроена безопасность вокруг ИИ.
Дырку в HF закрыли, серверы пересобрали, доступы обновили, подключили и своих специалистов, правоохранительные органы и возможно подключили тех, чьё имя нельзя называть... Как полезный вывод для вас — если у вас есть токены на HuggingFace — обновите их обязатеьлно, не откладывая, хз что там и как утекло. Автономные, наступательные инструменты, основанные на искусственном интеллекте, больше не являются теоретическими. Это снижает стоимость проведения широкой, терпеливой, многоэтапной кампании и работает со скоростью машины. Защита онлайн-платформы теперь означает рассматривать поверхность данных и модели как первоклассную поверхность атаки и использовать ИИ в защите, чтобы не отставать. Мы будем продолжать инвестировать туда и делиться тем, что узнаем.Жаль, что если взломают нас — у нас нет своего ИИ, который согласится разбираться в атаке, а не заблокирует нас как подозрительного...статья от самих HF об этом инциденте.