The post has been translated automatically. Original language: Russian
It would seem a simple question:
1 . What types of files do you know?
Many analysts are able to upload data to pandas from Excel and csv. But the question becomes much more interesting when we are asked to organize data storage in a distributed repository.
2. Do you know what Parquet or ORC is? Do you know how they are organized internally to make the most of their functionality?
3. You've probably heard of data compression. But which of these algorithms is possible (and is it necessary?) should I use it during the calculation period for shuffle or to upload the calculation results to the storage?
Alexey Dral tells a case from his own practice: how they achieved a more than 10-fold performance effect in a search engine only with the help of proper data Layout and the choice of compression algorithms. The analytical calculation based on the analysis of user behavior on the Internet over the past month took 4 days of cluster time. We managed to optimize these calculations for up to 4 hours on the same cluster.
Do you want to learn how to solve problems for which you are clearly looking for a career advancement? Then sign up for module 10. "Optimization of storage (Data Layout)".
Available in the framework of:
A practical course on Big Data
A separate part of the course "Part 3. RT, NoSQL, Data Layout"
A bonus for math lovers:
Overview of HDFS 3.0 and Higher Algebra (Galois fields and Reed-Solomon codes)

Useful information
Past issues:
Read first: Part 1 guarantees in IT projects
Previous issue: NoSQL Part 9, CAP theorem, ring of omnipotence and denormalization
Save and subscribe if you want to be in demand in IT.
When this post gets 1k+ views, 25+ likes or 10+ comments, we will post a checklist for preparing for an interview for the position of Data Engineer. Access to knowledge is in your hands.
BigData Team: the way you learn best
Казалось бы простой вопрос:
1️⃣ Какие типы файлов вы знаете?
Многие аналитики умеют грузить данные в pandas из Excel и csv. Но вопрос становится гораздо интереснее, когда нас просят организовать хранение данных в распределенном хранилище.
2️⃣ Вы знаете, что такое Parquet или ORC? Вы знаете как они организованы внутри, чтобы использовать их функционал по максимуму?
3️⃣ Вы наверняка слышали о сжатии данных. Но какие из этих алгоритмов можно (и нужно ли?) использовать в период вычислений для shuffle или для выгрузки результатов расчетов в хранилище?
Алексей Драль рассказывает кейс из собственной практики: как они в поисковой системе только с помощью правильной организации данных (Data Layout) и выбора алгоритмов сжатия добились более чем 10-кратного эффекта в производительности. Аналитический расчет по анализу поведения пользователей в Интернете за последний месяц занимал 4 дня кластерного времени. Нам удалось оптимизировать эти расчеты до 4-х часов на том же кластере.
Хотите научиться решать задачи, за которые вам явно светит карьерное повышение? Тогда записывайтесь на модуль 10. "Оптимизация хранилища (Data Layout)".
Доступно в рамках:
Практического курса по Big Data
Отдельной части курса "Часть 3. RT, NoSQL, Data Layout"
Бонус для любителей математики:
Обзор HDFS 3.0 и высшей алгебры (поля Галуа и коды Рида-Соломона)

Полезная информация
Прошлые выпуски:
Читать сначала: ч.1 гарантии в IT проектах
Предыдущий выпуск: ч.9 NoSQL, CAP теорема, кольцо всевластия и денормализация
Сохраните и подпишитесь, если хотите быть востребованным в IT
Когда данная публикация наберет 1k+ просмотров, 25+ лайков или 10+ комментариев, то мы выложим чек-лист по подготовке к собеседования на позицию Data Engineer. Доступ к знаниям — в ваших руках.
BigData Team: the way you learn best