The post has been translated automatically. Original language: Russian
The radiologist looks at the CT scan. There are hundreds of slices and thousands of details in front of him. He reviews dozens of such studies per day. By the end of the shift — fatigue, decreased concentration, cognitive load.
The computer vision algorithm looks at the same image. He doesn't get tired. He doesn't get distracted. Processes each pixel with the same accuracy — the first image per shift and the last one.
This does not mean that the algorithm is better than the doctor. This means that they are good at different things. And it is this combination that changes the medical diagnosis.
Why medical imaging is an ideal task for computer vision
Computer vision works best where there is a large amount of labeled data and a clear criterion for the correct answer.
Medical imaging is ideal: decades of CT scans, MRI scans, X-rays, and histological preparations with verified diagnoses. Millions of images where the correct answer is known. This is an ideal training dataset.
Convolutional neural networks (CNNs) are able to extract hierarchical features from images: from edges and textures at a low level to complex anatomical patterns at a high level. This is what makes them a powerful tool for analyzing medical images.
Where computer vision is already changing diagnostics
Oncology. The detection of skin cancer from dermatoscopic photographs is one of the first areas where algorithms have achieved the accuracy of a dermatologist. A 2017 Stanford study showed that CNN diagnoses melanoma with an accuracy comparable to an experienced specialist based on a sample of 130,000 images.
Screening of cervical cancer by cytological smears, detection of polyps during colonoscopy in real time, analysis of biopsy preparations for prostate cancer — everywhere algorithms are already working in clinical practice or are undergoing final tests.
Ophthalmology. Diabetic retinopathy is one of the leading causes of blindness in the world. Screening requires the analysis of fundus images by an experienced ophthalmologist. There are not enough specialists, especially in countries with developing medicine.
Google has developed an algorithm that detects diabetic retinopathy with 90% sensitivity and 98% specificity — surpassing the average ophthalmologist. The system has already been deployed in clinics in Thailand and India, where the shortage of specialists is critical.
Radiology. Pneumonia, pneumothorax, fractures on X—rays - algorithms detect them with high accuracy. Stanford's CheXNet detects pneumonia on an X-ray better than the average radiologist on a standard dataset.
An important caveat: "better than the average radiologist on a standard dataset" does not equal "better in real clinical practice." The real pictures are dirtier, the patients are more complex, and the context is richer. This is an honest limitation that is important to understand.
Pathology. The analysis of histological preparations is one of the most time—consuming processes in oncology. The pathologist spends hours examining tissue sections under a microscope. Computer vision algorithms can automatically segment cells, classify tissue types, and detect malignant changes — many times faster than humans.
Where does the algorithm see what a human is missing
This is the most interesting part. There are tasks where computer vision detects patterns that humans are physically unable to see.
Prediction of cardiovascular risk based on retinal imaging. It sounds incredible, but Google's algorithm has found that based on the characteristics of the blood vessels in the fundus image, it is possible to predict the patient's age, gender, diabetes, and risk of heart attack. The human eye does not see these patterns — they are statistically significant only in large samples.
Prediction of genetic mutations based on histological images of a tumor — without expensive genetic testing. The algorithm finds visual correlates of genetic changes that the pathologist is not trained to notice.
Technical challenges
There are several problems that engineers are solving right now.
Data quality and markup. Medical images from different clinics — different equipment, different protocols, different quality. A model trained on data from one clinic may not work well in another. Domain adaptation — adapting the model to new data sources is an active area of research.
Explainability. The doctor should understand why the algorithm made this particular diagnosis. Activation visualization methods — Grad-CAM, SHAP — show which areas of the image the model is looking at. But there is no complete explainability yet.
Rare diseases. The algorithm is good where there are many training examples. For rare pathologies, there is little data, and accuracy drops sharply. Few-shot learning and synthetic data augmentation partially solve the problem.
What does this mean for IT developers?
Medical computer vision is a specialized field with its own standards. DICOM, a medical image format, requires specific libraries and an understanding of metadata. The regulatory requirements for model validation are much stricter than in a conventional CV.
But the market is huge. It is estimated that the market for AI-based medical imaging solutions alone will exceed $20 billion by 2030. And this is one of the few areas where AI has already proven its clinical value with a solid evidence base.
Computer vision is not a substitute for a radiologist or pathologist. It makes their work more accurate, faster and more accessible where there are not enough specialists. This is a rare case when AI in medicine has gone from hype to real clinical results.
Интуитивно кажется, что в такой консервативной и зарегулированной отрасли, как медицина, открытый исходный код — это риск. Где гарантии качества? Кто отвечает за безопасность?
Реальность оказывается противоположной интуиции. В ряде ключевых областей медицинских технологий именно open source стал стандартом, которому доверяют больше, чем закрытым коммерческим альтернативам.
Почему открытость работает в пользу доверия именно в медицине
В обычном бизнес-софте закрытый код часто воспринимается как конкурентное преимущество — защита интеллектуальной собственности. В медицине логика смещается: прозрачность алгоритма часто важнее коммерческой защиты, потому что от точности этого алгоритма зависит здоровье пациентов.
Если код открыт — любой исследователь, регулятор или независимый эксперт может проверить, как именно работает алгоритм, найти ошибки, предложить улучшения. Закрытый код — это вопрос веры в добросовестность производителя. Открытый код — это возможность верификации.
Где open source уже стал стандартом отрасли
AlphaFold от Google DeepMind — система, предсказывающая трёхмерную структуру белков, совершила революцию в структурной биологии. DeepMind открыл исходный код AlphaFold и в партнёрстве с EMBL-EBI создал AlphaFold Protein Structure Database — открытую базу, которая сегодня содержит более 200 миллионов предсказанных структур белков. Это ускорило тысячи исследовательских проектов по всему миру — от разработки лекарств до понимания механизмов заболеваний. Закрытая версия такого масштаба влияния не дала бы.
3D Slicer — open source платформа для анализа и визуализации медицинских изображений, используемая в тысячах исследовательских и клинических проектов по всему миру. Разработанная изначально как магистерский проект в Surgical Planning Laboratory при Brigham and Women's Hospital и MIT Artificial Intelligence Laboratory в 1998 году, она стала фактическим стандартом благодаря открытости — за более чем 25 лет на её основе опубликовано множество научных работ, а скачиваний насчитывается миллионы.
OpenMRS — open source платформа электронных медицинских карт, изначально разработанная Regenstrief Institute для масштабирования лечения ВИЧ в Кении в 2004 году. Сегодня используется в более чем 80 странах, обслуживая миллионы пациентов — преимущественно в регионах с ограниченными ресурсами, где лицензии крупных коммерческих вендоров недоступны.
Биоинформатические инструменты. Большинство фундаментальных инструментов для анализа геномных данных — BLAST, Bioconductor, многие пакеты для анализа секвенирования — open source. Это позволило геномной революции происходить с участием тысяч исследовательских групп по всему миру, а не только тех, кто может позволить себе дорогие коммерческие лицензии.
Hugging Face модели для медицинского NLP. BioBERT, ClinicalBERT и десятки других специализированных языковых моделей для медицинского текста публикуются открыто — это позволяет исследователям и разработчикам по всему миру строить на их основе собственные решения, не начиная с нуля.
Почему это особенно важно для развивающихся технологических экосистем
Для стран с ограниченным венчурным капиталом и более молодой технологической инфраструктурой — как Казахстан и Центральная Азия в целом — open source медицинские инструменты создают уникальную возможность.
Команда без многомиллионных инвестиций может строить продукт на базе проверенной открытой технологии — будь это платформа визуализации, NLP-модель или система электронных карт — вместо того чтобы тратить годы и огромный бюджет на разработку базовой инфраструктуры с нуля.
Где остаются обоснованные ограничения
Открытость не означает отсутствие регуляторики. Программное обеспечение, влияющее на клинические решения, должно проходить ту же сертификацию независимо от того, открыт исходный код или закрыт.
Безопасность через открытость работает только при активном сообществе, которое действительно проверяет и поддерживает код. Заброшенный open source проект может быть опаснее поддерживаемого коммерческого продукта.
Коммерческая поддержка остаётся ценной. Многие медицинские учреждения выбирают гибридную модель — используют open source ядро, но платят за профессиональную поддержку от компаний, специализирующихся на конкретной платформе.
Что это значит для разработчиков
Знание ключевых open source инструментов и платформ в медицинской сфере — реальное конкурентное преимущество для команды, которая хочет быстро выйти на рынок без огромных первоначальных инвестиций в базовую инфраструктуру.
📌 В медицине прозрачность алгоритма часто важнее коммерческой защищённости кода. Open source доказал, что открытость не противоречит надёжности — а в ряде случаев именно она создаёт основу для того уровня доверия, который требует медицина.
Источники:
- AlphaFold Protein Structure Database — открытый доступ от Google DeepMind и EMBL-EBI: https://alphafold.ebi.ac.uk/
- AlphaFold — открытый исходный код, GitHub: https://github.com/google-deepmind/alphafold
- 3D Slicer — история и текущее состояние проекта, Wikipedia: https://en.wikipedia.org/wiki/3D_Slicer
- 3D Slicer — официальный сайт проекта: https://www.slicer.org/
- OpenMRS — официальный FAQ, использование в 80+ странах: https://openmrs.org/faq/