Decision acceptance deadline

25.06.26 (inclusive)

Form of award

contractual

Product status

MVP

Task type

ICT tasks

Сфера применения

Robotics

Область задачи

Neurotechnology and artificial Intelligence

Type of product

Software/ IS

Problem description

In the practical application of speech recognition and offline speech analysis technologies, a significant obstacle is the inconsistency of audio materials with the requirements for correct information processing. Dialogues between two speakers are often recorded in the same audio channel, which makes it difficult to automatically separate them for later analysis and shorthand. Existing ASR solutions for Russian and Kazakh languages do not provide stable diorization, especially in mixed speech (code-switching), which leads to a decrease in the accuracy and readability of transcriptions.

Expected effect

The result of the development will be a software module that provides: 1. Automatic separation of audio recordings into replicas of two speakers (diorization) with an accuracy of at least 90%; 2. Correct speech recognition in Russian and Kazakh, including mixed utterances; 3. Formation of a structured transcript with timestamps and speaker IDs (Speaker 1, Speaker 2); 4. Compatibility with existing data analysis and storage systems via API; 5. The ability to use any external or local STT service that provides the required recognition quality.

Full name of responsible person

Danchenko Maxim

Purpose and description of task (project)

Creation of an automatic speech recognition (ASR) module with diorization support for two speakers, ensuring accurate separation of phrases and correct shorthand of audio recordings in Russian and Kazakh. Task description: To develop a software mechanism that accepts an audio file as input and performs the following functions: 1. definition of speech sections and automatic separation into two speakers; 2. Speech recognition in Russian and Kazakh languages; 3. formation of a text file with speaker markup and timestamps; 4. Ensuring diorization accuracy of at least 90% and recognition accuracy of at least 85%; 5. Providing the finished result in a format suitable for subsequent analysis and integration.

Note