Price: 0
Number of applications: 7
07.09.26 (inclusive)
contractual
Finished product
ICT tasks
Robotics
Neurotechnology and artificial Intelligence
Software/ IS
In the practical application of speech recognition and offline speech analysis technologies, one of the key problems is the insufficient quality of the source audio data and the limited possibilities of their automatic processing. In particular, when recording dialogues between two or more participants, one audio channel is often used, which significantly complicates the automatic identification and separation of cues by speaker and, as a result, reduces the quality of subsequent transcription and analysis. The processing of Russian-Kazakh mixed speech (code-switching) is an additional difficulty. Existing ASR solutions do not always provide stable recognition and diarization in such conditions, which leads to errors in identifying speakers, reducing the accuracy of transcription and impairing its readability. Thus, a solution is required that can provide more accurate separation of speech by speakers and high-quality transcription of Russian-Kazakh dialogues, including when using single-channel audio recordings.
The result of the development is a software module for automatic processing and analysis of audio recordings of dialogues between two speakers, which improves the quality and efficiency of subsequent transcription and speech analysis. The module should provide: 1. Automatic diarization of audio recordings — identification of speech fragments and their distribution between two speakers with a target accuracy of at least 90%, including for single-channel recording. Russian Russian and Kazakh language speech recognition with a target accuracy of at least 85%, including the processing of mixed Russian-Kazakh speech (code-switching). 3. Formation of a structured transcript with preservation of the sequence of the dialogue, timestamps of the beginning and end of the remarks and speaker identifiers (Speaker 1, Speaker 2). 4. A flexible recognition architecture that allows the use of various STT/ASR services - both external cloud and local/offline solutions — depending on the requirements for quality, security, performance and data placement. 5. Integration with existing information systems through API and standard data exchange formats. 6. The possibility of autonomous operation in a closed or limited information circuit when using a local STT/ASR service, without mandatory transmission of audio data to external cloud systems. 7. Standardization of the processing result — the formation of a single structured data representation suitable for subsequent search, analysis, storage, automatic processing and transmission to other information systems. 8. Reducing the volume of manual processing of audio recordings due to automatic speaker detection, speech recognition and preparation of the finished transcript. Thus, the development will make it possible to create a vendor-independent mechanism for the diarization and preparation of a structured transcript that can use various speech recognition technologies and adapt to the requirements of a specific information system.
Torchik V.V.
Purpose and description of task (project)
As part of the project, it is necessary to develop a software mechanism that provides automatic processing of audio files without the need for manual pre-markup. The decision should accept an audio recording of the dialogue between the two speakers, including one recorded in the same audio channel, and generate a structured text result. The module's functionality should include: 1 Audio input file processing — reception and preprocessing of audio recordings, including normalization and selection of speech fragments. 2 Speech detection — automatically detects sections of an audio recording containing speech, eliminating or minimizing the effects of pauses, background noise, and non-speech fragments. 3 Diarization of two speakers — automatic detection and separation of the remarks of two participants in the dialogue, even when using single-channel recording. 4 Speech Recognition (ASR) — speech to text conversion with support for Russian and Kazakh languages. 5 Mixed speech processing (code-switching) — correct recognition of fragments in which the speaker switches from Russian to Kazakh and vice versa. 6 Formation of a transcript is the creation of a text file with sequential markup of replicas by speaker, timestamps of the beginning and end of each speech fragment and, if necessary, an indication of a specific language. 7. Preserving the structure of the dialogue is to ensure that the order of the text phrases corresponds to their actual sequence in the audio recording. 8 Result quality control — achievement of targets: accuracy of diarization — at least 90%; speech recognition accuracy is at least 85%. 9 Result formatting — providing a transcript in a structured format suitable for further analysis, storage, processing and integration with other information systems. 10 Integration capability — implementation of a software interface or other standard mechanism for transferring results to third-party systems and applications.