Price: 50000
Number of applications: 8
29.09.26 (inclusive)
Payment
Finished product
ICT tasks
Robotics
Information processing and transformation
Software/ IS
What is it about In our service (verification of packages of documents for the state The text is extracted from the downloaded PDF for further analysis. Most files have a text layer and are parsed locally. But some of them come in scans without a text layer — they are now recognized by the cloud-based Google Gemini 2.5 Flash. We want to get away from the cloud and do it locally, on our hosting.
Increasing the independence of the product from third-party services
Zhantas Asylbek
Purpose and description of task (project)
A task for the community: local OCR of PDF scans (RU/KK) instead of Gemini What is it about In our service (verification of packages of documents for the state text is extracted from the downloaded PDF files for further analysis. Most files have a text layer and are parsed locally. But some of them come in scans without a text layer — they are now recognized by the cloud-based Google Gemini 2.5 Flash. We want to get away from the cloud and do it locally, on our hosting. What needs to be done A module that accepts a PDF scan and returns the extracted text (plain text, UTF-8), verbatim and in the correct reading order. It should work offline, without external APIs. Important caveat on the type of documents This is not the recognition of passports, identification cards, complex forms or manuscripts. Scans are ordinary printed text documents: charters, certificates, contracts, business plans, letters. That is: mostly solid printed text, 1 column; machine font (not a manuscript); there may be simple tables, a stamp/signature at the bottom — but this is not the main content. This greatly simplifies the task: you don't need the "intelligence" of the multimodal LLM level for complex layout — you need reliable OCR of printed RU/KK text. Specificity Languages: Russian + Kazakh (ә ғ қ ң ө))))), sometimes in the same document. Quality: scans of different quality (distortion, noise, DPI 150-300). Critical: numbers, amounts, BIN/IIN (12 digits), dates — an error in a digit is unacceptable. Input: PDF up to 15 MB, 1-30 pages. Output: UTF-8 text, top—to-bottom / left-to-right order; tables - line by line. Integration: A PHP module or a local CLI/HTTP service called from PHP (CodeIgniter 4 / PHP 8.2 stack). Allowed stack Everything that is hosted: Tesseract 5 (rus + kaz LSTM), PaddleOCR / EasyOCR, docTR; rasterization of PDF→image (pdftoppm / Ghostscript); preprocessing (OpenCV / ImageMagick: deskew, denoise, binarize). Combinations are welcome. Acceptance criteria Benchmark: we will provide a set of ~150-200 real impersonal PDFs (RU/KK) with a reference text. Metrics: CER / WER vs. benchmark; Field-accuracy for critical fields (BIN, amounts, dates). The goal: accuracy is no worse than Gemini 2.5 Flash on the same set (we will publish its results as a baseline). Performance: less than N seconds per document (to specify the hardware), completely local. The format of the submission Repository with code + README (installation, dependencies, launch). A benchmark run script that outputs CER/WER and field-accuracy. A brief description of the pipeline. Deadline: 01.10.26