OCR Language Pack: which languages does text recognition support?
OCR Language Pack: which languages does text recognition support?
Your benefit
With the right language pack you get significantly better recognition results for non-Latin scripts — for example with contracts in Russian, Arabic, Hindi or Chinese. Choosing the wrong language leads to unreadable OCR output and therefore to faulty AI analysis results. A single click in the upload dialogue saves you a great deal of manual correction later on.
How it works
1. Select the language pack when uploading
When you upload a contract, the Add contracts [Verträge hinzufügen] dialogue opens. At the very top you will see the Text recognition language [Sprache der Texterkennung] dropdown with the default Latin script [Lateinische Schrift] and the note:
"The language pack determines which writing system is used for text recognition."
Click the dropdown and select the writing system your contract is written in.
2. Select a PDF or upload it by drag and drop
Once you have chosen the language, drag your PDF into the blue upload area (Select contracts in PDF format or drag them here [Verträge im PDF-Format auswählen oder hier reinziehen]) or click inside it to select a file. The OCR engine then processes the contract with the selected language pack.
3. Overview of the available language packs
ContractHero currently offers eight language packs. Each pack covers several languages that use the same writing system:
Language pack | Languages covered |
|---|---|
Latin script [Lateinische Schrift] (default) | English, German, French, Spanish, Portuguese, Italian, Dutch, Polish, Czech, Hungarian, Romanian, Swedish, Danish, Norwegian, Finnish, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian, Indonesian, Vietnamese, Turkish, Malay, Tagalog, Swahili, Afrikaans, Azerbaijani, Bosnian, Catalan, Welsh, Galician, Icelandic, Irish, Albanian, Maltese, Uzbek, Sundanese, Javanese |
Chinese & Japanese [Chinesisch & Japanisch] | Chinese (simplified & traditional), Japanese |
Cyrillic [Kyrillisch] | Russian, Ukrainian, Bulgarian, Serbian |
Arabic [Arabisch] | Arabic, Persian (Farsi), Urdu |
Devanagari (Hindi) | Hindi, Marathi, Nepali |
Korean [Koreanisch] | Korean |
Greek [Griechisch] | Greek |
Thai | Thai |
Frequently asked questions
What is the difference between OCR and AI contract analysis?
OCR converts the scanned PDF or image into machine-readable text — this is the prerequisite for the AI contract analysis to then recognise fields such as the contract start date, the contracting party or the notice period. Without clean OCR there is no clean AI analysis.
Do I always have to select the language pack manually?
No — for contracts in Latin script (all Western European languages, Turkish, Vietnamese, etc.) the default is already correct. You only need to change the dropdown if your contract is written in a different script (e.g. Russian in Cyrillic, or Arabic).
What happens if I select the wrong language pack?
OCR recognition then delivers text that is virtually unusable — for example if you process an Arabic document with Latin script [Lateinische Schrift]. You can delete the contract after the upload and upload it again with the correct language.
Can one contract use several language packs at the same time?
One language pack is applied per upload. For bilingual contracts (e.g. German + English), Latin script [Lateinische Schrift] is sufficient — both languages are included. For mixed scripts (e.g. German + Chinese in the same document), select the pack that covers the contractually relevant part.
Does the AI analysis work in all of these languages?
OCR text recognition covers all of the languages listed. The AI contract analysis (filling in the fields via a prompt) works reliably primarily in German and English — with other languages the result depends heavily on how your prompt is worded and on the language of the contract. Get in touch with us if you would like to analyse contracts in further languages systematically.
Where do I find the language pack for email import or XLS import?
For automatic imports (email, XLS), Latin script [Lateinische Schrift] is always used. If you regularly receive non-Latin contracts, a manual upload with a language selection is the more reliable route.
Does ContractHero recognise handwritten notes or signatures?
The OCR engine is optimised for printed text. Handwriting is sometimes recognised, but not reliably. We process signature blocks separately via the signature function (see Related articles).
Good to know
- The default is enough for 95% of use cases. If you mainly process European contracts, you never need to change anything in the dropdown.
- The same language for everything in a bulk upload. If you drag several PDFs into the dialogue at the same time, the selected language pack is applied to all of them. Upload batches with mixed languages separately.
- Language pack ≠ translation. The dropdown selects the writing system for text recognition — the contract is not translated into another language.
- Scan quality matters more than the pack. A poor scan in the correct language delivers worse results than a clean scan with the wrong pack. Aim for 300 dpi and high-contrast originals.
- We support you with the setup. If you regularly process contracts in non-Latin scripts, we will go through the setup together during onboarding.
Related articles
- How do I add contracts to ContractHero?
- How does contract analysis work?
- What are prompts and how do you write good prompts for the AI contract analysis?
Updated on: 08/26/2026
Thank you!
