
Who Actually Knows Georgian? A Big Test of 22 AI Models
We put 22 leading AI models to the test in Georgian: morphology, ergative constructions, idioms and literary style. About 90% fail on the ergative case — the winner is DeepSeek V4.1 Flash (92/100).
Is it “გიორგი მივიდა” or “გიორგიმ მივიდა”? For a Georgian speaker the answer is obvious — the first is correct, because the verb “მივიდა” is intransitive and, in the aorist, demands a nominative subject. Yet when we put that question to 22 leading AI models, most of them — around 90% — slipped up exactly there.
We ran one of the largest AI benchmarks ever for the Georgian language: 22 models, four disciplines, a 100-point system. The results show that Georgian is still a serious test for modern technology — but there are models that genuinely master it.
Why Georgian is a special test for AI
Most models are trained on English and Chinese data, while Georgian ranks among the low-resource languages. The Kartvelian family is isolated, and its grammar has features that universal models often fail.
First — agglutination and polypersonal verb agreement: a Georgian verb expresses subject and object at once (“დაგვახვედრეს”, “გამომიგზავნეს”). Second — split ergativity: in the aorist the subject of a transitive verb takes the ergative case (-მა), the intransitive one the nominative (-ი). Third — foreign-language calques (“იქნა მიღებული”, “ჩემს მიერ”) that fill much of the Georgian text online and that models learn as the “correct” forms.
The test: four disciplines
Each model was scored in four areas, under a single API, identical prompts and minimal temperature:
1. Morphology and ergative constructions (25 points). Three case-filling sentences — including “გუშინ ______ (გიორგი) სკოლაში დროზე მივიდა”. The correct answer is “გიორგი”, because the verb is intransitive.
2. Error detection and editing (35 points). A text with five hidden flaws: the Russian passive calque “იქნა მიღებული”, the lexical error “რამოდენიმე”, the tautology “ყველზე საუკეთესო”, the number agreement “ათი სტუდენტები” and a missing comma.
3. Idioms and phraseology (20 points). “კოვზი ნაცარში ჩაუვარდა”, “წყალი შეუყენა”, “თვალში ნაცრის შეყრა” — grasping figurative meaning and contextual examples.
4. Literary style (20 points). Free writing on the theme “language as the memory of a nation” — without calques, in high style.
The drama: how the models got “stuck” while thinking
In the first round, 14 of the 22 models returned an empty response — even though the status was 200 OK. The culprit was the internal chain of reasoning: a simple task takes 100-200 “thinking” tokens in English, but in Georgian that figure exploded to 1,200-2,400 tokens. On a 1,000-token budget they spent the whole limit on thinking and never reached an answer.
With a 2,500-token limit the picture changed: DeepSeek V4.1 Flash spent 1,242 tokens thinking and then wrote a flawless Georgian answer; Qwen3.8-Max needed 1,443 tokens. A special case is GLM-5.3-flash, which fell into an endless reasoning loop — it burned 2,494 tokens and never reached an answer.
The winners
The best result — 92/100 — came from DeepSeek V4.1 Flash: 100% accuracy on the ergative test, flawless editing (it found every flaw, including the Russian calque) and an average of just 6.4 seconds per answer. Second place goes to Qwen3.8-Max (89/100) — the strongest logical reasoning, but with a 28.5-second delay; third is Qwen3.8-Flash (87/100).
MiniMax-M3 is an interesting case: fourth overall (81/100), yet an absolute result in the literary test — 20/20. Its text is so natural that it is hard to tell from a human’s: “Language is not born as a mere form of words, but turns into the spiritual trace of lived experience, pain and joy accumulated over centuries.”
The fastest was NVIDIA Nemotron-3-Ultra-550B — 4.5 seconds — though a Russian Cyrillic word (“გენეτიკური”) crept into its text, a classic example of language “leakage” during multilingual training; it also invented a non-existent grammar rule while explaining one.
The full leaderboard (top-10)
1. DeepSeek V4.1 Flash — 92/100 (6.4 sec)
2. Qwen3.8-Max — 89/100 (28.5 sec)
3. Qwen3.8-Flash — 87/100 (22.1 sec)
4. MiniMax-M3 — 81/100 (12.6 sec)
5. NVIDIA Nemotron-3-Ultra-550B — 79/100 (4.5 sec)
6. Qwen3.8-Max-0902 — 78/100
7. Xiaomi MiMo v2.5 — 54/100
8. StepFun Step-3.7-Flash — 44/100
9. StepFun Step-3.5-Flash — 40/100
10. Ling 3.0 Flash — 36/100
The overall picture is clear: the gap between the top five and the rest is huge — a drop from 79 points to 54. Three models never made it to testing (two because of package limits, one was temporarily unavailable).
Which model for what
For everyday work — writing, editing, question answering — the best choice is DeepSeek V4.1 Flash: the best speed-price-quality ratio. For complex analysis and legal texts, Qwen3.8-Max goes deeper, while for marketing and creative material — MiniMax-M3, whose Georgian is the most “alive”.
The main conclusion is simple: pick a model for Georgian not by brand or size, but by real testing. That is exactly why we published the full results and methodology — so that anyone working with Georgian texts has a ready answer.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.