
BTU Publishes Open Georgian-Language Dataset for AI Training
Business and Technology University has published an open Georgian-language resource on GitHub. Prepared with Palitra Media's support, it holds 645,555 structured records — roughly 12 million language-model tokens.
Business and Technology University (BTU) has created an open Georgian-language resource for training and evaluating AI systems and published it on GitHub. Prepared with the support of Palitra Media, the release contains 645,555 structured records — a volume BTU describes as roughly 12 million language-model tokens.
Why Georgian needs its own language data
According to BTU, a language's future in the AI era depends on how well it is represented digitally. When a model knows a language only through scattered texts, it may produce awkward sentences, misread context and struggle with grammar, idioms and professional terminology. Georgian is especially demanding: much of its meaning is carried by verb forms, grammatical cases and context, while word order stays flexible.
The resource is therefore not an accidental collection of texts but a structured system covering twelve areas: word forms, verbs, sentences, cases, idioms, professional terminology, quantities, quotations, concepts, official data, media titles and examples of typical AI mistakes.
What the release contains
Palitra Media supplied broad contemporary Georgian material, which BTU converted into an open format suited to research and technology development. The university says the resource can support Georgian-language chatbots, translation, search, educational tools and digital services, leading to more natural customer support and clearer public services.
Open publication and its limits
Publishing on GitHub gives researchers, universities and developers in Georgia and abroad free access to the resource and raises the chance that Georgian appears in new studies, model-adaptation projects and evaluation tasks. BTU notes that publication alone does not mean commercial systems such as ChatGPT, Gemini or Claude automatically learn from the data: deliberate training, integration and testing are still required.
What comes next
The university frames the project as a matter of digital language sovereignty — the capacity to describe and develop Georgian technologically within global systems. According to BTU researchers, its main value is not the number of tokens but the fact that Georgian now has a shared digital foundation for different technologies. The resource is not a finished national AI model, and the next stage is expert review, testing across models and measurable evidence of improvement.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.