← Back
SiTech Team⏱️ 7 წთ. საკითხავი

AI System Helped Pakistani Judges Clear Backlogs at $38.50 Return Per Dollar Invested

AI System Helped Pakistani Judges Clear Backlogs at $38.50 Return Per Dollar Invested

An ETH Zurich and Imperial College study showed that a GPT-4 based AI system helped Pakistani judges achieve a 6.3% increase in case resolution.

The Crisis in Pakistan's Judicial System

Pakistan's justice system is one of the most overburdened in the world. By the end of 2024, 2.26 million cases were pending, 82% of which were stuck in trial courts. The country has fewer than 2 judges per 100,000 residents — compared to 22 in the EU and 30 in England. This crisis means millions of people wait years for justice, creating not just a social problem but an economic one as well. Unresolved cases stifle business activity, deter investment, and erode public trust in state institutions. The backlog represents an enormous drag on Pakistan's economic potential.

Building JudgeGPT: An AI System for the Courts

Researchers from ETH Zurich and Imperial College London built "JudgeGPT" on OpenAI's GPT-4, using Retrieval-Augmented Generation (RAG) technology. RAG fundamentally changes how AI works in specialized domains — rather than relying solely on training data, the model actively retrieves relevant information from a trusted database in real time. The system was grounded in 129,235 documents, including 128,292 court rulings and 943 Pakistani laws. This meant a judge could input the specifics of a case and receive an answer based on actual, current legislation and precedents — not just statistical patterns from the internet. The database was carefully curated and validated by legal experts to ensure accuracy and relevance.

The Largest Randomized Trial of AI in Government

This was the largest experimental evidence ever collected on AI boosting government productivity. 1,559 judges across 118 courts were divided into three groups: (1) JudgeGPT with targeted training, (2) JudgeGPT with a general seminar, and (3) a control group with no AI access at all. This gold-standard randomized controlled trial (RCT) ran for 40 weeks. Researchers tracked not just how often judges used the AI, but also the quality of rulings, appeal rates, working hours, and potential bias indicators. The study's methodology was rigorous enough to establish causation, not just correlation — a rare achievement in AI impact research.

Training: The Decisive Factor

The most striking finding was that AI access alone (with only a general seminar) did almost nothing. This result is particularly telling given that only about 25% of participants had ever used an LLM like ChatGPT — suggesting unfamiliarity was a major barrier. Judges who received targeted training used the AI 4x more: after 40 weeks, they logged ~60 sessions with 200+ prompts, compared to ~20 sessions and fewer than 50 prompts in the general seminar group. Targeted training included hands-on exercises, simulated case reviews, and personalized feedback loops that helped judges understand both the capabilities and limitations of the AI system. This demonstrates that technology deployment without proper training is not just suboptimal — it is essentially ineffective.

Productivity Gains That Matter

The results were significant: 1,848 extra cases resolved per year per district, representing a 6.3% bump. While this percentage may seem modest, in absolute terms it translates to thousands of additional resolved cases across the entire system. The appeal rate actually fell slightly per 1,000 resolved cases, indicating that AI-assisted decisions were not lower quality — if anything, they were better. When independent legal experts evaluated judge rulings in blind pairwise comparisons, AI-assisted rulings were rated better 59% of the time versus 42% in the control group. This quality improvement is a crucial counterargument to concerns that AI might sacrifice quality for speed.

Ethics, Bias, and Work-Life Balance

Importantly, working hours and work-life balance did not change — AI made work more efficient, not longer. This is a critical finding for overburdened judicial systems worldwide: AI enables judges to accomplish more within the same time frame rather than simply accelerating burnout. The study found no evidence of increased gender or religious bias in AI-assisted rulings. This suggests that, when properly designed and monitored, AI can actually reduce human bias rather than amplify it. The researchers specifically tested for bias across multiple dimensions and found the AI system to be at least as fair as human judges acting alone.

Unprecedented Economic Returns

The ROI of $38.50 per dollar invested is unprecedented for a government technology initiative. Even the conservative estimate puts it above $10 per dollar. The researchers note that these results were achieved using a pre-reasoning version of GPT-4, meaning today's reasoning models like GPT-4o, Claude 4, or Gemini 2.5 would likely yield even higher returns. The cost calculations included API usage, training expenses, infrastructure, and judge time — making it a comprehensive assessment of true economic impact. For resource-constrained judiciaries, this ROI figure makes a compelling case for pilot programs.

Challenges and Risks

Despite its success, the study surfaced several important challenges. Infrastructure: many courts had limited internet access, restricting AI usage. Training costs: targeted training requires time and resources, though the ROI clearly justifies the investment. Model bias: GPT-4 is predominantly trained on Western data, potentially creating friction with Pakistan's Islamic legal framework. The researchers monitored this carefully and found no significant issues, but it remains a consideration for any other jurisdiction adopting similar technology. Data quality and maintenance of the legal database was another ongoing challenge — laws and precedents change, and the system needs continuous updates to remain reliable.

Lessons for Georgia and the South Caucasus

Georgia's judicial system faces similar challenges: case backlogs, limited human resources, and slow processes. While the scale differs from Pakistan's, the core lessons are universal.

1. AI without training fails. The JudgeGPT study clearly showed that technology access without knowledge is useless. Georgia would need a structured training program for judges as part of any AI deployment — not just a one-day seminar, but ongoing practical workshops.

2. Legal data infrastructure is foundational. A RAG-based system needs high-quality, digitized legal databases. Georgia's Legislative Herald (matsne.gov.ge) already contains digitized legislation — a strong foundation for a similar initiative. Court decisions are also available digitally, though their structuring and systematization could be improved for AI consumption.

3. ROI potential is transformative. $38.50 per dollar invested suggests that AI in justice is economically highly beneficial. For Georgia, where public budgets are constrained, high-ROI investments are especially valuable. Even a $100,000 pilot project could yield millions of lari in economic benefits through faster case resolution.

4. Transparency and human oversight. Ensuring AI decisions are not a "black box" is critical. Involving judges in editing and finalizing AI-generated content preserves human control and accountability — particularly important in countries where trust in judicial institutions needs strengthening.

The High Council of Justice of Georgia and ongoing judicial reform programs could leverage these findings to design AI pilot projects in Georgian courts. This is particularly relevant as the government pursues its broader digital transformation agenda. Armenia and Azerbaijan face similar judicial challenges — the entire South Caucasus region could benefit from the lessons of Pakistan's experiment, perhaps through a coordinated regional approach to legal AI.

Conclusion

The JudgeGPT study represents a watershed moment for AI in government. It provides the first large-scale experimental evidence that AI can meaningfully boost government productivity and improve service quality for citizens. The key ingredients — training, infrastructure, and data quality — are universal. If Pakistan, with its 2.26 million pending cases, could make significant progress with AI at a $38.50 return per dollar invested, why shouldn't other countries — including Georgia — give it a try? The future of justice is already here; it is up to us to decide how to use it wisely.