Pakistan AI liability civil courts

JudgeGPT's Efficiency Gains in Pakistan Expose a Liability Gap the Courts Haven't Closed

A GPT-4 trial across 1,559 Pakistani judges cut backlogs 6.3%, but courts still have no rule for who answers when AI-assisted reasoning is wrong.

JudgeGPT: Gains and the Accountability Gap People of Internet Research · Pakistan 1,559 Judges in trial Roughly half of Pakistan's trial c… 6.3% Case resolution increase About 1,848 additional cases resol… ~1 in 5 Prompts showing AI delegation Judges asked the model to determin… $38.50 Return per dollar invested Estimated return on the training-p… peopleofinternet.com
JudgeGPT: Gains and the Accountability… People of Internet Research · Pakistan 1,559 Judges in trial 6.3% Case resolution increase ~1 in 5 Prompts showing AI delegation $38.50 Return per dollar invested peopleofinternet.com

Key Takeaways

Pakistan's trial courts are drowning: more than 2.2 million cases were pending as of late 2024, and the country has fewer than two judges per 100,000 residents against roughly 22 in the EU. Against that backdrop, a randomized field experiment covering 1,559 Pakistani trial judges — roughly half the country's trial bench, across 118 courts — tested whether a GPT-4 tool called JudgeGPT, built with retrieval-augmented search over 129,235 documents including 128,292 prior rulings, could help. Judges given the tool plus targeted training resolved 6.3% more cases, an estimated 1,848 extra cases a year, with no detectable drop in ruling quality and no increase in gender or religious bias in judicial language, according to IEEE Spectrum's August 2026 account of the study by researchers at ETH Zurich, the New Economic School, and Imperial College London.

That is a real result, not a pilot press release. But buried in the same study is the number that should worry anyone who cares about how judicial power is actually exercised: roughly one in five prompts sent to JudgeGPT showed what the researchers termed "substantial AI delegation" — judges asking the model to determine the outcome or generate the legal reasoning itself, rather than to summarize or locate precedent. Training reduced this, shifting usage toward editing and citation-checking. But it did not eliminate it, and no institution in Pakistan has yet answered the question that follows: when a delegated ruling turns out to be wrong, who is liable — the judge who signed it, the court that deployed the tool, or nobody?

The case for caution, stated fairly

The strongest version of the accountability worry is not speculative. Pakistan's own Supreme Court raised it first. In Ishfaq Ahmed v. Mushtaq Ahmed (PLD 2025 SC 582), decided April 11, 2025 in an otherwise ordinary landlord-tenant dispute that had dragged on for seven years, Justice Syed Mansoor Ali Shah used an 18-page opinion to warn that AI in courts brings "hallucinations, bias, opacity, and erosion of public trust," and that the right to a fair trial under Article 10-A of Pakistan's constitution must stay anchored in human judicial reasoning, not a model's output — a summary corroborated by Digital Rights Foundation's account of the ruling and by a peer-reviewed analysis of the case in the Sociology & Cultural Research Review, which flags an "accountability gap" that current safeguards don't close.

That concern deserves to be taken seriously on its own terms, independent of JudgeGPT's aggregate numbers. A 6.3% throughput gain measured across a trial population says nothing about the tail: the specific case where a judge, facing a docket of hundreds, accepted a hallucinated citation or a model-generated legal conclusion without the scrutiny the constitution presumes a human brings. Averages can mask exactly the failure mode critics are worried about, and unlike a wrongly denied loan or a mis-ranked search result, a wrongly decided civil judgment can cost someone their home, custody of a child, or years of their life in appeals. Courts are also unlike most institutions that adopt AI: their legitimacy rests specifically on the public's belief that a human being reasoned through the facts and the law, not on throughput.

Where the guidelines actually land

Following the Supreme Court's direction, the National Judicial (Policy Making) Committee — chaired through its automation arm by the same Justice Mazhar-led body — approved National Guidelines for the Use of Generative AI in Judicial Institutions at its 57th meeting on April 29–30, 2026. The guidelines state that AI "will assist — not replace" judicial decision-making and stress "explainability and accountability," per The News's coverage of the NJPMC announcement. That framing is right as far as it goes. But notably, neither the guidelines as reported nor the JudgeGPT study itself specify what happens procedurally when a delegated ruling is later found to be wrong — whether that triggers automatic appellate review, disciplinary referral, or nothing at all beyond the ordinary appeals process. "Accountability" is asserted as a value; it has not yet been operationalized as a rule.

This is the gap regulators should close, and it is a narrow, solvable one — not a reason to restrict the tool itself. The JudgeGPT trial's own design points to the fix: the judges who got targeted training, not just tool access, produced better-rated rulings (59% preferred in pairwise comparison, versus 42% for the untrained control group) and shifted away from delegation-heavy prompts. The researchers estimate a $38.50 return for every dollar spent on the program, largely because training, not restriction, is what changed behavior — a detail confirmed in reporting on the trial's cost-benefit analysis. That is the strongest evidence yet that the answer to judicial AI delegation is procedural, not prohibitive: mandate that any prompt crossing into outcome-determination or full legal-reasoning generation be logged and flagged for review, require judges to certify in writing that AI-drafted reasoning reflects their own independent analysis, and make that certification the point of individual liability if a ruling is later overturned for reliance on unverified AI output.

A blanket restriction on tools like JudgeGPT — the reflexive response some accountability advocates favor — would forfeit a proven backlog reduction in a judiciary that badly needs one, to guard against a failure mode that a disclosure-and-certification rule can address directly. Pakistan's courts do not have the luxury of choosing between speed and accountability; 2.2 million pending cases means the status quo already fails litigants on both counts. The NJPMC guidelines are a genuine first step. What they need next is not more caution in principle, but the specific liability rule — logged delegation, mandatory certification, a clear disciplinary and appellate consequence — that turns "accountability" from an aspiration into an enforceable standard.

Sources & Citations

  1. IEEE Spectrum: JudgeGPT Experiment
  2. Digital Rights Foundation: SC Urges Regulated AI Use
  3. Sociology & Cultural Research Review: Ishfaq Ahmed v. Mushtaq Ahmed analysis
  4. The News: NJPMC AI Guidelines for Judicial Institutions
  5. The Decoder: Pakistani Judiciary AI Cost-Benefit