Pakistan’s Courts Have an AI Guideline. What They Don’t Have Is a Liability Chain

Staff Report
14 Min Read

Summary

  • But a guideline that describes AI as safe without saying who is liable when it isn’t, and that says nothing about where sensitive judicial data physically sits once it leaves a courtroom, is not yet a framework a litigant can rely on.
  • Guidelines should tie any AI rollout to mandatory training on what the tool is reliable for, such as drafting and summarising, and what it is not, such as legal reasoning or case outcomes.
  • Conclusion Pakistan’s guidelines get the premise right: AI belongs in the courts as an assistive tool, not a decision-maker, and the backlog is severe enough that refusing the technology outright would be its own failure.
AI Generated Summary

By Muhammad Khan Dummar, Law Student, SAHSOL-LUMS

In March 2023, an Additional District and Sessions Judge in Phalia, a small court in Mandi Bahauddin district, asked ChatGPT whether a juvenile accused in a criminal case could get post-arrest bail. Judge Muhammad Amir Munir wrote in his order that the decision itself did not rest on what the chatbot told him. He used it to test the tool, not to decide the case. It was the first known instance of a Pakistani judge experimenting with AI on the bench, and it happened three years before the country had any rule telling him whether that was allowed.

The rule exists now. On 29 April 2026, the National Judicial Policy Making Committee issued National Guidelines for the Use of Artificial Intelligence in Judicial Institutions, describing AI as an assistive tool that supports judges without replacing them. Weeks earlier, the Supreme Court had already pointed in the same direction, though not as a central holding. In Ishfaq Ahmed v. Mushtaq Ahmed (2025 SCP 112), a tenancy dispute between two brothers that had dragged on for seven years, Mr. Justice Syed Mansoor Ali Shah used the delay itself as the occasion to recommend AI adoption, specifically for case allocation and bench assignment, as a way to curb arbitrary discretion in how cases get distributed among judges, while warning against any use that would compromise human judicial autonomy. A case analysis published by LUMS’s Shaikh Ahmad Hassan School of Law reads the ruling the same way: AI is welcome in the administrative machinery of a court, but the judgment stops well short of endorsing it anywhere near reasoning or decision-making. That is exactly the line the guidelines gesture at without yet enforcing. Both moves respond to a real crisis. The Law and Justice Commission of Pakistan recorded 2,270,584 pending cases nationwide as of 30 June 2025, and Justice Syed Mansoor Ali Shah has noted that even the Supreme Court alone carries close to 56,000, despite the bench growing to 24 judges. A quarter of sanctioned judicial posts across the country sit vacant. Delay in Pakistan is not incidental. It is closer to the operating condition of the system.

Whether AI actually helps is no longer a guess. A randomized field trial run with Pakistan’s Federal Judicial Academy, using a purpose-built assistant called JudgeGPT across roughly half the country’s trial judges and eighty percent of district courts, found that AI access paired with targeted training raised annual case resolution by about 1,848 cases, a 6.3 percent increase, with no measurable drop in writing quality. The finding that matters most for policy is easy to miss in that headline number: access alone did not produce the gain. Judges who got the tool without training used it less, and less well, drifting toward open-ended legal questions that are expensive to verify rather than the narrower tasks, like drafting and summarising, where a language model is actually reliable. Training decided whether the tool helped or sat unused. That is precisely the argument this piece is going to make about the guidelines themselves: the document exists, but whether it does anything depends on details it has not yet worked out.

Read against that backdrop, the NJPMC guidelines get the big call right and leave two smaller ones dangerously open.

The first gap is accountability. The guidelines promise explainability and safeguards against bias, but they do not say, in operational terms, who answers for an AI-assisted error: the judge, the court, or the vendor supplying the model. India’s own draft AI regulations for courts, released by the Supreme Court’s AI Committee on 3 June 2026, go a step further on exactly this point, and name two separate mechanisms. Regulation 52 gives any party harmed by a prohibited use of AI the right to file an application with the court where the system was used and be heard before it passes orders. Regulation 46 goes further still, requiring every contract between a court and a private AI vendor to include mandatory indemnity clauses protecting the court from harm caused by defects in the vendor’s system, and a clear contractual allocation of liability between the court and the vendor when something goes wrong. Pakistan’s guidelines describe a value and expect adherence to it in principle. India’s draft names an actual door a litigant can knock on, and a separate chain of who pays when a vendor’s system is at fault. That difference is not cosmetic. It is the difference between a document with real repercussions when things go wrong and one that reads well in a press release.

This is not a hypothetical concern. U.S. courts have spent the past year working through exactly this failure mode. In Connecticut, a federal judge warned a solo practitioner that an eye-catching sanction might be needed to stop the pattern after his brief was found to contain three fabricated citations. In Kansas, five attorneys on a single patent brief were all found to have violated their duty to verify citations, even though only one of them had actually used the AI tool that invented the case law. The lesson is blunt: once the veracity of a citation goes unchecked and a fabricated one reaches a filing, courts are holding everyone who signed it responsible, not just whoever typed the prompt. Pakistan’s guidelines do not yet say whether that principle applies here, or how.

The second gap is where the data actually lives. Courts do not process ordinary information. A case file can contain medical records, financial disclosures, identity documents, and testimony from minors and other vulnerable witnesses. If judges or court staff run any of that through a public AI tool built by a foreign company, that material may leave Pakistan’s jurisdiction entirely, processed on servers governed by a different country’s laws. OpenAI does not currently list Pakistan among the regions where it offers data residency, and its own documentation notes that certain processing logs can be retained for up to thirty days regardless of the customer’s settings. None of this makes foreign AI tools unusable. It does mean that using them on sensitive judicial material without a residency and retention agreement in place is a real exposure, not a theoretical one, and the April guidelines do not address it at all.

There is a deeper reason both gaps matter, beyond drafting. Legal scholarship on “human in the loop” systems, most recently a 2026 article by Elisabeth Paar in the German Law Journal, has started challenging the assumption that keeping a judge nominally in charge of an AI-assisted decision is automatically safer than full delegation. Her argument is that oversight tends to degrade on its own: humans monitoring an automated system for long enough become what she calls rubber stampers, losing the vigilance and independent judgment the oversight was supposed to preserve in the first place. Pakistan’s guidelines lean entirely on the reassurance that a judge remains the ultimate decision-maker. That reassurance is doing more work than it can actually bear unless it is backed by something concrete, training requirements, verification duties, and a named party who answers when the system gets it wrong.

None of this is an argument against using AI in Pakistan’s courts. The JudgeGPT trial shows real gains are available, and the backlog is severe enough that leaving those gains on the table would be its own kind of failure. But a guideline that describes AI as safe without saying who is liable when it isn’t, and that says nothing about where sensitive judicial data physically sits once it leaves a courtroom, is not yet a framework a litigant can rely on. Before the next phase of rollout, the NJPMC should amend its guidelines to do two specific things: name who is accountable, in the same operational detail India’s draft regulations already attempt, and require that any AI system touching case-sensitive material meet an explicit data residency and retention standard.

Judge Munir’s 2023 experiment worked precisely because he treated the chatbot’s answer as something to test against his own judgment, not something to rely on. Three years and one national guideline later, Pakistan still has not written down, in a way anyone could actually enforce, what happens the day a judge trusts the wrong answer instead.

Recommendations

The NJPMC does not need to start over. It needs to close two specific holes before the next phase of rollout, and address two more that follow directly from them.

  1. Name a liability chain, not just a value. Amend the guidelines to specify who answers when an AI-assisted error reaches a filing or order: the judge who used the tool, the court that deployed it, or the vendor whose system produced the error. India’s draft Regulations 52 and 46 are a usable template: a filing mechanism for harmed parties, and mandatory indemnity language in every vendor contract.
  2. Set a data residency and retention standard before, not after, adoption. Any AI system touching case files with medical records, financial disclosures, or testimony from minors should be certified against an explicit residency and retention rule: where the data is processed, how long logs are kept, and under whose jurisdiction, before it is approved for judicial use.
  3. Make training a condition of access, not an add-on. The JudgeGPT trial found that untrained access made outcomes worse, not neutral: judges drifted toward exactly the open-ended, hard-to-verify uses that are riskiest. Guidelines should tie any AI rollout to mandatory training on what the tool is reliable for, such as drafting and summarising, and what it is not, such as legal reasoning or case outcomes.
  4. Build in a verification duty, not just a supervision principle. Given Paar’s “rubber stamper” concern, nominal human oversight is not self-sustaining. The guidelines should require affirmative, documented verification steps, akin to the citation-checking duty now being enforced in Connecticut, rather than assuming a judge’s presence in the loop is protection enough.

Conclusion

Pakistan’s guidelines get the premise right: AI belongs in the courts as an assistive tool, not a decision-maker, and the backlog is severe enough that refusing the technology outright would be its own failure. But a document that describes AI as safe without saying who is liable when it isn’t, and that says nothing about where sensitive judicial data ends up once it leaves a courtroom, is a statement of intent rather than a framework anyone can enforce. Judge Munir’s 2023 experiment worked because he treated the chatbot as something to test, not trust. The guidelines now need to make that same caution binding for everyone who follows him, with a named party to answer to, and a rule about where the data goes, when someone doesn’t.

We welcome your contributions! Submit your blogs, opinion pieces, press releases, news story pitches, and news features to opinion@minutemirror.com.pk and minutemirrormail@gmail.com
Share This Article
Leave a Comment

Leave a Reply

Your email address will not be published. Required fields are marked *