The vast potential and multifaceted uses of AI make it almost impossible for international courts and tribunals to disregard and remain isolated from technological developments. Artificial intelligence can aid in summarising and organising jurisprudence and generate ‘draft paragraphs which could serve as a starting point’ for the Court’s decisions and memos and provide courts and tribunals with information about the evidence gathering processes and/or provide summaries of testimonies, translate or transcribe interviews, and extract key facts and details from extensive documentary evidence. More importantly, AI addresses the longstanding issues of human resources and costs, significantly aiding the work of international justice. Yet, these technologies are far from perfect and even further from being wholly explained or understood by their developers. Hence, their implementation should be carried out while maintaining the highest possible standards and practices. In this line, this blog post seeks to analyse potential scenarios for the use of AI at the ICC, highlight the risks they would pose, and provide solutions and strategies to minimise them.
AI at the International Criminal Court
At the International Criminal Court, the Office of the Prosecutor (‘OTP’) has been a precursor of its adoption through the development of ‘Project Harmony’. In 2023, the OTP announced the launch of ‘OTPLink’, an advanced evidence submission platform that processes data with AI and machine learning, as well as OTP eDiscovery, a cloud-based eDiscovery software and eVault, an electronic system for the preservation of evidence. Thus, improving the capacity of the OTP to review field-based evidence as well as to analyse, process, and transcribe different audio and video materials.
In relation to its current use by other organs of the Court, as well as more broadly by the OTP – as there is limited publicly available information – we can only speculate on the potential uses of these technologies. Among others, AI could be utilised during fact-finding stages by facilitating facial or speech recognition, within digital evidence such as videos or images, thus, aiding in the identification of victims and perpetrators of international crimes. AI can also serve in the processing, categorisation, and analysis of data, e.g., by helping spot similar modus operandi in crimes, identifying structural similarities across different criminal organisations, and pointing out evidentiary gaps.
Discrimination and bias risks
First, as stressed by McIntyre and Vialle, the use of AI could increase the risk of discrimination against vulnerable or marginalised groups. In this sense, AI may replicate or amplify existing biases. Since AI learns based on the information of data it is provided and how the model is trained, if biased data or training is given to the algorithm, there is a high risk for its output to be intentionally or unintentionally biased. For instance, a system such as the OTP e-Discovery might run the risk of downplaying the likelihood of gender-based crimes against less visible groups, such as men, intersex individuals, or LGBTQ+ persons, when categorising data and analysing it. A similar scenario might occur when utilising AI to transcribe or translate videos or audios or conduct speech analysis, as the model, if trained with native English or western speakers, might have issues in adequately functioning with less represented categories, excluding or misinterpreting common slang or accents.
Similarly, the model or system can suffer from interaction bias or biases arising from its interactions with humans. For instance, as investigators from the OTP use and interact with the model, it might become more prone to finding information that leads to convictions rather than to investigative or exculpatory evidence.
A key strategy to reduce these risks is to ensure that the data used for its training is representative of the population, or in other words, it contains a relevant number of all the attributes and sub-classes of the parameters. Therefore, it is crucial for any system developed to adequately review the datasets and adjust the training of the algorithm multiple times. Likewise, it is essential to guarantee that the system’s development and training are free from intentional and unintentional biases. To do so, experts (from both the technical AI side and international criminal law) must be consulted at all stages of the model’s design, training, implementation, and deployment (this last refers to the monitoring stage). Another potential strategy to guarantee transparency and reliance on the model is to require the potential developers of the system to provide the Court with model cards or a sort of ‘nutritional label’ which would allow the Court to compare the ‘candidate models for deployment across not only traditional evaluation metrics but also along the axes of ethical, inclusive, and fair considerations.’
Gradual or sudden loss of control
A second risk is the gradual or sudden loss of control over the model, which refers to the problem of controlling what an advanced AI does. A gradual loss of control will occur, for example, when AI systems are increasingly given more tasks and responsibilities. Initially, the system might just be used to automate certain administrative activities. However, as time goes on, the Court’s legal officers might come to rely too heavily on the system. In this line, they might be tempted to overlook or miss minor issues in the cataloguing of evidence or other areas, as they benefit from reduced workloads and improved efficiency. For example, a legal officer at the OTP may initially use AI merely to sort, label, and organise large volumes of evidentiary material. However, as the legal officer becomes more confident in the system, they may stop thoroughly checking the classification or whether relevant items have been omitted. Accordingly, what was initially meant to be merely an administrative support tool may progressively shape substantive assessments, not because its designers/users have granted the system decision-making authority, but rather because human actors increasingly defer to its outputs.
On the other hand, a sudden loss of control may occur when an AI system produces an unexpected or unreliable output at a critical stage of the ICC’s work, and the Court is unable to detect, explain, or correct the problem before it affects legal decision-making. A plausible scenario is that a system used by the OTP to sort, label, and prioritise large volumes of evidentiary material suddenly misclassifies relevant evidence, overlooks exculpatory material, mistranslates witness statements, or gives undue importance to unreliable open-source information. For example, an AI tool may group together videos, witness statements, and social media posts as supporting the same incident, even though some of the material comes from different locations or dates. If legal officers rely on this output when developing their case theory, selecting witnesses, or deciding which evidence to review first, the system may influence the direction of the investigation before its error is noticed.
In this situation, the loss of control is not dramatic, but procedural. The ICC may still have human officials formally responsible for the decision, but those officials may no longer have meaningful control over the reasoning process if they cannot understand why certain evidence was prioritised, why other material was excluded, or how the system reached its conclusions. The risk is especially serious in international criminal proceedings because the evidentiary record is often massive, multilingual, fragmented, and politically sensitive. Once an AI-generated classification or summary becomes embedded in the prosecution’s internal analysis, correcting the error may be difficult, particularly if later decisions have already been built on it.
To mitigate these risks, a robust governance framework should be established, combining technical, legal, and institutional safeguards. This includes embedding human-in-the-loop mechanisms for critical decisions, implementing strict oversight and auditability standards, and enforcing legal limits on AI’s scope of authority. Additionally, institutions like the ICC should adopt clear policies defining non-delegable human responsibilities and ensure regular external reviews to assess AI’s role, performance, and compliance with ethical and legal norms such as the EU AI Act. Hence, the OTP should not delegate to AI systems evidence review or classification, and in the case of transcriptions of interviews, it should guarantee a constant review by experts. Similarly, the ICC should not delegate tasks such as drafting decisions or reviewing confidential documents to the system. This is particularly important if one considers the OTP’s desire to become a technological hub and to fully utilise new technologies.
Towards responsible AI at the ICC
Overall, AI’s novel and varied applications can significantly contribute to the work and enforcement of international criminal justice. At the outset, the ICC should ensure its compliance with existing regulatory frameworks – specifically the EU AI Act – thereby guaranteeing risk classification, transparency, accountability, and human oversight in any AI use. Nevertheless, this is likely insufficient given the heightened risks of international criminal proceedings. In this sense, the Court should develop its own internal policies for the design, application, and improvement of any AI systems within the Court. Although a comprehensive governance model is outside the scope of this blog post, as emphasised in the previous sections, the Court should rely on strict human-in-the-loop review mechanisms, e.g., by introducing verification protocols or auditing mechanisms to test the systems for bias or errors. The Court should also clearly limit which functions will be delegated to AI. Similarly, the Court should invest in adequate training for legal officers; for instance, the ICC could have specialised AI officers within each organ of the Court who receive additional training to understand the capabilities and limitations of these tools. Ultimately, responsible AI at the ICC should not be seen as an obstacle to innovation, but as a condition for ensuring that AI strengthens, rather than undermines, fairness, accountability, and the rule of law.
