In the Italian judiciary system, Public Prosecutors’ Offices still rely on heterogeneous and partially paper-based document workflows, where crime reports and related attachments are often printed, manually signed and annotated, scanned, and finally stored as image-based files. As a consequence, a significant portion of prosecutorial documentation remains only partially machine-readable, limiting the effectiveness of digital case-management systems. In this context, Optical Character Recognition (OCR) and Named Entity Recognition (NER) are enabling technologies that transform unstructured, non-searchable judicial documents into computationally usable legal information. This paper analyzes the technical, organizational, and legal challenges associated with OCR-based processing of crime reports, in the Italian and foreign jurisdictions, and identifies the main methodological requirements for governance-aware NER models used in judicial environments. These include layout-aware document analysis, legal-domain adaptation, human-in-the-loop validation, and privacy-aware processing mechanisms that support pseudonymization and controlled access to sensitive data. Finally, the paper discusses the broader topic of computer-aided judicial digitalization, highlighting the need for reliable, privacy-aware pipelines capable of processing documents at scale in contemporary criminal justice systems.
Bridging digitalization gaps: OCR and NER for rrime report automation in the Italian criminal justice system
Deidda, Nicola
;Fumera, Giorgio;Giacinto, Giorgio;
2027-01-01
Abstract
In the Italian judiciary system, Public Prosecutors’ Offices still rely on heterogeneous and partially paper-based document workflows, where crime reports and related attachments are often printed, manually signed and annotated, scanned, and finally stored as image-based files. As a consequence, a significant portion of prosecutorial documentation remains only partially machine-readable, limiting the effectiveness of digital case-management systems. In this context, Optical Character Recognition (OCR) and Named Entity Recognition (NER) are enabling technologies that transform unstructured, non-searchable judicial documents into computationally usable legal information. This paper analyzes the technical, organizational, and legal challenges associated with OCR-based processing of crime reports, in the Italian and foreign jurisdictions, and identifies the main methodological requirements for governance-aware NER models used in judicial environments. These include layout-aware document analysis, legal-domain adaptation, human-in-the-loop validation, and privacy-aware processing mechanisms that support pseudonymization and controlled access to sensitive data. Finally, the paper discusses the broader topic of computer-aided judicial digitalization, highlighting the need for reliable, privacy-aware pipelines capable of processing documents at scale in contemporary criminal justice systems.| File | Dimensione | Formato | |
|---|---|---|---|
|
_ARES26__Iris.pdf
embargo fino al 18/08/2027
Descrizione: AAM
Tipologia:
versione post-print (AAM)
Dimensione
1.97 MB
Formato
Adobe PDF
|
1.97 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.



