In the Italian judiciary system, Public Prosecutors’ Offices still rely on heterogeneous and partially paper-based document workflows, where crime reports and related attachments are often printed, manually signed and annotated, scanned, and finally stored as image-based files. As a consequence, a significant portion of prosecutorial documentation remains only partially machine-readable, limiting the effectiveness of digital case-management systems. In this context, Optical Character Recognition (OCR) and Named Entity Recognition (NER) are enabling technologies that transform unstructured, non-searchable judicial documents into computationally usable legal information. This paper analyzes the technical, organizational, and legal challenges associated with OCR-based processing of crime reports, in the Italian and foreign jurisdictions, and identifies the main methodological requirements for governance-aware NER models used in judicial environments. These include layout-aware document analysis, legal-domain adaptation, human-in-the-loop validation, and privacy-aware processing mechanisms that support pseudonymization and controlled access to sensitive data. Finally, the paper discusses the broader topic of computer-aided judicial digitalization, highlighting the need for reliable, privacy-aware pipelines capable of processing documents at scale in contemporary criminal justice systems.

Bridging digitalization gaps: OCR and NER for rrime report automation in the Italian criminal justice system

Deidda, Nicola
;
Fumera, Giorgio;Giacinto, Giorgio;
2027-01-01

Abstract

In the Italian judiciary system, Public Prosecutors’ Offices still rely on heterogeneous and partially paper-based document workflows, where crime reports and related attachments are often printed, manually signed and annotated, scanned, and finally stored as image-based files. As a consequence, a significant portion of prosecutorial documentation remains only partially machine-readable, limiting the effectiveness of digital case-management systems. In this context, Optical Character Recognition (OCR) and Named Entity Recognition (NER) are enabling technologies that transform unstructured, non-searchable judicial documents into computationally usable legal information. This paper analyzes the technical, organizational, and legal challenges associated with OCR-based processing of crime reports, in the Italian and foreign jurisdictions, and identifies the main methodological requirements for governance-aware NER models used in judicial environments. These include layout-aware document analysis, legal-domain adaptation, human-in-the-loop validation, and privacy-aware processing mechanisms that support pseudonymization and controlled access to sensitive data. Finally, the paper discusses the broader topic of computer-aided judicial digitalization, highlighting the need for reliable, privacy-aware pipelines capable of processing documents at scale in contemporary criminal justice systems.
2027
978-3-032-35576-8
978-3-032-35575-1
Optical Character Recognition; Named Entity Recognition; Justice System; Privacy-Preserving Data Processing; Crime Reports Classification
File in questo prodotto:
File Dimensione Formato  
_ARES26__Iris.pdf

embargo fino al 18/08/2027

Descrizione: AAM
Tipologia: versione post-print (AAM)
Dimensione 1.97 MB
Formato Adobe PDF
1.97 MB Adobe PDF   Visualizza/Apri   Richiedi una copia

I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11584/491385
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact