Document analysis is an interdisciplinary field of research that lies at the intersection of computer vision and natural language processing (NLP), with the aim of extracting, interpreting and processing information contained in structured and unstructured documents. These documents can take many forms, including digital texts, scanned images, historical documents or legal documents, often characterised by a high degree of heterogeneity from both a visual and semantic point of view. Document analysis is therefore a complex problem that requires the integration of visual and linguistic techniques to obtain an accurate and usable representation of the information content. In recent years, this field has undergone profound changes thanks to the advent of Large Language Models (LLMs), capable of understanding and generating natural language with a high level of accuracy, and subse- quently Vision-Language Models (VLMs), which extend these capabilities by integrating visual and textual information within a single model. These advances have led to a paradigm shift from traditional approaches, enabling complex tasks such as information extraction, classification, topic modelling and question answering to be tackled in a more robust and effective man- ner. Document analysis techniques, while having different objectives, are highly interconnected and are frequently used in conjunction with each other within complex pipelines aimed at transforming raw documents into struc- tured knowledge. The development of such pipelines is particularly important in highly complex application contexts, such as medical-legal ones. In the health- care sector, hospitals have to manage large volumes of documentation, espe- cially in cases of alleged medical malpractice, which involve the production and analysis of numerous clinical and legal documents. Similarly, the legal domain is characterised by a considerable amount of documents, such as compensation letters, party-appointed technical consultants (CTP), court- appointed technical consultants (CTU), defence briefs, expert reports and judgements. Manual analysis of these documents is costly, prone to errors and difficult to scale, making it necessary to adopt automatic support sys- tems. Equipping hospitals and relevant institutions with advanced document analysis systems not only supports the management and resolution of med- ical malpractice cases, but also creates structured digital archives of cases handled. These archives are a valuable resource, as they allow data to be prepared and organised for subsequent quantitative and qualitative analy- sis, facilitating the application of data mining and text mining techniques. This approach also opens up the possibility of identifying recurring pat- terns, supporting epidemiological and legal studies, and improving data- driven decision-making processes. Another critical aspect of document analysis concerns the quality of the documents themselves. Real-world documents are often affected by imper- fections such as blurring, noise, scanning artefacts or degradation of the original medium, which compromise the performance of optical character recognition (OCR) systems and, more generally, information extraction al- gorithms. For this reason, document restoration is a fundamental step in document analysis pipelines, as it improves the visual quality of documents and, consequently, the reliability of the information extracted. In this context, this thesis aims to address the problem of document analysis in the medical-legal field through an integrated approach. In par- ticular, document restoration algorithms are developed to reduce the effects of blurring and noise; a complete document analysis pipeline applied to the forensic domain is designed and implemented; a dataset for topic modelling based on civil and criminal judgments of the Italian supreme court is also constructed, with the aim of analysing recurring themes. Finally, the thesis presents the development of a multimodal graph-based information retrieval system capable of retrieving relevant information from a ruling and auto- matically responding to queries of interest.

Artificial Intelligence methods and techniques for text understanding and risk estimation, predictive models for healthcare claims management / Matteo Marulli, M.B.. - (2026).

Artificial Intelligence methods and techniques for text understanding and risk estimation, predictive models for healthcare claims management

Matteo Marulli
;
Vilma Pinchi
2026

Abstract

Document analysis is an interdisciplinary field of research that lies at the intersection of computer vision and natural language processing (NLP), with the aim of extracting, interpreting and processing information contained in structured and unstructured documents. These documents can take many forms, including digital texts, scanned images, historical documents or legal documents, often characterised by a high degree of heterogeneity from both a visual and semantic point of view. Document analysis is therefore a complex problem that requires the integration of visual and linguistic techniques to obtain an accurate and usable representation of the information content. In recent years, this field has undergone profound changes thanks to the advent of Large Language Models (LLMs), capable of understanding and generating natural language with a high level of accuracy, and subse- quently Vision-Language Models (VLMs), which extend these capabilities by integrating visual and textual information within a single model. These advances have led to a paradigm shift from traditional approaches, enabling complex tasks such as information extraction, classification, topic modelling and question answering to be tackled in a more robust and effective man- ner. Document analysis techniques, while having different objectives, are highly interconnected and are frequently used in conjunction with each other within complex pipelines aimed at transforming raw documents into struc- tured knowledge. The development of such pipelines is particularly important in highly complex application contexts, such as medical-legal ones. In the health- care sector, hospitals have to manage large volumes of documentation, espe- cially in cases of alleged medical malpractice, which involve the production and analysis of numerous clinical and legal documents. Similarly, the legal domain is characterised by a considerable amount of documents, such as compensation letters, party-appointed technical consultants (CTP), court- appointed technical consultants (CTU), defence briefs, expert reports and judgements. Manual analysis of these documents is costly, prone to errors and difficult to scale, making it necessary to adopt automatic support sys- tems. Equipping hospitals and relevant institutions with advanced document analysis systems not only supports the management and resolution of med- ical malpractice cases, but also creates structured digital archives of cases handled. These archives are a valuable resource, as they allow data to be prepared and organised for subsequent quantitative and qualitative analy- sis, facilitating the application of data mining and text mining techniques. This approach also opens up the possibility of identifying recurring pat- terns, supporting epidemiological and legal studies, and improving data- driven decision-making processes. Another critical aspect of document analysis concerns the quality of the documents themselves. Real-world documents are often affected by imper- fections such as blurring, noise, scanning artefacts or degradation of the original medium, which compromise the performance of optical character recognition (OCR) systems and, more generally, information extraction al- gorithms. For this reason, document restoration is a fundamental step in document analysis pipelines, as it improves the visual quality of documents and, consequently, the reliability of the information extracted. In this context, this thesis aims to address the problem of document analysis in the medical-legal field through an integrated approach. In par- ticular, document restoration algorithms are developed to reduce the effects of blurring and noise; a complete document analysis pipeline applied to the forensic domain is designed and implemented; a dataset for topic modelling based on civil and criminal judgments of the Italian supreme court is also constructed, with the aim of analysing recurring themes. Finally, the thesis presents the development of a multimodal graph-based information retrieval system capable of retrieving relevant information from a ruling and auto- matically responding to queries of interest.
2026
Marco Bertini, Vilma Pinchi
ITALIA
Matteo Marulli, Marco Bertini, Vilma Pinchi
File in questo prodotto:
File Dimensione Formato  
tesi_dottorato_mmarulli_2026.pdf

accesso aperto

Descrizione: Tesi dottorato ciclo 38 matteo marulli
Tipologia: Tesi di dottorato
Licenza: Creative commons
Dimensione 16.56 MB
Formato Adobe PDF
16.56 MB Adobe PDF

I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificatore per citare o creare un link a questa risorsa: https://hdl.handle.net/2158/1488953
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact