Machine learning models for Document Layout Analysis (DLA) rely on large-scale datasets such as DocBank and many others. However, automatic generation of these datasets from PDF or LaTeX sources often introduces errors in region annotations, including missing captions, misplaced regions and overlapping elements. Such inconsistencies negatively impact both training quality and evaluation reliability, propagating noise or bias through downstream models. In this work we present a preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents. The proposed system builds a relational model of article structure, integrating geometric (e.g. near, below, lower page), sequential (e.g. follows, precedes), and semantic (e.g. disjoint, subclass) relations among layout elements. These relations can be learned directly from document data, enabling automatic detection and correction of violations, including semantic and hierarchical inconsistencies. Error typologies that can be detected include those derived from class reading order (e.g. title -> author -> abstract) and proximity, which are found to be the most frequent. The detection of such violations highlights missing or misplaced components and guide automated correction that we perform through LLMs. The resulting datasets provide a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models, with reduced noise and bias inherent in automatically generated annotations.

Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs / Lorenzo Massai, S.M.. - ELETTRONICO. - (2026), pp. 0-0. (Information and Research science Connecting to Digital and Library science Modena February 19-20, 2026).

Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs

Lorenzo Massai
;
Simone Marinai
2026

Abstract

Machine learning models for Document Layout Analysis (DLA) rely on large-scale datasets such as DocBank and many others. However, automatic generation of these datasets from PDF or LaTeX sources often introduces errors in region annotations, including missing captions, misplaced regions and overlapping elements. Such inconsistencies negatively impact both training quality and evaluation reliability, propagating noise or bias through downstream models. In this work we present a preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents. The proposed system builds a relational model of article structure, integrating geometric (e.g. near, below, lower page), sequential (e.g. follows, precedes), and semantic (e.g. disjoint, subclass) relations among layout elements. These relations can be learned directly from document data, enabling automatic detection and correction of violations, including semantic and hierarchical inconsistencies. Error typologies that can be detected include those derived from class reading order (e.g. title -> author -> abstract) and proximity, which are found to be the most frequent. The detection of such violations highlights missing or misplaced components and guide automated correction that we perform through LLMs. The resulting datasets provide a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models, with reduced noise and bias inherent in automatically generated annotations.
2026
Proceedings of IRCDL 2026
Information and Research science Connecting to Digital and Library science
Modena
February 19-20, 2026
Lorenzo Massai,Simone Marinai
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificatore per citare o creare un link a questa risorsa: https://hdl.handle.net/2158/1484172
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact