Machine learning models for Document Layout Analysis (DLA) rely on large-scale datasets such as DocBank and many others. However, automatic generation of these datasets from PDF or LaTeX sources often introduces errors in region annotations, including missing captions, misplaced regions and overlapping elements. Such inconsistencies negatively impact both training quality and evaluation reliability, propagating noise or bias through downstream models. In this work we present a preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents. The proposed system builds a relational model of article structure, integrating geometric (e.g. near, below, lower page), sequential (e.g. follows, precedes), and semantic (e.g. disjoint, subclass) relations among layout elements. These relations can be learned directly from document data, enabling automatic detection and correction of violations, including semantic and hierarchical inconsistencies. Error typologies that can be detected include those derived from class reading order (e.g. title -> author -> abstract) and proximity, which are found to be the most frequent. The detection of such violations highlights missing or misplaced components and guide automated correction that we perform through LLMs. The resulting datasets provide a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models, with reduced noise and bias inherent in automatically generated annotations.
Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs / Lorenzo Massai, S.M.. - ELETTRONICO. - (2026), pp. 0-0. (Information and Research science Connecting to Digital and Library science Modena February 19-20, 2026).
Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs
Lorenzo Massai
;Simone Marinai
2026
Abstract
Machine learning models for Document Layout Analysis (DLA) rely on large-scale datasets such as DocBank and many others. However, automatic generation of these datasets from PDF or LaTeX sources often introduces errors in region annotations, including missing captions, misplaced regions and overlapping elements. Such inconsistencies negatively impact both training quality and evaluation reliability, propagating noise or bias through downstream models. In this work we present a preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents. The proposed system builds a relational model of article structure, integrating geometric (e.g. near, below, lower page), sequential (e.g. follows, precedes), and semantic (e.g. disjoint, subclass) relations among layout elements. These relations can be learned directly from document data, enabling automatic detection and correction of violations, including semantic and hierarchical inconsistencies. Error typologies that can be detected include those derived from class reading order (e.g. title -> author -> abstract) and proximity, which are found to be the most frequent. The detection of such violations highlights missing or misplaced components and guide automated correction that we perform through LLMs. The resulting datasets provide a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models, with reduced noise and bias inherent in automatically generated annotations.I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



