We describe one tool for Table of Content (ToC) identification and recognition from PDF books. This task is part of ongoing research on the development of tools for the semi-automatic conversion of PDF documents in the Epub format that can be read on several E-book devices. Among various sub-tasks, the ToC extraction and recognition is particularly useful for an easy navigation of book contents. The proposed tool first identifies the ToC pages. The bounding boxes of ToC titles in the book body are subsequently found in order to add suitable links in the Epub ToC. The proposed approach is tolerant to discrepancies between the ToC text and the corresponding titles. We evaluated the tool on several open access books edited by University Presses that are partner of the OAPEN EcontentPlus project

Table of contents recognition for converting PDF documents in e-book formats / S. MARINAI; E. MARINO; G. SODA. - STAMPA. - (2010), pp. 73-76. (Intervento presentato al convegno ACM Symposium on Document Engineering 2010 tenutosi a Manchester (UK)) [10.1145/1860559.1860576].

Table of contents recognition for converting PDF documents in e-book formats

MARINAI, SIMONE;MARINO, EMANUELE;SODA, GIOVANNI
2010

Abstract

We describe one tool for Table of Content (ToC) identification and recognition from PDF books. This task is part of ongoing research on the development of tools for the semi-automatic conversion of PDF documents in the Epub format that can be read on several E-book devices. Among various sub-tasks, the ToC extraction and recognition is particularly useful for an easy navigation of book contents. The proposed tool first identifies the ToC pages. The bounding boxes of ToC titles in the book body are subsequently found in order to add suitable links in the Epub ToC. The proposed approach is tolerant to discrepancies between the ToC text and the corresponding titles. We evaluated the tool on several open access books edited by University Presses that are partner of the OAPEN EcontentPlus project
2010
Proceedings of the 10th ACM symposium on Document engineering
ACM Symposium on Document Engineering 2010
Manchester (UK)
S. MARINAI; E. MARINO; G. SODA
File in questo prodotto:
File Dimensione Formato  
Marinai-DocEng10.pdf

accesso aperto

Descrizione: Paper
Tipologia: Versione finale referata (Postprint, Accepted manuscript)
Licenza: Tutti i diritti riservati
Dimensione 130.44 kB
Formato Adobe PDF
130.44 kB Adobe PDF

I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificatore per citare o creare un link a questa risorsa: https://hdl.handle.net/2158/397150
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 21
  • ???jsp.display-item.citation.isi??? 13
social impact