We describe one tool for Table of Content (ToC) identification and recognition from PDF books. This task is part of ongoing research on the development of tools for the semi-automatic conversion of PDF documents in the Epub format that can be read on several E-book devices. Among various sub-tasks, the ToC extraction and recognition is particularly useful for an easy navigation of book contents. The proposed tool first identifies the ToC pages. The bounding boxes of ToC titles in the book body are subsequently found in order to add suitable links in the Epub ToC. The proposed approach is tolerant to discrepancies between the ToC text and the corresponding titles. We evaluated the tool on several open access books edited by University Presses that are partner of the OAPEN EcontentPlus project
Table of contents recognition for converting PDF documents in e-book formats / S. MARINAI; E. MARINO; G. SODA. - STAMPA. - (2010), pp. 73-76. (Intervento presentato al convegno ACM Symposium on Document Engineering 2010 tenutosi a Manchester (UK)) [10.1145/1860559.1860576].
Table of contents recognition for converting PDF documents in e-book formats
MARINAI, SIMONE;MARINO, EMANUELE;SODA, GIOVANNI
2010
Abstract
We describe one tool for Table of Content (ToC) identification and recognition from PDF books. This task is part of ongoing research on the development of tools for the semi-automatic conversion of PDF documents in the Epub format that can be read on several E-book devices. Among various sub-tasks, the ToC extraction and recognition is particularly useful for an easy navigation of book contents. The proposed tool first identifies the ToC pages. The bounding boxes of ToC titles in the book body are subsequently found in order to add suitable links in the Epub ToC. The proposed approach is tolerant to discrepancies between the ToC text and the corresponding titles. We evaluated the tool on several open access books edited by University Presses that are partner of the OAPEN EcontentPlus projectFile | Dimensione | Formato | |
---|---|---|---|
Marinai-DocEng10.pdf
accesso aperto
Descrizione: Paper
Tipologia:
Versione finale referata (Postprint, Accepted manuscript)
Licenza:
Tutti i diritti riservati
Dimensione
130.44 kB
Formato
Adobe PDF
|
130.44 kB | Adobe PDF |
I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.