Publication

Two complementary techniques for digitized document analysis

Jan 1, 1988 · 5 authors · 3 topics

Abstract

Two complementary methods are proposed for characterizing the spatial structure of digitized technical documents and labelling various logical components without using optical character recognition. The top-down method segments and labels the page image simultaneously using publication-specific information in the form of a page-grammar. The bottom-up method naively segments the document into rectangles that contain individual connected components, combines blocks using knowledge about generic layout objects, and identifies logical objects using publication-specific knowledge. Both methods are based on the X-Y tree representation of a page image. The procedures are demonstrated on scanned and synthesized bit-maps of the title pages of technical articles.

Showing the abstract — retrieve the full paper via the Exa API.

Authors

George NagyJunichi KanaiMukkai S. KrishnamoorthyMathews ThomasMahesh Viswanathan

Topics

Handwritten Text Recognition TechniquesAlgorithms and Data CompressionImage Retrieval and Classification Techniques

About

PublishedJan 1, 1988
Citations44
References22

Powered by the Exa API