Abstract
Lottg alld eolni)licated seltteltces prov(: to b(: a. stumbling block for current systems relying on N[, input. These systenls stand to gaill frolil ntethods that syntacti<:aHy simplily su<:h sentences. ']b simplify a sen= tence, we nee<t an idea of tit(." structure of the sentence, to identify the <:omponents to be separated out. Obviously a parser couhl be used to obtain the complete structure of the sentence. ][owever, hill parsing is slow a+nd i)rone to fa.ilure, especially on <:omph!x sentences. In this l)aper, we consider two alternatives to fu]l parsing which could be use<l for simplification. The tirst al)l)roach uses a Finite State Grammar (FSG) to pro-dn<:e noun and verb groups while the second uses a Superta.gging model to i)roduce dependency linkages. We discuss the impact of these two input representations on the simplification pro(:ess. Reasons for Text Simplification l,ong and <:oml)licatcd sentences prove to be a stumlJing block for <'urrent systems which rely on natural language input. 'l'lmsc systems stand to gain from metho<ls that preprocess such sentences so as to make them simpler. Consider, for examph;, the following sentence: (l) 7'he embattled Major government survived a crucial 'vole on coal pits closure as its last-minute concessions curbed the extent of 'lbry revolt over an issue that generated u'ausual heat in the l]ousc of Commons and brought the miners to London streets. Such sentences are not uncommon in newswire texts. (]ompare this with the multi-sentence version which has been manually simplified: (2) The embatlled Major governmcnl survived a crucial vote o'u coal pits closure. Its last:minute conccssious curbed the cxlenl o]" *On leave fl'om the National Centre for Soft, ware Techno]ogy, (lulmohar (?ross Road No. 9, Juhu, Bombay 4:0(/ (149, India Tory revolt over the coal-miue issue. Th.is issue generaled unusual heat in the ltousc of Commons. II also brought the miners to London streels. If coml>lex text can be made simph'x, senten(-es beconae easier to process, both for In:Ograms and humans. Wc discuss a simplification process which identifies components of a sentence that may be separated out, and transforms each of these into frec-sta,ding simpler sentences. (]learly, some mmnees of meaning from the original text may be lost in the simplification process. Simplitication is theretbre inappropriate for texts (such as legal docunlents) where it is importa.nt not to lose any nuance. I|owew;r, one c.~tl] COilceive of several areas of natural language processing where such simplitication would be of great use. This is especially true in dolnains such as Inachine translation, which commonly have a manual post-processing stage, where semantic and pragmatic repairs may be <'arried out if ne<;essary. • Parsing: Syntactically <:omplex sentence's arc likely to generate a large number of parses, and may cause parsers to fail altogether. Resolving ambiguities in attachment of constituents is non-trivial. This ambiguii, y is reduced for simpler sentences sin<'e they involve fewer constituents. 'Fhus simpler sentences lead to faster parsing and less parse aml)iguity. Once the i>arses for the simpler sentences are obtained, the subparses can be assembled to form a full parse, or left as is, depending on the application. • Machine Translation (MT): As in the parsing case, simplification results in simpler scntential structures and reduced ambiguity. As argued in (Chandrasekar, 1994) , this conld lead to improvements in the quality of machine translation. • Information Retrieval: IR systems usually retrieve large segments of texts of which only a part n]ay bc reh~'wml,. Wit|, simplified texts, it is possible to extract Sl>eCific phrases or simple sentences of relevance in response to queries. • Summarization: With the overload of information that people face today, it would be very helpful to have text summarization tools that; reduce large bodies of text to the salient minimum. Simplification can be used to weed out irrelevant text with greater precision, and thus aid in summarization. • Clarity of Text: Assembly/use/maintenance manuals must be clear and simple to follow. Aircraft companies use a Simplified English for maintenance manuals precisely for this reason (Wojcik et M., 1993) . However, it is not easy to create text in such an artificially constrained language. Automatic (or semi-automatic ) simplification could be used to ensure that texts adhere to standards. We view simplification as a two stage process. The first stage provides a structural representation for a sentence on which the second stage applies a sequence of rules to identify and extract the components that can be simplified. One could use a parser to obtain the complete structure of the sentence. If all the constituents of the sentence along with the dependency relations are given, simplification is straightforward, ttowever, full parsing is slow and prone to failure, especially on complex sentences. To overcome the limitations of full parsers, researchers have adopted FSG based approaches to parsing (Abney, 1994; Hobbs et al., 1992; Grishman, 1995) . These parsers are fast and reasonably robust; they produce sequences of noun and verb groups without any hierarchical structure. Section 3 discusses an FSG based approach to simplification. An alternative approach which is both fast and yields hierarchical structure is discussed in Section 4. In Section 5 we compare the two approaches, and address some general concerns for the simplification task in Section 6.
Showing the abstract — retrieve the full paper via the Exa API.