Publication

The acquisition of English speech rhythm by adult Chinese ESL and EFL learners

Aug 1, 2003 · 1 author · 17 topics

Abstract

THE ACQUISITION OF ENGLISH SPEECH RHYTHM BY ADULT CHINESE ESL AND EFL LEARNERS by Te-fang Hua Chairperson of the Supervisory Committee: Professor Ann M. Peters Department of Linguistics Mandarin Chinese speakers are frequently reported by ESL professionals to speak English in a syllable-timed rhythm. However, little empirical evidence is available to physically characterize their speech rhythm in English. In view of the paucity of information available on this issue, the current study compares speech samples of Taiwan Mandarin (TM) and English speakers with respect to their difficulties in producing English rhythm by analyzing three well-attested correlates of stress in English, duration, intensity, and pitch. The Participants in this study were 10 native speakers of English, 10 TM speakers learning English as a Second Language (ESL), and 10 TM speakers learning English as a Foreign Language (EFL). The subjects were requested to read two prosodically diverse sets of sentences, with Type A featuring a single strong syllable or two widely spaced strong syllables and Type B featuring a regular alternation between strong and weak syllables. The results showed that the TM ESL and EFL speakers experienced difficulties with Type A but not with Type B rhythm. For Type A sentences, the TM speakers produced relatively shorter, softer, and lower-pitched strong syllables and relatively longer, louder, and higher-pitched weak syllables than the English speakers. The combination leads to less duration, intensity, and pitch differentiation between the strong and the weak syllables. Additionally, the TM speakers produced fewer levels of stress than the English speakers did. Increased proficiency and exposure is correlated with positive changes in the use of duration, intensity, and pitch as correlates for stress. The current study strongly challenges using "syllable-timing" as a cover tenn in describing the speech rhythm of TM speakers because they were apparently able to manage at least one type of English stress-timing well. We propose multiple parameters under the traditional rhythmical category "stress-timing" by building in possible language-specific variations as to the number of unstressed syllables permitted within a foot and the number of prosodically weak syllables within a higher prosodic domain. TABLE OF CONTENTS Acknowledgements iv Abstract vi List of Tables xiv List of Figures xxi CHAPTER 1: Introduction 1 CHAPTER 2: Characteristics of rhythm in English and in Chinese 5 2.1 Defining rhythm 5 2.2 Speech rhythm in English 6 2.2.1 Characteristcs of English speech rhythm 6 2.2.1.1 Tendency toward isochrony 6 2.2.1.2 Metrical representation of stress 10 2.2.1.3 The acoustic correlates of stress in English 12 2.2.2 Acquisition of English rhythm 15 2.3 Characteristics of speech rhythm in Mandarin Chinese 23 2.3.1 Foot strucure in Mandarin 23 2.3.2 Full-toned versus neutral-toned syllables 25 2.3.3 Idiosyncrasies of lexical rhythm across dialects 30 2.3.4 The Phonetic correlates of stress in Mandarin Chinese 33 2.3.5 Are BM and TM mora-timed, syllable-timed, or foot-timed? 35 2.4 Potential difficulties Taiwan Mandarin and Beijing Mandarin speakers might have with the production of English speech rhythm 36 CHAPTER 3: The present study 40 3.1 Purpose of the study 40 3.2 Research questions 40 3.2.1 Do Taiwan mandarin speakers have difficulties with the duration, intensity, or pitch of strong syllables, weak syllables, or both? 42 3.2.2 Do Taiwan mandarin speakers produce less differentiation in duration, intensity, or pitch between strong and weak syllables than English speakers? 43 3.2.3 Do Taiwan Mandarin speakers correlate duration, intensity, and pitch with the target stress patterns in an English-like way? 44 3.2.4 Do TM speakers produce more English-like duration, intensity, and pitch patterns with improved proficiency and exposure to English? 45 3.2.5 Are duration, intensity, and pitch coordinated as correlates of stress for Taiwan Mandarin speakers? 46 3.3 Method 47 3.3.1 Subjects 47 3.3.2 Materials 48 3.3.3 Procedures 53 3.3.4 Analyses 54 3.3.4.1 Acoustical analyses 54 3.3.4.2 Normalization of data 63 3.3.4.3 Statistical analyses 68 CHAPTER 4: Results and discussion (1): Duration 70 4.1 Duration patterns of Type A sentences 71 4.1.1 Absolute duration patterns 71 4.1.2 Relative durations Patterns 78 4.1.3 Significant differences between NS, ESL and EFL 85 4.1.3.1 Differences between NS and ESL vs. differences between NS and EFL 86 4.1.3.2 Significant differences between ESL and EFL 87 4.1.4 Correlation of duration patterns between speaker groups 91 4.1.5 Test reliability 92 4.1.6 Summary of results for Type A sentences 93 4.2 Duration patterns of Type B sentences 94 4.2.1 Absolute duration of syllables 94 4.2.2 Relative duration patterns 101 ix 4.2.3 Significant differences between NS, ESL and EFL 108 4.2.3.1 Differences between NS and ESL vs. differences between NS and EFL 110 4.2.3.2 Significant differences between ESL and EFL 111 4.2.4 Correlation of duration patterns between speaker groups 114 4.2.5 Test Reliability 114 4.2.6 Summary of duration results for Type B sentences 115 4.3 Discussion of the duration in Type A and Type B sentences 117 4.3.1 Do TM speakers have difficulty lengthening strong syllables, shortening weak syllables, or both? 117 4.3.2 Do TM speakers use duration as a correlate of stress? 121 4.3.3 Do TM speakers produce smaller duration contrasts between strong and weak syllables than English speakers? 126 4.3.4 Do TM speakers produce duration patterns that are closer to the target speech rhythm with improved proficiency and increased exposure to English? 133 4.3.5 Summary for the discussion of duration 134 CHAPTER 5: Results and discussion (2): Intensity 136 5.1 Intensity patterns of Type A sentences 137 5.1.1 Intensity patterns in dB 137 5.1.2 Relative intensity patterns 144 5.1.3 Significant differences between NS, ESL and EFL 151 5.1.3.1 Differences between NS and ESL vs. differences between NS and EFL 153 5.1.3.2 Significant differences between ESL and EFL 154 5.1.4 Correlation of intensity patterns between speaker groups 157 5.1.5 Test reliability 158 5.1.6 Summary of intensity results for Type A sentences 159 5.2 Intensity patterns of Type B sentences 161 5.2.1 Intensity patterns in dB 161 5.2.1.1 Relative intensity patterns 167 x 5.2.2 Significant differences between NS, ESL and EFL 174 5.2.2.1 Differences between NS and ESL vs. differences between NS and EFL 176 5.2.3 Significant differences between ESL and EFL 177 5.2.4 Correlation of intensity ratios between speaker groups 178 5.2.4.1 Test Reliability 179 5.2.5 Summary of intensity results for Type B sentences 180 5.3 Discussion of the intensity in Type A and Type B sentences 182 5.3.1 Do TM speakers have difficulty with the intensity of strong syllables, the intensity of weak syllables, or both? 182 5.3.2 Do TM speakers use intensity as a correlate of stress? 186 5.3.3 Do TM speakers produce smaller intensity contrasts between strong and weak syllables than English speakers? 189 5.3.4 Do TM speakers with improved proficiency and increased exposure to English produce intensity patterns that are closer to the target speech rhythm? 195 5.3.5 Summary for the discussion of intensity 197 CHAPTER 6: Results and discussion (3): Pitch 199 6.1 F0 patterns of Type A sentences 200 6.1.1 Fo frequency in Hz 201 6.1.2 Fo frequency in semitone ratios 208 6.1.3 Significant differences between NS, ESL and EFL 215 6.1.3.1 Differences between NS and ESL vs. differences between NS and EFL 217 6.1.3.2 Significant differences between ESL and EFL 218 6.1.4 Correlation of pitch patterns between speaker groups 220 6.1.4.1 Test Reliability 221 6.1.5 Summary of pitch results for Type A sentences 222 6.2 Pitch patterns of Type B sentences 224 6.2.1 Fo frequency in Hz 224 6.2.2 FO frequency in semitone ratios 231 xi 6.2.3 Significant differences between NS, ESL and EFL 238 6.2.3.1 Differences between NS and ESL vs. differences between NS and EFL 240 6.2.3.2 Significant differences between ESL and EFL 241 6.2.4 Correlation of pitch patterns between speaker groups 242 6.2.5 Test reliability 242 6.2.6 Summary of results for Type B sentences 243 6.3 Discussion of the pitch in Type A and Type B sentences 244 6.3.1 Do TM speakers have difficulty with the pitch of strong syllables, the pitch of weak syllables, or both? 244 6.3.2 Do TM speakers use pitch as a correlate of stress? 249 6.3.3 Do TM speakers produce smaller pitch contrasts between strong and weak syllables than English speakers? 253 6.3.4 Do TM speakers with improved proficiency and increased exposure to English produce pitch patterns that are closer to the target speech rhythm? 260 6.3.5 Summary for the discussion of pitch 262 CHAPTER 7: Coordination among duration, intensity, and pitch 265 7.1 The coordination of duration, intensity, and pitch of syllables in Non-final positions 267 7.1.1 Significant differences in duration, intensity, and pitch 267 7.1.2 The variations in duration, intentisy, and pitch and stress 270 7.1.3 Differentiation between strong and weak syllables in non-final position .. 274 7.2 Duration, intensity, and pitch of syllables in final position 280 7.2.1 Significant differences in duration, intensity, and pitch between pairs of subject groups · 281 7.2.2 The variations in duration, intentisy, and pitch and stress 282 7.2.3 Differentiation between strong and weak syllables in final position 286 7.3 Summary 289 CHAPTER 8: Conclusion 291 8.1 Findings and implications 291 8.2 Strengths and limitations of the currenty study 296 8.2.1 Strengths 296 8.2.2 Limitations 297 8.3 directions for further research 301 8.3.1 Production of speech rhythm 301 8.3.2 Perception of stress 304 8.3.3 A quantitaive representation of speech rhythm , 306 8.3.4 Rhythm and speech processing 308 8.3.5 Pedagogy for teaching speech rhythm 309 Appendix A: List of Experimental Sentences 313 A.l Type A Sentences 313 A.2 Type B Sentences 314 Appendix B: Segmentation Criteria 315 B.l Type A Sentences 315 B.2 Type B Sentences 321 Bibliography 327 LIST OF TABLES Table 2.1 Tones in Mandarin Chinese 27 Table 3.1 Profile of the three subject groups 47 Table 3.2 Labels, stress patterns, and syllable numbers of test sentences 53 Table 3.3 The Hertz to Semitone conversion chart 66 Table 3.4 Mathematic derivation of semitones from frequency in Hz 67 Table 4.1 Mean syllable durations in ms for Type A sentences 72 Table 4.2 Mean syllable duration as its percentage of the total sentence duration for Type A sentences 79 Table 4.3 Student's t-test scores for duration (%) of individual syllables between pairs of groups for Type A sentences 85 Table 4.4 Number of strong and weak syllables with duration (%) significantly different from NS in non-final vs. final positions for Type A sentences ........ 86 Table 4.5 Number of strong and weak syllables with duration (%) significantly different between EFL and ESL in non-final vs. final position for Type A sentences 88 Table 4.6 Number of strong and weak syllables with duration (%) significantly different between EFL and ESL speakers categorized as content and function in non-final vs. final position 89 Table 4.7 Pearson Product-moment Correlation Coefficients for mean syllable duration between groups for Type A sentences 91 Table 4.8 Test-retest reliability for syllable duration from three productions for Type A sentences 92 Table 4.9 Mean syllable durations in ms for Type B sentences 95 Table 4.10 Mean syllable duration as its percentage of the total sentence duration for Type B sentences 102 xiv Table 4.11 Student's t-test scores for durations (%) of individual syllables between groups for Type B sentences 109 Table 4.12 Number of strong and weak syllables with durations (%) significantly different from NS in non-final vs. final positions for Type B sentences ...... 110 Table 4.13 Number of strong and weak syllables with duration (%) significantly different between EFL and ESL in non-final and final positions for Type B sentences 112 Table 4.14 Pearson Product-moment Correlation Coefficients for mean syllable duration between pairs of groups for Type B sentences 114 Table 4.15 Test-retest reliability for syllable duration from three productions of Type B sentences 115 Table 4.16 Number of strong and weak syllables with duration (%) significantly different from NS in non-final vs. final positions for Type A and Type B sentences 117 Table 4.17 Number of strong, weakly stressed, and unstressed syllables with duration (%) significantly different from NS in non-final vs. final position for Type A sentences 118 Table 4.18 Group average duration (%) of strong and weak syllables in non-final and final positions for Type A and Type B sentences 120 Table 4.19 Group average duration (%) of strong, weakly stressed, and unstressed syllables in non-final and final positions for Type A sentences 121 Table 4.20 Number of strong and weak syllables classified as "+LRPS" or "-LRPS" in non-final and final positions for Type A and Type B sentences 122 Table 4.21 Number of "+LRPS" vs "-LRPS" strong, weakly stressed, vs. unstressed syllables in non-final vs. final position for Type A sentences 125 Table 4.22 Average duration (%) of strong, weakly stressed, and unstressed syllables in non-final and final positions for Type A and Type B sentences 126 Table 4.23 Duration contrasts (%) between strong and weak syllables in non-final position for Type A sentences 127 Table 4.24 Duration contrasts (%) between strong and weak syllables in final position for Type A sentences 128 xv Table 4.25 Duration contrasts (%) between strong and unstressed syllables in non final and final position for Type A sentences 130 Table 4.26 Duration contrasts (%) between stressed and unstressed syllables in non final and final position for Type A sentences 131 Table 4.27 Duration contrasts (%) between strong and weak syllables in non-final positions for Type B sentences 131 Table 4.28 A comparison of duration results between ESL and EFL. 133 Table 5.1 Group Average intensity in dB of individual syllables for Type A sentences 138 Table 5.2 Group Average intensity ratios of individual syllables for Type A sentences 145 Table 5.3 Student's t-test scores for intensity ratios of individual syllables between groups for Type A sentences 152 Table 5.4 Number of syllables with intensity ratios significantly different from NS as strong or weak in non-final and final positions for Type A sentences ..... 153 Table 5.5 Number of strong and weak syllables with intensity ratios significantly different between EFL and ESL in non-final vs. final positions for Type A sentences 154 Table 5.6 Number of strong and weak syllables with intensity ratios significantly different between EFL and ESL speakers categorized as content and function in non-final vs. final positions for Type A sentences 155 Table 5.7 Pearson Product-moment Correlation Coefficients for mean syllable intensity ratios between groups for Type A sentences 157 Table 5.8 Test-retest reliability for intensity ratios from three productions for Type A sentences 159 Table 5.9 Group Average intensity in dB of individual syllables for Type B sentences 162 Table 5.10 Group Average intensity ratios of individual syllables for Type B sentences 168 Table 5.11 Number of syllables with intensity ratios significantly different from NS in categories of strong and weak in non-final and final positions for Type B sentences 176 Table 5.12 Number of syllables with intensity ratios significantly different between EFL and ESL in categories of strong and weak in non-final and final positions for Type B sentences 178 Table 5.13 Pearson Product-moment Correlation Coefficients for mean intensity ratios between pairs of groups for Type B sentences 179 Table 5.14 Test-retest reliability for intensity ratios from three productions of Type B sentences 180 Table 5.15 Number of strong and weak syllables with intensity ratios significantly different from NS in non-final vs. final positions for Type A and Type B sentences 182 Table 5.16 Number of strong, weakly stressed, and unstressed syllables with intensity ratios significantly different from NS in non-final vs. final position for Type A sentences 183 Table 5.17 Group average intensity ratios of strong and weak syllables in non-final and final positions of Type A and Type B sentences 184 Table 5.18 Group average intensity ratios of strong, weakly stressed, and unstressed syllables in non-final and final position for Type A sentences 185 Table 5.19 Number of strong and weak syllables classified as "+IRPS" or "-IRPS" in non-final and final position for Type A and Type B sentences 186 Table 5.20 Number of "+IRPS" vs "-IRPS" strong, weakly stressed, vs. unstressed syllables in non-final vs. final position for Type A sentences 188 Table 5.21 Average intensity ratios of strong, weakly stressed, and unstressed syllables in non-final and final positions for Type A and Type B sentences 189 Table 5.22 Intensity contrasts (ratio) between strong and weak syllables in non-final positions for Type A sentences 190 Table 5.23 Intensity contrasts (ratio) between strong and weak syllables in final positions for Type A sentences 192 Table 5.24 Intensity contrasts (ratios) between strong and unstressed syllables in non final and final position for Type A sentences 193 Table 5.25 Intensity contrasts (ratios) between stressed and unstressed syllables in non-final and final position for Type A sentences 193 Table 5.26 Intensity contrasts (ratio) between strong and weak syllables in non-final positions for Type B sentences 194 Table 5.27 A comparison of intensity results between ESL and EFL 196 Table 6.1 Group Average Fo in Hz of individual syllables for Type A sentences 202 Table 6.2 Group Average Fo in semitone ratios of individual syllables for Type A sentences 209 Table 6.3 Student's t-test scores for semitone ratios of individual syllables between NS, ESL and EFL for Type A sentences 216 Table 6.4 Number of syllables with semitone ratios significantly different from NS as strong and weak in non-final and final positions for Type A sentences... 217 Table 6.5 Number of strong and weak syllables with semitone ratios significant different between EFL and ESL in non-final vs. final positions for Type A sentences 218 Table 6.6 Pearson Product-moment Correlation Coefficients for mean syllable semitone ratios between groups for Type A sentences 221 Table 6.7 Test-retest reliability for Fo frequency in semitone ratios from three productions of Type A sentences 222 Table 6.8 Group Average Fo in Hz of individual syllables for Type B sentences 225 Table 6.9 Group average Fo (Hz) of strong content words vs. weak function in non final and final positions for Type B sentences 231 Table 6.10 Group Average Fo in semitone ratios of individual syllables for Type B sentences 232 Table 6.11 Student's t-test scores for semitone ratios of individual syllables between groups for Type B sentences 239 Table 6.12 Number of strong and weak syllables with semitone ratios significantly different from NS in non-final vs. final positions for Type B sentences ...... 240 Table 6.13 Number of strong and weak syllables with semitone ratios significantly different between EFL and ESL in non-final vs. final positions for Type B sentences 241 Table 6.14 Pearson Product-moment Correlation Coefficients for mean semitone ratios between groups for Type B sentences 242 Table 6.15 Test-retest reliability for syllable semitone ratios from three productions of Type B sentences 243 Table 6.16 Number of strong and weak syllables with semitone ratios significantly different from NS in non-final vs. final positions for Type A and Type B sentences 245 Table 6.17 Number of strong, weakly stressed, and unstressed syllables with semitone ratios significantly different from NS in non-final vs. final position for Type A sentences 246 Table 6.18 Group average semitone ratios of strong and weak syllables in non-final and final positions of Type A and Type B sentences 247 Table 6.19 Group average semitone ratios of strong, weakly stressed, and unstressed syllables in non-final and final position for Type A sentences 248 Table 6.20 Number of strong or weak syllables classified as "+PRPS" or "-PRPS" in final and non-final position for Type A sentences 249 Table 6.21 Number of "+PRPS" vs "-PRPS" strong, weakly stressed, vs. unstressed syllables in non-final vs. final position for Type A sentences 251 Table 6.22 Average semitone ratios of strong, weakly stressed, and unstressed syllables in non-final and final positions for Type A and Type B sentences 252 Table 6.23 Group average semitone ratios of strong vs. weak syllables in non-final position for Type A sentences 254 Table 6.24 Group average semitone ratios of strong vs. weak syllables in final position for Type A sentences 256 Table 6.25 Pitch contrasts (ratios) between strong and unstressed syllables in non final and final position for Type A sentences 258 Table 6.26 Pitch contrasts (ratios) between stressed and unstressed syllables in non final and final position for Type A sentences 258 xix Table 6.27 Group average semitone ratios of strong vs. weak syllables in non-final position for Type B sentences 259 Table 6.28 A comparison Fo results between ESL and EFL 261 Table 7.1 Number of strong and weak non-final syllables for Type A sentences with significant greater or smaller duration, intensity, or semitone ratios than those of NS 268 Table 7.2 Number of strong and weak non-final syllables with duration, intensity, or semitone ratios significantly greater or smaller than those of NS for Type B sentences 269 Table 7.3 Number of strong and weak non-final syllables classified as Raised or Lowered for Type A sentences 271 Table 7.4 Number of strong and weak non-final syllables classified as "+SRPS" or "-SRPS" for Type B sentences 273 Table 7.5 Duration, intensity, and semitone ratios of strong versus weak syllables in non-final positions for Type A sentences 274 Table 7.6 Duration, intensity, and semitone ratios of strong versus weak syllables in non-final positions for Type B sentences 278 Table 7.7 Number of strong and weak syllables in final position of Type A sentences classified as significantly different in duration, intensity, and/or pitch 281 Table 7.8 Number of strong and weak syllables in final position of Type B sentences classified as significantly different in duration, intensity, and/or pitch 282 Table 7.9 Number of strong and weak final syllables classified as "+SRPS" or "- SRPS" for Type A sentences 283 Table 7.10 Number of strong final syllables classified as "+SRPS" or "-SRPS" for Type B sentences 285 Table 7.11 Duration, intensity, and semitone ratios of strong versus weak syllables in final positions for Type A sentences 287 LIST OF FIGURES Figure Figure 3.1 The sampling of the Extreme Point Fosfor the utterance Jane is not my mom 60 Figure 3.2 The corresponding Extreme Point Fos on the original pitch track for the utterance Jane is not my mom 61 Figure 3.3 The Middle Fos, the Fosat Peak Intensity, the Mean Fos, and Extreme Point Fos for the utterance Jane is not my mom 62 Figure 4.1 Mean syllable durations (ms) for sentences Al through A7 77 Figure 4.2 Mean syllable durations (%) for sentences Al through A7 84 Figure 4.3 Distribution of significant differences between ESL and EFL in terms of the spread of difficulties to ESL and/or EFL 90 Figure 4.4 Mean syllable durations (ms) for sentences B1 through B7 100 Figure 4.5 Mean syllable durations (%) for sentences B1 through B7 107 Figure 4.6 Distribution of significant differences in syllable duration between ESL and EFL speakers for Type B sentences 113 Figure 4.7 Number of strong and weak non-final syllables classified as "+LRPS" or "-LRPS" for Type A and Type B sentences 123 Figure 4.8 Number of strong and weak final syllables classified as "+LRPS" for Type A and Type B sentences 124 Figure 4.9 Group average duration (%) of strong vs. weak syllables in non-final position for Type A sentences 127 Figure 4.10 Group average duration (%) of strong vs. weak syllables in final position for Type A sentences 129 Figure 4.11 Group average duration (%) of strong vs. weak syllables in non-final position for Type B sentences 132 Figure 5.1 Group Average peak syllable intensity in dB for sentences Al through A7 143 xxi Figure 5.2 Group Average intensity ratios for sentence Al through A7 150 Figure 5.3 Distribution of significant differences in intensity ratios between ESL and EFL in terms of the spread of difficulties to EFL and/or EFL 156 Figure 5.4 Group Average peak syllable intensity in dB for sentences Bl through B7. 167 Figure 5.5 Mean syllable intensity ratios for sentences B1 through B7 173 Figure 5.6 Number of strong and weak syllables classified as "+IRPS" or "-IRPS" for Type A and Type B sentences 187 Figure 5.7 Group average intensity ratios of strong vs. weak syllables in non-final position for Type A sentences 191 Figure 5.8 Group average intensity ratios of strong vs. weak syllables in final position for Type A sentences 192 Figure 5.9 Group average intensity ratios of strong vs. weak syllables in non-final position for Type B sentences 194 Figure 6.1 Group Average FoFrequency (Hz) for sentences Al through A7 207 Figure 6.2 Group Average FoFrequency in semitone ratios for sentence Al through A7214 Figure 6.3 Distribution of significant differences in semitone ratios between ESL and EFL in terms of the spread of difficulties to ESL and/or EFL 220 Figure 6.4 Group Average FoFrequency (Hz) for sentences Bl through B7 230 Figure 6.5 Group Average FO Frequency in semitone ratios for sentences B1 through B7237 Figure 6.6 Number of strong and weak syllables classified as "+PRPS" or "-PRPS" for Type A and Type B sentences 250 Figure 6.7 Group average semitone ratios of strong vs. weak syllables in non-final position for Type A sentences 255 Figure 6.8 Group average semitone ratios of strong vs. weak syllables in final position for Type A sentences 257 Figure 6.9 Group average semitone ratios of strong vs. weak syllables in non-final position for Type B sentences 259 Figure 7.1 Percentage of strong non-final syllables classified as "+SRPS" and weak non-final syllables classified as "-SRPS" for Type A sentences 271 xxii Figure 7.2 Percentage of strong non-final syllables classified as "+SRPS" and weak non-final syllables classified as "-SRPS" for Type B sentences 273 Figure 7.3 Duration, intensity, and pitch contrasts between strong and weak non-final syllables for Type A sentences for NS, ESL, and EFL speakers 277 Figure 7.4 Duration, intensity, and pitch contrasts between strong and weak non-final syllables for Type B sentences for NS, ESL, and EFL speakers 280 Figure 7.5 Percentage of strong final syllables classified as "+SRPS" and "-SRPS" final syllables classified as "+SRPS" for Type A sentences 284 Figure 7.6 Percentage of strong final syllables classified as "+SRPS" for Type B sentences 286 Figure 7.7 Duration, intensity, and pitch contrasts between strong and weak syllables in final position for Type A sentences for NS, ESL, and EFL speakers....... 289 CHAPTER 1: INTRODUCTION As fundamental as rhythm is in the perception and production of speech, the rhythm of many languages has not been positively determined for native speakers, much less even considered for non-native speakers. To date, how second language learners acquire the rhythm of a new language remains largely an empirical question. Mandarin Chinese speakers, for example, are frequently reported by professionals teaching and researching English as a second language to experience difficulties with English speech rhythm. Their rhythm has been informally described as staccato or syllable-timed. However, little empirical evidence is available to physically characterize what it is that makes their speech rhythm different from that of the native speakers of English. What makes this inquiry particularly interesting is that although learning to produce a new speech rhythm seems one of the most difficult tasks for second language learners, it is one of the first linguistic features discovered by children learning their first language. It seems that prosodic aspects of language, which are usually acquired at an early stage for children learning their first language, may be especially difficult to break away from for second language learners. Psycholinguists have long noted that newborn infants can discriminate languages differing in rhythmic structure, even when the speech itself has been low-pass filtered, namely, when only suprasegmental information is available (Mehler et aI., 1988; Moon et aI., 1993). At 9 months of age, infants already show reliable preference for unfamiliar words with a rhythmic pattern that is dominant in their parental language (Jusczyk, Cutler, & Rendanz, 1993). There is also evidence of rhythmic constraints on children's early speech. Numerous studies on the omissions of syllables in children's production of polysyllabic words reported that English-speaking children develop a strong preference for the Strong-Weak metrical footing at an early stage of language acquisition (Allen & Hawkins, 1980; Klein, 1978, 1981; Wijnen et aI., 1994; Gerken, 1994, 1996). Furthermore, speech rhythm plays an essential role in the processing of speech. Over the past two decades, there is increasing evidence that native listeners draw heavily on rhythmic units as the most natural and efficient way of segmenting continuous speech into words (Mehler et aI., 1981; Cutler et aI., 1986; Cutler and Norris, 1988; Cutler and Butterfield, 1992; Otake et aI., 1993; Cutler, 1994). Results of these studies generally suggest that in a stress-timed language such as English, native listeners segment at the beginning of a stressed syllable, whereas in a syllable-timed language such as French, segmentation takes place at each syllable. In other words, the units of rhythm, perceptual prominence, and segmentation tend to converge for a given language. This means that placing stress on the right syllables is essential to the comprehension of speech. For second language learners, deviance in speech rhythm is not only a key element for the detection of a foreign accent, but it also negatively affects the intelligibility of one's speech (Adams, 1979; Wang, 1987; Anderson-Hsieh, Johnson, & Koehler, 1992). If the speaker makes a stress mistake within a word, the listener may not be able to understand the word, even if all the individual sounds are pronounced correctly. Or if the speaker places stress on the wrong word(s) within a sentence, the listener may misunderstand the meaning of the sentences, even if all the individual sounds are pronounced correctly. Studies have shown that native speakers of English are less tolerant of prosodic deviance than of segmental errors when they are asked to judge intelligibility, acceptability of accent, or overall pronunciations of foreigners' speech (James, 1976; Johansson, 1978; Anderson-Hsieh, Johnson, & Koehler, 1992). Although the rhythmic structure of English has been more extensively researched than most other languages, what we know about the second language acquisition of English speech rhythm to date reveals only the tip of an iceberg. Little empirical evidence is available in the form of acoustic measurements which physically characterize the difficulties non-native speakers might experience with English speech rhythm. There is still considerable speculation regarding which parameters (fundamental frequency, intensity, or duration) cause difficulty in speech rhythm for non-native speakers of English. Yet if the foreign learner wants to acquire the stress-timed rhythm of English with any degree of success, it is apparent that they must know the means to produce English stress and in particular, what it is that distinguishes stressed from unstressed syllables in the stream of speech. Previous research in this area is generally limited to the examination of the duration of isolated words in the speech of a small number of non-native speakers drawn from a single proficiency level (see Chapter Two for a full review). In view of the paucity of information available on the second language acquisition of speech rhythm and the limitationsrof the previous research, the current study compares Taiwan Mandarin and English speakers with respect to their phonetic realization and distribution of stress by analyzing all three well-attested correlates of stress in English and using longer stretches of speech produced by a relatively larger number of subjects representing two English proficiency levels. Results of this study are expected to contribute toward our understanding of the acquisition of second language prosody, particularly how second language learners develop the new rhythm, including using a phonetic process reserved for marking a different set of phonemic contrasts in their native language. In this case, the pitch correlate of stress in English is also used for signaling phonemic pitch contrasts in Chinese Mandarin. Results of the study may also be used for comparing first and second language acquisition of speech rhythm and for investigating the ways in which speakers of various native languages are similar or different in their use of duration, intensity, and pitch as correlates for stress in English. In addition, the current results may have implications for the construction of a theoretical model of the rhythm of language. Second language rhythms provide additional data that might help reveal any natural constraints on the structure or development of a speech rhythm. The study will possibly interest psycholinguists researching the link between rhythm and speech segmentation. The results of the current study are expected to reveal rhythmic patterns of TM ESL/EFL learners, which may affect their strategy for segmenting speech in English. Although the link between rhythm and speech segmentation in the native language has been an active area of investigation, whether a 3 similar link exists in the second language is generally unknown. In addition, the current research methods, including our approaches for sampling, measuring, and normalizing duration, intensity, and pitch, may become useful for researchers who want to conduct experiments on the prosodic aspects of language. The results are expected to raise the awareness of second language professionals about the kinds of difficulties that learners might have in learning to produce a new speech rhythm. In effect, the study hopes to encourage second language teachers and researchers to reexamine the ways in which stress and rhythm are introduced in current classroom practices. The current findings are expected to have implications for the development of teaching materials appropriate for TM speakers learning English as a second language. The organization of the chapters will be as follows. Chapter One is a general introduction. Chapter Two reviews the phonetics and the phonology of rhythm in English and in Mandarin. Based on the review, a comparison is made between these two languages to make preliminary predictions about the types of difficulties TM speakers might have with English speech rhythm. With these predictions in mind, Chapter Three presents the research questions and details regarding the methodological design ofthe current study. Chapters Four through Six each focus on analysis of a different phonetic cue, in the order duration, intensity, and pitch. Results from both acoustical and statistical analyses are presented. Immediately following the presentation of results is a discussion of the results, which directly addresses the research questions. Chapter Seven combines results from duration, intensity, and pitch to examine possible interactions. Chapter 8 summarizes results and draws conclusions for the whole study. Finally the limitations and strengths of the study are discussed along with proposals for further research on the second language acquisition of rhythm. CHAPTER 2: CHARACTERISTICS OF RHYTHM IN ENGLISH AND IN CHINESE The organization of this chapter is as follows. In section 2.1, I discuss speech rhythm in general. In section 2.2, I discuss the characteristics of speech rhythm in English and review previous research on the second language acquisition of English speech rhythm. In section 2.3, I discuss the characteristics of speech rhythm in Beijing Mandarin (BM) and Taiwan Mandarin (TM). In section 2.4, I project possible difficulties TM speakers might have with English speech rhythm based on similarities and differences between TM and English with regard to their characteristics of speech rhythm. 2.1 DEFINING RHYTHM Speech rhythm is defined as the phonological representations of isochrony, i.e., "the process of providing phonological representations of time with phonetic substance (i.e., duration)" (Podesva, 2003). Rhythm is defined as a durational phenomenon although in some languages (e.g., English), stress may serve to define rhythm (to be discussed in section 2.2). Most linguists discuss rhythm in terms of languages being 'foot-timed' (also referred to as 'stress-timed'), 'syllable-timed', or 'mora-timed'. The distinction is based on the notion of isochrony such that duration is perceived to remain constant from one foot to the next in a foot-timed language (e.g., English), from one syllable to the next in a 'syllable-timed' language (e.g., French), and from one mora to the next in a 'mora-timed' language (e.g., Japanese) (Pike, 1945; Abercrombie, 1967; see McCawley, 1968 for 'mora-timing'). It can be illustrated by the following three sentences: (1) (English): This is absolutely rediculous. (French): C'est absolument ridicule. (Japanese):Kore-wa zettai-ni bakageteiru. The perceptual impression that stresses in English, syllables in French, and moras in Japanese seem to occur at equal intervals of time is perhaps one of the first triggers for proposing these three different timing units. However, it is important to note that isochrony is only intended by native speakers as a mental target. Phonetic isochrony in speech production is rarely confirmed empirically (for evidence against perfectly isochronous stress, see Classe, 1939; O'Connor, 1968; Uldall, 1971; Lea, 1975; Faure et aI., 1980; for evidence against perfectly isochronous syllables, see Delattre, 1966; Olsen, 1972; Pointon, 1980; Balasubramanian, 1980; Wenk and Wioland 1982, for evidence against perfectly isochronous moras, see Beckman 1982). Factors such as the intrinsic duration of a segment, the number of segments in a mora/syllable/foot, the number of syllables in a foot, the position of a mora/syllable/foot in an utterance (i.e., final ones tend to be longer than non-final ones) do matter (Dauer, 1982). However, if one could keep all of these factors under control, it may be possible to approximate phonetic isochrony. 2.2 SPEECH RHYTHM IN ENGLISH I first discuss the characteristics of English speech rhythm in section 2.2.1. Then I review previous research on the second language acquisition of English speech rhythm in section 2.2.2. 2.2.1 CHARACTERISTCS OF ENGLISH SPEECH RHYTHM 2.2.1.1 Tendency toward isochrony English is commonly referred to as an archetypal stress-timed language. This view originates from the impression that stressed syllables in English to occur at more or less equal intervals of time as illustrated in (2). (2) The 'teacher is 'interested in 'buying some 'books. (Pike, 1945, p.34) Other things being equal, the more syllables there are in an interstress interval, the shorter they tend to be (Gimson, 1962; Halliday, 1970; Allen, 1975). Compare the following examples: (3) a. Ken's here. b. Kenny's here. c. Kennedy's here. Although each of them contains a different number of syllables in the first foot, the length of the foot does not increase proportionately to the number of syllables. For example, sentence (3c) has twice the number of syllables as sentence (3a), but it does not take twice as long to say as (3a). Apparently, the addition of a syllable to a foot does not take fixed amount of time to utter. Instead, the durations of syllables can be stretched or compressed according to the number of syllables within the foot. In addition, other things being equal, the more syllables there are in an interstress interval, the shorter the stressed syllable tends to be. Lehiste (1972) studied the timing of words spoken in isolation compared to the timing of the same words with monosyllabic or bisyllabic suffixes. Two native speakers of American English produced base words such as stick, speed, sleep, and shade in isolation, and then the same words followed by monosyllabic suffixes such as -y and -er and bisyllabic suffixes such as -ily, and -iness. Lehiste found that her subjects consistently compressed the duration of the base words according to the number of the syllables added to them, supporting the hypothesis that in reality the foot is the domain of timing. One of the major consequences of stress-timing is thus the wide variation in syllable length. If a speaker intends to say each stretch of varying numbers of syllables within similar time limits, stretches with more syllables must be spoken faster than stretches with fewer syllables. It seems that the more syllables there are in an interval, the more compressed the duration of the unstressed syllables tends to become as well. This can cause the unstressed syllables to undergo various reductions or modifications. For example, consonants or vowels in unstressed syllables can be reduced or completely omitted in English. (4) a./p:;)red/ [phred] 'parade' b./p:;)lisl [phlis] 'police' c. If:;)netIksl [fnethIkhs] 'phonetics' d. lay so hIm! [ay so 1m] 'I saw him.' e. lay wII got [ayl go] 'I will go.' It is important to note that although there is a tendency toward isochronous stresses in the production of English (i.e., the more syllables there are in a foot, the shorter they become), stresses in English are by no means physically isochronous in speech production. Physical evidence for perfect isochrony is slight. Many experimental studies have disproved the existence of phonetic isochrony in English speech production. Classe (1939) measured interstress intervals in English and found that they are measurablely longer when they contain more syllables. Bolinger (1965) examined the recordings of two lengthy sentences read by six English speakers. Accents were marked and the intervals between the accents were measured. His results did not support isochronic stresses, either. O'Connor (1965) recorded a strictly rhythmic limerick in English spoken with a tap of the hand at each stress. However, even under such favorable conditions, isochrony could not be demonstrated. In a follow-up study (1968), he used a set of seven utterances, each containing three monosyllabic feet, with the first and third foot held constant, but with the second varied in segmental length from three to nine segments. Duration measurements showed that the variable foot had a clear tendency to greater duration as the number of segments increased. He found no evidence for equal interstress intervals in English. Uldall (1971) studied the recording of a passage from "The North Wind and the Sun" read by David Abercrombie. The 45-second passage was divided into metrical feet, whose durations ranged from 260 to 870 ms. Similar to what was found in previous studies, he also reported that the average duration of all feet varied considerably according to the number of syllables. The results showed that the lengths of the inter stress intervals were not regular. Lea (1980) found a linear relationship between the number of intervening unstressed syllables and the duration of the inter-stress intervals. The results seriously contradict the claim that the number of intervening unstressed syllables has no effect on the duration of the interval. Faure, Hirst, and Chafcouloff (1980) had two British subjects record a number of sentences. The recordings were played to three English phoneticians to identify the stresses. The durations of the inter-stress intervals were then measured. They concluded that stressed syllables in English were not even roughly isochronous because the inter-stress intervals varied considerably in length with the number of syllables. Results from such studies show that it is extremely difficult to confirm isochrony of stresses in production by measuring interstress intervals instrumentally. However, evidence shows that native listeners seem to perceive stresses as more isochronous than they actually are. Lehiste (1975; 1977) found that English speakers tend to perceive isochrony in utterances which are not actually isochronous. In one of her experiments, she played four noise-filled intervals to a group of 30 listeners. Of the four sequences of intervals, three were of equal length, and the fourth was either longer or shorter. The duration of the intervals corresponded to the range observed in actual productions of metric feet in another production experiment. The subjects were first asked to identify the longest intervals, and on a second presentation, to identify the shortest intervals. Results showed that in order to obtain significant agreement that a given interval was actually longer or shorter than the rest, an increment between 30-100 milliseconds was needed. Increments smaller 30 milliseconds were never reliably detected, which accounts for most of the differences observed in the production experiment. The results suggest that some of the differences between the lengths of intervals are actually below the perceptual threshold. In fact, the just noticeable differences in actual speech may be even larger than those obtained for synthesized noise since there is independent evidence that listeners have better discrimination of length with non-speech than with speech materials (Fujisaki, Nakamura& Imoto, 1973). 2.2.1.2 Metrical representation of stress The perception of timing in English centers around stress. To understand the characteristics of English rhythm, it is important to understand how stress is organized in English. In English, stress is a property of words and phrases. Therefore, the citation form of the verb "perMIT" has its stress on the second syllable and the noun "PERmit" has its stress on the first syllable. The requirement of stress often exempts function words, which are usually realized as stressless in English. For example, English the is usually pronounced as unstressed [5e] in connected speech, but stressed as [5i] or [51\] when pronounced in isolation. When individual words are strung together to form connected speech, the underlying stress patterns of the words undergo readjustments and a new hierarchy of prominence is formed. For example, when considered separately, both words "baby" in (5a) and "sitter" in (5b) share the same Strong-Weak rhythmic pattern. But when they are connected into a single utterance in (5c), their stress patterns are reorganized in a hierarchical way. (5) x x x a.baby x x x b. sitter x x x x x x x c. baby sitter Following Metrical Theory, stress is expressed, not as features, but as relative prominence. Prominence is represented by a metrical grid, which depicts a sequence of beats that vary in strength. (6) x x x x x x x Look at Annie. The grid columns are intended to depict the temporal structure of the beats. The height of the grid columns represents the relative prominence. Intrinsically, various levels of rhythm may be present in an utterance at one time. That is, a given utterance may contain several nested hierarchical rhythmic patterns. For example, given the six-syllable word reconciliation, which has the stress pattern shown in (7), a native speaker of English will find it natural to tap on one, two, three, or six syllables, but not to tap on four or five syllables (Hayes, 1995). | | (7) | x | | | [ | a | ] | 1 taps | | A | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | | x | x | | | [re | a | ] | 2 taps | | B | | | | x | x x | | | [re | ci a | ] | 3 taps | | C | | | | x x | xxx x | | | [reconciliation] | | | 6 taps | | D | | reconciliation By definition for any rhythmic structure that has multiple levels, any beat on a higher level must also serve as a beat on all lower levels. And stresses on any level are nested within the lower levels on the hierarchy. For example, in (7), level D includes all stresses at levels A, B, C; level C includes all stresses at levels A, B, etc. The metrical representation of stress is consistent with three major phonological characteristics of stress (Kager, 1995; Hayes, 1995). First, stress is "culminative." In English, there is a single strongest syllable at the foot level and at the intonational phrase level. For example, the four-syllabic word babysitter defines two stress feet at the lowest level. The constituents at the lower level are grouped together to form larger units at higher levels. At any given level of the prosodic hierarchy, there are elements of greater and lesser prominence but there is only a single strongest stressed syllable within each domain. Second, stress is rhythmically distributed. In English, equal timing of stresses tends to occur at multiple levels, as depicted in (8) below. (8) x x x x x x x x x x x x x x x x xx xx xx twenty-seven Mississipi legislators (Hayes, 1995, p.28) Third, stress is hierarchical. In English, there are multiple levels of stresses. At least three levels of stress are present in English, (1) primary stress, (2) secondary stress, and (3) no stress. (9) 2 1 3 abstraction This is not to suggest that there is a direct mapping between the strength of the acoustic signals and the levels of prominence on a metrical grid. The location in which a syllable occurs in an utterance, for example, could have an effect on its acoustic realization. For instance, the pitch contours over a syntactic unit generally follow a downward trend in most languages. This general pitch lowering is often known as pitch declination. Thus it is common for a high pitch accent earlier in a sentence to have a higher fundamental frequency than a high pitch accent later in a sentence. Not only does the overall pitch drift lower, the range within which pitch varies also narrows toward the end of the utterance. Sometimes the range can be so narrow that it is difficult to distinguish a high pitch accent from a low pitch accent. Furthermore, the intensity of speech also tends to decrease toward the end of an utterance. As a result, it is common for a stressed syllable earlier in an utterance to have a greater intensity than a stressed syllable that occurs later in the utterance. 2.2.1.3 The acoustic correlates of stress in English Rhythmic beats in English are signaled by recurrences of stressed syllables, which are usually longer, louder, and higher in pitch than the unstressed ones (Ladefoged, 1993). 12 Lieberman (1960) recorded 16 speakers of American English producing 24 word pairs in which a change of grammatical category from noun to verb is commonly associated with a shift of primary stress from the first syllable to the second. An example for this type of words is 'progress versus pro'gress. The results indicated that 90% of their stressed syllables had higher pitch, 87% had greater intensity, and 66% had a greater duration. Although duration has the most direct impact on the temporal aspect of speech among other phonetic cues, it is highly susceptible to variations in segmental composition across syllables. This may explain in part why the percentage of stressed syllables having a greater duration turns out to be much lower than the percentage of stressed syllables having higher pitch or greater intensity in Lieberman's study. Although stressed syllables in English are often pronounced with high pitch, they can also be uttered with low pitch. The shapes of the pitch contours on stressed syllables are often used to signal different meaning (Pierrehumbert and Hirschberg, 1990). For example, high pitch can be used to signal certainty of the information whereas low pitch can be used to signal uncertainty. An example is given in (9) to illustrate that a change of intonation could lead to a change in intonation meanings. Note that these intonation contours are transcribed following the ToBI conventions, l where H* stands for high pitch accent, L* for low pitch accent, H- for high phrasal tone, and H% for high boundary tone. (10) May 1 interrupt you? H* H-H% (I have confidence that it is okay.) L* H-H% (I am uncertain if it is okay.) Even though that duration, intensity, and pitch all contribute to signaling stress in English, they are not equally important with regards to the perception of stress (Lehiste, 1970). Among these three acoustic correlates of stress, native listeners of English seem to rely most heavily on pitch and least heavily on intensity. 1 The ToBI framework is a set of community-wide standard conventions for transcribing the intonation and prosodic structure of spoken utterances in a language variety. As intonation and prosody vary from language to language, there are language specific variations of ToBI systems. Fry (1955) evaluated the relative importance of duration and intensity in the perception of stress. The material used was 125 synthesized test words created from five pairs of disyllabic words, where a difference of rhythm was associated with a difference of grammatical function, noun vs. verb. They included object, subject, digest, permit, and contract. The duration and intensity ratios of the first and the second vowels of each test word were varied independently in five steps within ranges based on the actual productions of these words by 12 native speakers of English. 100 subjects were asked to listen to these synthesized test words and to identify the accented syllable in each case. The results showed that when duration and intensity varied in the same direction, there was excellent agreement among the listeners. That is, listeners agreed that a syllable was stressed when it was both long and loud, and unstressed when it was both short and soft. The results also suggested that both duration and intensity were linked to the perception of stress. The number of noun judgments increased with the increase of duration ratio and the increase of intensity ratio of Vowel One and Vowel Two. However, the duration ratios of the two vowels had a stronger influence on the perception of stress than the intensity ratios. The whole range of variations in intensity contributed toward a 29% increase of "noun" judgments whereas the variations in duration contributed toward an increase of 70% of "noun" judgments. Fry (1958) took pitch into account in two additional experiments. In the first experiment, he combined the five duration ratios between the first and the second vowels of the disyllabic word subject with step changes of fundamental frequency, while holding intensity constant throughout both syllables. With each duration ratio, the pitch of one syllable was varied in eight steps while the other syllable was held constant at a low pitch. A total of 80 different synthesized patterns of the test word was presented to 41 subjects for stress judgments. The results confirmed pitch as a factor in the perception of stress. In all duration contexts, when two syllables differed in fundamental frequency, the syllable with the higher Fo was more likely to be judged as stressed than the syllable with the lower Fo. In a subsequent experiment, Fry combined the five duration ratios with variations in fundamental frequency over one syllable while keeping the intensity ratios of the two vowels constant at all times. In each case, the fundamental frequency of one syllable featured either a linear or a curvilinear change within a fixed range while the frequency of the other syllable was held level at either the upper or the lower bound of the predetermined pitch range. For the linear pattern, the Fochanged continuously throughout the vowel and for the curvilinear pattern, the Fochange took place in the second half of the vowel duration. The purpose was to simulate common intonation patterns on disyllabic words. 76 English listeners were subjected to the stress judgment test of the 80 synthesized versions of the test word subject. The results showed that when two syllables differ in tone shapes, the contour-tone syllable is more likely to be perceived as stressed than the level-tone syllable. In all duration patterns that contained both contour and level tone syllables, a significantly greater number of contour tone syllables were judged as stressed than level tone syllables. Whether the tonal inflection is linear or curvilinear, rising or falling does not change the judgment of stress. The results suggested that the pitch cue might outweigh the duration cue in the perception of stress in English. This finding was supported by Bolinger (1958), who indicated that a syllable could stand out from others by increasing or decreasing its pitch level, as well as by including corners or sharp points of the pitch curve. 2.2.2 ACQUISITION OF ENGLISH RHYTHM Given that there are more second language learners of English than any other language in the world, it is not surprising that most of what is known about the acquisition of a second speech rhythm comes from research on English as a target language. However, the majority of these studies were focused solely on the duration parameter of rhythm and the studies were generally to determine whether the inter-language produced isochronous stresses. Despite variability in the research methods and the native language backgrounds of the subjects, overall, the results of these studies showed that non-native speakers of English produced less duration differentiation between stressed and unstressed syllables and stressed more syllables than did native speakers of English. Adams and Munro (1978) studied the placement and the correlates of stress in the connected utterances of native and non-native speakers of English. Eight Australian English native and eight non-native graduate teachers of English of various Asian languages classified as syllable-timed were asked to read aloud 12 short passages, including two nursery rhymes, three excerpts of verse with a strongly metrical rhythm, five fabricated equivalents of these items, and one passage each of colloquial and literary prose. In the perception part of the experiment, 10 naiVe native listeners of Australian English were asked to listen to the recordings and identify the syllables that they perceived as "prominent." The results showed that the non-native speakers stressed more syllables than did native speakers of English. In addition to stressing syllables that were stressed by native speakers, non-native speakers also placed stress on a number of syllables that were not stressed by native speakers, including prepositions and conjunctions. The native speakers stressed an average of 52 syllables, while the non native speakers stressed an average of 82 syllables. In the production part of the experiment, the fundamental frequency (Hz), intensity (dB), and duration (ms) of the individual syllables were measured. Fundamental frequency is defined as the number of complete cycles of variations in air pressure per second caused by the opening and closing movements of the vocal cords. The intensity of a sound is relative to the size of variations in air pressure. It is usually measured in decibels (dB). Overall, little evidence was found to suggest that native and non-native speakers of English used different phonetic cues to signal stress at the sentence level. Stressed syllables were associated with increased duration, a greater fall of amplitude from its peak, and a wider range of fundamental frequency for both groups of subjects. It is noteworthy that although very little difference was shown between native and non native speakers in the duration of their stressed syllables, the non-native speakers' unstressed syllables were generally longer than those of the native subjects. From Table V given by Adams and Munro (1978, p.142), it is possible to figure out that the mean length of stressed syllables was 293 ms for the native speakers and 307 for the non-native speakers, whereas that of unstressed syllables was 226 ms for the native speakers and 277 ms for the non-native speakers. This study concluded that the real difference in stress production between native and non-native speakers of English lies not in the mechanisms employed to signal stress, but the distribution of stress. This was one of the few studies of rhythm that took perception into consideration and that examined all three phonetic correlates of stress. However, it had several limitations. (1) All of their non-native speakers were highly proficient English learners, whose stress production was potentially less likely to differ significantly from English native speakers than that of less proficient learners. (2) The great majority of their materials were composed of highly rhythmical texts, which could potentially induce a regular alternating rhythm out of the speakers and therefore underestimate the differences between native and non-native speakers. (3) The study did not specify the individual first languages represented among the non-native speakers nor did it explain by what criteria these languages were classified as syllable-timed or stress timed. It seems over-simplifying to assume that speakers from these different languages form a natural group. (4) The results were statistically analyzed based on the absolute measurements of fundamental frequency, intensity, and duration. Individual variations over pitch range, volume of speech, and speech rates were not taken into consideration. (5) Final and non-final syllables were not separated in the comparison between stressed and unstressed syllables. Given that final syllables are usually specially marked to signal the sentence boundary, it is desirable to analyze final and non-final syllables separately so that non-native speakers' difficulties with marking boundaries do not become confounded with their difficulties with stress. Observations of non-native speakers of English over a number of years convinced Taylor (1981) that rhythm is perhaps one of the most widespread difficulties among foreign learners of English. In an informal survey, a recording of reading aloud and unscripted speech prompted by questions was made of 49 experienced teachers of English with 23 different native languages. Of the 49 subjects, 25 were judged to have "considerable difficulties with rhythm". Of these 25, 15 were judged to have produced syllable-timed rhythm instead of stress-timed rhythm. The rhythm type of Ll was found to affect the rhythm in L2 to some extent. Of these 15 speakers, 11 were native speakers of syllable-timed languages. Insufficient length differentiation between stressed and unstressed vowels seems the basic cause of syllable-timed rhythm among non-native speakers of English, particularly the inability to use reduced forms in unstressed syllables. None of the 15 subjects who were judged to produce syllable-timed rhythm reduced unstressed vowels properly whereas all of those that did produce reduced forms produced acceptable English rhythm. This study raised a very interesting question: Do English L2 learners transfer Ll rhythmic patterns into L2 or do they all start out with a syllable-timed rhythm regardless of their LIs? However, the methodology of this survey raised some serious doubts. First, there are no indications of what these 23 languages were and what criteria the author used to classify them into the two rhythmic types. Given that there is a lack of general consensus about how the distinction should be drawn, the classification of languages into rhythmic types seems to be an empirical question of its own. Second, there were no specified criteria or procedures as to how the subjects were judged to produce a stress timed or a syllable-timed rhythm and by whom. Third, apart from duration and vowel quality, the effect of pitch and intensity on speech rhythm was largely ignored in the interpretation of why syllable-timed rhythm or stress-timed rhythm was perceived among learners of English. Bond and Fokes (1985) studied the timing of words with zero, monosyllabic, and disyllabic suffixes. Three pairs of ESL learners, native speakers of Thai, Malaysian, and Japanese, were asked to read the base words stick, speed, sleep, and shade in isolation and with monosyllabic suffixes -y and -er and disyllabic suffixes-ily and -iness. Each subject said each word printed on cards three times and proceeded through the cards three times. The relative duration of the base word to the duration of the entire word were measured and compared with the results from two English native speakers in an earlier study (Lehiste, 1972). Instead of compressing the base words in proportion to the number of syllables in the suffixes as the English native speakers did, the non-native speakers produced about the same amount of duration of on the base words, regardless of the number of suffixed syllables. There is variability between speakers of the same language or within the same speaker. They concluded that non-native speakers' difficulties were with consistency and proportion of timing. The results suggested that English learners did not observe the tendency toward isochrony of feet to the same degree nor as consistently as native speakers did. Bond and Fokes (1985) also focused on the timing aspect of rhythm. Although pitch is the most important cue for the perception of stress in English, it has been largely unexplored in the study of rhythm. It would be interesting to see how the two speakers of Thai (a tone language) and the two speakers of Japanese (a pitch accent language) differed in their pitch patterns from native English speakers. However, one needs to exercise caution when using isolated words in the study of rhythm. Speakers are more likely to enunciate clearly when asked to read vocabulary removed from context, which makes the rhythm less natural. The small numbers of subjects from each language also makes the results hard to interpret. Given that individual variations are usually greater among non-native than among native speakers, it is desirable to have a large number of subjects in the sample to increase the reliability and generalizability of the data. Mandarin speakers are frequently reported by English teachers to experience difficulties with speech rhythm in English (Chang, 1987). Their rhythm has been informally characterized as staccato or syllable-timed. However, this area of research has been little explored. Juffs (1990) studied the stress errors of 19 Mandarin-speaking first-year undergraduates from Hunan Agricultural College. A recording was made of each student reading a lOS-word passage chosen from their English textbook. Ten polysyllabic lexical items were chosen for analysis of stress placement and syllable structure. Defining word stress as primarily manifested by pitch height and tonic stress as primarily manifested by pitch movement, he found, based on his own perception as a native speaker of RP, i.e., the standard British English dialect, that Mandarin learners seemed to be using pitch movement to realize not only tonic accent, but also word stress. He suggested that 19 Mandarin speakers might have interpreted English stress as tonal due to the fact that tone is such a salient feature in their first language. In addition, Mandarin speakers sometimes produced stress on the incorrect syllables, such as saying red'skinned for 'redskinned or cont'nent for 'continent. He further proposed that there might be a link between errors in stress placement and errors in syllable structure. Taking the word 'continent as an example, he found that several students produced it with stress on the final syllable. His hypothesis was that Mandarin speakers might have syllabified it as con-ti-nent. According to his analysis of the first syllable, the vowel Ial is a weak vowel in Hunan dialect and the In! is not counted as a consonant in the same way as it is in Mandarin. This leaves the third syllable as the only heavy syllable eligible for stress assignment. He concluded that Mandarin speakers' errors with stress in English productions occurred not only in placement but also in the phonetic process used to achieve stress. Juffs drew a number of very interesting conclusions: (1) Mandarin speakers did not seem to be making tonal distinctions between word stress and sentence stress and they tended to place pitch accent on both. (2) They might have interpreted English stress as Mandarin tone. (3) They sometimes misplaced stress on polysyllabic words. And (4) their difficulties with English syllable structure may be linked to their difficulties with stress placement in English. However, Juffs' study suffers from a number of limitations. First, his observations were based on very limited supporting data. In support of each of the above-mentioned claims, only a very small portion of the results were analyzed and discussed. Second, his analysis was limited to the pitch aspect of stress and did not examine duration and intensity, the other phonetic correlates of stress. Without knowledge of what other phonetic cues Mandarin speakers actually use to achieve stress in English, it is difficult to determine (1) whether Mandarin speakers make other phonetic distinctions between word and sentence stress or (2) whether Mandarin speakers actually interpret English stress as a purely tonal event like Mandarin tones. Third, the distinction between pitch movement (sentence stress) vs. pitch height (word stress) was based on the judgment of a single native listener. Because the perception of whether or not there is a pitch movement could be subject to individual variation even among native listeners, it would be desirable to do acoustic measurements in addition to establishing inter-rater reliability by collecting judgements from several trained native English listeners. Fourth, the notion that word stress correlates most closely with pitch height and sentence stress correlates with pitch movement seems to be an overstatement. In English, stress is often but not always realized with higher pitch. Moreover, sentence stress, or accent can assume various tonal shapes, including both level tones and contour tones, depending on the meaning one wants to convey. Fifth, although the materials were taken from the subjects' English textbook, they may still have made errors in stress placement during the recording. To increase the reliability of the results, it would be desirable to collect multiple reading samples from each subject. Finally, Juffs' study only examined stress errors on individual lexical items. Although misplacement of stress in polysyllabic words may contribute in part to their difficulties with English rhythm, such errors largely reflect difficulties with the rhythmic patterns within the words. The data did not adequately address rhythmical problems at the sentence level, which involves not only the prominence relations among syllables within words, but also the prominence relations among words. To tap into the latter aspect of speech rhythm, one would have to look at longer stretches of speech. Anderson-Hsieh and Venkatagiri (1994) examined syllable duration and pausing in the reading samples of three low intermediate and three high intermediate Mandarin ESL speakers and three English native speakers. For syllable duration, they compared native and non-native speakers on two types of syllables representing two durational extremes: (1) tonic syllables witli full vowels before a clause or sentence boundary, and (2) unstressed monosyllabic function words with reduced vowels in non-clause-final positions. Their results showed that although the tonic syllables were longer in duration than the unstressed syllables for all three groups, the tonic syllables were relatively much longer than the unstressed syllables for English native speakers and the high intermediate group than for the low intermediate group. Ratios of tonic syllable and unstressed syllable duration to total sentence duration are 0.37 vs. 0.09 for English speakers, 0.37 vs. 0.08 for the high intermediate group, and 0.24 vs. 0.12 for the low intermediate group. It appears that the low intermediate speakers produce less differentiation between the two types of syllables. The findings of Anderson-Hsieh and Venkatagiri's study have two important implications. First, the durational contrast between stressed and unstressed syllables could constitute a major area of difficulty for Mandarin speakers. Second, the acquisition of stress-timing is possible as proficiency improves, but Mandarin speakers may start out with a more syllable-timed rhythm at early stages. However, we need to be very cautious with the interpretation and generalization of results found for such limited samples. There are also issues that are not adequately addressed in this study. Overall, I found Anderson-Hsieh and Venkatagiri's study limited in three ways. First, other key components of speech rhythm in English, such as pitch and intensity, were excluded from the analyses. Because Mandarin speakers have to make the transition from using pitch for tones to using pitch for stress, the investigation of pitch as a correlate of stress in speech is particularly significant. Second, the sample size was very small. Third, the research design made it impossible to disentangle stress from syllable position as two potential variables affecting duration. Based on the durational underdifferentiation between these two types of syllables, one can not generalize that Mandarin speakers would have the same problems with other types of syllables in other positions. In summary, previous research in this area generally manifests a number of limitations: (1) Pitch and intensity have seldom been taken into consideration despite their proven influence on the perception of stress. (2) The number of subjects has usually been very small. Although acoustic analyses take a tremendous amount of time, it would improve the reliability and generalizability of the results to include a larger of number of subjects. (3) Rhythm has seldom been analyzed at the sentence level or over longer stretches of speech. Most studies examined rhythm of isolated words or highly selected syllables. (4) Few studies have considered possible developmental stages of speech rhythm by comparing learners of different proficiency levels. (5) Specifically, little is known about how non-native speakers realize stress and rhythm in the new language when their native language signals stress in a different way from their second language or 22 when the phonetic cues for stress in the second language are used for different major linguistic functions in their native language. 2.3 CHARACTERISTICS OF SPEECH RHYTHM IN MANDARIN CHINESE Rhythm is one of the least researched areas in Chinese phonology. Little empirical evidence is available to confirm the classification of Mandarin dialects into any of the three recognized categories of speech rhythm (foot-timed, syllable-timed, mora-timed). In this section, I discuss the characteristics of speech rhythm in Beijing Mandarin and Taiwan Mandarin based on recent analyses of their foot structure, the contrasts between full-toned (heavy and stressed) vs. neutral-toned (light and unstressed) syllables, and the phonetic realization of stress. 2.3.1 FOOT STRUCURE IN MANDARIN Duanmu (1999, 2001) proposes that Mandarin consists of left-dominant disyllabic feet, where prominence falls on the first syllable, which must be filled by a heavy (i.e., bimoraic) syllable. A heavy syllable has either a long vowel or a coda, whereas a light (monomoraic) syllable has a short vowel and no coda. Furthermore, he assumes the mora to be the tone bearing unit in Mandarin and that full tones are sequences of two level tones (Duanmu, 1990). Following from these assumptions, heavy (i.e., bimoraic) syllables can carry full tones, but light syllables, which are monomoraic, cannot. The central notion of his proposal is that a minimal word has two levels of metrical structure: the moraic level (which he calls the 'M-foot') and the syllabic level (which he calls the 'S-foot'). He uses 'M-foot' to refer to a bimoraic foot and 'S-foot' to refer to a disyllabic foot. Based on the fact that only the second, but not the initial syllable of a Mandarin disyllabic word can be pronounced with a neutral tone, he proposes that disyllabic feet in Mandarin are left-dominant (see also Lin, 2001). Because a disyllabic foot in Mandarin has prominence on the first syllable, the first syllable must be an 'M-foot'. The second syllable of a foot can be either an 'M-foot' or an unfooted single mora. According to Duanmu's analysis, a metrically acceptable foot in Mandarin consists of either two heavy syllables (heavy-heavy), as shown in example (11a), or a heavy followed by a light syllable (heavy-light), as shown in example (lIb), where foot boundaries are shown in parentheses, 0' is used to represent syllable, and J.l is used to represent mora. (11) a. ( 0' 0' ) (J.lJ.l) (J.lJ.l) heavy heavy b. ( 0' (J.lJ.l) heavy 0') J.l light S-feet M-feet One kind of evidence for disyllabic feet in Mandarin comes from poetry. (12) shows how syllables fall into disyllabic feet. When only one syllable is available to form a foot, a silent beat, represented by 0, is inserted after the syllable to keep the beat. Aside from being realized silently, 0 can also be realized as lengthening on the preceding syllable. (12)(Zhang Lao) (San 0), (WG wen) (nz 0), Zhang Lao San, I ask you, 'Zhang Lao San, let me ask you' (n'i de) (jia xiang) (zed mi) (ti 0) you Gen. home town at which place 'Where is your home town?" (Duanmu, 2001, p.122). (13) shows a similar example from a popular nursery rhyme in Taiwan Mandarin. (13) (san lun) (che 0), (pao de) (kuai 0), three wheel car run fast 'A three-wheel wagon is running fast.' (shang mian) (zuo ge) (lao tai) (tai 0) top surface sit one old lady lady 'With an old lady sitting on the top.' Supporting evidence for disyllabic feet in Mandarin also comes from the fact that there exists a strong preference in all Chinese dialects for an expression to have at least two syllables (Duanmu, 1999). Examples (14) through (17) show that a monosyllabic expression is usually prefixed or suffixed with a semantically redundant syllable. The symbol '*' is used to indicate an ill-formed expression. (14) Personal address *Wang 'Wang' xiaoWang 'Little Wang' la03 Wang 'Old Wang' | *Luo | | LuoCheng | Shanghai | | | | --- | --- | --- | --- | --- | --- | | 'Los Angeles' | | Los Angeles | top City | ocean | | | | | 'Los Angeles' | 'Shanghai' | | | (15) Place names (16) Country names *de deguo Germany Germany Country 'Germany' 'Germany' rlben 'Japan' | *sun | sunZl | | | sunnyu | | | --- | --- | --- | --- | --- | --- | | grandson | grandson | Nominalizer | | grand daughter | | | 'grandson | 'grandson' | | | ,granddaughter' | | (17) Empty morphemes _ v 2.3.2 FULL-TONED VERSUS NEUTRAL-TONED SYLLABLES Following Duanmu's analysis, the contrast between a heavy (bimoraic) and a light (monomoraic) syllable reflects a contrast between a stressed and an unstressed syllable, which must be manifested through the contrast between a full-toned and a neutral-toned syllable. But if we assume that the mora is the tone-bearing unit, the moraic account of foot structure already implies that Mandarin disyllabic feet are formed by sequences either of two full-toned syllables or a full-toned followed by a neutral-toned syllable. It seems redundant to assume an additional stress contrast between the two types of tone, i.e., to assume that stress is a rhythmic primitive for Mandarin. Duanmu (2001) argues for stress as an independent construct in Mandarin by using it to solve the 'word length problem' (p.124). In Mandarin, there are monosyllabic and disyllabic synonyms as shown in (19). | (19)Monosyllabic | | | Disyllabic Gloss | | | --- | --- | --- | --- | --- | | | su~m (garlic) | | 'garlic' da-suan (big garlic) | | | | zhong | (plant) | 'to plant' zhong-zhi (plant plant) | | | The | examples | in (20) | illustrate the word-length constraint | problem in BM and TM, where | | it | is acceptable | to have | a monosyllabic verb (V) followed | (0) by a disyllabic object (20b), | but unacceptable to have a disyllabic V followed by a monosyllabic 0 (20d). Gloss 'to plant garlic' | | a.zhong | | , 'to plant garlic' suan | | | --- | --- | --- | --- | --- | | | b.zhong | | da-suan | | | | c. zhong-zhi | | da-suan | | (20) Verb Object , suan d.*zhong-zhi According to Duanmu, the word length problem can be resolved if one adopts the notion of stress in Mandarin, assuming that a word with more stress should not have fewer syllables than a word with less stress and assuming that 0 must be more stressed than V. Based on these assumptions, 0 may not be shorter than V in a V-O. This analysis provides one solution to the problem. However, it is not necessary to resort to the notion of stress to solve this problem. An alternative is to postulate that V must not be heavier (counting moras) than O. This is to illustrate how difficult it is to justify the reality of stress in Mandarin. In this study, I follow Duanmu's (2001) assumption that Chinese has stress. But it is important to be aware of the fact that there has not been strong evidence or a clear definition for stress in TM or BM. Despite the fact that it is difficult to justify stress in both TM and BM, it is uncontroversial that a neutral-toned syllable is prosodically weaker than a full-toned syllable. A neutral-toned syllable is relatively shorter in duration than a full-toned syllable and is variable in pitch. In order to illustrate the impact of the alternation 26 between full-toned and neutral-toned syllables on the speech rhythm, I now review the general characteristics of full and neutral tones in Mandarin. Dialectal variations regarding the use of neutral tones between BM and TM will be discussed later. There are four full tones and one neutral tone in Mandarin. Among the four full tones, one is a level tone and the other three are contour tones. Table 2.1 Tones in Mandarin Chinese Tone number Description Pitch Tone letter Example Gloss 1 high level 55 1 Ma 'Mother' 2 high rising 35 1 Ma 'Hemp' 3 low falling rising 214 A Ma 'Horse' 4 high falling 51 Ma 'Scold' 5 neutral variable Ma Particle A popular view is that when a syllable is completely unstressed, its default tone disappears and becomes reduced to a neutral tone. When in isolation, the neutral tone is a low plateau close to the bottom of the speaker's pitch range and its duration is relatively short (Chao, 1968; Tseng, 1981). In connected speech, the pitch level of a neutral tone syllable is variable, depending on the pitch value of the preceding tone, except for those in sentence-final position whose pitch level is determined by the intonation (Chao, 1956; Cheng, 1973, Tseng, 1981). When non-final, the pitch contour of a neutral tone is falling following Tone 1, 2, and 4, and rising following Tone 3 (Gao, 1980; Dreher and Lee, 1966). In an acoustic study conducted by Gao (1980), the tonal value of a neutral tone syllable at its starting point is lower than the ending point of a preceding Tone 1 and 2, higher than that of a preceding Tone 4, and continues upward from the pitch movement of a preceding Tone 3. His instrumental analyses indicated the following starting tones for the neutral tone following the four citation tones. 3 after Tone 1 3 after Tone 2 4 after Tone 3 2 after Tone 4 27 In the case of two or more successive neutral tone syllables, Chao (1933) proposed that the pitch level of each is determined by the immediately preceding tone. Thus if the pitch level of the first neutral tone is low, any following neutral tone syllables would also have low pitch. For example, (21) kimjian 'to see' kimjian Ie 'to have seen' kanjian Ie mei you 'to have or have not seen' J J J J J J J However, in running speech, the pitch level of the neutral tone syllables following the first neutral tone syllable is not so much determined by the preceding neutral tone. Instead, it is heavily influenced by the intonation contour of the sentence (Shen, 1994). A neutral tone syllable may occur under two conditions: (1) In polysyllabic words where any full-toned syllable except for the initial one may become reduced to a neutral toned syllable, and (2) in functional categories (Chao, 1932; Shen, 1994). For example, tone loss happens to the second syllable of the disyllabic word pianyi 'inexpensive' (pianyi, if said without tonal neutralization) and the middle syllable of the tri-syllabic word laogudong 'old conservative person' (laogudong without tonal neutralization). Syllables that are toneless are often found in the following grammatical categories of words. (22) Categories Grammatical particles Plural marker Perfective marker Progressive marker Possessive marker Generic classifier Personal enclitic Examples a, ba, ne, ma, ye men as in women 'we', tamen 'they' Ie as in chIle 'have eaten', shulle 'is already asleep' zhe as in kanzhe 'watching', zouzhe 'walking' de as in wade 'my', n'ide 'your' ge as in y{gerin 'a person', y{gemeng 'a dream' zi as in lauzi 'father', ha{zi 'child' Disyllabic words, which constitute the majority of the Mandarin lexicon, may consist of sequences of two full tones or a full tone followed by a neutral tone. According to Chao, the majority of disyllabic words belong to this category, such as xianzai 'now' andjlzer 'egg'2, where the second syllable cannot lose its tone. Trochaic disyllabic words are fewer in number but higher in average frequency of occurrence. They account for words such as mianhua (cotton flower) 'cotton' and y'iba 'tail', where the second syllables have lost their inherent lexical tones 3 due to lack of stress. According to Chao, when words or phrases have three or more syllables, the final syllable has primary stress, the first secondary, and the medial tertiary stress as in huashengtang 'peanut candy', xiiishuobiidao 'stuff and nonsense,4. Chao's position is supported by Yan and Lin (1988), who found that the last syllable in an isolated trisyllabic word is longest. However, their results were later challenged by Duanmu (2001) and Lin (2001) in that the longer duration of the final syllable may be attributed to final lengthening. Wang and Wang (1993) measured the duration of polysyllabic words ranging from one to four syllables placed within a carrier sentence and found that the first syllable tends to be the longest. Their results provide supporting evidence for Lin's assertion that Mandarin polysyllabic words have main stress on the initial syllable. Stress is also used in BM to distinguish minimal pairs of disyllabic words that have identical segmental and underlying tonal composition and differ only in their stress 2 Jizer 'egg' is an idiosyncratic expression used in Beijing Mandarin. In Taiwan Mandarin, 'egg' is referred to as jidan. 3 In Taiwan Mandarin, both syllables in mianhua 'cotton' and y'iba 'tail' are spoken with full tones mianhuii and y'ibii. 4 Xiiishuobiidao 'stuff and nonsense' has a different expression in Taiwan Mandarin, namely hushuobiidao. patterns. There are approximately 200 such minimal pairs in Beijing Mandarin (Chen, 1984; Shen, 1993). A few examples are given in (23). | (23) | | sh'ifii | | 'right | or wrong' | | | | --- | --- | --- | --- | --- | --- | --- | --- | | | | shljei | | 'quarrel' | | | | | | | matou | | 'horse's | head' | | | | | | matou | | 'pier' | | | | | | | shengq'i | | 'get | angry' | | | | | | shengqi | | 'vitality' | | | | | | | day'i | | 'main | point' | | | | | | dayi | | 'careless' | | | | 2.3.3 IDIOSYNCRASIES OF LEXICAL RHYTHM ACROSS DIALECTS There are idiosyncrasies across various dialects of Chinese Mandarin when it comes to the rhythmic patterns of polysyllabic words. Unlike Beijing Mandarin where any syllable except the first may be pronounced with a neutral tone, Taiwan Mandarin seldom reduces any syllables in polysyllabic words to neutral tones. As a result, most of the neutral tone syllables in Taiwan Mandarin are grammatical morphemes (see examples in 22 above). Many of the common trochaic disyllabic words in Beijing Mandarin are spoken with full tones on both syllables in Taiwan Mandarin. | | Beijing | | Mandarin | | Taiwan Mandarin | Gloss | | | --- | --- | --- | --- | --- | --- | --- | --- | | | zhldao | | | | zhldao | 'know' | | | | ben3shi | | | | bensh'i | 'ability' | | | | m{ngbai | | | | m{ngbdi | 'understand' | | (24) Disyllabic words The stress rule that typically applies to polysyllabic words of three or more syllables in Beijing Mandarin is often non-existent in Taiwan Mandarin. In Beijing Mandarin, tone loss usually happens to the middle syllable of a tri-syllabic word. For example, the tri-syllabic word laushouzhang 'old senior officer' is usually said with a middle neutral-tone syllable laushouzhang, laugudong 'old conservative person' as lagudong, and guoluke 'passerby' as guoluke (Chao, 1948; Shen, 1994). In contrast, words like these are spoken with full tones on all syllables in Taiwan Mandarin, with appropriate tone sandhi 5 • 5 In Mandarin Chinese, when a 3rd tone is followed another 3 rd Tone, the first one is usually realized as a rising tone. | | Beijing | | Mandarin | Taiwan | | | Mandarin | Gloss | | | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | | laushouzhang | | | laushouzhang | | | | 'old | | senior officer' | | | | laugudong | | | laugudong | | | | 'old | | conservative person' | | | | guoluke | | | guolitke | | | | 'passerby' | | | | | | | Moreover, | stress | does | | not | seem to | be contrastive | | in Taiwan Mandarin. | No minimal | | pairs | | of words | are | distinguished | | | solely | on the | basis | of stress. The minimal | pairs of words | | distinguished | | | by contrastive | | | tones | in Beijing | Mandarin | | as shown in | (26) are homophones | (23)Tri-syllabic words with full tones on both syllables in Taiwan Mandarin. | (26) | | Stressed-based | | minimal | | | pairs | | | | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | | | Beijing | Mandarin | | Taiwan | | Mandarin | | | Gloss | | | | | shlfii | | | sh'ifii | | | | | 'right or wrong' | | | | | shlfei | | | shlfii | | | | | 'quarrel' | | | | | matou | | | matou | | | | | 'horse's head' | | | | | matou | | | matou | | | | | 'pier' | | | | | shengq'i | | | shengq'i | | | | | 'get angry' | | | | | shengqi | | | shengqi | | | | | 'vitality' | | | | | day'i | | | day'i | | | | | 'main point' | | | | | dayi | | | day'i | | | | | 'careless' | | | | | It is | possible | for both | | Beijing | | and Taiwan | | Mandarin to have | two consecutive | It is possible for both Beijing and Taiwan Mandarin to have two consecutive neutral-toned syllables in the middle or at the end of a sentence. For example, | (27) | | jiejie | de shu | | | 'older | sister's | book' | | | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | | | gege | de pengyou | | | 'older | brother's | friend' | | | | | | | nz | kanzhe ba | | | 'You'll | see.' | | | | | | | | However, | it | is very | rare | | in Taiwan | Mandarin | | to have three | consecutive neutral | However, it is very rare in Taiwan Mandarin to have three consecutive neutral toned syllables except when the last one is a sentence final particle. For example, (28) daoshlhou, fangzi jiush'i baba de Ie. 'At that time, the house will have belonged to Daddy.' Overall, it is more common to have two or more consecutive neutral tone syllables in BM than in TM due to the large number of neutral-toned syllables in polysyllabic words in BM. Long stretches of consecutive neutral tone syllables in BM are often interrupted by full tones in TM as described above. (29) BM: kanjian Ie 'to have seen' TM: kanjian Ie BM: kanjian Ie mei you 'to have or have not seen' TM: kanjian Ie mei you BM: daban Ie 'to have dressed up' TM: daban Ie BM: TM: daban Ie mei you daban Ie mei you 'to have or have not dressed up' In summary, although both BM and TM allow the same types of foot structure (i.e., heavy-heavy and heavy-light), a heavy-light foot is much more common in BM than in TM because tonal neutralization is much less frequent in TM. 2.3.4 THE PHONETIC CORRELATES OF STRESS IN MANDARIN CHINESE Chao (1968, p.35), based on his auditory impressions, first described tonic stress in Mandarin Chinese as "primarily an enlargement in pitch range and time duration and only secondarily in loudness." His position later gained support from several instrumental analyses (Howie, 1976; Tseng, 1981; Lin, Yan, & Sun, 1983; Shen, 1994). When stressed, high level tones are pushed higher, rising tones rise higher, dipping tones dip lower, and falling tones start higher and drop deeper. Moreover, the stressed syllables also have longer duration to realize the extended Po movement. Shen (1993) analyzed the 33 natural speech of a female Beijing Mandarin speaker and the reiterated speech of three Beijing Mandarin speakers and reported a duration ratio of approximately 3:2 and a difference of nearly 8 dB between stressed and unstressed vowels of the same quality. In addition, contrastive stress is primarily achieved by changes in duration. Note that there are two kinds of stress that need to be distinguished. One is lexical stress, which is associated with the presence or absence of tones. The other is emphatic stress, which is achieved by increasing pitch range, length, and loudness. To illustrate this contrast, an example is given in Figure 2.1 showing a female Mandarin speaker saying wobaba 'my father' first without and second with emphatic stress on the middle syllable. afflo 4Stll~' '- ~~~~~"""'~" ,~'~-,~"" ",~, ""~~""",, '~'~"" "~~,,~~,~~~,,, ....9 ';'. '" "~"'~~ "~,," I 40(J1'~~""~~ ""',', ,,-~'" ,~""-,-,," ~~"" '''~~''''''''''''''''''~~-'',", ,,~~~~,-,,~-"-'" "~" ~'I .9 . 35tl "-"",,~,,, ,-~, -'" '~-"'~-""', ","~, ,~~, ,,,--,,-,,,~,",~, " "'-~-'" ,'-~,,~,~~,~~~, , ~'. """""'"'''''' ,~~~'" I -.... .''"'' - 2001;"""""-""""'" ",~~~""",,'"",-~"""""""-~~"",~~"""""",--~"",,,"""",""""--""""""""""'---,""-""~'-'""""'i""~>""'~"'--'~"I !5tll~'=·0/6"'·-'~~"-"~b'"a."-~"-;i~b~a.~'-"'~-I-"""~~-~~"--'-~il-'-'--~;'··'wy"'''=.·----T"ba.'~--------:ib;a-;'''---I Figure 2.1 Lexical stress vs. emphatic stress in Mandarin Chinese The middle syllables in both utterances in Figure 2.1 carry lexical stress but the middle syllable in the second utterance also carries emphatic stress. Because both are lexically stressed, both were said with full tones. However, the emphatically stressed syllable was said with a much broader pitch range (192Hz vs. 50Hz) with the Fo contour starting at 385Hz, peaking at 476Hz, and ending at 284Hz. Compare this with the same syllable without emphatic stress, whose Fo starts at 290Hz, peaks at 295Hz, and ends at 34 245Hz. The emphatically stressed syllable is also slightly longer (238 ms vs. 208 ms) and louder (peak intensity 56 dB vs. 51 dB) than its non-emphatically stressed counterpart. Although pitch serves as a major correlate of emphatic stress in BM and TM, Fo information is not necessary in order for stress to be perceived. In her perception tests, Shen (1993) asked four Beijing Mandarin speakers to listen to natural and reiterant speech under three different conditions: (1) Low-pass-filtered, with a cutoff frequency of 400 Hz to eliminate segmental information, (2) Monotoned, with the Fo of the already filtered speech held constant at 135 Hz to eliminate Fo variations, and (3) Fixed intensity at a constant 60 dB so the only remaining cue is length. The subjects were then asked to circle syllables that they perceived as stressed. The results showed that the Chinese speakers perceived stress equally well under all three conditions. Neither the presence of Fo nor intensity variations changed the stress judgments significantly. She also found much stronger correlations between duration and stress judgments than between intensity and stress judgments, suggesting that duration is a more important perceptual cue of stress than intensity for Mandarin Chinese speakers. 2.3.5 ARE BM AND TM MORA-TIMED, SYLLABLE-TIMED, OR FOOT-TIMED? From this perspective, how do the rhythms of BM and TM be categorized? Are they mora-timed, syllable-timed, or foot-timed? First, they cannot be mora-timed because evidence from poetry as shown in (11) and (12) indicates that BM and TM are syllable counting (i.e. each foot must consist of two syllables, regardless of the second syllable being heavy or light), not mora-counting. Second, they cannot be syllable-timed, because both BM and TM maintain a distinction between heavy and light syllables. Third, if BM and TM are foot-timed, given that a foot can consist of either four moras (heavy-heavy), or three moras (heavy-light), I would expect their heavy-heavy feet and heavy-light feet to show approximately equal lengths. That is, the initial heavy syllable in a heavy-light foot would tend to be longer than that in a heavy-heavy foot. 2.4 POTENTIAL DIFFICULTIES TAIWAN MANDARIN AND BEIJING MANDARIN SPEAKERS MIGHT HAVB WITH THE PRODUCTION OF ENGLISH SPEECH RHYTHM In order for a second language learner to produce a native-like rhythm in English, at least three important aspects of speech rhythm must be learned. First, one must place stress cues on the appropriate syllable(s). Second, one must employ the same phonetic cues to signal stress as native speakers. Although there are no invariant phonetic correlates for stress and no two stressed syllables are realized in exactly the same way, there are strong tendencies for stressed syllables to be longer, louder, and/or broader in pitch range. And third, one must produce native-like differentiation between stressed and unstressed syllables so that the rhythmical beats can be distinctively perceived. It is true that stress isochrony does not only mean lengthening stressed syllables and shortening unstressed ones. It means producing approximately equal time between stresses. However, because isochrony is only a mental target and phonetic isochrony is rarely established empirically in production, I will focus our investigation based on the three proposed criteria. The assumption is that if an ESL or EFL learner is capable of achieving all three, it would be a good indication that s/he is able to produce near native-like English speech rhythm. While the speech rhythm of English has been widely researched, little empirical evidence is directly available to either characterize or classify the speech rhythm of Taiwan Mandarin or Beijing Mandarin into any of the three recognized rhythmic categories. However, I might be able to glean some ideas about the potential difficulties TM and BM speakers might have with the production of English speech rhythm from what I have learned about the two languages in sections 2.2 and 2.3. First, TM speakers should be able to use duration, intensity, and pitch as correlates of stress in English with at least some degree of success because the same correlates are used to signal stress in TM (see Figure 2.1). However, because of the strong predisposition to pitch in their first language, TM speakers may rely more on pitch in producing the distinction between stressed and unstressed syllables than on duration and intensity. If this prediction is true, I would expect TM speakers to produce similar or greater pitch contrasts but smaller duration and intensity contrasts between stressed and unstressed syllables than English speakers do. Second, because most syllables are produced with full tones in Mandarin, TM speakers might tend to assign tones to syllables. If they do, they might have difficulty with reducing unaccented syllables, including those that bear some stress and those that are totally unstressed. Given that toned syllables must be spoken with a minimal length in order for the tones to be sufficiently realized, which could limit the possible variations in duration across syllables. Also because accented syllables must also be stressed i~ English, by assigning tones to syllables, TM speakers would likely be perceived as assigning stresses to syllables. Note that the distinction between accented and unaccented syllables is not necessarily realized within a foot in English because not all stressed syllables bear pitch accents. This would suggest that TM speakers' difficulty with English speech rhythm (assuming that there are as many levels of rhythm as levels of stress) may be realized at the foot level (when the distinction between accented and unaccented syllables coincide with the distinction between stressed and unstressed syllables) or at a level above the foot level (when the distinction is between accented syllables and unaccented weakly stressed or unstressed syllables). (30) x (Swwww) x x x x x x x Mom made itfor me. Consider the example in (30), where the main stressed syllable mom, shown in boldface, is stronger in relation to all the other syllables (made, it, for, me) in the sentence. The distinction between syllables that bear main stress and those that do not is hereafter defined as 'Strong' (main-stressed) syllables versus 'Weak' (secondary-stressed and unstressed) syllables on a metrical grid in the sense that main stressed syllables are relatively stronger than the others. At this level of the metrical grid, it is possible in English to have stretches of relatively weaker and tonally unspecified syllables, which could be potentially difficult for TM speakers. TM Speakers, in particular, may experience greater difficulties with the production of English speech rhythm than BM speakers do. Although both BM and TM consist of left-dominant disyllabic feet, alternating full-toned (stressed) and neutral-toned (unstressed) syllables in polysyllabic words are much less common in TM than in BM. From this perspective, TM may sound syllable-timed because most of the syllables in polysyllabic words are produced with full tones. In addition, TM speakers may sound relatively syllable-timed when they speak English because they may tend to produce most syllables with full tones. This makes TM and English a more interesting pair of languages to compare in terms of rhythm. However, it is also possible that TM speakers may have little problem with the alternating stress pattern in English given that a Strong-Weak foot structure is legitimate in TM. If this is true, TM speakers would have little trouble realizing the alternating strong and weak syllables in English, particularly when the weak syllables consist of unstressed function words as they usually are in TM. Realizing the hierarchical organization of stresses in English could be a challenge for TM speakers. It is likely that TM speakers could produce a more English-like rhythm when words are produced in isolation as compared with the phrase or sentence level where reorganization of prominence relations is required. For example, they might have less trouble producing correct stress patterns for elevator and operator as two isolated words than as a single phrase. Instead of assigning greater prominence to the first than to the second word as shown in (31a), they might assign equal prominence to both words as shown in (31b). (31) a. s w / \ / \ S wSw /\ /\ /\ /\ SwSwSwSw elevator operator | | /\ /\ | /\ /\ | /\ /\ /\ /\ | | | --- | --- | --- | --- | --- | | | SwSwSwSw | | SwSwSwSw | | | | elevator | operator | elevator operator | | b. S S / \ / \ S wSw If this were true, they would produce stress on all syllables bearing main lexical stress at the same level. If these predictions are correct, it appears that unaccented ('weak') syllables (including those that bear some stress and those that are completely unstressed) would pose a great challenge to Taiwan Mandarin speakers. And rhythmic patterns of connected speech would be more difficult than words in isolation. Because TM allows alternating full-toned and neutral-toned syllables when the neutral-toned syllables are also grammatical morphemes, TM speakers may be able to produce alternating strong and weak (in this case, unstressed grammatical morphemes) syllables with some degree of success. Thus potentially, sentences that contain long stretches of weak syllables may be rhythmically more difficult than those that contain alternating weak syllables. CHAPTER 3: THE PRESENT STUDY 3.1 PURPOSE OF THE STUDY The current study investigates Taiwan Mandarin (TM) speakers of two different proficiency levels with respect to their difficulties in producing English speech rhythm by analyzing three well-attested correlates of stress in English - duration, intensity, and pitch - of Strong and Weak syllables in non-final versus final position over complete sentences. The purpose is to correct some of the limitations of previous research on the acquisition of English speech rhythm (as noted in 2.2.2) and to provide physical evidence that would help identify difficulties, if any, in two key components of speech rhythm: (1) the physical realization of stress, i.e., whether TM speakers are able to use duration, intensity and pitch effectively in the discrimination between syllables that have more stress versus syllables that have less stress, and (2) the distribution of stress, i.e., whether TM speakers are able to focus stress on the appropriate syllables. ### 3.2. RESEARCH QUESTIONS Based on Anderson-Hsieh and Vekatagiri's findings and the characteristics of English and Mandarin speech rhythms as reviewed in 2.2 and 2.3, I have six general predictions. More formally they are: Prediction 1 TM speakers should be able to use duration, intensity, and pitch as correlates of stress in English with at least some degree of success. Among the three variables, TM speakers may rely more on pitch in producing stress distinction than on duration and intensity. Prediction 2 TM speakers might tend to assign tones to syllables and therefore have difficulty with reducing unaccented syllables, including those that bear some stress and those that are totally unstressed. Prediction 3 TM speakers may produce less differentiation between strong and weak syllables than do English speakers. Prediction 4 TM speakers may have little problem realizing the alternating strong and weak syllables in English, particularly when the weak syllables consist of unstressed function words. Prediction 5 Realizing the hierarchical organization of stresses in English could be a challenge for TM speakers. Prediction 6 TM speakers who have a higher overall proficiency in English may produce a more English-like rhythm than those who have a lower proficiency in English. With these general predictions in mind, the current study will aim at addressing five research questions. In the following five subsections, 3.2.1 through 3.2.5, we will discuss specific predictions for each question and the kinds of data needed to answer them. 3.2.1 DO TAIWAN MANDARIN SPEAKERS HAVE DIFFICULTIES WITH THE DURATION, INTENSITY, OR PITCH OF STRONG SYLLABLES, WEAK SYLLABLES, OR BOTH? If Prediction 2 is correct, reducing weak syllables is expected to be more difficult for TM speakers than strengthening strong syllables, where 'Strong' refers to syllables that bear main stress (including nuclear-pitch-accented and non-nuclear pitch-accented) and 'Weak' refers to those that are unaccented (including those that bear some stress and those that are totally unstressed). TM speakers may find it difficult to reduce weak syllables in English partly because they may tend to assign tones to syllables, and pitch accents on syllables are often perceived as stressed by English speakers. In order to test this prediction, we will examine their strong and weak syllables in two types of sentences. Type A sentences feature long stretches of weak syllables and Type B sentences feature a highly regular rhythmic pattern of alternating strong and weak syllables. We will compare TM and English speakers with respect to the distribution of identifiable differences in strong and weak syllables in these two types of sentences. Specifically, we will focus on those differences in each of the three variables that usually weaken the contrasts between strong and weak syllables, namely, relatively shorter, softer, lower-pitched strong syllables versus relatively longer, louder, and higher-pitched weak syllables. Strong syllables which are longer, louder, and higher-pitched than those of English speakers or weak syllables which are shorter, softer, and lower-pitched are not considered to be difficulties because they strengthen rather than weaken the target rhythmic pattern by broadening the contrast between strong and weak syllables. As supporting evidence, we will also compare TM and English speakers with respect to the relative duration, intensity, and pitch of the strong versus weak syllables.. 3.2.2 DO TAIWAN MANDARIN SPEAKERS PRODUCE LESS DIFFERENTIATION IN DURATION, INTENSITY, OR PITCH BETWEEN STRONG AND WEAK SYLLABLES THAN ENGLISH SPEAKERS? If Prediction 3 is correct, I expect TM speakers to experience difficulties producing a native-like differentiation between strong and weak syllables. Furthermore, if Prediction 1 is also correct, TM speakers are expected to have greater difficulty realizing duration and intensity contrasts than realizing pitch contrasts between strong and weak syllables. It might be easier for TM speakers to discover a correlation between pitch and stress in English because they are predisposed by their first language to be sensitive to variations in pitch. If this is true, I expect them to produce a pitch differentiation between strong and weak syllables, which is similar to or perhaps greater than that of English speakers. There are two ways of measuring duration, intensity, and pitch contrasts between strong and weak syllables. Each approach tells only part of the story. One is to measure the contrasts between target (i.e., expected) strong and weak syllables. This provides a general index of how far TM speakers are from being native-like in a particular sentence. The other is to measure the contrasts between actual strong and weak syllables produced by TM and English speakers. This tells us whether TM speakers produce a greater or smaller differentiation between their strong and weak syllables than English speakers, regardless of any differences that they might have from English speakers in terms of the distribution of stress. However, the latter distinction is hard to determine based on physical measurements of duration, intensity, and pitch of syllables alone. As an alternative to the latter approach, we will try to identify possible factors (i.e., lexical categories such as content vs. function words) which might better predict the placement of strong vs. weak syllables in the speech of TM speakers when the expected strong and weak syllables do not match the actual strong and weak syllables. 3.2.3 DO TAIWAN MANDARIN SPEAKERS CORRELATE DURATION, INTENSITY, AND PITCH WITH THE TARGET STRESS PATTERNS IN ANENGLISH-LIKE WAY? Even when the same sentences are presented to TM and English speakers within a context that strongly favors a particular stress pattern, there is no guarantee that TM speakers will produce the expected stress pattern. If Prediction 5 is correct, TM speakers are expected to have difficulty realizing the prominence relations among syllables of a sentence in production. Without assuming that TM speakers intend to produce the same stress patterns as English speakers, the current research question is indeed twofold. The first dimension to this question is whether TM speakers put main stress on the right syllable. The other dimension to this question is whether TM speakers correlate duration, intensity, and pitch with stress, no matter how different their stress patterns might be from those of English speakers. For the latter, the goal is to identify possible variable(s) that might influence the distribution of stress in the speech of TM speakers. For example, given that toneless syllables are usually function morphemes in Taiwan Mandarin (see 2.3.2), do they tend to associate lack of stress with monosyllabic function words and stress with monosyllabic content words in English? We want to examine the extent to which variations in duration, intensity, and pitch correlate with the target stress patterns in Type A and Type B sentences. First we will trace the fluctuations of duration, intensity, and pitch across syllables by counting the number of strong and weak syllables classified as relatively longer, louder, and higher pitched versus those that are relatively shorter, softer, and lower-pitched than their preceding syllables. In cases where TM and English speakers differ in their placement of stress, we will investigate whether other variable(s) (e.g., function words vs. content words) better correlate with the placement of stress for TM speakers. 3.2.4 DO TM SPEAKERS PRODUCE MORE ENGLISH-LIKE DURATION, INTENSITY, AND PITCH PATTERNS WITH IMPROVED PROFICIENCYAND EXPOSURE TO ENGLISH? If Prediction 6 is correct, TM speakers are expected to produce a more English-like duration, intensity, and pitch patterns as their overall proficiency in English and exposure to an English-speaking environment improve. In order to compare two proficiency groups of TM speakers with respect to their degree of success in achieving the target rhythmic patterns using duration, intensity, and pitch, we will examine the following two aspects of their stress patterns: (1) the physical realization of stress, and (2) the placement of stress. We mentioned in Chapter Two that some of the prosodic cues used to mark rhythmic impulses in English are employed for other linguistic functions in TM. Although increased pitch range is used to signal emphatic stress (See Figure 2.1), variations in pitch within a syllable are primarily used for discriminating lexical tones in TM (see Table 2.1). And duration is often linked with the contrast between full-toned and neutral-toned syllables. We will investigate how successful TM speakers are in correlating pitch and duration with prominence in English at different stages of acquisition. Second, as reviewed in 2.3.3, we have seen that at the lexical level stress is a feature of polysyllabic words in English but in Taiwanese Mandarin polysyllabic words are often spoken with full tones on all syllables. At the sentence level, any lexically stressed syllables in English have the potential for bearing sentential stress. It is possible to pronounce the same English sentence with a variety of stress patterns depending on which syllable(s) are emphasized. Therefore, in order to produce native-like English speech rhythm with any success at the sentence level, TM speakers must also be able to focus stress on the appropriate syllables. Answers to these two research questions will be drawn from a wide variety of sources. First, we will compare TM learners of English of two different proficiency levels with respect to the basic statistics of the relative duration, intensity, and pitch of individual syllables, induding standard deviations, ranges and means where appropriate. The information will give a general picture of how native-like the two groups of TM speakers are with respect to each variable. Second, we will compare the two groups of TM speakers with respect to the number of strong syllables which are shorter, softer, and lower-pitched and the number of weak syllables which are longer, louder, and higher pitched than those of English speakers. The results will provide crucial information about the degree of difficulty each group has with the duration, intensity, and pitch of strong versus weak syllables. Third, we will compare the two groups of TM speakers with respect to the number of strong versus weak syllables whose duration, intensity, or pitch values rise or fall from their preceding syllables. With this information, we will be able to gauge the extent to which their duration, intensity, and pitch patterns conform to the target stress patterns. Fourth, we will compare the two groups of TM speakers with respect to the correlation coefficients of their duration, intensity, and pitch across the syllables of the sentences with those of English speakers. A higher correlation will indicate a more native-like pattern. Finally, we will compare the two groups of TM speakers and the English native speakers with respect to the relative degree of differentiation between strong and weak syllables in each of the three variables. The results will provide information about the extent to which TM speakers produce native like differentiation between strong and weak syllables at different stages of acquisition. 3.2.5 ARE DURATION, INTENSITY, AND PITCH COORDINATED AS CORRELATES OF STRESS FOR TAIWAN MANDARIN SPEAKERS? After we have analyzed how successful TM speakers are in using each of the three variables for achieving the target stress patterns, an important question to ask is how well the three variables are coordinated among TM speakers. The purpose of this investigation is to identify any similarities and/or differences between TM and English speakers about the ways and the extent to which duration, intensity, and pitch coordinate with one another in the realization of stress. Three different measures will be used for comparing TM and English speakers with respect to the coordination among these three variables. First, we will compare duration, intensity, and pitch in terms of the number and type of differences each shows between English native speakers and the two groups of TM speakers. Second, we will compare 46 duration, intensity, and pitch with respect to their variations across the syllables of the sentences. Specifically, we will count the number of strong versus weak syllables whose duration, intensity, and pitch values rise or lower from their preceding syllables, assuming that syllables that are longer, louder, and higher-pitched are more stressed and syllables that are shorter, softer, and lower-pitched are less stressed. Third, we will compare duration, intensity, and pitch in terms of the differentiation between strong and weak syllables for each of the three subject groups. The results will provide information about whether the three variables present similar problems to TM speakers, whether the three variables tend to rise and fall together according to stress, and whether variations in each of the variables are consistent with one another in creating a native-like differentiation between strong and weak syllables. Last but not least, if Prediction 4 is correct, I expect TM speakers to show little trouble with strengthening strong syllables, weakening syllables, producing a native-like contrasts between strong and weak syllables, correlating duration, intensity, and pitch with the target stress patterns, and coordinating duration, intensity, and pitch in the production of sentences that feature alternating strong and weak syllables in English, particularly when the weak syllables consist of unstressed function words. 3.3 METHOD 3.3.1 SUBJECTS | Table 3.1 | Profile of | the three | subject | | | | groups | | | | | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | Group | | | Number | | | | | Age | Onset Age | of English Time | in US | | | | | All | | M | | F | | (years) | Instruction | (years) (years) | | | | Control | English Native | 10 | | 5 | | 5 | | 24.8 | - | | - | | | Experimental | TMESL | 10 | | 5 | | 5 | | 27.6 | 11.9 | 3.8 | | | | | TMEFL | 10 | | 5 | | 5 | | 18.9 | 11.5 | | - | | The participants in this study were 10 native speakers of English, 10 TM ESL learners, and 10 TM EFL learners. The English speakers were either graduate or undergraduate students at the University of California in Irvine or at the University of Hawai'i and all came from the US west coast. None of them were raised as bilinguals or considered themselves as fluent in any foreign language. The average age of the English native speakers was 24.8. The two groups of TM speakers were distinguished from each other along two criteria: (1) exposure to an English-speaking environment, and (2) general proficiency. The members of the ESL group were 10 TM learners of English from Taiwan, who were enrolled in a variety of graduate programs at the University of Hawai'i. Nine out of the 10 speakers provided age information. The average length of stay in the United States was three years nine months. The average age of the TM ESL speakers was 27.6 and the average onset age of English instruction was 11.9 years old. The members of the EFL group were 10 TM learners in Taiwan, who were all freshmen students enrolled at the Tamkang University in Taipei and none of them had ever been to any English-speaking countries prior to the experiment. The average age of the TM EFL speakers was 18.9 and the average onset age of English instruction was 11.5 years old. Each of the three subject groups consisted of five males and five females. 3.3.2 MATERIALS The materials include an Agreement to Participate Form, a short questionnaire, an instruction sheet, and two stacks of test cards. The current project and its official Agreement to Participate Form were reviewed and approved by the Committee on Human Studies of the University of Hawai'i. The Form described specifically what the subjects were asked to do for the experiment, including the kinds of data that were collected and the ways in which these data were collected, stored and used. Two separate questionnaires were developed for each language group, with one written in English and the other in Chinese. The content of the questionnaires generally overlapped except for a few questions, which were relevant only to specific language 48 groups. For instance, the English speakers were asked about their home state and whether they had lived in any non-English speaking countries. The TM speakers were asked about their age at the onset of English instruction, their length of stay in the United States or any other English-speaking countries, and their TOEFL scores, if available. The instructions were also presented in the subject's native language, with the English version printed one side of a piece of laminated A4 paper and the Chinese version on the other side. The purpose was to ensure that the participants understood the procedures clearly. The two versions were identical in terms of content. Both included step-by-step directions, followed by a few practice examples in English. The English examples were designed to familiarize the subjects with the type of test stimuli that they would be asked to read and to preclude any priming effect as a precaution that TM speakers might tend to process the stimuli using their own native language after reading the instructions in Chinese. The test stimuli comprise two prosodically diverse sets of sentences with seven tokens in each set. Type A sentences feature long stretches of weak syllables. Each of these sentences is narrow-focused and consists of either a single strong syllable or two widely spaced strong syllables. The classification of 'Strong' versus 'Weak' is defined at the top level of the metrical grid where a syllable is either 'Strong' (bears pitch accent) or 'Weak' (bears no pitch accent). The distinction is made based on the prediction that TM speakers may experience difficulty in reducing unaccented syllables (see discussion in section 2.4). The weak syllables include both unaccented syllables that bear some degree of stress and totally unstressed syllables. The example in (32) shows how syllables are classified into 'Strong' (S) or 'Weak' (W) on a metrical grid for a Type A sentence, where the syllables it, for, me are weak in relation to made, which is weak in relation to Jane, which bears main stress. This results in one strong and four weak syllables. (32) S w w w w x x x x x x x x Jane made it for me Four subsets of rhythmic patterns are represented for Type A sentences. Sentences Al and A2 have the five-syllable rhythmic pattern Swwww, sentences A3 and A4 have the seven-syllable rhythmic pattern SwwwwwS, sentences A5 and A6 have the seven syllable rhythmic pattern wSwwwww, and sentence A7 has the eight-syllable rhythmic pattern SwwwwwSw. Type B sentences are broad-focused and feature a highly regular rhythmic pattern of alternating strong and weak syllables where each strong syllable is immediately preceded and followed by a weak syllable. Three different rhythmic patterns are represented in the Type B sentences. Sentences BI through B4 have the six-syllable iambic (in a poetic sense) pattern wSwSwS, sentences B5 and B6 have the seven-syllable trochaic pattern SwSwSwS, and sentence B7 has the eight-syllable iambic pattern wSwSwSwS. Furthermore, the alternating rhythmic pattern featured in Type B sentences was reinforced by choosing lexical categories so that all strong syllables are content words, which are more likely to bear stress than function words, which make up the weak syllables. Each test sentence was embedded within a brief context designed to help the subjects arrive at the same interpretation of the test sentences so as to elicit the target rhythmic patterns. Examples of these two types of test sentences with their contexts are given in (33) and (34) (see Appendix A for a full list of all sentences used in the experiment). The test sentences are shown in italics and the expected strong syllables are shown in boldface. In the actual experiment, the subjects were not made aware of the purpose of the study, nor of which sentences were being tested, nor which syllables were targeted as strong (i.e., No syllables were shown in italics or boldface in the actual experiment). (33) Type A This is a beautiful bag! Did Mary make it for you? No, Jane made itfor me. (34) Type B Well, I didn't know what happened. But anyway, Mom and Dad were mad at Jim. Efforts were made to keep the lead-in sentences casual in order to encourage a natural speech style. For examples, we tried to incorporate colloquial expressions such as Well, Anyway, and Sure into the contexts for the test sentences. However, we deliberately avoided using formulaic expressions, (i.e., everyday phrases that learners pick up as chunks) such as How are you?, Where are you going?, Where're you from ?, That's what I thought., I donna., See you later., etc. Non-native speakers can often sound near native like when saying these commonly used expressions even though they might have problems with English rhythm in their non-formulaic speech. Several criteria were taken into consideration in the construction of the experimental sentences. First, to minimize controversies over syllabification, the majority of the test sentences were composed of monosyllabic words. English speakers do not always agree about how syllables should be divided in multi-syllabic words. Treiman and Danis (1988) found that native speakers of English disagreed about where to place the syllable boundary when asked to segment words such as honest, happy, letter, and ladder into syllables. Syllabification in an interlanguage could be even more complicated. Some of the phonological processes that help define syllable boundaries may not be present in the interlanguage. For instance, the aspiration of stops in English often indicates they are syllable-initial. However, TM speakers are known to produce aspirated stops in contexts where American English speakers normally would not, such as the second Itl in the word tomato, or the Itl in kitty. Second, efforts were made to choose sequences of words that would make the segmentation reliable. We deliberately avoided sequences of highly similar segments at syllable boundaries, such as "wan[t t]o," where it is almost impossible to determine the boundary between the two sounds within the bracket. Where possible, we also tried to avoid vowel-vowel or vowel-glide sequences because the transition between vowels is so 51 gradual that it is very difficult to pinpoint the boundary. In addition, we preferred combinations of vowels and stops or vowels and fricatives at word boundaries because they involve distinct changes of acoustic events to help with the segmentation. Third, we tried to use words that would generate clear pitch tracks. For instance, we tried to use words with as many voiced segments as possible because pitch information for voiceless sounds is very difficult to extract. We also preferred stops to fricatives, which are relatively more difficult to extract pitch from due to their high-frequency noise. In general, the more sonorant a sound is, the better it is for pitch extraction. Fourth, we tried to avoid sounds and sound sequences that are known to be difficult for TM speakers because the purpose of the study was not to test the pronunciations of sounds. We tried to avoid English words with sounds that are not present in TM, such as the interdental fricatives 18,01, the postvocalic alveolar lateral approximant Ill, and the voiced palato-alveolar fricative 131. Efforts were also made to avoid words with heavy initial or final consonant clusters because there are no consonant clusters in TM except for consonants followed by glides. Fifteen sentences were originally formulated for each sentence type, and the seven that elicited the most consistent results in an informal pretest with 10 English native speakers were selected as the final test sentences. The final test sentences were then reviewed by three assistant professors teaching EFL at three different universities in Taiwan and informally pretested with two freshmen students from each of their classes to make sure that the sentences were appropriate for their proficiency levels. The 14 test sentences are given in Table 3.2, where the target strong words are shown in boldface to show the rhythmic patterns because it is hard to infer the expected rhythm when the test sentences are removed from their lead-in sentences. Note that the actual test sentences were presented to the subjects without boldfacing. A complete list of the sentences used in this study including the lead-in sentences can be found in Appendix A. Table 3.2 Labels, stress patterns, and syllable numbers of test sentences | Type | Item | | Test Sentence | | Rhythm | | | Syllable | Total | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | A | Al | | Jim wrote it with | me | Swwww | | | 5 | | | | A2 | | Jane made it for | me | Swwww | | | 5 | | | | A3 | | You like me to | wear the jeans | SwwwwwS | | | 7 | | | | A4 | | You want me to | bring the wine | SwwwwwS | | | 7 | | | | A5 | | My mom made | the lemon pie | wSwwwww | | | 7 | | | | A6 | | The old man gave | it to me | wSwwwww | | | 7 | | | | A7 | | Jane's the one wearing | the blue dress | SwwwwwSw | | | 8 | | | B | BI | | I need it back by | noon | wSwSwS | | | 6 | | | | B2 | | They play with | dad and laugh | wSwSwS | | | 6 | | | | B3 | | We learn to read | and write | wSwSwS | | | 6 | | | | B4 | | I met with John | at noon | wSwSwS | | | 6 | | | | B5 | | Mom and dad were | mad at Jim | SwSwSwS | | | 7 | | | | B6 | | John is good at | baking bread | SwSwSwS | | | 7 | | | | B7 | | I need a ride to | work at once | wSwSwSwS | | | 8 | | 3.3.3 PROCEDURES All subjects were recorded individually in a quiet place using an Audio-Technica ATR20 unidirectional microphone connected to a Toshiba laptop computer. The utterances were digitally recorded using Speech Analyzer v. 1.06 at 16 bits with a sampling frequency of 1l025Hz. The subjects were instructed to speak into the microphone and hold the microphone close to their mouths. Prior to recording, subjects were first given an opportunity to read the consent form and sign it if they agreed to participate in the experiment. The experiment did not proceed without their consent. After that, they were requested to fill out a short questionnaire in their native language, which collected anonymously coded biographic information. Upon completion of the agreement form and the questionnaire, the subjects were presented with instructions printed on a piece of laminated A4 size paper written in their native language and were given the opportunity to read the English practice sentences, which resembled the actual stimuli. The subjects were then request@d to read from two stacks of laminated color-coded cards, with Type A sentences in one stack and Type B sentences in the other. They were 53 instructed to say the sentences naturally at their normal speed as if they were talking to a friend. The cards with the sentences were presented to the subjects in random order to minimize the possibility of any ordering effect. Once all seven sets of sentences in each stack had been read, the subjects were instructed to shuffle the cards. The same procedure was repeated twice so that each card was read three times by each speaker. After the recording for one stack of cards was completed, the same procedure was repeated for the other stack of cards. Half of the subjects within each group read Type A sentences first and the other read Type B sentences first to counterbalance any potential sequencing effect. 3.3.4 ANALYSES 3.3.4.1 Acoustical analyses The speech data were segmented into syllables and subsequently analyzed for duration, intensity, and pitch using PitchWorks v. 4.5. The duration, peak intensity, and Extreme Point Fo' innovatively defined as the maximum Fo of a syllable whose pitch rises most of the time, or the minimum Fo of a syllable whose pitch drops most of the times, were obtained for each syllable of each test sentence (See 3.3.4.1.3 for a full description of Extreme Point Fo). The total number of utterances analyzed was 630 and the total number of syllables analyzed was 4140 for each sentence type. 3.3.4.1.1 Measurement of duration The durations of the syllables were measured based on the placement of the syllable boundaries. Given that the test sentences were largely composed of monosyllabic words, syllable boundaries usually coincided with word boundaries. The segmentation of the syllables generally followed three basic principles. First, wherever possible, the syllable boundary was placed at a distinct acoustic event that could be reliably and consistently identified. For instance, the end of a stop can be measured at the sudden burst of energy from the stop release, while the onset of a labiodental fricative can be measured at the onset of a quasi-random wave on the waveform and the onset of a high frequency noise region on the spectrogram, both of which are characteristic of fricatives. Second, when it was difficult to determine syllable boundaries, the location of a syllable boundary was placed close to the end of a real word or morphological boundary for practical reasons. For instance, the boundary between need and it in sentence B1 is placed at the onset of dark formant bands of the second vowel. In this case, it could be difficult to determine whether to syllabify the intervocalic consonant with the first or the second syllable if it was pronounced as a flap. However, sometimes the morphological boundary perfectly matched the syllable boundary. For instance, the syllable boundary between with and John in sentence B4 could be placed at the onset of silence due to the closure of /d3/. And third, in cases where there was no visible acoustic event to help, the syllable boundary could be determined via reasonable arbitrary rules that could be reliably applied. For example, the syllable boundary between the two syllables mo[m mjade in sentence A5 was arbitrarily placed at the midpoint between the end of formant bands of the first vowel [0] and the onset of formant bands of the second vowel [e]. Note that in doing so, it is assumed that the [m] coda of the first syllable and the [m] onset of the second syllable have equal lengths, which may not always be true. 3.3.4.1.2 Measurement of intensity The intensity of a syllable was measured at its intensity maximum. The measurement of peak intensity was preferred over the measurement of the average intensity of a syllable because the measurement of peak intensity minimizes segmental variation as a potential variable. Due to the fact that intensity is highly sensitive to the phonation type, i.e., whether the segment is a vowel or a fricative, etc., a syllable that contains voiceless stops and fricatives (e.g., socks), is likely to have lower average intensity than a syllable that contains only vowels (e.g., ah). Unlike average intensity, which takes into consideration the intensity of all segments in the syllable, peak intensity considers only the loudest portion of a syllable, which usually occurs within the vowel. Therefore, we chose peak intensity over average intensity as our measure of intensity as a correlate of stress. 3.3.4.1.3 Measurement of pitch Measuring the Fo value of a syllable is often not as straightforward as measuring duration and intensity. In a language like English where contour pitches on accented syllables are 55 common, determining an Fo value that best represents a syllable can be a great challenge. As part of an effort to develop a suitable approach for measuring the Fo value of a syllable in English, we first review three existing approaches. MiddleFo This approach is sometimes used in experimental studies of intonation. With this approach, the Fo is measured right in the middle of each syllable. The biggest advantage of this method is that it can be reliably and easily applied. However, it has two major disadvantages. First, the location of Middle Fo can vary due to segmental variations. For instance, the middle Fo of a syllable with a heavy initial consonant cluster (e.g., splash) wi11likely occur earlier within the syllable than in a syllable that has no initial consonants at all (e.g., ash). Even when two syllables have identical pitch contours, we are likely to come up with two different Fo measurements. Although the middle of the syllable is likely to fall inside the vowel, it can be anywhere within the vowel. Since pitch values within a vowel can vary rapidly, Middle Fo could fall either in the lower portion or the higher portion of a pitch contour depending on the segmental composition of the syllable. Second, it leads underestimation of a syllable's full Fo range. For instance, in a rising pitch accent L+H*, the Fo range of the syllable actually goes as high as the peak of the H*, which usually occurs toward the end of the syllable's pitch track. However, the middle of the syllable would probably occur earlier. As a result, Middle Fo is likely to be Fo value that is somewhat lower in this case. F0 at peak intensity With this approach, the Fo value is measured at the maximum intensity of each syllable. Compared with the first approach, peak intensity Fo is a more dependable reference point because it is less likely to be influenced by segmental variations. In addition, the maximum intensity does not necessarily coincide with the point where the Fo excursion reaches its maximum. Another problem with this approach is that the peak intensity can correspond to several Fo values within a syllable. One would have to arbitrarily decide which Foto choose, or as we decided to do in the example at the end of this section, average all the Fos that correspond to the intensity peak. MeanFo With this approach, the Fo value is measured by averaging all Fo values sampled within the boundaries of the syllable. This approach is quite valid for languages where there is little pitch movement inside a syllable. For languages that contain both level and contour pitch accents, however, measuring the mean Focan be problematic. By averaging the high and low elements of a contour pitch accent, we flatten its actual pitch range. For instance, the mean Foof an H* is likely to appear greater than that of an L+H* because the low element of an L+H* will bring the average Fodown. So how can we measure the pitch of a syllable in English? For our purposes, a desirable way of measuring the pitch of a syllable should achieve three goals. First, it should show as closely as possible where pitch prominence actually occurs. Second, it should provide a good estimate of the physical range of pitch variations over the utterance. Third, it should be in crucial ways consistent with our knowledge of the characteristics of English pitch accents. As complicated as pitch is, it is difficult to achieve all of the above by sampling a single Fovalue for each syllable. A fourth goal is that when the sampled Fos of each individual syllable are connected, the result approximates the contour of the original pitch track as closely as possible. With these principles in mind, the current study proposes an alternative approach for sampling the Fo. named the Extreme Point F0(EPF0), Extreme Point F0 This approach measures Foat the Fomaximum or minimum within a syllable. The purpose is to capture the shape and the range of the pitch contour across the utterance as closely as possible. When the syllable is accented, the approach is aimed at sampling the upper bound of an H* in pitch accents such as H* or L+H*, or the lower bound of an L* in pitch accents such as L* or L*+H. And when the syllable is unaccented, this approach is designed to sample the maximum or minimum pitch of the syllable relative to the next 57 tonal event (See Pierrehumbert, 1980; Ladd, 1996 for a review of the phonetics and phonology of English intonation). Three sampling rules were developed to serve this purpose. Rule #1 Ifthe pitch values ofa syllable rise most ofthe time, pitch will be measured at the Fa maximum. In English, a pitch accent can consist of either a level or a contour tone. A level pitch accent can be either High or Low. The High and Low of a pitch accent generally refer to the ceiling and the floor of the speaker's pitch range. The common types of pitch accent in English include H*, L*, L*+H, L+H*, etc. The asterisk "*,, is a convention of the TOBI system to mark the weighted tone (Ladd, 1996). For instance, the pitch accent L+H*, a very common type of pitch accent in English, has more weight on the high than on the low (e.g. Jane in Figure 3.1). The Fo values of an L+H* usually rise rapidly throughout. Even when the pitch accent is a monotone H*, it usually appears to be slightly rising because of the transition from the preceding syllable. With this rule, the EPFo should sample the upper bound of the H*. When the syllable in question does not carry any pitch accent, this rule directs EPFo to sample the upper bound of a tonal interpolation between a low tone and a high tone. Rule #2 If the pitch values of a syllable fall most of the time, pitch will be measured at the Fo minimum. This rule is intended to capture the L* in pitch accents such as L* or an L*+H. In an L*+H, the Fo values usually remain low within most of the syllable with a relatively late onset of the Fo rise. Therefore, if the pitch values of a syllable show a predominant fall, it contains a weighted low tone. Even when the pitch accent is a monotone L*, .it usually appears slightly falling due to the tonal transition from the previous syllable. When the syllable in question does not carry any pitch accent (e.g., is in Figure 3.1), this rule directs the EPFo to sample the lower bound of a tonal interpolation between a high tone and a low tone. Rule #3 Ifthe Fa drop occurs on the final syllable ofa declarative utterance and is preceded by an Fa rise, pitch will be measured at the Fa maximum. Phrase or sentence-final syllables have to be treated differently because they are also marked by phrasal and boundary tones in English. For instance, when a predominant Fo drop is observed on the final syllable of a declarative sentence and the Fo drop is preceded by an Forise, it is usually an L+H* or an H* followed by a low phrasal tone and a low boundary tone (e.g., Mom in Figure 3.1). In that case, this rule directs the EPFoto sample the upper bound of the high pitch accent rather than the trailing low phrasal or boundary tones. With gracious help from Dr. Jonas Chen, a computer program was written to automatically sample the EPFos of individual syllables from the segmented pitch log files for all utterances generated using PitchWorks. Figure 3.1 illustrates how the Extreme Point Fos are measured. The sentence was produced by a male English native speaker. The complete context within which the target sentence was embedded is given below. The target sentence is shown in italics and the target strong syllables are shown in boldface. (23) I think you're mistaken. Jane is not my mom. She is my sister. ..yll .. bln jane is not my mop Il:In~ L+HiI L+HiI L+H* L-L~ E:pr iI *I iI *l • ·~~jr. TIlT II ."w,rrrrrrr r ....., I,f'/'/ """.._\. I,//~~"""-~ r~""'\·\ r/'"///- ~" /" V··'~~--~'c .•__, r'~~~\- ~_r 250 200 -_. • 150 ...._! !I... •• ..... 10. ......... • ••• • • .10............ ......- 100 ........ ...........- I I I I I 50 '"" 200 400 SOD 800 1000 Figure 3.1 The sampling of the Extreme Point Fosfor the utterance Jane is not my mom. Three syllables, Jane, not, and mom, are assigned L+H* pitch accents for this sentence. The Fos of these syllables were measured at the Fa maxima to capture the weighted H*, following sampling rule #1 for the syllables Jane and not, and rule #3 for the final syllable mom. The pitch movements within the unaccented syllables, is and my, generally decline due to the tonal interpolation between the preceding high tone and the following low tone. For these syllables, the EPFs were sampled at the Fa minima within the syllable following rule #2. The vertical lines immediately to the right of the asterisks in the "EPF" row mark the point in time where each EPF was sampled. Figure 3.2 shows what the sampled Fos look like on the original pitch track. The labeled points mark the Fa values and the points in time where the Extreme Point Fos were sampled. 170 160 150 140 N' e. .r:. 130 u := 0. 120 110 100 90 <:l 143 • ~~~~~~~~~~~~~~~~~~~~~ ti ti ~. Time (Ms) Figure 3.2 The corresponding Extreme Point Fos on the original pitch track for the utterance Jane is not my mom. The EPFoapproach is not only innovative but also theoretically sound. To show its strength, we compare this approach against the three earlier described approaches for measuring Fowith respect to how closely each approach corresponds to the pitch accents and the pitch contours of the original pitch track. 170 160 150 140 N e. ..r::: 130 ii: 120 110 100 90 jane is not my mom Syllable -t-MidFO _dBMax MeanFO "ooq)'I!'"'''' EPFO Figure 3.3 The Middle Fos, the Fos at Peak Intensity, the Mean Fos, and Extreme Point Fos for the utterance Jane is not my mom. A few observations can be made based on the results shown in Figure 3.3. First, all three earlier approaches, including the Middle Fo' Peak Intensity Fo' Mean Fo' fail to represent the high pitch accent on the final syllable. Both Middle Foand Peak Intensity Fo of the final syllable reflect instead the lengthy low phrasal and low boundary tones at the end of the declarative statement. The Mean Foof the final syllable averages the high weighted tone of the pitch accent with the lengthy low phrasal and boundary tones, therefore underestimating the height of the high tone. EPFo is the only approach that successfully passes this test. Second, Middle Fo' Peak Intensity Fo' and Mean Foeach may underestimate the range of the pitch excursion of a contour pitch accent (e.g., Jane). Mean Fois particularly susceptible to flattening the contour pitch accents by averaging out the high and the low elements. We can see in this example that the Mean Foconsistently lowers the upper bounds of the high pitch accents and raises the lower extremes. Among the four approaches, EPFomost authentically represents both the pitch accents and the pitch contour of the original pitch track. 3.3.4.2 Normalization of data After the duration, intensity, and pitch of individual syllables were measured, the absolute measurements were then normalized within each utterance to control for variations in speech rate, loudness of speech, and pitch ranges across utterances. The absolute duration, intensity, and pitch data are valuable in showing the actual differences between subject groups. For instance, by using the absolute duration in milliseconds to compare the groups, one can see whether TM learners generally take longer to produce each utterance. However, caution must be taken when one uses absolute duration, intensity, and pitch to compare groups because these differences often heavily reflect individual variations in speech rate, loudness, and pitch ranges. In order to eliminate these factors as potential variables, the absolute duration, peak intensity, and Extreme Point Fomeasurements were normalized within each utterance before any further computations or comparisons were made. After all, as far as rhythm is concerned, it is the relative prominence relationships among syllables of an utterance that count. 3.3.4.2.1.1 Normalization of duration There are two ways of normalizing duration within an utterance. The first approach arbitrarily assigns the longest syllable a value of one and converts others into proportions. The second approach divides the duration of individual syllables by the duration of the entire utterance and shows them as percentages of the length of the entire utterance. The second approach is advantageous over the first in two ways. First, only the second approach truly eliminates speech rate as a variable. The second approach does exactly that. After the normalization procedure, the duration of all utterances becomes 1.00 so that there are no more rate differences across utterances. However, with the first approach, even after normalization, the durations of entire utterances are bound to vary and therefore so are the speech rates. Second, the duration ratio obtained using the first approach is biased by the duration of a single syllable, whose length is an independent random variable. For example, suppose two people both produce the greatest lengthening 63 on the final syllable. One speaker produces a moderate final lengthening while the second speaker produces considerably greater final lengthening. With the first approach, the duration ratios of the other syllables will appear much smaller for the second speaker than for the first with other things being equal. By taking the duration of all syllables into consideration, the second approach is less susceptible to the duration variations of a single syllable. Third, the first approach is computationally more demanding. As the longest syllable may vary across utterances, one will have to manually compute the duration ratio for each syllable of each utterance. With the second approach, the duration percentages can be more easily computed. Consequently, the current study normalized duration within an utterance by applying the second approach. 3.3.4.2.1.2 Normalization of intensity The normalization of intensity across syllables within an utterance is less complicated. The intensity ratio of a syllable was simply obtained by dividing the obtained peak intensity for that syllable by the maximum intensity of the utterance. In doing so, we eliminate variations in loudness of speech as a potential variable across utterances. 3.3.4.2.1.3 Normalization of pitch The normalization of pitch among syllables within an utterance proceeded in two steps. First, the Fo frequency in Hz was converted into semitones to make the units of comparison perceptually equal. The reason for doing this is that a given absolute difference in Hz is perceived differently at different pitch ranges. It takes a larger difference in Hz to achieve the same amount of perceptual difference at a higher pitch range than at a lower pitch range. Semitones take such perceptual differences into consideration, so that a difference of one semitone at a higher pitch range is perceived the same as a difference of one semitone at a lower pitch range. After the Fo frequency in Hz was converted into semitones, the semitone value of each syllable was calculated as a ratio of the highest semitone value in that utterance. The concept of a semitone originates from music theory. Each of the 12 notes in an octave is one semitone higher than its adjacent note toward the lower end of the pitch 64 scale. The ratio of frequencies between successive notes is a constant. When we multiply the frequency of a note (or any frequency) by this constant 12 times, we end up doubling that frequency. Therefore, this constant ratio equals the twelfth root of 2, or approximately 1.0595. This constant ratio is called a semitone. In the current study, we converted frequency data in Hz into semitone units. The procedure took shape in two steps. First, we established a Hertz to Semitone Conversion Scheme. The 'Note' to 'Frequency' to 'Semitone' conversion chart for music is expanded to be inclusive of the Fo range of ordinary speech data. The chart starts with "C" at 32 Hertz because this is the closest "C" to the estimated lowest frequency sample in the speech data. This C is then arbitrarily set to be the zero of the semitone scale. In the rightmost column, you will find the corresponding semitone value for the notes at various frequencies. As the frequency ratio between successive notes is constant at one semitone, the semitone scale is designed so that a difference of one on the scale means exactly a difference of one semitone. Table 3.3 The Hertz to Semitone conversion chart Note Name Frequency Semitone (Hertz) Scale C 32*2°112 32 0 C# 32*21/12 34 1 D 32*2 2112 36 2 D# 32*23112 38 3 ...... ...... C 32*i 2l12 64 12 C# 32*213112 68 13 D 32*i 4l12 72 14 D# 32*215112 66 15 ...... ...... C 32*i 4l12 128 24 C# 32*225/12 136 25 D 32*2 26/12 144 26 D# 32*227112 152 27 ...... ...... C 32*236112 256 36 C# 32*i 7 /12 272 37 D 32*238/12 288 38 D# 32*239112 305 39 ...... ...... Next, we conducted a series of mathematical derivations for converting frequency in Hz into semitones. Given that X represents the frequency in Hz and Y represents the corresponding semitone level, the relationship between the two variables can be written as follows. (35) X =Frequency in Hz Y =Semitone X =32*2 Y/1Z The relationship between X and Y should become clearer with the help of Table 3.4. Table 3.4 Mathematic derivation of semitones from frequency in Hz Frequency Mathematic Semitone (Hertz) Derivation 32 32*2°112 0 64 32*i 2l12 12 128 32*2 24112 24 256 32*2 36112 36 Based on the relationship between X and Y, the following step-by-step mathematic derivations were performed to obtain y. (36) Step 1. X =32*2 YIIZ Step 2. X/32 =2Y/IZ Step 3. logiX/32) =Y/12 Step 4. 12* logz(X/32) =Y Step 5. Y =12* logz(X/32) Now that we know how Y can be mathematically derived from X, we can convert any frequency into a semitone. Once frequency is converted into semitones, we can describe the pitch differences among syllables in semitone units. 3.3.4.3 Statistical analyses A set of four statistical analyses were performed to provide supporting evidence for answering the research questions. First of all, general comparisons were drawn from the basic statistics, including means, standard deviations, and ranges of both the absolute and the normalized data. Second, Student's t-tests were conducted for each individual syllable between each pair of the three speaker groups to identify significant differences. We focused on the number of strong versus weak syllables that were significantly different in non-final versus final positions between English native speakers and the two groups of TM speakers. The results provide crucial information about the degree of difficulty TM speakers experience with each of the three variables, the distribution of difficulties between strong vs. weak syllables in non-final vs. final positions, and any differences between the two groups of TM learners of English. Third, Pearson Product-moment correlation coefficients were obtained for the duration, intensity, and pitch of each sentence between pairs of the three subject groups. The results provide an estimate of how closely the duration, intensity, or pitch ratios between pairs of the three groups co vary across syllables of a sentence. Fourth, Test-retest reliability indices were calculated for the three productions of each sentence for each group. The 'results measure how consistent the subjects were with their duration, intensity, and pitch ratios across the three productions of the test sentences. Three reliability indices were obtained for each sentence for each group. The significant level for all statistical tests were set at p<O.05. Because multiple statistical tests were conducted on the same set of data, the probability of obtaining spurious statistical significance increases. If I were to use the Bonferoni adjustment for my t-tests, given that there were 828 of them, an alpha level of 0.0000603865 (i.e., 0.05 divided by the total 828 tests) will be necessary in order to interpret each of my t-tests such that the risk of finding a significant difference by chance remains at 0.05. However, it would be very difficulty to establish statistical differences with such a small P-value. Therefore, Bonferoni adjustment was not performed on any of the statistical tests. CHAPTER 4: RESULTS AND DISCUSSION (1): DURATION This chapter reports and analyzes results from the duration data of Type A and Type B sentences. English, TM ESL, and TM EFL speakers are compared with respect to their use of duration as a correlate of stress in English. Results of the two rhythmically diverse sets of sentences are reported separately and then combined for more detailed analyses. The focus of the current investigation will be the extent to which these two types of rhythmic patterns differ in terms of the types and degrees of difficulty in duration presented to TM speakers and how such differences may help us understand the source of any difficulties. The duration data of Type A are reported in Section 4.1 and those for Type B sentences in Section 4.2. These two sections are organized into five subsections each. The first pair of sections, 4.1.1 and 4.2.1, report basic statistics and patterns revealed by the absolute durations of syllables in ms. The data are valuable in showing the absolute durational differences between subject groups. However, caution should be taken when one compares absolute durational differences across groups because these differences often heavily reflect individual variations in speech rates. Therefore, interpretations of the absolute duration data are limited to revealing general patterns. The second pair of sections, 4.1.2 and 4.2.2, report basic statistics and general patterns revealed from relative durations. The relative duration of a syllable is obtained by converting its absolute duration into a percentage of the duration of the entire utterance. The purpose is to eliminate speech rate as a potential variable. In the third pair of sections, 4.1.3 and 4.2.3, direct syllable-to-syllable comparisons are made between speaker groups based on their relative durations. Both similarities to and significant differences from the duration patterns of English speakers are reported. The fourth pair of sections, 4.1.4 and 4.2.4, report correlation coefficients of overall duration patterns between speaker groups for each test sentence to provide additional information about how close to target ESL and EFL speakers are. The fifth pair of sections, 4.1.5 and 4.2.5, 70 report test reliability information for all test sentences to provide information about how consistent NS, ESL, and EFL speakers are with their duration patterns across three productions of the same sentence. 4.1 DURATION PATTERNS OF TYPE A SENTENCES This section reports duration results of Type A sentences, which feature long stretches of weak syllables. Each sentence contains either one single strong syllable or two widely spaced strong syllables. Four rhythmic patterns are represented. Sentences Al and A2 feature the five-syllable rhythmic pattern Swwww, sentences A3 and A4 feature the seven syllable rhythmic pattern SwwwwwS, sentences A5 and A6 feature the seven syllable rhythmic pattern wSwwwww, and sentence A7 has the eight-syllable rhythmic pattern SwwwwwSw. The total number of utterances analyzed is 630 and the total number of syllables analyzed is 4140. 4.1.1 ABSOLUTE DURATiON PATTERNS This section summarizes results based on the absolute durations of the individual syllables. A group average duration in milliseconds was obtained for each syllable of every test sentence. Overall observations based on the lengthening and shortening of syllables across sentences and the basic statistics, including means, standard deviations, ranges of absolute duration, and the number of syllables uttered per second, are reported. Table 4.1 Mean syllable durations in ms for Type A sentences Syllable Mean SD Range SylJ sec wrote it with me 162.33 108.00 193.33 208.00 257.67 204.00 221.67 271.67 325.33 271.00 256.67 295.00 made it for 171.33 100.67 174.33 241.33 123.00 205.33 276.33 293.67 192.67 240.67 You like me to wear the 162.67 220.00 123.00 101.00 167.00 89.00 189.67 266.67 188.00 193.00 219.67 133.00 193.67 299.00 233.33 272.67 280.67 161.00 You want me to bring the 170.67 138.00 151.33 91.67 243.33 65.00 197.00 229.67 219.67 192.67 281.00 105.67 227.00 267.33 257.00 239.67 351.67 151.33 M made the Ie mon 152.67 197.00 65.33 119.33 156.33 182.67 268.00 110.00 144.67 190.67 214.00 302.00 192.33 182.33 231.00 The old man gave it to 117.00 189.00 173.33 105.00 101.00 129.33 269.00 264.00 231.33 177.67 109.67 178.67 316.33 320.00 311.67 277.00 155.67 Jane's the one wear .ing the 49.00 195.67 175.00 115.00 39.67 91.67 311.00 223.(i7 149.33 77.00 209.67 369.00 272.67 214.00 126.00 The syllables are generally shortest for NS, longest for EFL and in between for ESL. NS produce the shortest syllables across the board except for the final syllable in sentence A3. EFL produce the longest syllables among all subject groups except for the first syllable in A2, and the final syllables in sentences A3, A4, and A5. The mean syllable duration for all Type A sentences averages 177.38 ms for NS, 232.75 ms for ESL, and 273.28 IDS for EFL. The mean syllable duration reflects the flip side of speech rates, which are fastest for NS, slowest for EFL, and in between for ESL. NS on average speak at the rate of 5.64 syllables per second, ESL at the rate of 4.31 syllables per second, and EFL at the rate of 3.67 syllables per second. Advancement in English proficiency yields positive improvement on overall speech rates for TM speakers. The variation of absolute duration among syllables of a sentence tends to be greatest for ESL, smallest for EFL, and in betweenfor NS. NS and ESL on average produced very similar standard deviation of absolute duration for Type A sentences. The average standard deviation of absolute duration for all Type A sentences is 80.61 ms for NS, 80.06 ms for ESL, and 70.06 ms for EFL. The ranges between the longest and the shortest syllables of the sentences are generally broadest for ESL, narrowest for EFL, and in between for NS. The average range of syllable duration for all Type A sentences is 228.71 ms for NS, 240.48 ms for ESL, and 195.29 ms for EFL. Among the three subject groups, EFL speakers produce syllables that are longest but least variable in length. (a) Mean syllable duration (ms) for sentences Al (Swwww) 400 | | 400 350 | A .. " | | --- | --- | --- | | | 300 | .. "A " '- .. .. .. | | ....... 1/1 .§. c :8 l! ::I 0 | 250 200 150 | .. .. .. .. .. ",,- ". " ....... NS • --ESL .... EFL .. | | | 100 | | | | 50 0 | | 100 Jim wrote it with me Syllable (b) Mean syllable duration (ms) for sentence A2 (Swwww) 350 250 | | 350 ..300 | | --- | --- | | | t1 , ... & | | E NS--c 0 i ... ::I 0 | 250 200 • --ESL .... EFL .. 150 100 | | | 50 0 | 100 Jane made it for me Syllable (c) Mean syllable duration (ms) for sentence A3 (SwwwwwS) -r--------------------------, 400 +------------------------1IIIIO------f 350 +-------------------------<If------f 300 +-------->it<------------------¥----I "iii I' ..... .. .. .. .... r;;;;.~NSl S 250 +-__~"___"'c--_-----'''--~'''-----..------'----''=_I_--______j I- • NS l$ --ESL 200 +--------.ItiIII"---,~-~-____"Io=_--__-=---'l-----"lr____,f+_-------f .. .... EFL :::l c 150 +-------''''-----------'~------~~'"""-~'!EU'-----I -1----------~~~----~y:..-.----~ 50 +--------------------------1 O+------,.----r----,...------,-----,----,....----I You like me to Syllable wear the jeans (d) Mean syllable duration (ms) for sentence A4 (SwwwwwS) 500 450 400 350 , " "iii 300 E .IF--c ...... 250 .. .. 0 :::l 200 c 150 100 50 0 • NS --ESL ...... EFL You want me to Syllable bring the wine (e) Mean syllable duration (ms) for sentence A5 (wSwwwww) 400 Ui" 250 100 | | 400 350 | .. | | | | | --- | --- | --- | --- | --- | --- | | Ui" NS--r:: | 300 250 E 0 200 | | | | • --II-ESL | | .... c | III 150 | | | | EFL 'Ill - - | | | | | | | | | | 100 | | | | | | | 50 0 | | | | | | (f) | Mean syllable 400 | My mom duration | made the Ie Syllable (ms) for sentence A6 (wSwwwww) | man pie | | | | 350 | | | | | | | -...300 ..... 250 III §. r:: 0 200 ..I! c 150 100 50 0 | ... , | - .. .. -. | I• | --II-ESL EFL - 'Ill - | | | | | | | | 400 • NS --II-ESL - 'Ill - EFL it to me The old man gave Syllable (g) Mean syllable duration (ms) for sentence A7 (SwwwwwSw) • NS --ESL - ... - EFL | | 450 | | | --- | --- | --- | | | | | | 'iii' g ;: III... d | 400 '\ 350 300 250 200 150 100 50 0 | --ESL - ... - EFL | the blue dress Jane's the one wear ing Syllable Figure 4.1 Mean syllable durations (ms) for sentences Al through A7 Five observations can be made based on the results summarized in Table 4.1 and Figure 4.1. First, syllables are noticeably longer for ESL and EFL speakers than for NS across the sentences. Generally speaking, EFL speakers produce the longest syllables across groups. Second, the longest syllables of the sentences are usually the first strong syllables or the final syllables. The longest syllables are the first strong syllables in three sentences AI, A2, and A7 for NS and ESL speakers, and in three sentences AI, A5, and A7 for EFL speakers. The longest syllables are the final syllables in three sentences A3, A4, and A5 for NS, in four sentences A3, A4, A5, and A6 for ESL speakers, and in A2, A3, A4, and A6. Third, regardless of major differences in speech rates, syllables appear to lengthen and shorten in similar fashion across groups. The syllables that are lengthened by English speakers are also lengthened by TM speakers, but the syllables that are shortened by English speakers are not necessarily shortened by TM speakers. For example, EFL speakers lengthen the weak syllable made in sentence A2, and the weak syllable gave in sentence A6, both of which are shortened by English speakers. Fourth, in addition to strong syllables and final syllables, weak syllables, especially weak content words, are sometimes lengthened. Here content words are defined as words that denote object, property, or action, including nouns, adjectives, adverbs, and verbs. For instance, the weak syllables wear in sentence A3, bring in sentence A4, man in sentence A6, and one in sentence A7 are lengthened by all groups. Fifth, other factors, such as the number of segments in a syllable and the inherent length of a syllable, could also affect the actual length of a syllable. For example, fewer number of segments may explain why the target strong syllable you in sentence A3 appears relatively shorter than the immediately following target weak syllable like. Additionally, the larger number of syllables and the presence of inherently longer segments such as fricatives may explain why the syllable with in Al and the syllable for in A2 are relatively longer than the preceding weak syllable it in the speech of NS. 4.1.2 RELATIVE DURATIONS PAITERNS This section reports results based on the relative durations of the syllables of a sentence. The relative duration of a syllable is calculated as its percentage of the lengths of the entire sentence. To eliminate variations in speech rates across utterances, the lengths of the syllables are normalized within each utterance first. After that, a group average duration (%) is obtained for each syllable of a sentence by averaging the durations (%) of that syllable in the 30 utterances. These basic statistics, including means, standard deviations, and ranges, are summarized in Table4.2. Table 4.2 Mean syllable duration as its percentage of the total sentence duration for Type A sentences GrOll S lIable Al wrote it with me NS 17.55 11.63 20.97 22.20 ESL 20.02 15.78 17.53 20.98 EFL 21.19 17.80 16.90 19.67 A2 made it for NS 19.03 11.08 19.24 ESL 21.05 10.49 17.93 EFL 20.97 22.22 14.43 18.23 A3 You like me to wear the NS 12.84 17.33 9.76 7.91 13.21 7.02 ESL 11.71 16.61 11.79 12.08 13.72 8.36 EFL 10.62 16.23 12.61 14.71 15.30 8.76 A4 You want me to bring the NS 14.16 11.32 12.47 7.50 20.07 5.36 ESL 11.62 13.48 12.92 11.30 16.89 6.35 EFL 12.07 14.31 13.68 12.75 18.94 8.10 AS M made the Ie mon NS 12.38 15.93 5.29 9.71 12.61 ESL 12.02 17.54 7.15 9.61 12.55 EFL 12.12 16.99 10.78 10.07 12.98 A6 The old gave it to NS 10.04 16.23. 14.88 9.01 8.64 ESL 8.83 17.99 17.63 15.65 11.97 7.37 EFL 9.35 16.62 16.80 16.57 14.53 8.21 A7 Jane's the one wear .ing the blue dress NS 3.43 13.73 12.30 8.08 2.78 16.90 18.21 ESL 4.70 15.77 11.76 7.77 3.97 16.80 17.29 EFL 9.00 15.85 11.79 9.24 5.37 15.03 14.63 AI·7 Ion est shortest NS 26.88 7.49 ESL 23.98 8.50 EFL 21.02 10.26 Six observations can be made based on the results summarized in Table 4.2. First, the longest syllables of the sentences are usually the first strong syllables or the final syllables. The longest syllables are the first strong syllables in three sentences for NS (AI, A2, A7), ESL (AI, A2, A7) andEFL (AI, AS, A7) and they are the final syllables in three sentences for NS (A3, A4, AS) and in four sentences for ESL (A3, A4, AS, A6) and EFL (A2, A3, A4, A6). Second, the longest syllables of the sentences are relatively 79 longer for NS than for ESL and EFL. The relative duration of the longest syllables of all sentences averages 26.88% for NS, 23.98% for ESL speakers, and 21.02% for EFL speakers. Third, the shortest syllables of the sentences are all unstressed function words, such as it, the, to, except for Ie of EFL speakers in sentence A5. The unstressed function words with reduced vowels, especially it, the, to, are particularly short in the speech of NS. Fourth, the shortest syllables tend to be relatively shorter for NS than for ESL and EFL speakers. The shortest syllables of the sentences are relatively shortest for NS in five sentences AI, A3, A4, A5, and A7. The relative duration of the shortest syllables of all sentences averages 7.49% for NS, 8.50% for ESL speakers, and 10.26% for EFL speakers. As a direct result of observations Two and Four, NS produce the widest ranges of relative duration for all Type A sentences. The average range of duration percentages is 19.39% for NS, 15.49% for ESL, and 10.75% for EFL speakers. Fifth, the standard deviations of duration percentages are greatest for NS and smallest for EFL speakers for all Type A sentences. The average standard deviation of duration percentages of all sentences is 6.85% for NS, 5.57% for ESL, and 3.88% for EFL speakers. Sixth, English speakers produce greater lengthening on final syllables than ESL and EFL speakers. The final syllables of NS are relatively longer than those of ESL and EFL for all Type A sentences. Figure 4.2 shows the duration patterns of Type A sentences in terms of percentages of the lengths of the entire sentences. (a) Mean syllable duration (%) for sentence Al (Swwww) -r--------------------------, 30 +--------------------------1 25 +------"!'......,~---------------------I 20 +--------":~~_-------_F=---_____c:..,,_~ii'_----j c: o !15 15 +--------------"'~----"'=----~'----------------I • NS -- ESL - .. - EFL 10 +----------------------------j 5+--------------------------1 0+-----..,.-----..,-----...,.....------,------1 Jim wrote it Syllable with me (b) Mean syllable duration (%) for sentence A2 (Swwww) -r--------------------------, 30 +--------------------------1 • NS --ESL - .. - EFL 25 +-----.... ~,_____----------------_::J __----j 20 +----==-------'''=*-_---------::;;;;;c:;;;;orttP-'''---------I c: o 15 15 +----------~-__"__,.:__""-~~---------l 10 +-------------1---------------1 5+--------------------------l 0+-----..--------,,-------.,.------,------1 Jane made it Syllable for me (c) Mean syllable duration (%) for sentence A3 (SwwwwwS) 35 30 25 20 c 0 :;::: l! :I 15 c 10 5 0 You like me to Syllable wear the jeans • NS --ESL .... EFL (d) Mean syllable duration (%) for sentence A4 (SwwwwwS) T"""------------------------..., 30 +----------------------------=:----1 +-----------------------6------1 20 +-----------------..------I--.IIr-----1 c o IS 15 +-----::::------.----------...£-,f-'--~---_I_----l 10 +---------------"'~-_+---------'Il~_4'--------1 5+-------------------~~----_:_I O+-----,.-----r----r----,.---....-----,,....----:-I • NS --ESL - .. - EFL You want me to Syllable bring the wine (e) Mean syllable duration (%) for sentence A5 (wSwwwww) 35 30 25 20 c:: 0 :;:; lIS... 6 15 10 5 0 My mom made the Syllable Ie mon pie • NS --ESL • 'lIt.. EFL (f) Mean syllable duration (%) for sentence A6 (wSwwwww) 35 30 25 -- 20 c:: 0 i... 6 15 10 5 0 The old man gave Syllable it to me • NS --ESL • 'lIt.. EFL strong or weak, weakly stressed or unstressed. However, TM speakers produce relatively shorter final syllables than English speakers. 4.1.3 SIGNIFICANT DIFFERENCES BETWEENNS, ESL AND EFL T-tests were performed on individual syllables to determine whether or not the obtained differences in relative duration between pairs of groups were statistically significant or likely by chance. Results of the t-test scores for the duration percentages of individual syllables between pairs of groups are shown in Table 4.3. Table 4.3 Student's t-test scores for duration (%) ofindividual syllables between pairs of groups for Type A sentences | Al | | | | | | | | | with | | me | | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | NS-ESL | | | | | | | | ---:2::-=.47:: | | 078*-:---+--:0:=:.5=-=87"8---1-----+---+----11 | | | | | NS-EFL | | | | | | | | -=2~.7:..::5:..::6_*----1f-'1~.4.:,:5:....:4-_I_---+-----+--__j1 | | | | | | | ESL-EFL | | | | | | | | | 0.416 | | 0.616 | | | | A2 | | | | | | | | | for | | me | | | | NS-ESL | | | | | | | | | 1.234 | | 1.440 | | | | NS-EFL | | | | | | | | | 0.951 | | 1.118 | | | | ESL-EFL | | | | | | | | | 0.282 | | 0.416 | | | | A3 | | | | | | | | | to | | | | | | NS-ESL | | | | | | | | | | | | | | | NS-EFL | | | | | | | | | | | | | | | ESL-EFL | | | 0.715 | | | | | | | | | | | | A4 | | | You | | | | | | | | | | | | NS-ESL | | | 2.645* | | | | | | | | | | | | NS-EFL | | | 2.649* | | | | | | | | | | | | ESL-EFL | | | 0.475 | | | | | | | | | | | | AS | | | My | | | mom | | | | | | | | | NS-ESL | | | 0.549 | | | 0.028 | | | | | | | | | NS-EFL | | | 0.293 | | | 0.477 | | | | | | | | | ESL-EFL | | | 0.130 | | | 0.353 | | | | | | | | | A6 | | | The | | | old | | | | | | | | | NS-ESL | | | 1.233 | | | 1.429 | | | | | | | | | NS-EFL | | | 0.641 | | | 0.400 | | | | | | | | | ESL-EFL | | | 0.503 | | | 1.411 | | | | | | | | | A7 | | | Jane's | | | the | | | wear | | | | | | NS-ESL | | | 2.685* | | | 2.814* | | | 0.784 | | | | | | NS-EFL | | | | | | | | | 0.689 | | | | | | ESL-EFL | | | 2. | | | | bO~.0;,;7,,;;0=~0;;.;.04~2=~~~=='=.;;;;;;;~=~~~= | | | | | | | *p<.05 (tcrlt=2.101), **p<.Ol(tcrlt=2.878), df=18, two-tailed In the next two subsections we will examine which syllables are found to be significantly different between groups in tenns of duration percentages. Section 4.1.3.1 reports and compares differences between native and non-native speakers. Results provide crucial infonnation as to the kinds of difficulties TM speakers have with duration as a correlate of stress, as well as in what ways ESL and EFL are similar or different in their problems with duration. Section 4.1.3.2 examines significant differences between ESL and EFL. Results in this section provide infonnation about the changes that might have taken place from EFL to ESL in the way duration correlates with stress. 4.1.3.1 Differences between NS and ESL vs. differences between NS and EFL Out of 46 syllables compared for Type A sentences, a total of 19 syllables was found to be significantly different between NS and ESL. They comprise two relatively shorter strong syllables in non-final position found in sentences A4 (you) and A7 (Jane's), 12 relatively longer weak syllables in non-final position in sentences Al (it), A3 (me, to, the), A4 (want, to, the), A5 (the), A6 (it) and A7 (the, one, the), three relatively shorter weak syllables in non-final position in sentences Al (with), A4 (bring), and A6 (man), one relatively shorter strong syllable in final position in sentence A3 (jeans), and one relatively shorter weak syllable in final position in sentence A5 (pie). Table 4.4 Number of strong and weak syllables with duration (%) significantly different from NS in non-final vs. final positions for Type A sentences Position Stress EFL 0 =,,;;;2~d=~= * Shaded cells dilute the contrast, unshaded ones do not. Weak Total (k=5) differ Lon er shorter o 1 19 o 2 27 A total of 27 syllables was found to be significantly different between NS and EFL. They comprise six relatively shorter strong syllables in non-final position in sentences Al (Jim), A2 (Jane), A3 (you), A4 (you), A7 (Jane's, blue), 15 relatively longer weak syllables in non-final position in sentence Al (wrote, it), A2 (it), A3 (me, to, wear, the), 86 A4 (want, to, the), AS (the), A6 (man), and A7 (the, one, the), two relatively shorter weak syllables in non-final position in sentences Al (with) and A6 (it), two relatively shorter strong syllables in final position in sentences A3 (jeans) and A4 (wine), and two relatively shorter weak syllables in final position in sentences AS (pie) and A7(dress). ESL produced the same types of difficulties with duration as a correlate of stress as EFL. Significant differences between NS and ESL and significant differences between NS and EFL highlight three types of difficulties, (1) relatively shorter strong syllables in non-final positions (you in A4 and Jane's in A7), (2) relatively longer weak syllables in non-final positions (it in AI, me, to, the in A3, want, to, the in A4, the in AS, it in A6, the, one, the in A7), and (3) relatively shorter final syllables (jeans in A3). Despite having the same types of problems, ESL speakers experience slightly less difficulty with duration than EFL speakers. A smaller number of syllables are found to be significantly different between NS and ESL than between NS and EFL. Compared with EFL speakers, ESL speakers produce fewer relatively shorter strong syllables and relatively longer weak syllables in non-final positions. EFL speakers produce a very high rate of relatively shorter strong syllables. Six of their eight strong syllables in non-final positions are significantly shorter than those of NS. ESL speakers have slightly less difficulty with lengthening final syllables than EFL speakers have. In particular, both ESL and EFL show little difficulty lengthening final unstressed function words. None of their final function words in sentences AI, A2, and A6 are significantly different from those of NS. 4.1.3.2 Significant differences between ESL and EFL This section reports significant differences in relative duration between ESL and EFL. Syllables that are found to be significantly different between these two groups are further examined under four categories: strong syllables in non-final position, weak syllables in non-final position, strong syllables in final position, and weak syllables in final position. The purpose is to identify patterns in changes that might have taken place due to improved proficiency and increased exposure to English. Table 4.S Number of strong and weak syllables with duration (%) significantly different between EFL and ESL in non-final vs. final position for Type A sentences Position Non-Final Final Total Stress Strong Strong (k=8) (k=2) T e Ion er shorter Ion er Shorter EFLvs.ESL 2 ESUNS only EFUNSonly 1 1 1 EFUESUNS 1 1 1 ..ESUEFL only 1 Of the 46 syllables compared, the relative duration of 17 syllables was found to be significantly different between ESL and EFL. Of these 17 syllables, EFL speakers produced two relatively shorter strong syllables in non-final position in A7 (Jane's, blue), 10 relatively longer weak syllables in non-final position in A2 (it), A3 (to, wear), A4 (bring, the), AS (the), A6 (it), A7 (the, -ing, the), two relatively shorter strong syllables in final position in A3 (jeans) and A4 (wine), and three relatively shorter weak syllables in final position in AS (pie, me, dress). It appears that EFL speakers differ from ESL speakers in ways that are higWy similar to the ways ESL and EFL differ from NS. While ESL and EFL speakers both produce relatively shorter strong syllables, relatively longer weak syllables in non-final position, and relatively shorter strong and weak syllables in final position than NS, EFL speakers produce more instances of each of these difficulties than ESL speakers. The results suggest positive improvement from EFL to ESL on the lengthening of strong syllables, shortening of weak syllables, and final lengthening. Table 4.6 Number of strong and weak syllables with duration (%) significantly different between EFL and ESL speakers categorized as content and function in non-final vs. final position | Position | | | Non-Final | | | Final | | | --- | --- | --- | --- | --- | --- | --- | --- | | Stress | | (k=8) Strong | Weak | (k=31) | (k=2) Strong | (k=5) Weak | | | Length | | (k=2) Shorter | Longer | (k=1O) | (k=2) Shorter | (k=2) Shorter | | | Svl.Tvoe | | content function I | content | function I | content function I | content I | function | | Number | | 2 I | 2 | 8 I | 2 I | 2 I | | ESL speakers have less difficulty shortening weak function words than EFL speakers. Of the 10 relatively longer weak syllables produced by EFL speakers, eight are weak function words. They are the syllable it in sentences A2 and A6, the syllable to in sentences A3, the syllable -ing in sentence A7, and the syllable to in sentences A3, A4, AS, and A7. We also notice that both of the relatively shorter strong syllables in non-final position and all of the relatively shorter strong and weak final syllables are content words. These results suggest that TM learners may start out having greater difficulty with shortening unstressed function words, lengthening strong content words, and lengthening final syllables at earlier stages of acquisition. Not all of the differences between ESL and EFL speakers translate into difficulties with speech rhythm. Simply because EFL and ESL differ from each other on the relative duration of one syllable does not mean either of them produce the syllable significantly different from English speakers. For example, the syllable me in sentence A6 is found to be significantly different between ESL and EFL, but not between NS and ESL or NS and EFL. This is a typical example of differences between ESL and EFL that do not amount to difficulties for ESL or EFL. Differences that are interpreted as difficulties are syllables that are significantly different between NS and ESL and/or between NS and EFL. When a syllable is found to be significantly different both between NS and ESL and between NS and EFL, it implies a difficulty common to both ESL and EFL. When a syllable is found to be significantly different between NS and ESL but not between NS and EFL, it implies a difficulty unique to ESL speakers. Similarly, when a syllable is found to be significantly different between NS and EFL but not between NS and ESL, it implies a difficulty unique to EFL speakers. Figure 4.3 Distribution ofsignificant differences between ESL and EFL in terms of the spread of difficulties to ESL and/or EFL Among the 17 syllables that are found to be significantly different between ESL and EFL, nine suggest difficulties common to both ESL and EFL, five suggest difficulties unique to EFL speakers, one suggests difficulties unique to ESL speakers, and two indicate difficulties to neither ESL nor EFL speakers. These results suggest that whenever a significant difference in syllable duration is found between ESL and EFL, it is more likely for EFL than for ESL to be the group having difficulty. It also suggests that if ESL is having difficulty, EFL will very likely manifest this difficulty as well, but the reverse will not apply. 4.1.4 CORRELATION OF DURATIONPATTERNS BETWEEN SPEAKER GROUPS Pearson Product-moment correlation coefficients were obtained to estimate how closely syllable duration within the same sentence covaries between pairs of subject groups. Correlation tests were performed between NS and ESL, NS and EFL, and ESL and EFL respectively for all seven type A sentences. Complete results of the correlation tests are summarized in Table 4.7. Table 4.7 Pearson Product-moment Correlation Coefficients for mean syllable duration between groups for Type A sentences Sentence Al A2 A3 A4 AS A6 A7 *P<.05, **p<.Ol NS andESL 0.8767 0.9233* r NS and EFL ESL and EFL 0.6459 0.9268* .....:0:.;,;.6::.::9.,::8::..3 0.9181* 0.8614* There are two major observations. First, the duration patterns of ESL and EFL speakers tend to covary closely. Correlation coefficients between ESL and EFL are found to be significant at p<O.05 in all sentences. Second, the correlations of mean syllable duration are consistently stronger between NS and ESL than between NS andEFL. In addition, a slightly larger number of significant correlations was established between NS and ESL than between NS and EFL speakers. Significant correlations were established at p<O.05 in six sentences between NS and ESL and five sentences between NS and EFL. The results lead to two suggestions. First, ESL and EFL speakers are quite similar with their duration patterns. Second, the duration of syllables tends to vary in a more similar fashion between NS and ESL speakers than between NS and EFL speakers. #### 4.1.5. TEST REUABILITY This section summarizes results of test-retest reliability between the three administrations of all seven Type A sentences. Three reliability indices were obtained for each sentence for each group. High test-retest reliability indicates speakers are consistent with their duration patterns. Table 4.8 Test-retest reliability for syllable duration from three productions for Type A sentences Sentence Administration Test-retest r NS Al 1Sl and 2 00 0.9870** ISland 3'" 0.9860** 0.9703** . 2 00 and 3 rd 0.9993** 0.9721** A2 1Sl and 2 nd 0.9506** 0.9899** 0.9674** ISland 3 n1 0.9827** 0.9726** 0.9699** 2 00 and 3'" 0.9894** 0.9871** 0.8969** A3 1Sl and 2 nd 0.9952** 0.9872** 0.9613** ISland 3 rd 0.9981** 0.9750** 0.9698** 2 nd and 3 n1 0.9974** 0.9880** 0.9778** A4 1Sl and 2 nd 0.9957** 0.9948** 0.9972** ISland 3 rd 0.9962** 0.9881** 0.9875** 2 nd and 3 rd 0.9879** 0.9914** 0.9905** AS 1Sl and 2 nd 0.9927** 0.9975** 0.9831** ISland 3'" 0.9967** 0.9828** 0.9382** 2 nd and 3 n1 0.9971** 0.9831** 0.9576** A6 1Sl and 2 nd 0.9854** 0.9956** 0.9892** ISland 3'" 0.9906** 0.9703** 0.9854** 2 00 and 3 rd 0.9914** 0.9561** 0.9415** A7 1Sl and 2 nd 0.9910** 0.9872** 0.9986** ISland 3 n1 0.9978** 0.9927** 0.9944** 2 nd and 3 rd 0.9822** 0.9799** 0.9648** *p<.05, **p<.Ol Results of the test-retest reliability check indicate that speakers of all groups are highly consistent with their duration patterns in all three productions of Type A sentences. All but one correlation coefficient are statistically significant at p<O.05. In fact, all but three of the obtained correlations are significant at p<O.Ol. Significant correlation could not be established between the first and the third productions of sentence Al for EFL speakers at p<O.05. 4.1.6 SUMMARY OF RESULTS FOR TYPE A SENTENCES The absolute lengths of the syllables are on average shortest for NS, longest for EFL and in between for ESL. Despite differences in speech rates, NS, ESL and EFL speakers generally lengthen and shorten very similar syllables. The longest syllables of the sentences are usually the first strong syllables or the final syllables although weakly stressed syllables are sometimes lengthened. The longest syllables of the sentences are relatively longer and the shortest syllables are relatively shorter for NS than for ESL and EFL speakers. NS produce the widest ranges and the greatest standard deviations of duration percentages in all Type A sentences, as opposed to EFL speakers, who produce the narrowest ranges and the smallest standard deviations of duration percentages. Lengthening of the final syllables is common for speakers in all groups, but TM speakers do not produce as much lengthening as English speakers. ESL and EFL show the same types of difficulties with duration as a correlate of stress. Significant differences between NS and ESL and between NS and EFL indicate that ESL and EFL sometimes produce relatively shorter strong syllables, relatively longer weak syllables in non-final positions, and relatively shorter final syllables than native English speakers. ESL speakers evidence less difficulty with duration than EFL speakers. A larger number of syllables are found to be significantly different between NS and EFL than between NS and ESL. Most of all, ESL speakers produce fewer instances of each type of the identified difficulties. Relatively longer weak syllables account for the majority of differences between TM and English speakers. Relatively longer weak syllables produced by EFL speakers also account for the majority of differences between ESL and EFL speakers. All of the significant differences between ESL and EFL speakers indicate shorter strong syllables, longer weak syllables; and shorter final syllables on the part of EFL speakers. In addition, the majority differences between ESL and EFL indicate difficulties common to both ESL and EFL, or difficulties unique to EFL. The duration patterns of ESL and EFL speakers are higWy correlated. Despite great similarities between these two groups of TM speakers, ESL speakers produce more 93 native-like duration patterns than EFL speakers. The duration of syllables covaries more closely between NS and ESL than between NS and EFL. Speakers of all groups are highly consistent with their duration patterns across three productions of Type A sentences. All but one test-retest reliability index between two administrations of the same sentence fall below significance at p<O.05. 4.2 DURATION PATTERNS OF TYPE B SENTENCES This section reports results from the duration patterns for the Type B sentences, which feature a highly regular rhythmic pattern of alternating strong and weak syllables. Each strong syllable is immediately preceded and/or followed by an weak syllable, and vice versa. All sentences were embedded in broad-focused contexts to encourage speakers to introduce all words that carry content information as new and strong. The alternating stress pattern is further reinforced with lexical category. Researchers have long divided words into two lexical categories: content words vs. function words. Content words are words that denote objects, properties, or actions and function words are words that serve purely grammatical functions. For Type B sentences, all strong syllables are content words, which are more likely to bear stress than the function words and morphemes, which make up the weak syllables. Three different rhythmic patterns are represented in Type B sentences. Sentences B1 through B4 feature the six-syllable iambic rhythmic pattern wSwSwS, sentences B5 and B6 feature the seven-syllable trochaic rhythmic pattern SwSwSwS, and sentence B7 features the eight-syllable iambic rhythmic pattern wSwSwSwS. The total number of utterances analyzed is 630 and the total number of syllables analyzed is 4140. 4.2.1 ABSOLUTE DURATION OF SYLIABLES This section summarizes results based on the absolute durations of the individual syllables. A group average duration in milliseconds was obtained for each syllable of every test sentence. Overall observations based on the lengthening and shortening of syllables across sentences and the basic statistics, including means, standard deviations, ranges of absolute durations, and speech rates, are reported. Table 4.9 Mean syllable durations in ms for Type B sentences | | | | | | | | | | llable S | | | | | | sec | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | Bl | | I | | need | | it | | back | | b | | | | | | | NS | | 108.00 | | 151.67 | | 132.00 | | 243.00 | | 110.00 | | | | | | | ESL | | 142.33 | | 218.00 | | 189.00 | | 289.67 | | 138.67 | | | | | | | EFL | | 162.00 | | 307.33 | | 313.00 | | 338.67 | | 196.00 | | | | | | | B2 | | They | | play | | with | | dad | | and | | | | | | | NS | | 108.67 | | 231.00 | | 118.33 | | | | 147.00 | | | | | | | ESL | | 124.33 | | 301.00 | | 154.33 | | | | 207.67 | | | | | | | EFL | | 186.00 | | 338.00 | | 228.33 | | | | 294.67 | | | | | | | B3 | | We | | learn | | to | | read | | and | | | | | | | NS | | 127.33 | | 237.00 | | 118.33 | | 173.33 | | 119.33 | | | | | | | ESL | | 135.00 | | 324.67 | | 167.00 | | 287.00 | | 187.67 | | | | | | | EFL | | 199.00 | | | | 269.33 | | 304.00 | | 242.33 | | | | | | | B4 | | I | | met | | with | | John | | at | | | | | | | NS | | 91.00 | | 180.67 | | 132.67 | | 280.33 | | 114.67 | | | | | | | ESL | | 141.33 | | 257.67 | | 197.33 | | 354.67 | | 188.33 | | | | | | | EFL | | 167.67 | | | | 289.67 | | 404.33 | | 230.67 | | | 337.67 | | | | B5 | | Mom | | | | dad | | were | | mad | | | at | | | | NS | | 247.33 | | | | 202.67 | | 110.33 | | 245.67 | | | 81.00 | | | | ESL | | 300.67 | | | | 264.33 | | 168.67 | | 316.67 | | | 147.00 | | | | EFL | | 294.67 | | | | 286.00 | | 292.67 | | 389.00 | | | 204.33 | | | | B6 | | John | | is | | 000 | | at | | bake | | | .ing | | | | NS | | 239.33 | | 136.67 | | 192.67 | | 157.67 | | 184.67 | | | 142.67 | | | | ESL | | 281.33 | | 177.67 | | 265.00 | | 234.67 | | 237.33 | | | 163.00 | | | | EFL | | 280.33 | | 218.67 | | 348.67 | | 302.67 | | 292.67 | | | 201.67 | | | | B7 | | I | | need | | a | | ride | | to | | | work | | | | NS | | 92.67 | | 157.00 | | 91.67 | | 274.67 | | 111.67 | | | 240.67 | | | | ESL | | 117.33 | | 208.00 | | 117.33 | | 312.00 | | 150.67 | | | 285.00 | | | | EFL | | 152.67 | | 283.67 | | 205.00 | | | | 215.67 | | | 314.33 | | | | 1-7 | | | | | | | | | | | | | | | | | NS | | | | | | | | | | | | | | | | | ESL | | | | | | | | | | | | | | | | | EFL | | | | | | | | | | | | | | | | | | | The | syllables | | | are | generally | | | shortest | | | for | NS, longest for EFL | and in between for | | ESL. | | NS | produce | | | the | shortest | | syllables | | | | across | the board. EFL speakers | produce the | | longest | | | syllables, | | except | | for | the | final | | syllables | | | in sentences B3, B4, | B5, B6, and B7. The | | average | | | syllable | | length | | for | Type | | B | sentences | | | is 185.49 ms for NS, | 243.45 ms for ESL, | | and | | 285.56 | | ms | for | EFL | | speakers. | | | Speech | | | rates, which are the flip | side of the mean | | syllable | | | lengths, | | average | | | fastest | | for | | NS, | slowest | for EFL, and in | between for ESL | | speakers. | | | NS | on | average | | | uttered | | 5.40 | | syllables | | per second, ESL speakers | 4.13 syllables | Mean SD Range Syll per second, and EFL speakers 3.51 syllables per second. TM speakers' speech rates improve with advancement in English proficiency. The variations in absolute duration among syllables of a sentence tend to be greatest for ESL, smallest for EFL, and in between for NS. The average standard deviation of syllable lengths for all Type B sentences is 83.49 ms for NS, 96.64 ms for ESL, and 73.27 ms for EFL speakers. The ranges between the longest and the shortest syllables of a sentence are generally widest for ESL, narrowest for EFL, and in between for NS. The average range of syllable lengths for all Type B sentences is 214.33 ms for NS, 258.95 ms for ESL, and 204.17 ms for EFL speakers. EFL speakers produce syllables that are longest but least variable in length across groups. (a) Mean syllable duration (ms) for sentences Bl (wSwSwS) • NS --ESL .. 'll" EFL | (b) Mean | syllable duration (ms) for sentence 450 | B2 (wSwSwS) | | --- | --- | --- | | .s c 0 :;::l l! :J Q | 400 350 300 250 200 150 100 50 | --ESL EFL .. 'll" | 'iii" • NS --ESL .. 'll" EFL .s 250 c 0 :;::l l! 200 :J Q 150 0 and laugh They play with dad Syllable (c) Mean syllable duration (ms) for sentence B3 (wSwSwS) 400 350 , ... fA. ... , 300 ... .1.'- 250 'iii" E..... • NS 200 --ESL .. 'It. .. EFL :::J a 150 100 50 a We learn to read and write Syllable (d) Mean syllable duration (ms) for sentence B4 (wSwSwS) (e) Mean syllable duration (ms) for sentence B5 (SwSwSwS) 500 450 400 350 '0' 300 ...... E S 250 g 200 150 100 50 a Mom and dad were Syllable mad at Jim • NS --ESL ...... EFL (f) Mean syllable duration (ms) for sentence B6 (SwSwSwS) , .. ... if .. 450 400 350 300 '0' .§. 250 c 0 :;::l l! 200 ::l Q 150 100 50 a John is good at Syllable bake -ing bread • NS --ESL ...... EFL (g) Mean syllable duration (ms) for sentence B7 (wSwSwSwS) Figure 4.4 Mean syllable durations (ms) for sentences Bl through B7 Four observations can be made based on the results summarized in Table 4.9 and Figure 4.4. First, syllables are generally longest for EFL speakers, shortest for NS, and in-between for ESL speakers. Second, NS, ESL, and EFL speakers generally lengthen and shorten the same syllables. The lengthening and shortening of the syllables generally coincides with the alternation between the target strong and weak syllables. Syllables are usually lengthened when strong and shortened when weak except for the final syllable laugh of sentence B2, the syllable dad in sentence B5, and the syllable bake in sentence B6 in the speech of EFL speakers. Third, although EFL speakers lengthen and shorten syllables according to stress most of the time, they sometimes produce a succession of syllables of similar length. For example, the successive syllables need it back in Bl, the syllables mon and dad were in B5, and the syllables at bake in sentence B6 are all of similar length. Fourth, the longest syllables of the sentences are generally the final syllables for NS and ESL. The final syllables are the longest syllables in 6 sentences for NS and ESL, but in only three sentences for EFL speakers. For EFL speakers, the longest syllable can be either the final syllable or the first or second strong syllable of the sentences. 4.2.2 RELATNE DURATIONPAITERNS This section reports results based on the relative durations of the syllables of a sentence. The length of each syllable was converted into a percentage of the lengths of the entire sentence. An average group duration percentage was obtained for each syllable of each sentence. The basic statistics, including means, standard deviations, and ranges, are summarized in Table 4.10. Table 4.10 Mean syllable duration as its percentage of the total sentence duration for Type B sentences | Type | B | sentences | | | | | | | | | | | | | | | | | | | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | GrOll | | | | | | | | | | | | | | | | | | | | | | Bl | | I | | | need | | | it | | | | | | | | | | | | | | NS | | 10.11 | | | 14.19 | | | 12.32 | | | | | | | | | | | | | | ESL | | 10.43 | | | 15.87 | | | 13.67 | | | | | | | | | | | | | | EFL | | 9.47 | | | 18.12 | | | 18.31 | | | | | | | | | | | | | | B2 | | They | | | play | | | with | | | | | | | | | | | | | | NS | | 9.66 | | | 20.34 | | | 10.40 | | | | | | | | | | | | | | ESL | | 8.69 | | | 20.99 | | | 10.70 | | | | | | | | | | | | | | EFL | | 10.96 | | | 20.14 | | | 13.42 | | | | | | | | | | | | | | B3 | | We | | | learn | | | to | | | | read | | | and | | | | | | | NS | | 12.18 | | | 22.71 | | | 11.33 | | | | 16.60 | | | 11.47 | | | | | | | Ii--=E=S=L=----r:-9~.2..;;..1_ | | | | | 21.90 | | | 11.30 | | | | 19.36 | | | 12.70 | | | | | | | EFL | | 11.80 | | | | | | 15.94 | | | | 18.32 | | | 14.55 | | | | | | | B4 | | I | | | met | | | with | | | | John | | | at | | | | | | | NS | | 7.98 | | | 15.77 | | | 11.62 | | | | 24.58 | | | 10.02 | | | | | | | Ij-:E=.;S::.::L=---+-=-9..:..:.1..;;..3_+-=16:..:..;.7~8--+-=1=2.:..=.80=-- | | | | | | | | | | | | 23.10 | | | 12.07 | | | | | | | EFL | | 9.42 | | | 19.24 | | | 15.99 | | | | | | | 13.05 | | 19.26 | | | | | B5 | | Mom | | | and | | | dad | | | | were | | | mad | | at | | | | | NS | | 17.78 | | | 8.13 | | | 14.50 | | | | 7.85 | | | 17.62 | | 5.83 | | | | | ESL | | 16.10 | | | 11.07 | | | 14.05 | | | | 8.99 | | | 16.94 | | 7.88 | | | | | EFL | | 13.54 | | | 13.42 | | | 13.10 | | | | 13.41 | | | 17.77 | | 9.35 | | | | | B6 | | John | | | is | | | | 000 | | | at | | | bake | | -ing | | | | | NS | | 17.60 | | | 10.05 | | | 14.18 | | | | 11.57 | | | 13.65 | | 10.62 | | | | | ESL | | 15.69 | | | 9.97 | | | 14.62 | | | | 13.25 | | | 13.37 | | 9.24 | | | | | EFL | | 13.76 | | | 10.79 | | | 17.38 | | | | 15.09 | | | 14.40 | | 9.84 | | | | | B7 | | I | | | need | | | a | | | | ride | | | to | | work | | | | | NS | | 6.76 | | | 11.40 | | | 6.60 | | | | 19.84 | | | 8.04 | | 17.32 | | | | | ESL | | 6.96 | | | 12.42 | | | 6.97 | | | | 18.39 | | | 8.86 | | 16.80 | | | | | EFL | | 7.31 | | | 13.52 | | | 9.76 | | | | | | | 10.29 | | 15.14 | | | | | Bl-7 | | | | | | | | | | | | | | | | | | | | | | NS | | | | | | | | | | | | | | | | | | | | | | ESL | | | | | | | | | | | | | | | | | | | | | | EFL | | | | | | | | | | | | | | | | | | | | | The variations in duration (%) among syllables of a sentence are on average greatest for NS, smallest for EFL, and in between for ESL for Type B sentences. The average standard deviation of duration (%) for all Type B sentences is 6.90% for NS, 6.14% for ESL, and 4.03% for EFL speakers. The ranges between the longest and the shortest syllables of a sentence are on average widest for ESL, narrowest for EFL, and in between for NS. The average range of duration (%) for all Type B sentences is 17.60% 102 for NS, 16.34% for ESL, and 11.21% for EFL speakers. The longest syllables of the sentences are relatively longer for NS than for ESL and EFL. The relative duration of the longest syllables of all Type B sentences averages 26.40% for NS, 25.10% for ESL speakers, and 20.90% for EFL speakers. EFL speakers produce syllables that are longest but least variable in length across groups. Figure 4.5 shows the duration patterns of Type B sentences in terms of the percentages of the lengths of the syllables to the lengths of the entire sentences. (a) Mean syllable duration (%) for sentence Bl (wSwSwS) +----------------------JII'----1 +----------------------#1----1 10 +------1I111F---------------_I-------I 20 +--------------J~~,__---____.1I~'--------1 r~;::::::NSl -fA. .. .. 1-. NS c ........ o --ESL 5 15 +------;r~~ ...........o;;;;:_-_7Jr_-----'~-___j------1 .. 'li." EFL 5-1-------------------------1 0-1----,..----.....------,,..-------.-----.----4 need it back Syllable by noon (b) Mean syllable duration (%) for sentence B2 (wSwSwS) ,..---~-------------------__, 25 -1---------------.1-+-------------1 20 ....... e... • NS c 0 15 --ESL ::J .. 'li." EFL Q 10 5+------------------------1 0-1----.,..----.....----,..-------.-----.----4 They play with dad Syllable and laugh (c) Mean syllable duration (%) for sentence B3 (wSwSwS) +----------------------......--1 20 ...... e... • NS c 0 15 --ESL .. ::l 'll" EFL c 10 5-1-------------------------1 O+----...,.-----,----..,.-----,...-----r-----I We learn to read and write Syllable (d) Mean syllable duration (%) for sentence B4 (wSwSwS) 35....-------------------------., -1----------------------__--1 25 +--------------__-------1-..,-------1 ......it 20 +--------.------------,;r:;r-------'tc-------E''---------.----I '~~ml c 1-. NS o f' --ESL 5 15 -I-----I-~~~""_:_-=-_I_---------'~-__#__JL-----I - 'll" EFL 10 +-----,""'#--------------- ------1 5-1-------------------------1 0+----...,.-----,----..,----..,.-----,...----1 met with John Syllable at noon (e) Mean syllable duration (%) for sentence B5 (SwSwSwS) 25 -1-------------------------11111--1 20 ....... :.l! e.... • NS c 0 15 _-ESL . ... EFL :J a 10 5+--------------------=-------1 0+---....,...---.,...---....,...---.,.---....,...---.,.----1 Mom and dad were mad at Jim Syllable (f) Mean syllable duration (%) for sentence B6 (SwSwSwS) _.._-------------------------.. | | 25 | +-------------------------1 | | --- | --- | --- | | | 20 | | | ....... | | | ....... .. :.l! e.... • NS c .. i 15 , --ESL I :J I . ... EFL a 10 5+---------------------------1 bake -ing bread O+---..,....-----.,----r----.,...---.....---..,----! John is good at Syllable (g) Mean syllable duration (%) for sentence B7 (wSwSwSwS) 20 +---------_...-----------_IIIIIIIt--t '0' 15 +----------I----"'f-------I-A-'l~--_I_____:&--I r:::e:::::r;ffil. c o _-ESL :e ...... EFL ::I c 10 +--------..-I-------'~___"\1tr_I_---------'.....~---4 5+------------------------1 O+-----r---.,...--...,.---"..----,.----r---...,.-----! need a ride to Syllable work at once Figure 4.5 Mean syllable durations (%) for sentences B1 through B7 Five observations can be made based on the results summarized in Table4.l0 and Figure 4.5. First, NS and ESL produce highly similar duration patterns, with strong syllables lengthened and weak syllables shortened. Not only do they lengthen and shorten the same syllables, but they do so with very similar duration percentages. Second, the longest syllables of the sentences are all final syllables for NS and ESL speakers except for sentence B2, where the second strong syllable dad is longest for all groups. For EFL speakers, the longest syllables are final syllables in three sentences B1, B5, and B6, and the first or the second strong syllables in other sentences. Third, the non-final strong syllables of EFL and ESL speakers are sometimes longer and sometimes shorter than those of NS, but their final strong syllables are always much shorter than those of NS speakers. This suggests that they might have difficulty with final lengthening. Fourth, EFL speakers generally produce relatively longer weak syllables than NS. 19 of the 22 weak syllables for Type B sentences are relatively longer for EFL speakers than for NS. All of these weak syllables are monosyllabic function words. Fifth, the relative duration of the syllables varies within the narrowest ranges for EFL speakers. The relatively narrower ranges are largely due to the substantially less lengthening of their final syllables, as well as their relatively longer weak syllables. 4.2.3 SIGNIFICANT DIFFERENCES BETWEENNS, ESLAND EFL T-tests were performed on individual syllables to determine whether or not an obtained difference in relative duration between pairs of speaker groups was statistically significant or likely by chance. Results of the t-test scores for the duration percentages of individual syllables between groups are listed in Table 4.11. Table 4.11 Student's t-test scores for durations (%) of individual syllables between groups for Type B sentences Group Bl I NS-ESL 0.531 0.861 NS-EFL ESL-EFL 1.610 They B2 NS-ESL 1.495 NS-EFL 1.473 2.476 ESL-EFL B3 NS-ESL NS-EFL ESL-EFL I B4 NS-ESL 1.629 1.589 NS-EFL ESL-EFL 0.354 B5 Mom 1.687 NS-ESL NS-EFL ESL-EFL B6 NS-ESL NS-EFL ESL-EFL | NS-ESL | | | | | | --- | --- | --- | --- | --- | | NS-EFL | | | | | | ESL-EFL | | 0.561 | ~0;.,;;.3;,;;5~4=~~~===!:.,;;;;;;;;.;".=~,;,;..;,~= 1.5 | | B7 *p<.05 (tcril=2.10l), **p<.Ol(tcril=2.878), df=18, two-tailed In the following two subsections, we will examine which syllables are significantly different between groups in terms of duration percentages. Section 4.2.3.1 reports and compares differences between NS and ESL and differences between NS and EFL. Results provide crucial information as to the kinds of difficulties TM speakers have with duration as a correlate of stress,as well as the ways in which ESL and EFL speakers are similar or different with respect to their problems with duration. Section 4.2.3.2 examines significant differences between ESL and EFL speakers. Results in this section will provide information about the changes that might have taken place between EFL and ESL speakers in the way duration correlates with stress, if there are any changes. 4.2.3.1 Differences between NS and ESL vs. differences between NS and EFL A much smaller number of syllables are significantly different between NS and ESL than between NS and EFL. Of the 46 syllables compared for Type B sentences, nine were significantly different between NS and ESL speakers and 30 between NS and EFL speakers. Table 4.12 Number of strong and weak syllables with durations (%) significantly different from NS in non-final vs. final positions for Type B sentences Position Non-Final Final Stress Strong Weak Strong Weak Total (k=17) (k=22) (k=7) (k=O) differ Ion er shorter Ion er shorter Ion er Shorter Lon er shorter 1 2 9 IrE=FL===---+--4"-- 30 *Shaded cells dilute the contrast, unshaded ones do not. Nonetheless, not all of the differences between TM and English speakers are disruptive of the speech rhythm of TM speakers. Relatively shorter strong syllables and relativelylonger weak syllables produced by TM speakers are considered differences that weaken the contrast between strong and weak syllables. Relatively longer strong syllables and relatively shorter weak syllables may not be damaging as far as speech rhythm is concerned because of the greater contrast between strong and weak syllables. With this distinction in mind, we found that six of the nine significant differences between NS and ESL constitute differences that dilute the contrast between strong and weak syllables. They include four relatively longer weak syllables in non-final position in sentences B1 (it), B5 (and, at), and B6 (at), and two relatively shorter strong syllables in final position in sentence B4 (noon) and B5 (Jim). Of the 30 syllables that were significantly different between NS and EFL speakers, 26 constitute differences that weaken the contrast between strong and weak syllables. They include five relatively shorter strong syllables in non-final position in sentences B1 (back), B3 (learn), B5 (Mom), B6 (John), and B7 (work), 14 relatively longer weak syllables in non-final position in sentences B1 (it), B2 (with, and), B3 (to, and), B4 (with, at), B5 (and, were, at), B6 (at), and B7(a, to, at), and seven relatively shorter strong syllables in final position in sentences B1 (noon), B2 (laugh), B3 (write), B4 (noon), B5 (Jim), B6 (bread), and B7(once). ESL and EFL share similar types of difficulties with duration, but to different degrees. Both ESL and EFL speakers produce (1) relatively longer weak syllables in non final positions, and (2) relatively shorter strong syllables in final positions. In particular, EFL speakers produce a very high rate of relatively longer weak syllables in non-final positions and relatively shorter final strong syllables. ESL speakers produce fewer instances of each than EFL speakers. Of the 22 weak syllables in non-final positions, 14 are relatively longer than those of NS for EFL speakers, as opposed to two for ESL speakers. And all of the seven final syllables produced by EFL speakers are relatively shorter than those of NS, but only two produced by ESL speakers are relatively shorter than those of NS. Besides having greater difficulties than ESL speakers within these two problem types, EFL speakers sometimes produce relatively shorter strong syllables in non-final position, which we do not see among ESL speakers. The results show that ESL speakers are able to produce native-like duration patterns except for a small number of relatively longer non-final weak syllables and relatively shorter final strong syllables. The results also suggest that less proficient EFL learners haver difficulty shortening weak syllables and lengthening final syllables, but they can improve with increased proficiency and exposure to the target language. 4.2.3.2 Significant differences between ESL and EFL This section reports significant differences in relative duration between ESL and EFL. Syllables that are found to be significantly different between these two groups are further examined under four categories: strong syllables in non-final position, weak syllables in non-final position, strong syllables in final position, and weak syllables in final position. The purpose is to identify changes that might have taken place due to improved proficiency and increased exposure to English. Table 4.13 Number of strong and weak syllables with duration (%) significantly different between EFL and ESL in non-final and final positions for Type B sentences Position Stress T EFLvs.ESL ESUNSonly EFUNSonly EFUESUNS ..ESUEFL only

Showing the abstract — retrieve the full paper via the Exa API.

Authors

Te-Fang Hua

Topics

EFL/ESL Teaching and LearningSecond Language Acquisition and LearningTHE ACQUISITION OF ENGLISH SPEECH RHYTHM BYADULT CIDNESE ESL AND EFL LEARNERSA DISSERTATION SUBMITTED TO THE GRADUATE DIVISION OF THE UNIVERSITY OF HAWAI'I IN PARTIAL FULFILLMENT OF THE REQUIREMENTS FOR THE DEGREE OF DOCTOR OF PIDLOSOPHY IN LINGUISTICSAugust 2003ByTe-fang HuaDissertation Committee:Ann M. Peters, ChairpersonPatercia J. DoneganKenneth L. RehgJames Dean BrownMartha Crosby©Copyright2003byTe-fang Hua

About

PublishedAug 1, 2003
TypeDissertation
Citations1

Powered by the Exa API