Publication

More on the Validity and Reliability of C-test Scores: A Meta-Analysis of C-test Studies

Jan 1, 2019 · 1 author · 3 topics

Abstract

Hundreds of C-test studies have been published since Klein-Braley's (1981) dissertation work in Duisburg, Germany (Grotjahn, 2016). C-tests are popular because many claim they are easy to develop, administer, and score. C-tests are widely used, and C-test studies vary in crucial ways: C-test scores are used to support different decisions, C-test users interpret scores as measuring different language constructs, and researchers construct and develop C-tests in many ways. Variations across C-test studies can pose unique challenges for C-test use. I report the results of a random-effects meta-analysis of C-test study information. I collect information about the types of decisions C-test scores are used to support, correlation coefficients and score reliabilities to shed light on what C-test scores measure, and information about steps Ctest users took to construct and develop their C-tests. Studies were retrieved from five major search channels. Inclusion-exclusion decisions were made during eligibility and prescreening stages. The main study-coding phase involved five qualified coders, rigorous coder-training procedures (see Stock, 1994) , the double coding of all studies, use of a FileMaker Pro 17 coding form, and assessing coder reliability. In addition to information about descriptive statistics, correlational analyses, and score-reliability estimates, the coding team gathered information about study features, language setting, participants, how C-test scores were used, and how C-tests were constructed and developed; variables in these categories were the basis for subgroup analyses. Score reliabilities were corrected for measurement artifacts, and correlation coefficients were corrected for attenuation before analysis. Following study screening, 239 studies were included in the dataset. Results show that, when study effects are grouped by criterion construct, C-test scores correlate most strongly with scores on other general language proficiency tests (r = .94 [.87; .97] ). Too few study effects were available to examine the relationship between decisions made on the basis of C-test scores and the magnitude of correlation coefficients. Key findings regarding C-test construction and development steps are that reliabilities are higher when users explore alternate deletion rules, use a series of dashes to indicate deleted letters or syllables, use alternate-answer scoring schemes, analyze scores with factor models, and order C-test texts randomly. There was little evidence of bias in study results. v ACKNOWLEDGMENTS My biggest thanks go to my committee members-Jeff Connor-Linton, John Norris, and Meg Malone-and John Davis during my first few years. Thanks for sticking things out with me despite my very many shortcomings, missed deadlines, and in-and out-of-class screw-ups (e.g., redesigning websites, finishing conference presentation PPTs at the last possible minute, skipping out on webinars, etc.). Thanks for all the office chats and phone calls to keep me moving along, great classes, insightful comments on course papers and program documents, AELRC office coffee… the list is long! You have all changed my thinking in one way or another, both academically and in other ways, and I am keen to get back into the classroom to use what I have learned to improve learning experiences for real people in real programs. I am very grateful for all the help and support from several members of the AELRC team during this dissertation. Solid fist-bumps to Amy Kim, Ayşenur Sağdıç, Yasser Teimouri, Derek Reagan, and Brad Salen for all the helpful chats and emails. More thanks go to Meg Malone for being such a constant throughout this dissertation work, always responding to emails quickly, and entertaining ideas and questions. I am leaving the program with two pearls of practice and one of wisdom from Meg: write a little bit each day, stay positive, and "If you don't ask, the answer's always no." This dissertation work was generously supported (financially and otherwise) by the AELRC (project number: P229A180017). I also owe a mountain of gratitude to several professors outside of Georgetown. In particular, I want to thank Ji-Seung Yang, Gregory Hancock, and Mike Linacre for answering my many questions both via email and in person. I also want to thank Dave Wilson at George Mason University for sharing one of his FileMaker apps so that I could reverse engineer the coding app vi for this dissertation. Thanks to Sunyoung Lee-Ellis for sharing Winsteps control files so that I could start to learn about IRT via the Rasch model. Thanks to my wife, Liz, for just putting up with me throughout this whole process (an incredibly arduous task and a real testament to her love and patience). Thanks for letting me leave the house early in a grumpy flurry every single morning, drink a bit too regularly in the evenings, and miss more than a few social outings to put my feet up and watch X-Files and Firefly. Your turn here in the next year. My thanks to the inter-library loan (ILL) staff in Georgetown's Lauinger Library. Dana Aronowitz helped me procure C-test studies all around the globe and in an exceptionally timely fashion. We also got to having some fun conversations via email, thanks to ILL's documentrequest forms. While taking care of some family affairs in the summer of 2018, Dana mailed me hard copies of more than a dozen C-test studies. She made the study-retrieval process of this dissertation much less painful. I also want to thank a few other people for help along the way. My thanks to Alison Mackey for always asking me how I was doing and where I was at in my program whenever we ran into each other. Thanks to Lara Bryfonski for being a killer classmate and for always casting a second pair of eyes on whatever I was writing. Thanks to Jennie for the many chats, beanies, and Christmas cards. Thanks to Erin for always letting me drop in for a few minutes to ask about deadlines and get advice on program milestones these five years. Thanks to Nic Subtirelu for putting up with my being a little despondent in my last semester of writing. Thanks to Lourdes Ortega for helping me get back on my feet with a solid lunch and face-to-face in the faculty lounge after that damn grant was canceled. Thanks to Steve Tarlow at Biostat for helping me get my analyses squared away.

Showing the abstract — retrieve the full paper via the Exa API.

Authors

Todd McKay

Topics

Diverse Approaches in Healthcare and Education StudiesGrit, Self-Efficacy, and MotivationHealth and Wellbeing Research

About

PublishedJan 1, 2019
TypeReview
Citations7

Powered by the Exa API