Showing posts with label
student evaluation of teaching.
Show all posts
Showing posts with label
student evaluation of teaching.
Show all posts
WHAT ARE TOTAL SCALE SCORES?
The highest level of score summary is the total scale score across all items. It’s like the total score on a test, except there are no right and wrong answers on a scale. If the scale consists of 36 items, each scored 0–3, the following results might be reported:
Total scale score range = 0–108; Midpoint = 54
Mean/Median = 96.43/101, where N = 97
The continuum for interpretation would be the following:
Extremely Mean Extremely
Unfavorable Neutral 96.43 Favorable
0__________________54___________________108
Mdn
101
INTERPRETATION OF TOTAL SCORES: This score gives a global, or composite, rating that is only as high as the ratings in each of its component parts (anchors, items, and subscales). It represents one overall index of teaching performance, from Extremely Unfavorable to Extremely Favorable. In this example, the score is very favorable. However, the total score is usually a little less reliable and less informative than the subscale scores.
COMPARISON OF ITEM, SUBSCALE, AND TOTAL SCORE SCALES
A comparison of the quantitative scales at the previous levels for a total scale is shown here:
Score Level Ex Unfav Neutral Ex Fav
Item 0 ___________1____ 1.5 ____ 2___________3
Subscale 1(8 items) 0 ________________ 12 __________________24
Subscale 2(5 items) 0 ________________ 7.5 _________________15
Subscale 3(4 items) 0 _________________ 6 __________________12
Subscale 4(13 items)0 ________________19.5 ________________39
Subscale 5(6 items) 0 _________________ 9 __________________18
Total Scale(36 items)0 _______________ 54 __________________108
These levels of score reporting and interpretation are based on a single course. This information has the greatest value to you, your department chair, and the curriculum committee evaluating the course.
How does the total score compare to the global item scores given at the end of the scale? Which score should you use? Is there really any difference?
COPYRIGHT © 2010 Ronald A. Berk, LLC
WHAT ARE SUBSCALE SCORES?
This is the first level at which item scores can be summed. If the items on the total scale are grouped into clusters related to topics such as instructional methods, evaluation methods, and course content, the item scores can be summed to produce subscale scores. These scores should only be used for decision making if adequate validity and reliability evidence support that internal scale structure. Since each subscale contains a different number of items, the score range will also be different. This range must be reported to interpret the results.
The summary of subscale score results is derived from the item score results. The statistics are the same. We are just aggregating or summarizing the item-level data (0–3) into subscale item clusters. For example, here are results for three selected subscales:
Instructional Methods (IM)—13 items
Subscale score range = 0–39; Midpoint = 19.5
Mean/Median = 34.79/37.00, where N = 97
Evaluation Methods (EM)—5 items
Subscale score range = 0–15; Midpoint = 7.5
Mean/Median = 13.40/15.00, where N = 97
Course Content (CC)—8 items
Subscale score range = 0–24; Midpoint = 12
Mean/Median = 21.97/24.00, where N = 97
COMPUTATION OF SUBSCALE SCORES: The interpretation of subscale results is analogous to the item results; only the numbers are BIGGER. For example, instead of a 0–3 range and a midpoint of 1.5 for an item, each subscale has a range and midpoint based on its respective number of items. So, for the Instructional Methods (IM) subscale with 13 items, a "0" (SD) response to every item produces a sum of 0 for the subscale, and a "3" (SA) response to all 13 items yields a sum of 39.
INTERPRETATION OF SUBSCALE MEANS AND MEDIANS: The zero-base for all score interpretations is easy to remember: the worst, most unfavorable rating on any item, subscale, or total scale is "0." What changes is the upper score limit for the most favorable rating on each subscale because the number of items change. Again, for the IM subscale, the mean and median can be referenced to the upper limit of 39 and also the midpoint of 19.5 to locate the position on the continuum, as indicated below:
Extremely Mean Extremely
Unfavorable Neutral 34.79 Favorable
0___________________19.5____________________39
Mdn
37
The mean/median ratings on the IM subscale are very favorable.
The subscale results can pinpoint areas of strength and weakness. They may be used by your department chair or the promotion review committee to identify your teaching strengths across different courses. Subscale scores cannot direct you toward particular aspects of teaching that can be improved or changed. The item and anchor results described previously are intended to provide that detailed level of direction.
Finally, the next blog will examine total scores on the scale. What additional info do they provide beyond what we already know?
COPYRIGHT © 2010 Ronald A. Berk, LLC
SHOULD YOU USE THE ITEM MEAN OR MEDIAN? As mentioned previously, when the anchor distribution is negatively skewed, the mean will always be lower than the median, as it is for all three items displayed in the previous blog. The reason is that the mean is drawn toward the few extremely low ratings of SD. The mean is sensitive to extreme scores. (Statisticians’ Concern: For as long as statisticians can remember, the mean has always had an attraction for extreme scores. Granted, this affinity for outliers is not normal. However, statisticians have tolerated this abnormal relationship for years, but also felt compelled to create another index that is not so easily swayed: the median. You may now resume this paragraph already in progress.) Depending on the degree of skew, the bias in interpreting the mean can be significant or insignificant.
DIRECTION AND DEGREE OF ITEM MEAN BIAS: The problem is that the mean misrepresents the actual ratings in a negatively skewed distribution by portraying lower class ratings than actually occurred. This bias MAKES THE INSTRUCTOR APPEAR WORSE in teaching performance, on all of the items, than the students’ ratings indicate. This is not very desirable, especially if these results are used for summative decisions by your department chair or associate dean.
Although means are reported on most commercially published scales, it is strongly recommended that MEDIANS SHOULD BE REPORTED ALONG WITH THE MEANS. Although the median is less discriminating as an index, it is more accurate, more representative, and less biased than the mean for markedly skewed distributions. The lower the degree of skew, the more similar both measures will be. In a perfectly normal distribution, the mean and median are identical. However, keep in mind that ratings of faculty, administrators, courses, programs, and fast food are typically skewed. Therein lays the importance of picking the right index.
BOTTOM LINE RECOMMENDATION: Use both mean and median.
INTERPRETATION: PROFILE OF STRENGTHS AND WEAKNESSES: Since the item means/medians are based on the total N for the class, they can be compared. They display a profile of strengths and weaknesses related to the different teaching behaviors and course characteristics. On a 0−3 scale, means/medians above 1.5 indicate strengths; those below 1.5 denote weaknesses. The means/medians in conjunction with the anchor percentages provide meaningful diagnostic information on areas that might need attention. Again, this report is intended for the instructor's use primarily, although the results on course characteristics may have curricular implications.
Next, the meaning and uses of subscale and total scale scores will be discussed. They are basically summaries of the item scores. Hope you’re finding this stuff helpful. If not, let me know.
COPYRIGHT © 2010 Ronald A. Berk, LLC
WHAT ARE ITEM SCORES?
The next level is the item, where a statistic such as a mean or median is reported. Since most anchor distributions are usually negatively skewed and answers are on a ranked, or ordinal, scale, the median is the most appropriate measure of central tendency. However, given the range of distributions that can occur, you may see both the mean and median on your report form.
“WAIT!! Back up. How did you get from responses of SD, D, etc. to means and medians?” Great question! Glad you’re on the ball. First, you have to convert the “verbal” anchors into “numbers.”
(MEASUREMENT ALERT: Keep in mind that we started with a “qualitative scale” of verbal expressions of how students feel about each behavior and now we’re converting the words into a “quantitative scale” for the convenience of performing analysis of those feelings. Actually, this conversion involves an arbitrary numerical coding scheme.)
CREATE A ZERO-BASED NUMERICAL SCORE SCALE: For simplicity and interpretability, a zero-based scale is recommended, so that the most negative anchor, such as SD, would be coded as “0.” Zero-based scoring was originally recommended by Likert (1932), who created this scaling method. Then the other anchors would be coded in 1-point increments above 0.
Higher values weight more desirable or positive ratings higher than negative ones. SA or Strongly Agreeing with a desirable teaching behavior or course characteristic is weighted with the highest value of 3. An example of this coding for a 4-point, agree–disagree scale is shown below:
SD D A SA
0 1 2 3
The score range for this single item is 0 to 3. (Note: These score points will vary with the number of anchors and the base number on different scales. Yours may be one of these. Sometimes the number 1 is used as the base instead of 0. Although the number scale may be different, the final interpretation will be the similar.)
COMPUTATION OF ITEM MEANS AND MEDIANS: If you hate stat, this section may make you hurl. Skip it. (SIDEBAR: Over 30 years of teaching stat, I had lots of student hurlers.) For you interested nonhurlers, here are the simple computational definitions:
MEAN = the sum of all students’ scores to each item, divided by the number of students or N. This is the average score for an item, within the range of 0–3 for this example.
MEDIAN = the middle score, after all students’ scores are ranked from high to low.
An example report, based on the anchor data shown in the previous blog, is shown below:
SD D A SA N Mean Median
Statement 1 1.0% 3.1% 37.5% 58.6% 96 2.52 3.00
Statement 2 1.0 3.1 24.0 71.9 96 2.65 3.00
Statement 3 1.1 1.1 28.9 68.9 90 2.57 3.00
The median score of 3 means the typical student in the middle of the distribution rated those behaviors as SA. The means were slightly lower with ratings between A and SA. Those are very respectable scores. Of course, they are consistent with the anchor percentage distribution, where the highest percentages are concentrated on the A and SA anchors.
So which index should you use? Mean? Median? Or both? Ah ha! The statistical plot thickens. Stay tuned…
COPYRIGHT © 2010 Ronald A. Berk, LLC
How do you interpret the percentage distribution across the anchors? Is a skewed distribution a good or bad sign for teaching performance? How do any of these results relate to your teaching? Keep perusing, dear colleague.
SKEWED DISTRIBUTION: The overall response pattern shown in the previous blog indicates a negatively skewed distribution of responses, which is the most common outcome and typically a desirable one as well. It occurs when the majority of the responses are A and SA, but there is also a sprinkling of a few Ds and SDs. The extreme SD responses, or outliers, create the skew. Realistically, a few students might mark SD to every statement to express their desire to see you whacked, while the majority of satisfied customers will choose the two “Agree” anchors. These distributions occur in more than 90% of the courses I’ve reviewed at different institutions. It’s rare to receive ratings without any Ds or SDs. Other anchors may yield a different pattern of responses.
DIAGNOSTIC PROFILE: This anchor distribution information provides you with the most detailed profile of responses to a single item. The percentages reveal the degrees of agreement and disagreement with each statement. It is diagnostic of how the class felt about each behavior or characteristic. Examine each distribution carefully to pinpoint your strength behaviors (high percentage of SAs and As) and your weakness behaviors (relatively high percentages of SDs and Ds). You can then consider specific changes in your teaching, evaluation, or course behaviors to shift the distribution farther to the right, in the A–SA zone, the next time the class is taught.
The next blog will examine responses at the item level. Is that info of any value after reviewing the anchor distributions? I will solve that mystery!
COPYRIGHT © 2010 Ronald A. Berk, LLC
RESPONSE RATE WARNING: As you begin to analyze the results from your class evaluations, please consider the response rate in your interpretations:
1. For class sizes of 30 to Super Bowl attendance, it is desirable to have at least 70% response, preferably 80 or 90%, to assure reasonably representative ratings. Anything less may be biased (aka "evil") in some unknown direction. Since all responses are anonymous, there is no way to assess the degree and direction of response bias. Just be cautious in your inferences about your teaching behaviors.
2. For classes less than 30, especially seminars of 5 to 10 students, be particularly careful in your interpretations based on both a less than desirable response rate and inadequate number of responses.
BOTTOM LINE: When your results are used to guide teaching improvement, view your ratings as suggestive of possible areas for change rather than as conclusive.
WHAT’S THE 1ST LEVEL OF INTERPRETATION?
IT’S ANCHOR-WORLD! The first level of score reporting is anchor results for each item. The anchors are the response options on the scale, such as STRONGLY DISAGREE or DISAGREE. Usually the percentage of students picking each anchor is reported, item by item. An example is shown below for student rating scale items with agree–disagree anchors:
SD D A SA N
Statement 1 1.0% 3.1% 37.5% 58.6% 96
Statement 2 1.0 3.1 24.0 71.9 96
Statement 3 1.1 1.1 28.9 68.9 90
ANCHOR SCALE: These results can be reported for any word, phrase, or statement stimuli and any response anchors. The anchor abbreviations for "Strongly Disagree" (SD), "Disagree" (D), "Agree" (A), and "Strongly Agree" (SA) are listed horizontally, left to right, from unfavorable to favorable ratings. This is the same order as the original scale. These anchors measure the degree or intensity of your feeling toward each statement. The N is the number of students that responded to the statement.
Other anchors may ask you evaluate the quality of a behavior, how frequently a behavior occurs, the quantity or the extent to which a behavior occurs, or how a behavior in one course compares to a behavior in another course. There are a variety of possible anchors and number of anchors on the scales now in use.
PERCENTAGE RESPONSES: The percentages for the four anchors indicate the percentage distribution based on the actual N. When the statements are positive teaching behaviors or course characteristics, you would expect low percentages for the first two “Disagree” anchors and high percentages for the second two “Agree” anchors, with the highest for SA. The percentages taper off drastically from right to left, with tiny percentages for D and SD for all three items. (NOTE: The percentages are slightly different for statement 3 compared to statements 1 and 2, particularly the 1.1% to SD and D. This was due, in part, to the six fewer students who responded to that item [N = 90].)
The next blog will examine the meaning of these distributions in terms of a diagnostic profile of your teaching strengths and weaknesses. Stay on board.
COPYRIGHT © 2010 Ronald A. Berk, LLC
ADMINISTRATION OF STUDENT RATING FORMS
Yup! It’s that time of the year. The Cherry Blossoms are gone and many professors on the east coast are sneezing and wheezing their brains out from the sky-high pollen counts. They’re medicating themselves with megadoses of antihistamines like Benadryl, Zyrtec, Claritin, Allegra, Clarinex, Flonase, Nazonex, and Sneezwhizz.
This is the perfect time to administer those end-of-course student rating forms and interpret the scores. The results usually appear better when you’re drowsy and punchy from those medications. This month has the highest number of forms administered world-wide, except for December when the same professors are sneezing and wheezing from the common cold.
The forms may be administered in class or online, but the results have to be reported in some format. You may receive the results in a couple of days to several months, depending on your processing system. Of course, all of this happens in between commencement exercises and end-of-year parties. Hopefully, you can glance at the form results before your next course begins in the summer or fall.
YOUR "RATING ANGEL"
I am your Rating Angel for this next week. If you’re not sure how to interpret the scores or use the results, this blog series is for YOU! I want you to milk those ratings for all their worth, to squeeze every drip of information that can guide your teaching improvement.
If you already know how to interpret these scores, STOP reading this blog immediately. Disregard it and get back to work, class, or lunch. Stop fooling around and wasting time on this blog. You should be ashamed of yourself.
SOURCE ALERT
My Thirteen Strategies book goes into considerable detail on the how, why, computing, reporting, formatting, and who cares about those scores. This series will simply focus on the scores you can use to guide decisions about your teaching improvement.
GOAL OF BLOG SERIES
This blog series is intended to present the BerksNotes® version of student rating scale interpretation. Hopefully, by cutting to the chase, whatever that is, you’ll be able to get the most out of your scores in lickety-split time or faster. Hold on to your keyboards. My next blog will begin with an overview of the different types of scores.
COPYRIGHT © 2010 Ronald A. Berk, LLC
I’m sorry to interrupt my blog series on “How to Put Pizazz into Your PowerPoint Conference Presentations,” but an article was just published that might be useful to some of you.
Are you struggling with issues related to student ratings of teaching, peer evaluation, teaching portfolio, and which gifts to buy for your family for the holidays? Well, have I got a deal for you. I have 2 articles recently published on a faculty evaluation model that might interest you. They are an extension of the model described in my Thirteen Strategies to Measure Teaching book (see Stylus Publishing link in right margin), applied specifically to formative and summative decisions about faculty. The model can be used to evaluate teaching performance and professionalism. Here they are:
Berk, R. A. (2009a). Beyond student ratings: “A whole new world, a new fantastic point of view.” Teaching Excellence, 20(1).
Berk, R. A. (2009f). Using the 360° multisource feedback model to evaluateteaching and professionalism. Medical Teacher, 31, 1073–1080.
The 1st article above is a brief description of the model; the 2nd is a full-blown presentation of the model and lit review. Although the latter was written for professors and administrators in medical schools, all of the characteristics are generalizable to any discipline, department, school, or kingdom. The model is simple, straightforward, and easily applied to impress even accreditation reviewers of your evaluation plan in your self-study. An abstract of the article is given below:
This MT article provides an overview of the salient characteristics, research, and practices of the 360° MSF models in management/industry and clinical medicine. Drawing on that foundation, the model was adapted to the specific decisions rendered to evaluate faculty teaching performance and professional behaviors. What remains unchanged in every application is the original spirit of the model and its primary function:
Multisource Ratings→Quality Feedback→Action Plan to Improve→Improved Performance.
Although the ratings were intended for formative decisions, in many cases they have also ended up being used for summative decisions. All of these applications of the 360° MSF model have advantages and disadvantages. In fact, it is possible to distill several persistent and, perhaps, intractable psychometric issues in executing these models. The top 10 issues are described. Although much has been learned during the 80-year history of scaling, 60-year history of faculty evaluation, and 50-year history of the 360° MSF model in management/industry, a lot of work is still necessary to realize the true meaning of “best practices” in evaluating teaching and professionalism.
You can download the articles from my Website (www.ronberk.com) under Publications. They are intended for your own use and research purposes. Please do not distribute to family members or farm animals. The latter may eat them and get sick. Enjoy!
COPYRIGHT © 2009 Ronald A. Berk, LLC
As a continuation of the last blog on the first two reasons for not using global items for summative decisions about faculty, this blog describes the third and most important reason:
3. PROFESSIONAL AND LEGAL STANDARDS: One or 2 global rating item scores alone for major summative decisions about faculty performance are totally inadequate. That administrative practice violates national testing/scaling practices according to the Standards for Educational and Psychological Testing and EEOC Uniform Guidelines on Employee Selection Procedures, plus the rulings from a large corpus of court cases on this topic. Essentially, it's ILLEGAL to make such personnel decisions about faculty. Clearly, these are PERSONNEL decisions about us, not instructional or curriculum decisions. In the case of employee decisions like these, 1 or 2 items do not reflect an accurate assessment of the instructor's job behaviors. A total scale score based on, for example, 35 items defining effective teaching behaviors, or subscale scores on specific areas of teaching competency would satisfy those criteria. A long history of court cases on personnel decisions indicates that the instrument used for personnel decisions must be based on a comprehensive job analysis of the job’s tasks related to a person’s knowledge, skills, and abilities (KSAs). The behaviors listed as items on the total scale satisfy that standard for teaching effectiveness.
Although administrators have used global items in some form for decisions about faculty teaching performance for quite some time, those practices should stop. They have lawsuit written all over them. As noted above, important, possibly career-changing, individual personnel decisions are held to the highest standards professionally and legally, as they should be. If the instructor being violated is a minority or female, be prepared for an EEOC offensive. If you know an administrator who is engaging in such practices, the recommendation is “cease and desist.”
What’s the alternative? Use the total scale score or subscale scores for different areas of performance in conjunction with other measures, such as peer evaluations, self-ratings, and a dozen other possible sources of evidence (for further details, see my October 25, 2009 blog and Thirteen Strategies… Stylus Publisher link in right margin).
Please let me know your thoughts and observations on my recommendations in these blogs.
COPYRIGHT © 2009 Ronald A. Berk, LLC
Please consider the following 2 reasons why the global items may not be the best option for those summative decisions:
1. ACCURACY AND FAIRNESS: After your students have spent 45 hours in your course over the semester, does their rating of 1 item seem to accurately capture the sum total of all of the experiences in your classroom? There is no doubt that the item furnishes information about your performance and the course, but should it be used for summative, super-important decisions about your career? Is 1 item score of 0-4 a fair and reasonable base to infer overall performance? Could you accurately and fairly rate the performance of your administrative assistant, department chair, or dean with 1 item to truly evaluate his or her degree of effectiveness?
2. VALIDITY AND RELIABILITY: Think of student rating scales like the S.A.T. One item from the S.A.T. doesn't provide a valid and reliable score of verbal ability any more than 1 or 2 global items on a rating scale provide a valid and reliable measure of teaching performance. Although the former measures knowledge and the latter measures attitudes or opinions, the psychometric problem is virtually identical. Such data would be tantamount to giving high schoolers 1 or 2 verbal and quantitative items from the S.A.T. and making individual college admission decisions on those scores. Those scores are technically unreliable. One or 2 items are extremely unreliable compared to a summary score derived from 5 subscales or 35 items. Reliability coefficients should be in the .80s-.90s for individual decisions, although coefficients as low as the .60s are acceptable for group-based decisions in research. Typically, a single item does not yield a coefficient in the acceptable range.
The third reason, PROFESSIONAL AND LEGAL STANDARDS, will be covered in the next blog. It is the most important of the 3 reasons. Don’t miss it.
COPYRIGHT © 2009 Ronald A. Berk, LLC
I know what you’re thinking: “I thought you were done with this student rating stuff. Get off it already.” I know, but I rethought my thought after receiving my weekly report that more people read my blogs this past week than at any other time over the past 3 months. Maybe I hit a note; maybe not. Anyway, these topics keep popping up on listservs and workshops.
But that’s not REEEAALLY what you were thinking. It is: “What in the world is a global item?” For those of you who are not of this world and are just passing through, as I am, here is a profile of the global item:
1. It provides a general, broad-stroke indication of teaching performance. It's intended to be an omnibus item, representing the collective judgments on all other items. But it isn't. Only the total scale score does that.
2. It doesn’t address specific teaching and course characteristics.
3. It usually appears at the end of the rating scale and should not be summed with the scores of all other items.
Using a "Strongly Agree-Strongly Disagree" anchor response scale, a couple of examples are given below:
Overall, my instructor is a dirtbag.
Overall, I learned squat in this course.
Of course, you know I’m kidding. Better items are:
Overall, my instructor is a moron.
Overall, this course is putrid.
Usually, there are 1 to 3 items. Frequently, administrators, such as your department chair, associate dean, or emperor or empress, will be encouraged to use the ratings on those items to provide a simple, quick-and-dirty measure of your teaching performance. Those ratings, in lieu of the total scale or subscale scores, are used in conjunction with other information to arrive at summative decisions regarding merit pay, contract renewal for full-time and adjunct faculty, and promotion and tenure recommendations. Are these important decisions about your career and life. You bet!!
Do you want those decisions to be rendered on the basis of 1 or 2 items? “Sure, why not?” Are you kidding me? There are several logical, psychometric, and legal reasons why global items should NOT be used for summative decisions. They will be described in my next blog. Stay tuned for more fun from RatingWorld.
COPYRIGHT © 2009 Ronald A. Berk, LLC
After the last few post-Halloween blogs, you’re probably sick of this topic and the dribbles of info in each blog. Sorry about that. This is the final installment on the “you know what.”
What is the BOTTOMLINE, HANDS-DOWN and UP, BEST STRATEGY TO USE? (Note: You don’t find this type of excitement and suspense in every blog. Brace yourself. Here it comes.) ANSWER: Withhold the posting of final grades. If your registrar is super-efficient at posting grades online in a timely fashion, you’re sunk. That is usually not a problem, because he or she has a bazillion courses to process.
If students are given a 48-hour window following the final exam or project within which to complete the evals, post the grades online on your course Website ASAP. Tell the students in advance both in class and online when they will be posted. That’s the academic version of a “grade (as opposed to movie) trailer” or teaser. This quick post will be perceived as a monster-size carrot or hot fudge sundae to your students. (Note: Certainly there are some courses where students can compute their own final grades, but they need the scores or grades from the final exam or project.)
We can leverage this Net Generation’s characteristics to maximize the response rates we need. Why should they bother to rate our performance? Although there may be some legitimate intrinsic reasons, why take a chance? Go after the extrinsic ones. As noted in my August 2009 blog series, these Net Geners have a need for speed, immediate gratification, and quick feedback on performance, plus they are driven to achieve, feel pressure to succeed, and expect rapid responses from us on everything. Receiving grades “immediately” feeds into these characteristics.
Institutions engaging in this practice over the past few years, including my own, have reported 80–90%+ response rates for most all courses with buckets of typed comments to open-ended or unstructured questions to boot. It might be worth giving it a whirl at your institution.
Let me know your thoughts and experiences with this issue and any other techniques you have used successfully to boost response rates.
COPYRIGHT © 2009 Ronald A. Berk, LLC
Based on the Top 10 List of Strategies in my previous blog, which ones did you pick as MOST EFFECTIVE? Oh, you didn't read my previous blog. Shame, shame. Is this your 1st blog reading? Well, congrats! Welcome to Berk's BlogWorld. I hope you find it informative and fun!
For those of you who did peruse my Top 10, what did you pick? Let’s examine the evidence and see whether you get the prize. (SIDEBAR: “Yo, Blogboy, you didn’t say a prize was involved!” Oops, sorry. It was an afterthought.)
To date, the evidence indicates that students must believe (Do you believe? Kinda like Peter Pan!) that the results will be used for important decisions about faculty and courses in order for the online system to be successful. When faculty assign students in class to complete the evaluation with or without incentives, response rates are high. Finally, withholding early access to grades is supported by students as highly effective, but not too restrictive. This “withholding” approach has been very successful at raising response rates to the 90s at several institutions.
These strategies have boosted response rates into the 80s and 90s. It is a combination of elements, however, as suggested above, that must be executed properly to assure a high rate of students' return on the online investment. The research base and track record of online student ratings belies “low response rate” as an excuse for not implementing such a system in any institution.
While the combo strategy is recommended, there is one BEST technique among the top 10 that can really spike your response rate. It will be examined in the next blog.
Let me know if you have tried other strategies that have worked in your evaluation system.
COPYRIGHT © 2009 Ronald A. Berk, LLC