SHOULD YOU USE THE ITEM MEAN OR MEDIAN? As mentioned previously, when the anchor distribution is negatively skewed, the mean will always be lower than the median, as it is for all three items displayed in the previous blog. The reason is that the mean is drawn toward the few extremely low ratings of SD. The mean is sensitive to extreme scores. (Statisticians’ Concern: For as long as statisticians can remember, the mean has always had an attraction for extreme scores. Granted, this affinity for outliers is not normal. However, statisticians have tolerated this abnormal relationship for years, but also felt compelled to create another index that is not so easily swayed: the median. You may now resume this paragraph already in progress.) Depending on the degree of skew, the bias in interpreting the mean can be significant or insignificant.
DIRECTION AND DEGREE OF ITEM MEAN BIAS: The problem is that the mean misrepresents the actual ratings in a negatively skewed distribution by portraying lower class ratings than actually occurred. This bias MAKES THE INSTRUCTOR APPEAR WORSE in teaching performance, on all of the items, than the students’ ratings indicate. This is not very desirable, especially if these results are used for summative decisions by your department chair or associate dean.
Although means are reported on most commercially published scales, it is strongly recommended that MEDIANS SHOULD BE REPORTED ALONG WITH THE MEANS. Although the median is less discriminating as an index, it is more accurate, more representative, and less biased than the mean for markedly skewed distributions. The lower the degree of skew, the more similar both measures will be. In a perfectly normal distribution, the mean and median are identical. However, keep in mind that ratings of faculty, administrators, courses, programs, and fast food are typically skewed. Therein lays the importance of picking the right index.
BOTTOM LINE RECOMMENDATION: Use both mean and median.
INTERPRETATION: PROFILE OF STRENGTHS AND WEAKNESSES: Since the item means/medians are based on the total N for the class, they can be compared. They display a profile of strengths and weaknesses related to the different teaching behaviors and course characteristics. On a 0−3 scale, means/medians above 1.5 indicate strengths; those below 1.5 denote weaknesses. The means/medians in conjunction with the anchor percentages provide meaningful diagnostic information on areas that might need attention. Again, this report is intended for the instructor's use primarily, although the results on course characteristics may have curricular implications.
Next, the meaning and uses of subscale and total scale scores will be discussed. They are basically summaries of the item scores. Hope you’re finding this stuff helpful. If not, let me know.
COPYRIGHT © 2010 Ronald A. Berk, LLC
How do you interpret the percentage distribution across the anchors? Is a skewed distribution a good or bad sign for teaching performance? How do any of these results relate to your teaching? Keep perusing, dear colleague.
SKEWED DISTRIBUTION: The overall response pattern shown in the previous blog indicates a negatively skewed distribution of responses, which is the most common outcome and typically a desirable one as well. It occurs when the majority of the responses are A and SA, but there is also a sprinkling of a few Ds and SDs. The extreme SD responses, or outliers, create the skew. Realistically, a few students might mark SD to every statement to express their desire to see you whacked, while the majority of satisfied customers will choose the two “Agree” anchors. These distributions occur in more than 90% of the courses I’ve reviewed at different institutions. It’s rare to receive ratings without any Ds or SDs. Other anchors may yield a different pattern of responses.
DIAGNOSTIC PROFILE: This anchor distribution information provides you with the most detailed profile of responses to a single item. The percentages reveal the degrees of agreement and disagreement with each statement. It is diagnostic of how the class felt about each behavior or characteristic. Examine each distribution carefully to pinpoint your strength behaviors (high percentage of SAs and As) and your weakness behaviors (relatively high percentages of SDs and Ds). You can then consider specific changes in your teaching, evaluation, or course behaviors to shift the distribution farther to the right, in the A–SA zone, the next time the class is taught.
The next blog will examine responses at the item level. Is that info of any value after reviewing the anchor distributions? I will solve that mystery!
COPYRIGHT © 2010 Ronald A. Berk, LLC
RESPONSE RATE WARNING: As you begin to analyze the results from your class evaluations, please consider the response rate in your interpretations:
1. For class sizes of 30 to Super Bowl attendance, it is desirable to have at least 70% response, preferably 80 or 90%, to assure reasonably representative ratings. Anything less may be biased (aka "evil") in some unknown direction. Since all responses are anonymous, there is no way to assess the degree and direction of response bias. Just be cautious in your inferences about your teaching behaviors.
2. For classes less than 30, especially seminars of 5 to 10 students, be particularly careful in your interpretations based on both a less than desirable response rate and inadequate number of responses.
BOTTOM LINE: When your results are used to guide teaching improvement, view your ratings as suggestive of possible areas for change rather than as conclusive.
WHAT’S THE 1ST LEVEL OF INTERPRETATION?
IT’S ANCHOR-WORLD! The first level of score reporting is anchor results for each item. The anchors are the response options on the scale, such as STRONGLY DISAGREE or DISAGREE. Usually the percentage of students picking each anchor is reported, item by item. An example is shown below for student rating scale items with agree–disagree anchors:
SD D A SA N
Statement 1 1.0% 3.1% 37.5% 58.6% 96
Statement 2 1.0 3.1 24.0 71.9 96
Statement 3 1.1 1.1 28.9 68.9 90
ANCHOR SCALE: These results can be reported for any word, phrase, or statement stimuli and any response anchors. The anchor abbreviations for "Strongly Disagree" (SD), "Disagree" (D), "Agree" (A), and "Strongly Agree" (SA) are listed horizontally, left to right, from unfavorable to favorable ratings. This is the same order as the original scale. These anchors measure the degree or intensity of your feeling toward each statement. The N is the number of students that responded to the statement.
Other anchors may ask you evaluate the quality of a behavior, how frequently a behavior occurs, the quantity or the extent to which a behavior occurs, or how a behavior in one course compares to a behavior in another course. There are a variety of possible anchors and number of anchors on the scales now in use.
PERCENTAGE RESPONSES: The percentages for the four anchors indicate the percentage distribution based on the actual N. When the statements are positive teaching behaviors or course characteristics, you would expect low percentages for the first two “Disagree” anchors and high percentages for the second two “Agree” anchors, with the highest for SA. The percentages taper off drastically from right to left, with tiny percentages for D and SD for all three items. (NOTE: The percentages are slightly different for statement 3 compared to statements 1 and 2, particularly the 1.1% to SD and D. This was due, in part, to the six fewer students who responded to that item [N = 90].)
The next blog will examine the meaning of these distributions in terms of a diagnostic profile of your teaching strengths and weaknesses. Stay on board.
COPYRIGHT © 2010 Ronald A. Berk, LLC