Showing posts with label global ratings. Show all posts
Showing posts with label global ratings. Show all posts

Monday, November 9, 2009

Why Should GLOBAL ITEM SCORES Not Be Used for Summative Decisions? PART III

Page copy protected against web site content infringement by Copyscape

As a continuation of the last blog on the first two reasons for not using global items for summative decisions about faculty, this blog describes the third and most important reason:



3. PROFESSIONAL AND LEGAL STANDARDS: One or 2 global rating item scores alone for major summative decisions about faculty performance are totally inadequate. That administrative practice violates national testing/scaling practices according to the Standards for Educational and Psychological Testing and EEOC Uniform Guidelines on Employee Selection Procedures, plus the rulings from a large corpus of court cases on this topic. Essentially, it's ILLEGAL to make such personnel decisions about faculty. Clearly, these are PERSONNEL decisions about us, not instructional or curriculum decisions. In the case of employee decisions like these, 1 or 2 items do not reflect an accurate assessment of the instructor's job behaviors. A total scale score based on, for example, 35 items defining effective teaching behaviors, or subscale scores on specific areas of teaching competency would satisfy those criteria. A long history of court cases on personnel decisions indicates that the instrument used for personnel decisions must be based on a comprehensive job analysis of the job’s tasks related to a person’s knowledge, skills, and abilities (KSAs). The behaviors listed as items on the total scale satisfy that standard for teaching effectiveness.

Although administrators have used global items in some form for decisions about faculty teaching performance for quite some time, those practices should stop. They have lawsuit written all over them. As noted above, important, possibly career-changing, individual personnel decisions are held to the highest standards professionally and legally, as they should be. If the instructor being violated is a minority or female, be prepared for an EEOC offensive. If you know an administrator who is engaging in such practices, the recommendation is “cease and desist.”



What’s the alternative? Use the total scale score or subscale scores for different areas of performance in conjunction with other measures, such as peer evaluations, self-ratings, and a dozen other possible sources of evidence (for further details, see my October 25, 2009 blog and Thirteen Strategies… Stylus Publisher link in right margin).

Please let me know your thoughts and observations on my recommendations in these blogs.


COPYRIGHT © 2009 Ronald A. Berk, LLC

Sunday, November 8, 2009

Why Should GLOBAL ITEM SCORES Not Be Used for Summative Decisions? PART II

Page copy protected against web site content infringement by Copyscape

Please consider the following 2 reasons why the global items may not be the best option for those summative decisions:


1. ACCURACY AND FAIRNESS: After your students have spent 45 hours in your course over the semester, does their rating of 1 item seem to accurately capture the sum total of all of the experiences in your classroom? There is no doubt that the item furnishes information about your performance and the course, but should it be used for summative, super-important decisions about your career? Is 1 item score of 0-4 a fair and reasonable base to infer overall performance? Could you accurately and fairly rate the performance of your administrative assistant, department chair, or dean with 1 item to truly evaluate his or her degree of effectiveness?


2. VALIDITY AND RELIABILITY: Think of student rating scales like the S.A.T. One item from the S.A.T. doesn't provide a valid and reliable score of verbal ability any more than 1 or 2 global items on a rating scale provide a valid and reliable measure of teaching performance. Although the former measures knowledge and the latter measures attitudes or opinions, the psychometric problem is virtually identical. Such data would be tantamount to giving high schoolers 1 or 2 verbal and quantitative items from the S.A.T. and making individual college admission decisions on those scores. Those scores are technically unreliable. One or 2 items are extremely unreliable compared to a summary score derived from 5 subscales or 35 items. Reliability coefficients should be in the .80s-.90s for individual decisions, although coefficients as low as the .60s are acceptable for group-based decisions in research. Typically, a single item does not yield a coefficient in the acceptable range.


The third reason, PROFESSIONAL AND LEGAL STANDARDS, will be covered in the next blog. It is the most important of the 3 reasons. Don’t miss it.


COPYRIGHT © 2009 Ronald A. Berk, LLC

Saturday, November 7, 2009

Should GLOBAL ITEM SCORES from Student Evaluations Be Used for Important Faculty Decisions? PART I

Page copy protected against web site content infringement by Copyscape

I know what you’re thinking: “I thought you were done with this student rating stuff. Get off it already.” I know, but I rethought my thought after receiving my weekly report that more people read my blogs this past week than at any other time over the past 3 months. Maybe I hit a note; maybe not. Anyway, these topics keep popping up on listservs and workshops.


But that’s not REEEAALLY what you were thinking. It is: “What in the world is a global item?” For those of you who are not of this world and are just passing through, as I am, here is a profile of the global item:

1. It provides a general, broad-stroke indication of teaching performance. It's intended to be an omnibus item, representing the collective judgments on all other items. But it isn't. Only the total scale score does that.
2. It doesn’t address specific teaching and course characteristics.
3. It usually appears at the end of the rating scale and should not be summed with the scores of all other items.

Using a "Strongly Agree-Strongly Disagree" anchor response scale, a couple of examples are given below:


Overall, my instructor is a dirtbag.
Overall, I learned squat in this course.


Of course, you know I’m kidding. Better items are:


Overall, my instructor is a moron.
Overall, this course is putrid.


Usually, there are 1 to 3 items. Frequently, administrators, such as your department chair, associate dean, or emperor or empress, will be encouraged to use the ratings on those items to provide a simple, quick-and-dirty measure of your teaching performance. Those ratings, in lieu of the total scale or subscale scores, are used in conjunction with other information to arrive at summative decisions regarding merit pay, contract renewal for full-time and adjunct faculty, and promotion and tenure recommendations. Are these important decisions about your career and life. You bet!!


Do you want those decisions to be rendered on the basis of 1 or 2 items? “Sure, why not?” Are you kidding me? There are several logical, psychometric, and legal reasons why global items should NOT be used for summative decisions. They will be described in my next blog. Stay tuned for more fun from RatingWorld.


COPYRIGHT © 2009 Ronald A. Berk, LLC

Sunday, October 25, 2009

What Scores Should Be Reported from Student Ratings of Faculty?

Page copy protected against web site content infringement by Copyscape

Recently, I was involved in a spirited marathon discussion with a bunch of colleagues on technical issues related to student ratings of teaching performance. One big topic was: How do you report results for formative and summative decisions? I thought some of my bloggees might be interested in the options available. These options with report form examples appear in my Thirteen Strategies... book (see Stylus link to right).

In order to answer the question, you don't need to administer multiple rating forms. There are a lot of options with the results from just one form. It is possible to "have your cake.." with one form for both formative and summative decisions up to a point. The trick is how the results are analyzed and reported for each decision maker.

Psychometrically speaking, I recommend the following:
1. A structured scale with 4-6 subscales measuring separate constructs such as Class Organization, Teaching Methods, Evaluation Techniques, and so on. The faculty evaluation lit reports several major constructs based on factor analyses. These core teaching behaviors should be generic enough to apply to most courses and disciplines.
2. A separate section devoted to course-specific items each instructor might want to add should be included. This optional section might contain up to 10 items.
3. One to three global items may be included as well, although individual item alpha reliabilities are typically much lower than item aggregates, such as subscale or total scale scores.
4. An unstructured section containing 2-5 stimulus questions to which students can comment is also important. Loads of online administrations reveal students spent considerable time typing buckets of comments. Frequently those comments explain the responses to some of the structured item ratings. Both forms of evaluation are valuable and furnish complementary information on teaching performance.

Analysis-wise, the above structure permits results at the following levels:
1. anchor distribution of percentages
2. item statistics such a mean and median (almost all distributions are negatively skewed)
3. subscale statistics
4. total scale statistics
5. summary of comments by stimulus question

That's a lot of information. Faculty would benefit from 1-5. 1 and 2, in particular, provide valuable diagnostic info to revise teaching or course materials that will benefit their next course-load of students. It is formative feedback only in that sense. Other formative methods administered during the course should be considered. You already know about those options.
Summative decisions by department chairs, associate deans, etc. can be based on 3 and 4 and possibly the global item scores.

The above strategy is certainly not new, but it is the simplest to get the biggest bang from your student rating scale. Of course, it is only 1 of 14 sources of evidence you might use in measuring teaching performance. Multiple sources of evidence should be involved in summative (personnel) decisions about faculty contract renewal, merit pay, and promotion and tenure. After all, faculty careers are on the line.

If you're grappling with this issue, I hope these suggestions may be helpful.
COPYRIGHT © 2009 Ronald A. Berk, LLC