Showing posts with label student ratings. Show all posts
Showing posts with label student ratings. Show all posts

Monday, September 27, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Finale!”

Page copy protected against web site content infringement by Copyscape

Epilogue

Well, there it is. I bet you’re thinking: “History, schmistory. What was that all about?” I’m sure your eyeballs hurt from rolling so many times and that one time when you’re your contacts blew out. Despite this cutesy romp through “Student-Ratings World” and a staggering 873 books and thousands of articles, monographs, conference presentations, blogs, etc. on the topic, some behaviors remain the same. For example, even today, the mere mention of teaching evaluation to many college professors triggers mental images of the shower scene from Psycho, with those bloodcurdling screams. They’re thinking, “Why not just beat me now, rather than wait to see my student ratings again.” Hummm. Kind of sounds like a prehistoric concept to me (a little "Meso-Pummel" déjà vu).

Despite the progress made with deans, department heads, and faculty moving toward multiple sources of evidence for formative and summative decisions, student ratings are still virtually synonymous with teaching evaluation in the United States, which is now located in Canada. They are the most influential measure of performance used in promotion and tenure decisions at institutions that emphasize teaching effectiveness. This popularity not withstanding, maybe the ubiquitous student rating scale will fair differently in the next "Meso-Cutback Era" by 2020! I hope I can update this schmistory for you then.

References
Arreola, R. A. (2007). Developing a comprehensive faculty evaluation system (3nd ed.). San Francisco: Jossey-Bass.
Berk, R. A. (2006). Thirteen strategies to measure college teaching. Sterling,VA: Stylus.
Knapper, C. & Cranton, P. (Eds). (2001). Fresh approaches to the evaluation of teaching (New Directions for Teaching and Learning, No. 88). San Francisco: Jossey-Bass.
Me, I. M. (2003). Prehistoric teaching techniques in cave classrooms. Rock & a Hard Place Educational Review, 3(4), 10−11.
Me, I. M. (2005). Naming institutions of higher education and buildings after filthy rich donors with spouses who are dead or older. Pretentious Academic Quarterly,14(4), 326−329.
Me, I. M., & You, W. U. V. (2005). Student clubbing methods to insure teaching
accountability. Journal of Punching & Pummeling Evaluation, 18(6), 170−183.
Seldin, P. (Ed.). (2006). Evaluating faculty performance. San Francisco: Jossey-Bass.

I gratefully acknowledge the valuable feedback of Raoul Arreola, Mike Theall, Bill Pallett, and another student-ratings expert for reviewing the skimpy facts reported in this blog series. To ensure the anonymity of one of the reviewers, I have volunteered him for the Federal Witness Protection Program or the USA cable TV series In Plain Sight. I forget which.

COPYRIGHT © 2010 Ronald A. Berk, LLC 

Friday, September 24, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Responserate Era (2000–2010)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Responserate Era (2000–2010)
The first decade of the new millennium rode on the search engines of the previous decade. A bunch of publications kicked it off with half a dozen edited volumes on student ratings and faculty evaluation published by Jossey-Bass in their Teaching and Learning series (Ryan, 2000, #83; Lewis, 2001, #87; Knapper & Cranton, 2002, #88; Sorenson & Johnson, 2004, #96) and Institutional Research series (Theall, Abrami,& Mets, 2001, #109; Colbeck, 2002, #114).

There were no planet-shattering technical developments, although there were several software options created specifically for online administration and reporting. Only a trickle (OOPS! Sorry. This is just residual fluid from the previous metaphor.) of articles on midterm formative ratings and other topics appeared. Finally, three books hit the Amazon pages in the last half of the decade: Arreola’s (2007)14th edition of his popular work, Seldin’s (2006) 128th edited volume, and our hero’s (Moi, 2006) psychometric-humorous attempt (see References in next blog).

Most of the activity and discourse on student ratings concentrated on practical issues. There were several trends that continued from the previous decade:

(1) student ratings data were being supplemented with other data, particularly peer review of teaching and course materials and letters of recommendation by each professor’s mommy or daddy, for decisions about teaching effectiveness;
(2) institutions reviewed the quality of their tools to consider either a commercial package, such as THOUGHT, or develop their own “homegrown” scales with online reporting by Academic Management Systems or other support;
(3) the technical quality of many “homegrown” scales prompted me to coin the new psychometric term "putrid"; and
(4) the debate over paper-based vs. online administration grew with cost and student response rates as the deal-breakers. Online packages were available everywhere. The importance of response rates to the validity of student ratings became so critical to the adoption of an online system that this era was named the "Meso-Responserate Era."

We’ve come to the end of this series. I’ve run out of jokes. The last blog will be an epilogue of a few final thoughts and jokes.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Wednesday, September 22, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Unleaded Era (1990s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Unleaded Era (1990s)
The 1990s were like a nice, deep breath of fresh gasoline (which was $974.99 a gallon at the pump, $974.00 for cash only), hereafter referred to as the Meso-Unleaded Era. Little did anyone anticipate how this era would be trumped at the pump in the years ending the current decade. The use of student rating scales had now spread to Kalamazoo (known to tourists as “The Big Apple”) and faculty began complaining about their validity, reliability, and overall value for decisions about promotion and tenure (the scales, that is, not Kalamazoo) . This was not unreasonable, given the lack of attention to the quality of scales over the preceding 90 billion years.

(WEATHER ALERT: I interrupt this section to warn you of impending wetness in the next three paragraphs. You might want to don appropriate apparel. Don’t blame me if you get wet. You may now rejoin this section already in progress. END OF ALERT.)

This debate intensified throughout the decade with a torrential downpour of publications challenging and contributing to the technical characteristics of the scales, particularly a series of articles by William Cashin of IDEA at Kansas State University, which was located in New Hampshire at the time, and an edited work by Mike Theall and Jennifer Franklin (Student Ratings of Instruction,1990). As part of this debate, another steady stream of research flowed toward alternative strategies to measure teaching effectiveness, especially peer ratings, self-ratings, videos, alumni ratings, interviews, learning outcomes, teaching scholarship, and teaching portfolios.

This stream leaked into books by John Centra (Reflective Faculty Evaluation, 1993), Larry Braskamp and John Ory (Assessing Faculty Work, 1994), Peter Seldin (Improving College Teaching, 1995), and Raoul Arreola (1st and 2nd editions of Developing a Comprehensive Faculty Evaluation System, 1995, 2000), and an edited volume by Seldin and Associates (Changing Practices in Evaluating Teaching, 1999). They furnished a confluence of valuable resources for faculty and administrators to use to evaluate teaching.

This cascading trend was also reflected increasingly in practice. Although use of student ratings had peaked at 88% by the end of the decade, peer and self-ratings were on the rise over the rapids of teaching performance as my liquid metaphor came to a screeching halt.

My next blog will address the developments in the first decade of the new millennium of the Meso-Responserate Era.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, September 20, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Meta Era (1980s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Meta Era (1980s)
The 1980s were really booooring! The research continued on a larger scale, and statistical reviews of the studies (a.k.a. meta-analyses) were conducted by such authors as Cohen (1980, 1981), d’Apollonia and Abrami (1997), and Feldman (1989). Of course, this period had to be labeled the Meso-Meta Era.

Book-wise, Peter Seldin of Pace University in upstate Saskatchewan published his first of thousands of books on the topic, Successful Faculty Evaluation Programs (1980). Ken Doyle produced his second book on the topic, Evaluating Teaching (1981), four years later (Are you still awake?).

The administration of student ratings metastasized throughout academe. By 1988, their use by college deans spiked to 80%, with still only a paltry 14% of deans gathering evidence on the technical aspects of their scales.

That takes us to—guess what? The next to last era in this blog series. Whew.

The next blog covers the 1990s with the major contributions by names you will know as gas prices spiked during the “Meso-Unleaded Era.”

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, September 16, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Golden Era (1970s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Golden Era (1970s)
The 1970s were called the “Golden Age of Student Ratings Research,” named after a bunch of “senior” professors who were doing research while living in assisted learning communities. Obviously, this period had to be named the "Meso-Golden Era." The research on the relationships between student ratings and instructor, student, and course characteristics burgeoned during these years. As the evidence on the validity and reliability of scales mounted, the use of the scales increased across the land, beyond East Lansing, in Nebraska (home of the Maryland Tar Heels).

To punctuate this burgeoning, commercially-produced, nationally-normed, student rating scales emerged during this era. Their superior technical quality and custom report forms compared to many “homegrown” tools provided colleges with the option of shopping for their rating scales in Ratings Mall. The first two scales were SIR (Educational Testing Service, 1972) and IDEA (Kansas State University, 1975). IAS Online (University of Washington) and SEEQ (Centre for Educational Assessment, Perth, Western Australia) became available later.

Student rating item banks also popped up during the ‘70s which permitted faculty to hand pick statements from a catalog to build their own customized scales. These banks were called POP-UPS PICES (Purdue University, 1974) and CIEQ (University of Arizona, 1977). Entering this commercial rating scale derby was not without challenges. Most ventures without barely pronounceable 3–5 letter acronym names tanked immediately.

These burgeoning golden years also witnessed the first 25 books on the topic. The winning prize for first book went to Richard Miller for Evaluating Faculty Performance (1972). He also took the silver medal with Developing Programs for Faculty Evaluation (1974). The bronze went to Ken Doyle for his work on student evaluations, aptly titled Student Evaluation of Instruction (1975). This was followed by John Centra’s Determining Faculty Effectiveness (1979), which synthesized all of the earlier work and proffered guidelines for future practices.

At the end of the decade, one survey found that 55% of liberal arts college deans always used student ratings to assess teaching performance, but only a tiny 10% conducted any research on the quality of their scales. These statistics were reported just in the nick of time before the next blog, which will continue with the flow of research during the 1980s and meta-analyses of previous studies of the “Meso-Meta Era.”

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, September 14, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Boomer Era (1960s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Boomer Era (1960s)
The 1960s were rocked by student protests on campuses, the Vietnam War, and the Broadway musical Hair (based on the TV program The Brady Bunch). This “boomer” generation was blamed for everything during this period, which was called the "Meso-Boomer Era." These boomers demanded that college administrators give them a voice in educational decisions that affected them. They expressed their collective voice by screaming like banshees, sitting in the entrances of administration buildings, and writing and administering rating scales to evaluate their instructors. They even published the ratings in student newspapers so students could use them as a consumer’s guide to course selection. There is still residual evidence of that practice today at several institutions, especially in countries like Texas.

There were few centrally-administered rating systems in universities to evaluate teaching effectiveness. Most uses of student ratings by faculty were voluntary. In general, the quality of the scales was dreadful and their use as evaluation tools was fragmented, unsystematic, and arbitrary, kinda like this blog series.

As for research, there was only a smidgen; it was all quiet on the publication front. The researchers were busy in the reference sections of their university libraries praying that someone would invent Google so that they could work at Starbucks instead. They were also preparing for the next decade: “Meso-Golden Era.” My next blog will document the research contributions during the 1970s.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, September 12, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Remmers Era (1927–1959)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Remmers Era (1927–1959)
Between 1927 and 1959, student rating practices began. The first student rating scale was developed during this period by Herman Remmers of Purdue University in Billings, Montana, home of the Pittsburgh Penguins. Consequently, these years were known as the Meso-PenguinsRemmers Era. Dr. Remmers pretty much owned this era. In fact, the rating scale was named after him: the Purdue Rating Scale for Instructors. Remmers was also the author of the first publication on the topic in 1930, the reliability of the scale, and studies on the relationships of student ratings to student grades and alumni ratings.

For his pioneer work on student ratings, Remmers was given the title “HH Cool R Rating Man.” WROOONG! He’s a rapper. Dr. Remmers’ real title was “Father of Student Evaluation Research.”

Now I suppose you’re going to shout out, “Who’s the Mother?” Are you ready for the answer? I don’t think so. It’s Darken “Foxy” Bubble Answer-Sheet. Mrs. Answer-Sheet was a descendant from a high stack of Answer-Sheets. What a striking couple they made. Well, that seems to wrap up this 32-year era.

My next blog will scope out one of the most traumatic periods in history—“Meso-Boomer Era,” when Baby Boomers attended and protested at our finest institutions of higher learning as out-of-control students during the ‘60s. How many of you are in this category? Are you still out of control like me?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, September 9, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: The Meso-Pummel Era!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Pummel Era

A long time ago, at a university far, far away, there were no student rating scales. During prehistoric times, there was only one university (near present-day Detroit) (Me, 2003), which was actually more like a community college because research and the four-year liberal arts curriculum hadn’t been invented yet. This institution was called Cave University, named after its major donor: Harold University (Me, 2005).

Students were very concerned about the quality of teaching back then. In fact, they created their own method of evaluation to express their feelings. For example, if an instructor strayed from the syllabus, fell behind the planned schedule, wrote faulty test items, or cut class to watch one of those Jurassic Park movies, which he (they were all men) later regretted, the students would club him (Me & You, 2005). This practice prompted historians to call this period the Meso-Pummel Era.

Admittedly, this practice seemed a bit crude and excessive at the time, but it held faculty accountable for their teaching. Obviously, there was no need for tenure. There was a lot of faculty turnover as word of the teaching evaluation method spread to Grand Rapids, which was on the land currently occupied by South Dakota.

That takes us up to 1927. I skipped over 90 billion years because nothing happened that was relevant to this documentary.

References

Me, I. M. (2003). Prehistoric teaching techniques in cave classrooms. Rock & a Hard
Place Educational Review, 3(4), 10−11.

Me, I. M. (2005). Naming institutions of higher education and buildings after filthy rich
donors with spouses who are dead or older. Pretentious Academic Quarterly, 14(4), 326−329.

Me, I. M., & You, W. U. V. (2005). Student clubbing methods to insure teaching
accountability. Journal of Punching & Pummeling Evaluation, 18(6), 170−183.

My next blog will cover the “Meso-Remmers Era,” named after the famous Purdue University professor Peter Seldin. His contributions over the succeeding 32 years will be described.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, September 6, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: A Parody!”

Page copy protected against web site content infringement by Copyscape

WARNING: This new blog series contains buckets of humor, which may not be suitable for all readers. If you have the sense of humor of an avocado or, even worse, a cumquat, this parody is not for you, despite its trailblazing, earth-shattering, Pulitzer Prize-caliber contribution to the teaching evaluation literature. If you fit this description, Buhbye!

ANOTHER BORING HISTORY?
Usually the history of any serious topic triggers the gag reflex in most nonhistorians. However, this topic is different from most. Over the past year, there have been incendiary debates over student rating forms, online administration procedures, and their use and interpretation on several professional listservs, LinkedIn groups, professional blogs, and other electronic and walkie-talkie communications. The reverberations of these debates have been felt on college campuses as far away as Pandora University, where the topic has provoked the verbal equivalent of the Avatar-scale firefight. Rather than fan the flames of this combustible metaphor, I thought this blog series might provide a time-out, refreshing break on this contentious topic, while also shedding some energy-saving light on how this situation evolved.

FRACTURED, BUT FACTUAL TOO!
Do any of you remember “Fractured Fairy Tales” on The Rocky and Bullwinkle Show? “NO!” What about Boris and Natasha? What were you doing? Oh well, it doesn’t matter, youngin’ academic readers. These blogs are written in the same spirit as that cult, politically-satirical cartoon, just without the cult, politics, and cartoon. It is a parody with the bonus of actual events in the history of student ratings. You’ll get a few morsels of content within a humor context. (FACT ALERT: Most of the names, dates, book and scale titles, and survey statistics are correct.)

TWO READER OUTCOMES:
There are two primary outcomes of this series for you:

1. to get a handle on the significant academic activities, research, and major players in the unfolding of the student ratings debate, and
2. to elicit a chuckle or two, maybe a guffaw, in that process.

If you laugh at any time during this series, I hope you will experience one of these physical signs:

a. burst your guts,
b. rupture key internal organs,
c. wet yourself, or
d. spurt your latté or green tea through your nostrils all over your keyboard.

Anything less will be disappointing. (NOTE: Most of the references along my historical path have been omitted to permit more space for jokes. Please refer to my Thirteen Strategies… book for those references and lots more jokes.)

My next blog will begin our historical journey with an overview of the state of the art of student ratings. Hold onto your guts.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, June 3, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Normative Score Comparisons

Page copy protected against web site content infringement by Copyscape

CRITERION-REFERENCED SCORE INTERPRETATIONS
All of the preceding score interpretations compare a score rating to the score range and points along the scale. For example, any rating can be above or below the midpoint or even higher or lower on the respective scale in terms of degree of favorableness, such as illustrated in the previous blog. Cut-off scores can also be set for criterion-referenced interpretations.

NORM-REFERENCED SCORE INTERPRETATIONS
Alternatively or in addition to these score interpretations, ratings can be compared to scores by a norm group, such as those by instructors of similar courses and/or instructors in your department, school, or university. They can also be compared to regional or national norms of instructors who teach the same courses you teach. Commercially-produced scales, such as IDEA, offer those normative scores. These comparisons are called norm-referenced interpretations. The scores at the various levels are still the same; they’re just summarized and reported for different groups of instructors and courses.

STRUCTURED VS. UNSTRUCTURED RESULTS
Did you do really well? If not, you should be able to figure out why you didn't. The answers to the unstructured or open-ended questions at the end of your scale usually provide reasons to explain the responses to the structured student ratings. The structured and unstructured items yield complementary evidence of your performance. Now you probably have more information than you want. That’s why I’m here for you.

I hope the preceding 2 weeks of blogs on interpreting student rating results have helped a little to make sense out of the scores you received. At least, you may have a basic understanding of the possible scores that can be reported and how you can use them to improve your teaching.

Let me know if you have any questions about the material presented. Much more detail on these topics is covered in my Thirteen Strategies book.

HAPPY STUDENT RATINGS!!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, June 1, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Total Scale Score vs. Global Item Scores

Page copy protected against web site content infringement by Copyscape

HOW DOES TOTAL SCORE COMPARE TO GLOBAL ITEM SCORES?
On many scales, global items are included at the end. These items ask students to provide a summary rating of the instructor and/or course. They take on a variety of formats, but the purpose is the same. For example, using anchors ranging from Excellent to Poor, the following items might be given:

What was the overall quality of your instructor’s teaching?
What was the overall value of this course?

OR

using Strongly Agree to Strongly Disagree,

This is the worst instructor on the planet.
This course sucks.

LIMITATIONS: The item response to each of these items would be interpreted the same as any other item on the scale. The problem is that these items are not diagnostic for teaching improvement. They provide an overall rating.

So what’s the problem? Individual item responses, either percentage responses to anchors or item means/medians, are usually unreliable. When those responses are used to suggest areas for improvement, they serve as a guide. No major career-shattering decisions are being made. If global item responses are used for summative decisions by your department chair or the promotion committee, there is a lot more at stake.

RECOMMENDATION: Subscale or total scale scores that summarize the quality of teaching or the value of the course are usually more reliable, based on a collection of items measuring those characteristics, than just a single item. It is recommended that those scores be used in lieu of global item scores whenever possible for any summative decisions about teaching performance.

My final blog in this bloated series will briefly describe criterion- and norm-referenced interpretations of the scores previously defined.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Saturday, May 29, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Total Scale Level

Page copy protected against web site content infringement by Copyscape

WHAT ARE TOTAL SCALE SCORES?
The highest level of score summary is the total scale score across all items. It’s like the total score on a test, except there are no right and wrong answers on a scale. If the scale consists of 36 items, each scored 0–3, the following results might be reported:

Total scale score range = 0–108; Midpoint = 54
        Mean/Median = 96.43/101, where N = 97

The continuum for interpretation would be the following:

Extremely                                                Mean     Extremely
Unfavorable                  Neutral               96.43     Favorable
      0__________________54___________________108
                                                                     Mdn
                                                                      101

INTERPRETATION OF TOTAL SCORES: This score gives a global, or composite, rating that is only as high as the ratings in each of its component parts (anchors, items, and subscales). It represents one overall index of teaching performance, from Extremely Unfavorable to Extremely Favorable. In this example, the score is very favorable. However, the total score is usually a little less reliable and less informative than the subscale scores.

COMPARISON OF ITEM, SUBSCALE, AND TOTAL SCORE SCALES
A comparison of the quantitative scales at the previous levels for a total scale is shown here:

Score Level        Ex Unfav                    Neutral                         Ex Fav
Item                       0 ___________1____ 1.5 ____ 2___________3
Subscale 1(8 items) 0 ________________ 12 __________________24
Subscale 2(5 items) 0 ________________ 7.5 _________________15
Subscale 3(4 items) 0 _________________ 6 __________________12
Subscale 4(13 items)0 ________________19.5 ________________39
Subscale 5(6 items) 0 _________________ 9 __________________18
Total Scale(36 items)0 _______________ 54 __________________108

These levels of score reporting and interpretation are based on a single course. This information has the greatest value to you, your department chair, and the curriculum committee evaluating the course.

How does the total score compare to the global item scores given at the end of the scale? Which score should you use? Is there really any difference?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, May 27, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Subscale Level

Page copy protected against web site content infringement by Copyscape

WHAT ARE SUBSCALE SCORES?
This is the first level at which item scores can be summed. If the items on the total scale are grouped into clusters related to topics such as instructional methods, evaluation methods, and course content, the item scores can be summed to produce subscale scores. These scores should only be used for decision making if adequate validity and reliability evidence support that internal scale structure. Since each subscale contains a different number of items, the score range will also be different. This range must be reported to interpret the results.

The summary of subscale score results is derived from the item score results. The statistics are the same. We are just aggregating or summarizing the item-level data (0–3) into subscale item clusters. For example, here are results for three selected subscales:

Instructional Methods (IM)—13 items
     Subscale score range = 0–39; Midpoint = 19.5
     Mean/Median = 34.79/37.00, where N = 97

Evaluation Methods (EM)—5 items
     Subscale score range = 0–15; Midpoint = 7.5
     Mean/Median = 13.40/15.00, where N = 97

Course Content (CC)—8 items
     Subscale score range = 0–24; Midpoint = 12
     Mean/Median = 21.97/24.00, where N = 97

COMPUTATION OF SUBSCALE SCORES: The interpretation of subscale results is analogous to the item results; only the numbers are BIGGER. For example, instead of a 0–3 range and a midpoint of 1.5 for an item, each subscale has a range and midpoint based on its respective number of items. So, for the Instructional Methods (IM) subscale with 13 items, a "0" (SD) response to every item produces a sum of 0 for the subscale, and a "3" (SA) response to all 13 items yields a sum of 39.

INTERPRETATION OF SUBSCALE MEANS AND MEDIANS: The zero-base for all score interpretations is easy to remember: the worst, most unfavorable rating on any item, subscale, or total scale is "0." What changes is the upper score limit for the most favorable rating on each subscale because the number of items change. Again, for the IM subscale, the mean and median can be referenced to the upper limit of 39 and also the midpoint of 19.5 to locate the position on the continuum, as indicated below:

Extremely                                                     Mean       Extremely
Unfavorable                       Neutral                34.79      Favorable
       0___________________19.5____________________39
                                                                             Mdn
                                                                              37
The mean/median ratings on the IM subscale are very favorable.

The subscale results can pinpoint areas of strength and weakness. They may be used by your department chair or the promotion review committee to identify your teaching strengths across different courses. Subscale scores cannot direct you toward particular aspects of teaching that can be improved or changed. The item and anchor results described previously are intended to provide that detailed level of direction.

Finally, the next blog will examine total scores on the scale. What additional info do they provide beyond what we already know?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, May 20, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Anchor Level—Part 2

Page copy protected against web site content infringement by Copyscape

How do you interpret the percentage distribution across the anchors? Is a skewed distribution a good or bad sign for teaching performance? How do any of these results relate to your teaching? Keep perusing, dear colleague.

SKEWED DISTRIBUTION: The overall response pattern shown in the previous blog indicates a negatively skewed distribution of responses, which is the most common outcome and typically a desirable one as well. It occurs when the majority of the responses are A and SA, but there is also a sprinkling of a few Ds and SDs. The extreme SD responses, or outliers, create the skew. Realistically, a few students might mark SD to every statement to express their desire to see you whacked, while the majority of satisfied customers will choose the two “Agree” anchors. These distributions occur in more than 90% of the courses I’ve reviewed at different institutions. It’s rare to receive ratings without any Ds or SDs. Other anchors may yield a different pattern of responses.

DIAGNOSTIC PROFILE: This anchor distribution information provides you with the most detailed profile of responses to a single item. The percentages reveal the degrees of agreement and disagreement with each statement. It is diagnostic of how the class felt about each behavior or characteristic. Examine each distribution carefully to pinpoint your strength behaviors (high percentage of SAs and As) and your weakness behaviors (relatively high percentages of SDs and Ds). You can then consider specific changes in your teaching, evaluation, or course behaviors to shift the distribution farther to the right, in the A–SA zone, the next time the class is taught.

The next blog will examine responses at the item level. Is that info of any value after reviewing the anchor distributions? I will solve that mystery!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, May 17, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Overview

Page copy protected against web site content infringement by Copyscape

APPLICATION TO DIFFERENT FORMS
Although each of you is using a different rating form with different numbers of items and scores, those differences do not matter in score interpretation. Whether you’re using a commercial package, such as IDEA, SIR II, PICES, or CIEQ DU SOLEIL, or a “homegrown scale,” there are only so many score reporting possibilities for any form in Likert-type format. So my suggestions are generic and should be applicable to your form. I encourage you to consult the guidelines or manual for your reporting system for more specific information.

FIVE BASIC CATEGORIES OF RESULTS
There are 5 possible categories of results reported for most student rating forms:

1. anchor distribution of percentages
2. item statistics (mean and/or median)
3. subscale statistics (mean and/or median)
4. total scale statistics (mean and/or median)
5. summary of comments to open-ended questions

Your report form may not provide all of the above, but it should certainly give you at least 2 and 4.

WHAT DO FACULTY NEED?
That's a lot of information. You could use all of those results, however, 1 and 2, in particular, provide the most valuable diagnostic info to revise teaching or course materials that will benefit your next course-load of students. These are called formative decisions about teaching. Category 5 can explain the reasons for the ratings to 1 and 2.

WHAT DO ADMINISTRATORS NEED?
Summative decisions about annual contract renewal, merit pay, or promotion and tenure review by department chairs, associate deans, etc. can be based on 3 and 4 and possibly the global item scores.

This blog series will focus primarily on the faculty needs. My next blog will examine the 1st level of interpretation: ANCHOR-WORLD!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, May 16, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Anchor Level—Part 1

Page copy protected against web site content infringement by Copyscape

RESPONSE RATE WARNING: As you begin to analyze the results from your class evaluations, please consider the response rate in your interpretations:

1. For class sizes of 30 to Super Bowl attendance, it is desirable to have at least 70% response, preferably 80 or 90%, to assure reasonably representative ratings. Anything less may be biased (aka "evil") in some unknown direction. Since all responses are anonymous, there is no way to assess the degree and direction of response bias. Just be cautious in your inferences about your teaching behaviors.
2. For classes less than 30, especially seminars of 5 to 10 students, be particularly careful in your interpretations based on both a less than desirable response rate and inadequate number of responses.

BOTTOM LINE: When your results are used to guide teaching improvement, view your ratings as suggestive of possible areas for change rather than as conclusive.

WHAT’S THE 1ST LEVEL OF INTERPRETATION?

IT’S ANCHOR-WORLD! The first level of score reporting is anchor results for each item. The anchors are the response options on the scale, such as STRONGLY DISAGREE or DISAGREE. Usually the percentage of students picking each anchor is reported, item by item. An example is shown below for student rating scale items with agree–disagree anchors:

                     SD       D         A          SA        N
Statement 1  1.0%   3.1%   37.5%    58.6%    96
Statement 2  1.0     3.1     24.0       71.9      96
Statement 3  1.1     1.1     28.9       68.9      90

ANCHOR SCALE: These results can be reported for any word, phrase, or statement stimuli and any response anchors. The anchor abbreviations for "Strongly Disagree" (SD), "Disagree" (D), "Agree" (A), and "Strongly Agree" (SA) are listed horizontally, left to right, from unfavorable to favorable ratings. This is the same order as the original scale. These anchors measure the degree or intensity of your feeling toward each statement. The N is the number of students that responded to the statement.

Other anchors may ask you evaluate the quality of a behavior, how frequently a behavior occurs, the quantity or the extent to which a behavior occurs, or how a behavior in one course compares to a behavior in another course. There are a variety of possible anchors and number of anchors on the scales now in use.

PERCENTAGE RESPONSES: The percentages for the four anchors indicate the percentage distribution based on the actual N. When the statements are positive teaching behaviors or course characteristics, you would expect low percentages for the first two “Disagree” anchors and high percentages for the second two “Agree” anchors, with the highest for SA. The percentages taper off drastically from right to left, with tiny percentages for D and SD for all three items. (NOTE: The percentages are slightly different for statement 3 compared to statements 1 and 2, particularly the 1.1% to SD and D. This was due, in part, to the six fewer students who responded to that item [N = 90].)

The next blog will examine the meaning of these distributions in terms of a diagnostic profile of your teaching strengths and weaknesses. Stay on board.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Friday, May 14, 2010

WHAT ARE YOU DOING WITH YOUR STUDENT RATING FORM RESULTS?

Page copy protected against web site content infringement by Copyscape

ADMINISTRATION OF STUDENT RATING FORMS
Yup! It’s that time of the year. The Cherry Blossoms are gone and many professors on the east coast are sneezing and wheezing their brains out from the sky-high pollen counts. They’re medicating themselves with megadoses of antihistamines like Benadryl, Zyrtec, Claritin, Allegra, Clarinex, Flonase, Nazonex, and Sneezwhizz.

This is the perfect time to administer those end-of-course student rating forms and interpret the scores. The results usually appear better when you’re drowsy and punchy from those medications. This month has the highest number of forms administered world-wide, except for December when the same professors are sneezing and wheezing from the common cold.

The forms may be administered in class or online, but the results have to be reported in some format. You may receive the results in a couple of days to several months, depending on your processing system. Of course, all of this happens in between commencement exercises and end-of-year parties. Hopefully, you can glance at the form results before your next course begins in the summer or fall.

YOUR "RATING ANGEL"
I am your Rating Angel for this next week. If you’re not sure how to interpret the scores or use the results, this blog series is for YOU! I want you to milk those ratings for all their worth, to squeeze every drip of information that can guide your teaching improvement.

If you already know how to interpret these scores, STOP reading this blog immediately. Disregard it and get back to work, class, or lunch. Stop fooling around and wasting time on this blog. You should be ashamed of yourself.

SOURCE ALERT
My Thirteen Strategies book goes into considerable detail on the how, why, computing, reporting, formatting, and who cares about those scores. This series will simply focus on the scores you can use to guide decisions about your teaching improvement.

GOAL OF BLOG SERIES
This blog series is intended to present the BerksNotes® version of student rating scale interpretation. Hopefully, by cutting to the chase, whatever that is, you’ll be able to get the most out of your scores in lickety-split time or faster. Hold on to your keyboards. My next blog will begin with an overview of the different types of scores.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, November 2, 2009

Top 10 Secret Strategies to Get 80%+ Response Rates for Online Student Evaluations!

Page copy protected against web site content infringement by Copyscape

Here is a list of the Top 10 most effective strategies to get 80%+ response rates for online student ratings (Berk, 2006; Johnson, 2003; Sorenson & Reiner, 2003):


10. Faculty communicate online and face-face to students the importance of their input and how the results will be used


9. Ease of computer access for all students


8. Assurance of anonymity


7. Convenient, user-friendly online system


6. Provide easy-to-understand instructions on how to use the system


5. Administration and faculty strongly encourage students to complete forms


4. Faculty “assign” students to complete forms


3. Faculty threaten to smash students’ iPods with a Gallagher-type sledgehammer


3. System withholds students’ early access to final grades


2. Faculty provide extra credit or points


1. Faculty provide positive incentives (aka assorted bribes, such as requiring students to pick-up dry cleaning and take dog to the vet)

This is quite a laundry list of techniques. Is there any single practice that has a proven track record of success? Yup. That one is even more top-secret and a higher security “code chartreuse” technique than the others. It will be revealed in my next blog. Stay tuned.



COPYRIGHT © 2009 Ronald A. Berk, LLC

Monday, October 26, 2009

When Should Student Rating Scales Be Administered? Who Cares?

Page copy protected against web site content infringement by Copyscape
One critical psychometric issue neglected in the literature on administration of student rating scales and in wide-spread administration practices is the permissible window for completing the scales. This involves standardization of administration procedures. These procedures are required in the administration of all instruments, especially those used for personnel decisions in the Standards for Educational and Psychological TestingEmployment decisions about merit pay, pay cuts, promotion, demotion, and tenure must be based on evidence of teaching performance that meets the Standards and EEOC Uniform Guidelines on Employee Selection Procedures.

The validity of the responses, their comparability, and aggregate meaning based on group data hinge on WHEN the students complete the evals. The window must be NARROW, such as 48 hours immediately following the final exam or project submission. If there are items on the scale measuring student assessment procedures, administration practices, fairness, etc., the final exam must be completed before the scales can be administered. This STANDARDIZATION of scale administration is essential, whether online or paper-and-pencil, to insure the scores from all students have the same meaning. If a wide window is given where some students can complete rating scales before the final assessment and others after that assessment at their discretion, the validity of responses is shot to smithereens!

The issue is the link between the behaviors measured on the rating scale and the students' opportunity to render an accurate rating of each behavior? This is a validity concern. It is assumed that every student has had the same 45 hours (3-credit course) to observe those behaviors during the semester. If they miss a few classes, their evaluations should not be significantly affected.


COPYRIGHT © 2009 Ronald A. Berk, LLC

Sunday, October 25, 2009

What Scores Should Be Reported from Student Ratings of Faculty?

Page copy protected against web site content infringement by Copyscape

Recently, I was involved in a spirited marathon discussion with a bunch of colleagues on technical issues related to student ratings of teaching performance. One big topic was: How do you report results for formative and summative decisions? I thought some of my bloggees might be interested in the options available. These options with report form examples appear in my Thirteen Strategies... book (see Stylus link to right).

In order to answer the question, you don't need to administer multiple rating forms. There are a lot of options with the results from just one form. It is possible to "have your cake.." with one form for both formative and summative decisions up to a point. The trick is how the results are analyzed and reported for each decision maker.

Psychometrically speaking, I recommend the following:
1. A structured scale with 4-6 subscales measuring separate constructs such as Class Organization, Teaching Methods, Evaluation Techniques, and so on. The faculty evaluation lit reports several major constructs based on factor analyses. These core teaching behaviors should be generic enough to apply to most courses and disciplines.
2. A separate section devoted to course-specific items each instructor might want to add should be included. This optional section might contain up to 10 items.
3. One to three global items may be included as well, although individual item alpha reliabilities are typically much lower than item aggregates, such as subscale or total scale scores.
4. An unstructured section containing 2-5 stimulus questions to which students can comment is also important. Loads of online administrations reveal students spent considerable time typing buckets of comments. Frequently those comments explain the responses to some of the structured item ratings. Both forms of evaluation are valuable and furnish complementary information on teaching performance.

Analysis-wise, the above structure permits results at the following levels:
1. anchor distribution of percentages
2. item statistics such a mean and median (almost all distributions are negatively skewed)
3. subscale statistics
4. total scale statistics
5. summary of comments by stimulus question

That's a lot of information. Faculty would benefit from 1-5. 1 and 2, in particular, provide valuable diagnostic info to revise teaching or course materials that will benefit their next course-load of students. It is formative feedback only in that sense. Other formative methods administered during the course should be considered. You already know about those options.
Summative decisions by department chairs, associate deans, etc. can be based on 3 and 4 and possibly the global item scores.

The above strategy is certainly not new, but it is the simplest to get the biggest bang from your student rating scale. Of course, it is only 1 of 14 sources of evidence you might use in measuring teaching performance. Multiple sources of evidence should be involved in summative (personnel) decisions about faculty contract renewal, merit pay, and promotion and tenure. After all, faculty careers are on the line.

If you're grappling with this issue, I hope these suggestions may be helpful.
COPYRIGHT © 2009 Ronald A. Berk, LLC