Showing posts with label Thirteen Strategies to Measure College Teaching. Show all posts
Showing posts with label Thirteen Strategies to Measure College Teaching. Show all posts

Monday, September 27, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Finale!”

Page copy protected against web site content infringement by Copyscape

Epilogue

Well, there it is. I bet you’re thinking: “History, schmistory. What was that all about?” I’m sure your eyeballs hurt from rolling so many times and that one time when you’re your contacts blew out. Despite this cutesy romp through “Student-Ratings World” and a staggering 873 books and thousands of articles, monographs, conference presentations, blogs, etc. on the topic, some behaviors remain the same. For example, even today, the mere mention of teaching evaluation to many college professors triggers mental images of the shower scene from Psycho, with those bloodcurdling screams. They’re thinking, “Why not just beat me now, rather than wait to see my student ratings again.” Hummm. Kind of sounds like a prehistoric concept to me (a little "Meso-Pummel" déjà vu).

Despite the progress made with deans, department heads, and faculty moving toward multiple sources of evidence for formative and summative decisions, student ratings are still virtually synonymous with teaching evaluation in the United States, which is now located in Canada. They are the most influential measure of performance used in promotion and tenure decisions at institutions that emphasize teaching effectiveness. This popularity not withstanding, maybe the ubiquitous student rating scale will fair differently in the next "Meso-Cutback Era" by 2020! I hope I can update this schmistory for you then.

References
Arreola, R. A. (2007). Developing a comprehensive faculty evaluation system (3nd ed.). San Francisco: Jossey-Bass.
Berk, R. A. (2006). Thirteen strategies to measure college teaching. Sterling,VA: Stylus.
Knapper, C. & Cranton, P. (Eds). (2001). Fresh approaches to the evaluation of teaching (New Directions for Teaching and Learning, No. 88). San Francisco: Jossey-Bass.
Me, I. M. (2003). Prehistoric teaching techniques in cave classrooms. Rock & a Hard Place Educational Review, 3(4), 10−11.
Me, I. M. (2005). Naming institutions of higher education and buildings after filthy rich donors with spouses who are dead or older. Pretentious Academic Quarterly,14(4), 326−329.
Me, I. M., & You, W. U. V. (2005). Student clubbing methods to insure teaching
accountability. Journal of Punching & Pummeling Evaluation, 18(6), 170−183.
Seldin, P. (Ed.). (2006). Evaluating faculty performance. San Francisco: Jossey-Bass.

I gratefully acknowledge the valuable feedback of Raoul Arreola, Mike Theall, Bill Pallett, and another student-ratings expert for reviewing the skimpy facts reported in this blog series. To ensure the anonymity of one of the reviewers, I have volunteered him for the Federal Witness Protection Program or the USA cable TV series In Plain Sight. I forget which.

COPYRIGHT © 2010 Ronald A. Berk, LLC 

Friday, September 24, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Responserate Era (2000–2010)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Responserate Era (2000–2010)
The first decade of the new millennium rode on the search engines of the previous decade. A bunch of publications kicked it off with half a dozen edited volumes on student ratings and faculty evaluation published by Jossey-Bass in their Teaching and Learning series (Ryan, 2000, #83; Lewis, 2001, #87; Knapper & Cranton, 2002, #88; Sorenson & Johnson, 2004, #96) and Institutional Research series (Theall, Abrami,& Mets, 2001, #109; Colbeck, 2002, #114).

There were no planet-shattering technical developments, although there were several software options created specifically for online administration and reporting. Only a trickle (OOPS! Sorry. This is just residual fluid from the previous metaphor.) of articles on midterm formative ratings and other topics appeared. Finally, three books hit the Amazon pages in the last half of the decade: Arreola’s (2007)14th edition of his popular work, Seldin’s (2006) 128th edited volume, and our hero’s (Moi, 2006) psychometric-humorous attempt (see References in next blog).

Most of the activity and discourse on student ratings concentrated on practical issues. There were several trends that continued from the previous decade:

(1) student ratings data were being supplemented with other data, particularly peer review of teaching and course materials and letters of recommendation by each professor’s mommy or daddy, for decisions about teaching effectiveness;
(2) institutions reviewed the quality of their tools to consider either a commercial package, such as THOUGHT, or develop their own “homegrown” scales with online reporting by Academic Management Systems or other support;
(3) the technical quality of many “homegrown” scales prompted me to coin the new psychometric term "putrid"; and
(4) the debate over paper-based vs. online administration grew with cost and student response rates as the deal-breakers. Online packages were available everywhere. The importance of response rates to the validity of student ratings became so critical to the adoption of an online system that this era was named the "Meso-Responserate Era."

We’ve come to the end of this series. I’ve run out of jokes. The last blog will be an epilogue of a few final thoughts and jokes.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Wednesday, September 22, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Unleaded Era (1990s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Unleaded Era (1990s)
The 1990s were like a nice, deep breath of fresh gasoline (which was $974.99 a gallon at the pump, $974.00 for cash only), hereafter referred to as the Meso-Unleaded Era. Little did anyone anticipate how this era would be trumped at the pump in the years ending the current decade. The use of student rating scales had now spread to Kalamazoo (known to tourists as “The Big Apple”) and faculty began complaining about their validity, reliability, and overall value for decisions about promotion and tenure (the scales, that is, not Kalamazoo) . This was not unreasonable, given the lack of attention to the quality of scales over the preceding 90 billion years.

(WEATHER ALERT: I interrupt this section to warn you of impending wetness in the next three paragraphs. You might want to don appropriate apparel. Don’t blame me if you get wet. You may now rejoin this section already in progress. END OF ALERT.)

This debate intensified throughout the decade with a torrential downpour of publications challenging and contributing to the technical characteristics of the scales, particularly a series of articles by William Cashin of IDEA at Kansas State University, which was located in New Hampshire at the time, and an edited work by Mike Theall and Jennifer Franklin (Student Ratings of Instruction,1990). As part of this debate, another steady stream of research flowed toward alternative strategies to measure teaching effectiveness, especially peer ratings, self-ratings, videos, alumni ratings, interviews, learning outcomes, teaching scholarship, and teaching portfolios.

This stream leaked into books by John Centra (Reflective Faculty Evaluation, 1993), Larry Braskamp and John Ory (Assessing Faculty Work, 1994), Peter Seldin (Improving College Teaching, 1995), and Raoul Arreola (1st and 2nd editions of Developing a Comprehensive Faculty Evaluation System, 1995, 2000), and an edited volume by Seldin and Associates (Changing Practices in Evaluating Teaching, 1999). They furnished a confluence of valuable resources for faculty and administrators to use to evaluate teaching.

This cascading trend was also reflected increasingly in practice. Although use of student ratings had peaked at 88% by the end of the decade, peer and self-ratings were on the rise over the rapids of teaching performance as my liquid metaphor came to a screeching halt.

My next blog will address the developments in the first decade of the new millennium of the Meso-Responserate Era.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, September 20, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Meta Era (1980s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Meta Era (1980s)
The 1980s were really booooring! The research continued on a larger scale, and statistical reviews of the studies (a.k.a. meta-analyses) were conducted by such authors as Cohen (1980, 1981), d’Apollonia and Abrami (1997), and Feldman (1989). Of course, this period had to be labeled the Meso-Meta Era.

Book-wise, Peter Seldin of Pace University in upstate Saskatchewan published his first of thousands of books on the topic, Successful Faculty Evaluation Programs (1980). Ken Doyle produced his second book on the topic, Evaluating Teaching (1981), four years later (Are you still awake?).

The administration of student ratings metastasized throughout academe. By 1988, their use by college deans spiked to 80%, with still only a paltry 14% of deans gathering evidence on the technical aspects of their scales.

That takes us to—guess what? The next to last era in this blog series. Whew.

The next blog covers the 1990s with the major contributions by names you will know as gas prices spiked during the “Meso-Unleaded Era.”

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, September 16, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Golden Era (1970s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Golden Era (1970s)
The 1970s were called the “Golden Age of Student Ratings Research,” named after a bunch of “senior” professors who were doing research while living in assisted learning communities. Obviously, this period had to be named the "Meso-Golden Era." The research on the relationships between student ratings and instructor, student, and course characteristics burgeoned during these years. As the evidence on the validity and reliability of scales mounted, the use of the scales increased across the land, beyond East Lansing, in Nebraska (home of the Maryland Tar Heels).

To punctuate this burgeoning, commercially-produced, nationally-normed, student rating scales emerged during this era. Their superior technical quality and custom report forms compared to many “homegrown” tools provided colleges with the option of shopping for their rating scales in Ratings Mall. The first two scales were SIR (Educational Testing Service, 1972) and IDEA (Kansas State University, 1975). IAS Online (University of Washington) and SEEQ (Centre for Educational Assessment, Perth, Western Australia) became available later.

Student rating item banks also popped up during the ‘70s which permitted faculty to hand pick statements from a catalog to build their own customized scales. These banks were called POP-UPS PICES (Purdue University, 1974) and CIEQ (University of Arizona, 1977). Entering this commercial rating scale derby was not without challenges. Most ventures without barely pronounceable 3–5 letter acronym names tanked immediately.

These burgeoning golden years also witnessed the first 25 books on the topic. The winning prize for first book went to Richard Miller for Evaluating Faculty Performance (1972). He also took the silver medal with Developing Programs for Faculty Evaluation (1974). The bronze went to Ken Doyle for his work on student evaluations, aptly titled Student Evaluation of Instruction (1975). This was followed by John Centra’s Determining Faculty Effectiveness (1979), which synthesized all of the earlier work and proffered guidelines for future practices.

At the end of the decade, one survey found that 55% of liberal arts college deans always used student ratings to assess teaching performance, but only a tiny 10% conducted any research on the quality of their scales. These statistics were reported just in the nick of time before the next blog, which will continue with the flow of research during the 1980s and meta-analyses of previous studies of the “Meso-Meta Era.”

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, September 14, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Boomer Era (1960s)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Boomer Era (1960s)
The 1960s were rocked by student protests on campuses, the Vietnam War, and the Broadway musical Hair (based on the TV program The Brady Bunch). This “boomer” generation was blamed for everything during this period, which was called the "Meso-Boomer Era." These boomers demanded that college administrators give them a voice in educational decisions that affected them. They expressed their collective voice by screaming like banshees, sitting in the entrances of administration buildings, and writing and administering rating scales to evaluate their instructors. They even published the ratings in student newspapers so students could use them as a consumer’s guide to course selection. There is still residual evidence of that practice today at several institutions, especially in countries like Texas.

There were few centrally-administered rating systems in universities to evaluate teaching effectiveness. Most uses of student ratings by faculty were voluntary. In general, the quality of the scales was dreadful and their use as evaluation tools was fragmented, unsystematic, and arbitrary, kinda like this blog series.

As for research, there was only a smidgen; it was all quiet on the publication front. The researchers were busy in the reference sections of their university libraries praying that someone would invent Google so that they could work at Starbucks instead. They were also preparing for the next decade: “Meso-Golden Era.” My next blog will document the research contributions during the 1970s.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, September 12, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: Meso-Remmers Era (1927–1959)!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Remmers Era (1927–1959)
Between 1927 and 1959, student rating practices began. The first student rating scale was developed during this period by Herman Remmers of Purdue University in Billings, Montana, home of the Pittsburgh Penguins. Consequently, these years were known as the Meso-PenguinsRemmers Era. Dr. Remmers pretty much owned this era. In fact, the rating scale was named after him: the Purdue Rating Scale for Instructors. Remmers was also the author of the first publication on the topic in 1930, the reliability of the scale, and studies on the relationships of student ratings to student grades and alumni ratings.

For his pioneer work on student ratings, Remmers was given the title “HH Cool R Rating Man.” WROOONG! He’s a rapper. Dr. Remmers’ real title was “Father of Student Evaluation Research.”

Now I suppose you’re going to shout out, “Who’s the Mother?” Are you ready for the answer? I don’t think so. It’s Darken “Foxy” Bubble Answer-Sheet. Mrs. Answer-Sheet was a descendant from a high stack of Answer-Sheets. What a striking couple they made. Well, that seems to wrap up this 32-year era.

My next blog will scope out one of the most traumatic periods in history—“Meso-Boomer Era,” when Baby Boomers attended and protested at our finest institutions of higher learning as out-of-control students during the ‘60s. How many of you are in this category? Are you still out of control like me?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, September 9, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: The Meso-Pummel Era!”

Page copy protected against web site content infringement by Copyscape

A HISTORY OF STUDENT RATINGS: Meso-Pummel Era

A long time ago, at a university far, far away, there were no student rating scales. During prehistoric times, there was only one university (near present-day Detroit) (Me, 2003), which was actually more like a community college because research and the four-year liberal arts curriculum hadn’t been invented yet. This institution was called Cave University, named after its major donor: Harold University (Me, 2005).

Students were very concerned about the quality of teaching back then. In fact, they created their own method of evaluation to express their feelings. For example, if an instructor strayed from the syllabus, fell behind the planned schedule, wrote faulty test items, or cut class to watch one of those Jurassic Park movies, which he (they were all men) later regretted, the students would club him (Me & You, 2005). This practice prompted historians to call this period the Meso-Pummel Era.

Admittedly, this practice seemed a bit crude and excessive at the time, but it held faculty accountable for their teaching. Obviously, there was no need for tenure. There was a lot of faculty turnover as word of the teaching evaluation method spread to Grand Rapids, which was on the land currently occupied by South Dakota.

That takes us up to 1927. I skipped over 90 billion years because nothing happened that was relevant to this documentary.

References

Me, I. M. (2003). Prehistoric teaching techniques in cave classrooms. Rock & a Hard
Place Educational Review, 3(4), 10−11.

Me, I. M. (2005). Naming institutions of higher education and buildings after filthy rich
donors with spouses who are dead or older. Pretentious Academic Quarterly, 14(4), 326−329.

Me, I. M., & You, W. U. V. (2005). Student clubbing methods to insure teaching
accountability. Journal of Punching & Pummeling Evaluation, 18(6), 170−183.

My next blog will cover the “Meso-Remmers Era,” named after the famous Purdue University professor Peter Seldin. His contributions over the succeeding 32 years will be described.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, September 7, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: The State of the Art!”

Page copy protected against web site content infringement by Copyscape

State-of-the-Art of Student Ratings
There is more research on student ratings than any other topic in higher education. More than 2,500 publications and presentations have been cited over the past 90 years. Those ratings have dominated as the primary and, frequently, only measure of teaching effectiveness at colleges and universities for the past five decades. In fact, the evaluation of teaching has been in a metaphorical cul-de-sac with student ratings as the universal barometer of teaching performance. And, if you’ve ever been in a cul-de-sac or metaphor, you know what that’s like. OMGosh, it can be Stephen Kingish terrifying.

In surveys over the past decade, it was found that 86% of U.S. liberal arts college deans and 97% of department chairs use student ratings for summative decisions about faculty. Only recently has there been a trend toward augmenting those ratings with other sources of evidence and better metaphors (Arreola, 2007; Berk, 2006; Knapper & Cranton, 2001; Seldin, 2006).

So how in the ivory tower did we get to this point? Let’s trace the major historical events. Hold on to your online administration response rates. Here we go.

A History of Student Ratings
This history covers a timeline of approximately 100 billion years, give or take a day or two, ranging from the age of dinosaurs to the age of Conan O’Brien’s new cable TV show. Obviously, it’s impossible to squish every event that occurred during that period in this series. Instead, that span is partitioned into six major eras within which salient student-ratings activities are highlighted. A blog will be devoted to each of those eras.

References

Arreola, R. A. (2007). Developing a comprehensive faculty evaluation system (3nd ed.). San Francisco: Jossey-Bass.
Berk, R. A. (2006). Thirteen strategies to measure college teaching. Sterling,VA: Stylus.
Knapper, C., & Cranton, P. (Eds). (2001). Fresh approaches to the evaluation of teaching (New Directions for Teaching and Learning, No. 88). San Francisco: Jossey-Bass.
Seldin, P. (Ed.). (2006). Evaluating faculty performance. San Francisco: Jossey-Bass.

My 1st era blog will tackle prehistoric student ratings of the “Meso-Pummel Era.” How did cave men and women measure teaching performance? Their methods were a bit crude, but effective.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, September 6, 2010

“A FRACTURED, SEMI-FACTUAL HISTORY OF STUDENT RATINGS OF TEACHING: A Parody!”

Page copy protected against web site content infringement by Copyscape

WARNING: This new blog series contains buckets of humor, which may not be suitable for all readers. If you have the sense of humor of an avocado or, even worse, a cumquat, this parody is not for you, despite its trailblazing, earth-shattering, Pulitzer Prize-caliber contribution to the teaching evaluation literature. If you fit this description, Buhbye!

ANOTHER BORING HISTORY?
Usually the history of any serious topic triggers the gag reflex in most nonhistorians. However, this topic is different from most. Over the past year, there have been incendiary debates over student rating forms, online administration procedures, and their use and interpretation on several professional listservs, LinkedIn groups, professional blogs, and other electronic and walkie-talkie communications. The reverberations of these debates have been felt on college campuses as far away as Pandora University, where the topic has provoked the verbal equivalent of the Avatar-scale firefight. Rather than fan the flames of this combustible metaphor, I thought this blog series might provide a time-out, refreshing break on this contentious topic, while also shedding some energy-saving light on how this situation evolved.

FRACTURED, BUT FACTUAL TOO!
Do any of you remember “Fractured Fairy Tales” on The Rocky and Bullwinkle Show? “NO!” What about Boris and Natasha? What were you doing? Oh well, it doesn’t matter, youngin’ academic readers. These blogs are written in the same spirit as that cult, politically-satirical cartoon, just without the cult, politics, and cartoon. It is a parody with the bonus of actual events in the history of student ratings. You’ll get a few morsels of content within a humor context. (FACT ALERT: Most of the names, dates, book and scale titles, and survey statistics are correct.)

TWO READER OUTCOMES:
There are two primary outcomes of this series for you:

1. to get a handle on the significant academic activities, research, and major players in the unfolding of the student ratings debate, and
2. to elicit a chuckle or two, maybe a guffaw, in that process.

If you laugh at any time during this series, I hope you will experience one of these physical signs:

a. burst your guts,
b. rupture key internal organs,
c. wet yourself, or
d. spurt your latté or green tea through your nostrils all over your keyboard.

Anything less will be disappointing. (NOTE: Most of the references along my historical path have been omitted to permit more space for jokes. Please refer to my Thirteen Strategies… book for those references and lots more jokes.)

My next blog will begin our historical journey with an overview of the state of the art of student ratings. Hold onto your guts.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, July 18, 2010

“WHAT IS WEB 2.0? I DON’T EVEN REMEMBER 1.0! WHO CARES?”

Page copy protected against web site content infringement by Copyscape

DISCLAIMER: You know I’m not a techy. So why am I writing blogs on this topic? First, if I don’t know some of this Web terminology, maybe there are few of you who don’t either. Also, writing about topics I know nothing about helps me grow. I grow inaccurately, but I grow. Actually, I can still read and write and I know you will correct me if I’m wrong. I have your built-in accountability. This is another Berk’sNotes® version on the subject of Web terminology.

Berk’sNotes®
If you’re not familiar with Berk’sNotes® from my Thirteen Strategies… book and previous blogs, here’s a definition:

Berk’sNotes® = In the spirit of CliffsNotes®, it is an abbreviated version or brief synthesis of the most salient information or critical elements of a given topic. It’s designed for people like you who don’t have the time or inclination to do that synthesis yourself. 
MOTTO: “I synthesize the stuff so you don’t have to.”

WHAT HAPPENED TO WEB 1.0?
What did I miss? I honestly don’t even remember Web 1.0. Maybe I experienced what is known as the “Rip van Winkle Effect” or amnesia, or I’m just out of touch. Did you hear about 1.0 before 2.0? It’s probably just me. Maybe I was living in a storm drain at the time.

Anyway, the goal of these blogs is to set me straight and clarify what these terms mean so we can think about how the various tech tools can be leveraged in the classroom and our lives. If you’re interested in all of the sources on this topic, just Google the numbers in the title and the sources cited in my blogs.

Right now, I think we’re into Web 2.0 world, or, maybe, 3.0. I’m not sure. Maybe this blog will get us both up to speed. You probably have an inkling about these numbers. I used to have inklings, but they cleared up with a topical medication. I’m going to try and level our keyboards and inklings on Web 1.0, 2.0, and 3.0. Hang on. My next blog will examine Web 1.0.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, June 3, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Normative Score Comparisons

Page copy protected against web site content infringement by Copyscape

CRITERION-REFERENCED SCORE INTERPRETATIONS
All of the preceding score interpretations compare a score rating to the score range and points along the scale. For example, any rating can be above or below the midpoint or even higher or lower on the respective scale in terms of degree of favorableness, such as illustrated in the previous blog. Cut-off scores can also be set for criterion-referenced interpretations.

NORM-REFERENCED SCORE INTERPRETATIONS
Alternatively or in addition to these score interpretations, ratings can be compared to scores by a norm group, such as those by instructors of similar courses and/or instructors in your department, school, or university. They can also be compared to regional or national norms of instructors who teach the same courses you teach. Commercially-produced scales, such as IDEA, offer those normative scores. These comparisons are called norm-referenced interpretations. The scores at the various levels are still the same; they’re just summarized and reported for different groups of instructors and courses.

STRUCTURED VS. UNSTRUCTURED RESULTS
Did you do really well? If not, you should be able to figure out why you didn't. The answers to the unstructured or open-ended questions at the end of your scale usually provide reasons to explain the responses to the structured student ratings. The structured and unstructured items yield complementary evidence of your performance. Now you probably have more information than you want. That’s why I’m here for you.

I hope the preceding 2 weeks of blogs on interpreting student rating results have helped a little to make sense out of the scores you received. At least, you may have a basic understanding of the possible scores that can be reported and how you can use them to improve your teaching.

Let me know if you have any questions about the material presented. Much more detail on these topics is covered in my Thirteen Strategies book.

HAPPY STUDENT RATINGS!!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, June 1, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Total Scale Score vs. Global Item Scores

Page copy protected against web site content infringement by Copyscape

HOW DOES TOTAL SCORE COMPARE TO GLOBAL ITEM SCORES?
On many scales, global items are included at the end. These items ask students to provide a summary rating of the instructor and/or course. They take on a variety of formats, but the purpose is the same. For example, using anchors ranging from Excellent to Poor, the following items might be given:

What was the overall quality of your instructor’s teaching?
What was the overall value of this course?

OR

using Strongly Agree to Strongly Disagree,

This is the worst instructor on the planet.
This course sucks.

LIMITATIONS: The item response to each of these items would be interpreted the same as any other item on the scale. The problem is that these items are not diagnostic for teaching improvement. They provide an overall rating.

So what’s the problem? Individual item responses, either percentage responses to anchors or item means/medians, are usually unreliable. When those responses are used to suggest areas for improvement, they serve as a guide. No major career-shattering decisions are being made. If global item responses are used for summative decisions by your department chair or the promotion committee, there is a lot more at stake.

RECOMMENDATION: Subscale or total scale scores that summarize the quality of teaching or the value of the course are usually more reliable, based on a collection of items measuring those characteristics, than just a single item. It is recommended that those scores be used in lieu of global item scores whenever possible for any summative decisions about teaching performance.

My final blog in this bloated series will briefly describe criterion- and norm-referenced interpretations of the scores previously defined.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Saturday, May 29, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Total Scale Level

Page copy protected against web site content infringement by Copyscape

WHAT ARE TOTAL SCALE SCORES?
The highest level of score summary is the total scale score across all items. It’s like the total score on a test, except there are no right and wrong answers on a scale. If the scale consists of 36 items, each scored 0–3, the following results might be reported:

Total scale score range = 0–108; Midpoint = 54
        Mean/Median = 96.43/101, where N = 97

The continuum for interpretation would be the following:

Extremely                                                Mean     Extremely
Unfavorable                  Neutral               96.43     Favorable
      0__________________54___________________108
                                                                     Mdn
                                                                      101

INTERPRETATION OF TOTAL SCORES: This score gives a global, or composite, rating that is only as high as the ratings in each of its component parts (anchors, items, and subscales). It represents one overall index of teaching performance, from Extremely Unfavorable to Extremely Favorable. In this example, the score is very favorable. However, the total score is usually a little less reliable and less informative than the subscale scores.

COMPARISON OF ITEM, SUBSCALE, AND TOTAL SCORE SCALES
A comparison of the quantitative scales at the previous levels for a total scale is shown here:

Score Level        Ex Unfav                    Neutral                         Ex Fav
Item                       0 ___________1____ 1.5 ____ 2___________3
Subscale 1(8 items) 0 ________________ 12 __________________24
Subscale 2(5 items) 0 ________________ 7.5 _________________15
Subscale 3(4 items) 0 _________________ 6 __________________12
Subscale 4(13 items)0 ________________19.5 ________________39
Subscale 5(6 items) 0 _________________ 9 __________________18
Total Scale(36 items)0 _______________ 54 __________________108

These levels of score reporting and interpretation are based on a single course. This information has the greatest value to you, your department chair, and the curriculum committee evaluating the course.

How does the total score compare to the global item scores given at the end of the scale? Which score should you use? Is there really any difference?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, May 27, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Subscale Level

Page copy protected against web site content infringement by Copyscape

WHAT ARE SUBSCALE SCORES?
This is the first level at which item scores can be summed. If the items on the total scale are grouped into clusters related to topics such as instructional methods, evaluation methods, and course content, the item scores can be summed to produce subscale scores. These scores should only be used for decision making if adequate validity and reliability evidence support that internal scale structure. Since each subscale contains a different number of items, the score range will also be different. This range must be reported to interpret the results.

The summary of subscale score results is derived from the item score results. The statistics are the same. We are just aggregating or summarizing the item-level data (0–3) into subscale item clusters. For example, here are results for three selected subscales:

Instructional Methods (IM)—13 items
     Subscale score range = 0–39; Midpoint = 19.5
     Mean/Median = 34.79/37.00, where N = 97

Evaluation Methods (EM)—5 items
     Subscale score range = 0–15; Midpoint = 7.5
     Mean/Median = 13.40/15.00, where N = 97

Course Content (CC)—8 items
     Subscale score range = 0–24; Midpoint = 12
     Mean/Median = 21.97/24.00, where N = 97

COMPUTATION OF SUBSCALE SCORES: The interpretation of subscale results is analogous to the item results; only the numbers are BIGGER. For example, instead of a 0–3 range and a midpoint of 1.5 for an item, each subscale has a range and midpoint based on its respective number of items. So, for the Instructional Methods (IM) subscale with 13 items, a "0" (SD) response to every item produces a sum of 0 for the subscale, and a "3" (SA) response to all 13 items yields a sum of 39.

INTERPRETATION OF SUBSCALE MEANS AND MEDIANS: The zero-base for all score interpretations is easy to remember: the worst, most unfavorable rating on any item, subscale, or total scale is "0." What changes is the upper score limit for the most favorable rating on each subscale because the number of items change. Again, for the IM subscale, the mean and median can be referenced to the upper limit of 39 and also the midpoint of 19.5 to locate the position on the continuum, as indicated below:

Extremely                                                     Mean       Extremely
Unfavorable                       Neutral                34.79      Favorable
       0___________________19.5____________________39
                                                                             Mdn
                                                                              37
The mean/median ratings on the IM subscale are very favorable.

The subscale results can pinpoint areas of strength and weakness. They may be used by your department chair or the promotion review committee to identify your teaching strengths across different courses. Subscale scores cannot direct you toward particular aspects of teaching that can be improved or changed. The item and anchor results described previously are intended to provide that detailed level of direction.

Finally, the next blog will examine total scores on the scale. What additional info do they provide beyond what we already know?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, May 25, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Item Level—Part 2

Page copy protected against web site content infringement by Copyscape

SHOULD YOU USE THE ITEM MEAN OR MEDIAN? As mentioned previously, when the anchor distribution is negatively skewed, the mean will always be lower than the median, as it is for all three items displayed in the previous blog. The reason is that the mean is drawn toward the few extremely low ratings of SD. The mean is sensitive to extreme scores. (Statisticians’ Concern: For as long as statisticians can remember, the mean has always had an attraction for extreme scores. Granted, this affinity for outliers is not normal. However, statisticians have tolerated this abnormal relationship for years, but also felt compelled to create another index that is not so easily swayed: the median. You may now resume this paragraph already in progress.) Depending on the degree of skew, the bias in interpreting the mean can be significant or insignificant.

DIRECTION AND DEGREE OF ITEM MEAN BIAS: The problem is that the mean misrepresents the actual ratings in a negatively skewed distribution by portraying lower class ratings than actually occurred. This bias MAKES THE INSTRUCTOR APPEAR WORSE in teaching performance, on all of the items, than the students’ ratings indicate. This is not very desirable, especially if these results are used for summative decisions by your department chair or associate dean.

Although means are reported on most commercially published scales, it is strongly recommended that MEDIANS SHOULD BE REPORTED ALONG WITH THE MEANS. Although the median is less discriminating as an index, it is more accurate, more representative, and less biased than the mean for markedly skewed distributions. The lower the degree of skew, the more similar both measures will be. In a perfectly normal distribution, the mean and median are identical. However, keep in mind that ratings of faculty, administrators, courses, programs, and fast food are typically skewed. Therein lays the importance of picking the right index.

BOTTOM LINE RECOMMENDATION: Use both mean and median.

INTERPRETATION: PROFILE OF STRENGTHS AND WEAKNESSES: Since the item means/medians are based on the total N for the class, they can be compared. They display a profile of strengths and weaknesses related to the different teaching behaviors and course characteristics. On a 0−3 scale, means/medians above 1.5 indicate strengths; those below 1.5 denote weaknesses. The means/medians in conjunction with the anchor percentages provide meaningful diagnostic information on areas that might need attention. Again, this report is intended for the instructor's use primarily, although the results on course characteristics may have curricular implications.

Next, the meaning and uses of subscale and total scale scores will be discussed. They are basically summaries of the item scores. Hope you’re finding this stuff helpful. If not, let me know.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, May 23, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Item Level—Part 1

Page copy protected against web site content infringement by Copyscape

WHAT ARE ITEM SCORES?
The next level is the item, where a statistic such as a mean or median is reported. Since most anchor distributions are usually negatively skewed and answers are on a ranked, or ordinal, scale, the median is the most appropriate measure of central tendency. However, given the range of distributions that can occur, you may see both the mean and median on your report form.

“WAIT!! Back up. How did you get from responses of SD, D, etc. to means and medians?” Great question! Glad you’re on the ball. First, you have to convert the “verbal” anchors into “numbers.”

(MEASUREMENT ALERT: Keep in mind that we started with a “qualitative scale” of verbal expressions of how students feel about each behavior and now we’re converting the words into a “quantitative scale” for the convenience of performing analysis of those feelings. Actually, this conversion involves an arbitrary numerical coding scheme.)

CREATE A ZERO-BASED NUMERICAL SCORE SCALE: For simplicity and interpretability, a zero-based scale is recommended, so that the most negative anchor, such as SD, would be coded as “0.” Zero-based scoring was originally recommended by Likert (1932), who created this scaling method. Then the other anchors would be coded in 1-point increments above 0.

Higher values weight more desirable or positive ratings higher than negative ones. SA or Strongly Agreeing with a desirable teaching behavior or course characteristic is weighted with the highest value of 3. An example of this coding for a 4-point, agree–disagree scale is shown below:

SD    D    A    SA
 0     1    2     3

The score range for this single item is 0 to 3. (Note: These score points will vary with the number of anchors and the base number on different scales. Yours may be one of these. Sometimes the number 1 is used as the base instead of 0. Although the number scale may be different, the final interpretation will be the similar.)

COMPUTATION OF ITEM MEANS AND MEDIANS: If you hate stat, this section may make you hurl. Skip it. (SIDEBAR: Over 30 years of teaching stat, I had lots of student hurlers.) For you interested nonhurlers, here are the simple computational definitions:

MEAN = the sum of all students’ scores to each item, divided by the number of students or N. This is the average score for an item, within the range of 0–3 for this example.

MEDIAN = the middle score, after all students’ scores are ranked from high to low.

An example report, based on the anchor data shown in the previous blog, is shown below:

                      SD       D         A        SA       N    Mean   Median
Statement 1   1.0%   3.1%   37.5%   58.6%   96    2.52    3.00
Statement 2   1.0     3.1     24.0     71.9     96    2.65    3.00
Statement 3   1.1     1.1     28.9     68.9     90    2.57    3.00

The median score of 3 means the typical student in the middle of the distribution rated those behaviors as SA. The means were slightly lower with ratings between A and SA. Those are very respectable scores. Of course, they are consistent with the anchor percentage distribution, where the highest percentages are concentrated on the A and SA anchors.

So which index should you use? Mean? Median? Or both? Ah ha! The statistical plot thickens. Stay tuned…

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, May 20, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Anchor Level—Part 2

Page copy protected against web site content infringement by Copyscape

How do you interpret the percentage distribution across the anchors? Is a skewed distribution a good or bad sign for teaching performance? How do any of these results relate to your teaching? Keep perusing, dear colleague.

SKEWED DISTRIBUTION: The overall response pattern shown in the previous blog indicates a negatively skewed distribution of responses, which is the most common outcome and typically a desirable one as well. It occurs when the majority of the responses are A and SA, but there is also a sprinkling of a few Ds and SDs. The extreme SD responses, or outliers, create the skew. Realistically, a few students might mark SD to every statement to express their desire to see you whacked, while the majority of satisfied customers will choose the two “Agree” anchors. These distributions occur in more than 90% of the courses I’ve reviewed at different institutions. It’s rare to receive ratings without any Ds or SDs. Other anchors may yield a different pattern of responses.

DIAGNOSTIC PROFILE: This anchor distribution information provides you with the most detailed profile of responses to a single item. The percentages reveal the degrees of agreement and disagreement with each statement. It is diagnostic of how the class felt about each behavior or characteristic. Examine each distribution carefully to pinpoint your strength behaviors (high percentage of SAs and As) and your weakness behaviors (relatively high percentages of SDs and Ds). You can then consider specific changes in your teaching, evaluation, or course behaviors to shift the distribution farther to the right, in the A–SA zone, the next time the class is taught.

The next blog will examine responses at the item level. Is that info of any value after reviewing the anchor distributions? I will solve that mystery!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, May 17, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Overview

Page copy protected against web site content infringement by Copyscape

APPLICATION TO DIFFERENT FORMS
Although each of you is using a different rating form with different numbers of items and scores, those differences do not matter in score interpretation. Whether you’re using a commercial package, such as IDEA, SIR II, PICES, or CIEQ DU SOLEIL, or a “homegrown scale,” there are only so many score reporting possibilities for any form in Likert-type format. So my suggestions are generic and should be applicable to your form. I encourage you to consult the guidelines or manual for your reporting system for more specific information.

FIVE BASIC CATEGORIES OF RESULTS
There are 5 possible categories of results reported for most student rating forms:

1. anchor distribution of percentages
2. item statistics (mean and/or median)
3. subscale statistics (mean and/or median)
4. total scale statistics (mean and/or median)
5. summary of comments to open-ended questions

Your report form may not provide all of the above, but it should certainly give you at least 2 and 4.

WHAT DO FACULTY NEED?
That's a lot of information. You could use all of those results, however, 1 and 2, in particular, provide the most valuable diagnostic info to revise teaching or course materials that will benefit your next course-load of students. These are called formative decisions about teaching. Category 5 can explain the reasons for the ratings to 1 and 2.

WHAT DO ADMINISTRATORS NEED?
Summative decisions about annual contract renewal, merit pay, or promotion and tenure review by department chairs, associate deans, etc. can be based on 3 and 4 and possibly the global item scores.

This blog series will focus primarily on the faculty needs. My next blog will examine the 1st level of interpretation: ANCHOR-WORLD!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, May 16, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Anchor Level—Part 1

Page copy protected against web site content infringement by Copyscape

RESPONSE RATE WARNING: As you begin to analyze the results from your class evaluations, please consider the response rate in your interpretations:

1. For class sizes of 30 to Super Bowl attendance, it is desirable to have at least 70% response, preferably 80 or 90%, to assure reasonably representative ratings. Anything less may be biased (aka "evil") in some unknown direction. Since all responses are anonymous, there is no way to assess the degree and direction of response bias. Just be cautious in your inferences about your teaching behaviors.
2. For classes less than 30, especially seminars of 5 to 10 students, be particularly careful in your interpretations based on both a less than desirable response rate and inadequate number of responses.

BOTTOM LINE: When your results are used to guide teaching improvement, view your ratings as suggestive of possible areas for change rather than as conclusive.

WHAT’S THE 1ST LEVEL OF INTERPRETATION?

IT’S ANCHOR-WORLD! The first level of score reporting is anchor results for each item. The anchors are the response options on the scale, such as STRONGLY DISAGREE or DISAGREE. Usually the percentage of students picking each anchor is reported, item by item. An example is shown below for student rating scale items with agree–disagree anchors:

                     SD       D         A          SA        N
Statement 1  1.0%   3.1%   37.5%    58.6%    96
Statement 2  1.0     3.1     24.0       71.9      96
Statement 3  1.1     1.1     28.9       68.9      90

ANCHOR SCALE: These results can be reported for any word, phrase, or statement stimuli and any response anchors. The anchor abbreviations for "Strongly Disagree" (SD), "Disagree" (D), "Agree" (A), and "Strongly Agree" (SA) are listed horizontally, left to right, from unfavorable to favorable ratings. This is the same order as the original scale. These anchors measure the degree or intensity of your feeling toward each statement. The N is the number of students that responded to the statement.

Other anchors may ask you evaluate the quality of a behavior, how frequently a behavior occurs, the quantity or the extent to which a behavior occurs, or how a behavior in one course compares to a behavior in another course. There are a variety of possible anchors and number of anchors on the scales now in use.

PERCENTAGE RESPONSES: The percentages for the four anchors indicate the percentage distribution based on the actual N. When the statements are positive teaching behaviors or course characteristics, you would expect low percentages for the first two “Disagree” anchors and high percentages for the second two “Agree” anchors, with the highest for SA. The percentages taper off drastically from right to left, with tiny percentages for D and SD for all three items. (NOTE: The percentages are slightly different for statement 3 compared to statements 1 and 2, particularly the 1.1% to SD and D. This was due, in part, to the six fewer students who responded to that item [N = 90].)

The next blog will examine the meaning of these distributions in terms of a diagnostic profile of your teaching strengths and weaknesses. Stay on board.

COPYRIGHT © 2010 Ronald A. Berk, LLC