Showing posts with label BerksNotes. Show all posts
Showing posts with label BerksNotes. Show all posts

Sunday, July 18, 2010

“WHAT IS WEB 2.0? I DON’T EVEN REMEMBER 1.0! WHO CARES?”

Page copy protected against web site content infringement by Copyscape

DISCLAIMER: You know I’m not a techy. So why am I writing blogs on this topic? First, if I don’t know some of this Web terminology, maybe there are few of you who don’t either. Also, writing about topics I know nothing about helps me grow. I grow inaccurately, but I grow. Actually, I can still read and write and I know you will correct me if I’m wrong. I have your built-in accountability. This is another Berk’sNotes® version on the subject of Web terminology.

Berk’sNotes®
If you’re not familiar with Berk’sNotes® from my Thirteen Strategies… book and previous blogs, here’s a definition:

Berk’sNotes® = In the spirit of CliffsNotes®, it is an abbreviated version or brief synthesis of the most salient information or critical elements of a given topic. It’s designed for people like you who don’t have the time or inclination to do that synthesis yourself. 
MOTTO: “I synthesize the stuff so you don’t have to.”

WHAT HAPPENED TO WEB 1.0?
What did I miss? I honestly don’t even remember Web 1.0. Maybe I experienced what is known as the “Rip van Winkle Effect” or amnesia, or I’m just out of touch. Did you hear about 1.0 before 2.0? It’s probably just me. Maybe I was living in a storm drain at the time.

Anyway, the goal of these blogs is to set me straight and clarify what these terms mean so we can think about how the various tech tools can be leveraged in the classroom and our lives. If you’re interested in all of the sources on this topic, just Google the numbers in the title and the sources cited in my blogs.

Right now, I think we’re into Web 2.0 world, or, maybe, 3.0. I’m not sure. Maybe this blog will get us both up to speed. You probably have an inkling about these numbers. I used to have inklings, but they cleared up with a topical medication. I’m going to try and level our keyboards and inklings on Web 1.0, 2.0, and 3.0. Hang on. My next blog will examine Web 1.0.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, June 3, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Normative Score Comparisons

Page copy protected against web site content infringement by Copyscape

CRITERION-REFERENCED SCORE INTERPRETATIONS
All of the preceding score interpretations compare a score rating to the score range and points along the scale. For example, any rating can be above or below the midpoint or even higher or lower on the respective scale in terms of degree of favorableness, such as illustrated in the previous blog. Cut-off scores can also be set for criterion-referenced interpretations.

NORM-REFERENCED SCORE INTERPRETATIONS
Alternatively or in addition to these score interpretations, ratings can be compared to scores by a norm group, such as those by instructors of similar courses and/or instructors in your department, school, or university. They can also be compared to regional or national norms of instructors who teach the same courses you teach. Commercially-produced scales, such as IDEA, offer those normative scores. These comparisons are called norm-referenced interpretations. The scores at the various levels are still the same; they’re just summarized and reported for different groups of instructors and courses.

STRUCTURED VS. UNSTRUCTURED RESULTS
Did you do really well? If not, you should be able to figure out why you didn't. The answers to the unstructured or open-ended questions at the end of your scale usually provide reasons to explain the responses to the structured student ratings. The structured and unstructured items yield complementary evidence of your performance. Now you probably have more information than you want. That’s why I’m here for you.

I hope the preceding 2 weeks of blogs on interpreting student rating results have helped a little to make sense out of the scores you received. At least, you may have a basic understanding of the possible scores that can be reported and how you can use them to improve your teaching.

Let me know if you have any questions about the material presented. Much more detail on these topics is covered in my Thirteen Strategies book.

HAPPY STUDENT RATINGS!!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, June 1, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Total Scale Score vs. Global Item Scores

Page copy protected against web site content infringement by Copyscape

HOW DOES TOTAL SCORE COMPARE TO GLOBAL ITEM SCORES?
On many scales, global items are included at the end. These items ask students to provide a summary rating of the instructor and/or course. They take on a variety of formats, but the purpose is the same. For example, using anchors ranging from Excellent to Poor, the following items might be given:

What was the overall quality of your instructor’s teaching?
What was the overall value of this course?

OR

using Strongly Agree to Strongly Disagree,

This is the worst instructor on the planet.
This course sucks.

LIMITATIONS: The item response to each of these items would be interpreted the same as any other item on the scale. The problem is that these items are not diagnostic for teaching improvement. They provide an overall rating.

So what’s the problem? Individual item responses, either percentage responses to anchors or item means/medians, are usually unreliable. When those responses are used to suggest areas for improvement, they serve as a guide. No major career-shattering decisions are being made. If global item responses are used for summative decisions by your department chair or the promotion committee, there is a lot more at stake.

RECOMMENDATION: Subscale or total scale scores that summarize the quality of teaching or the value of the course are usually more reliable, based on a collection of items measuring those characteristics, than just a single item. It is recommended that those scores be used in lieu of global item scores whenever possible for any summative decisions about teaching performance.

My final blog in this bloated series will briefly describe criterion- and norm-referenced interpretations of the scores previously defined.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Saturday, May 29, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Total Scale Level

Page copy protected against web site content infringement by Copyscape

WHAT ARE TOTAL SCALE SCORES?
The highest level of score summary is the total scale score across all items. It’s like the total score on a test, except there are no right and wrong answers on a scale. If the scale consists of 36 items, each scored 0–3, the following results might be reported:

Total scale score range = 0–108; Midpoint = 54
        Mean/Median = 96.43/101, where N = 97

The continuum for interpretation would be the following:

Extremely                                                Mean     Extremely
Unfavorable                  Neutral               96.43     Favorable
      0__________________54___________________108
                                                                     Mdn
                                                                      101

INTERPRETATION OF TOTAL SCORES: This score gives a global, or composite, rating that is only as high as the ratings in each of its component parts (anchors, items, and subscales). It represents one overall index of teaching performance, from Extremely Unfavorable to Extremely Favorable. In this example, the score is very favorable. However, the total score is usually a little less reliable and less informative than the subscale scores.

COMPARISON OF ITEM, SUBSCALE, AND TOTAL SCORE SCALES
A comparison of the quantitative scales at the previous levels for a total scale is shown here:

Score Level        Ex Unfav                    Neutral                         Ex Fav
Item                       0 ___________1____ 1.5 ____ 2___________3
Subscale 1(8 items) 0 ________________ 12 __________________24
Subscale 2(5 items) 0 ________________ 7.5 _________________15
Subscale 3(4 items) 0 _________________ 6 __________________12
Subscale 4(13 items)0 ________________19.5 ________________39
Subscale 5(6 items) 0 _________________ 9 __________________18
Total Scale(36 items)0 _______________ 54 __________________108

These levels of score reporting and interpretation are based on a single course. This information has the greatest value to you, your department chair, and the curriculum committee evaluating the course.

How does the total score compare to the global item scores given at the end of the scale? Which score should you use? Is there really any difference?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, May 27, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Subscale Level

Page copy protected against web site content infringement by Copyscape

WHAT ARE SUBSCALE SCORES?
This is the first level at which item scores can be summed. If the items on the total scale are grouped into clusters related to topics such as instructional methods, evaluation methods, and course content, the item scores can be summed to produce subscale scores. These scores should only be used for decision making if adequate validity and reliability evidence support that internal scale structure. Since each subscale contains a different number of items, the score range will also be different. This range must be reported to interpret the results.

The summary of subscale score results is derived from the item score results. The statistics are the same. We are just aggregating or summarizing the item-level data (0–3) into subscale item clusters. For example, here are results for three selected subscales:

Instructional Methods (IM)—13 items
     Subscale score range = 0–39; Midpoint = 19.5
     Mean/Median = 34.79/37.00, where N = 97

Evaluation Methods (EM)—5 items
     Subscale score range = 0–15; Midpoint = 7.5
     Mean/Median = 13.40/15.00, where N = 97

Course Content (CC)—8 items
     Subscale score range = 0–24; Midpoint = 12
     Mean/Median = 21.97/24.00, where N = 97

COMPUTATION OF SUBSCALE SCORES: The interpretation of subscale results is analogous to the item results; only the numbers are BIGGER. For example, instead of a 0–3 range and a midpoint of 1.5 for an item, each subscale has a range and midpoint based on its respective number of items. So, for the Instructional Methods (IM) subscale with 13 items, a "0" (SD) response to every item produces a sum of 0 for the subscale, and a "3" (SA) response to all 13 items yields a sum of 39.

INTERPRETATION OF SUBSCALE MEANS AND MEDIANS: The zero-base for all score interpretations is easy to remember: the worst, most unfavorable rating on any item, subscale, or total scale is "0." What changes is the upper score limit for the most favorable rating on each subscale because the number of items change. Again, for the IM subscale, the mean and median can be referenced to the upper limit of 39 and also the midpoint of 19.5 to locate the position on the continuum, as indicated below:

Extremely                                                     Mean       Extremely
Unfavorable                       Neutral                34.79      Favorable
       0___________________19.5____________________39
                                                                             Mdn
                                                                              37
The mean/median ratings on the IM subscale are very favorable.

The subscale results can pinpoint areas of strength and weakness. They may be used by your department chair or the promotion review committee to identify your teaching strengths across different courses. Subscale scores cannot direct you toward particular aspects of teaching that can be improved or changed. The item and anchor results described previously are intended to provide that detailed level of direction.

Finally, the next blog will examine total scores on the scale. What additional info do they provide beyond what we already know?

COPYRIGHT © 2010 Ronald A. Berk, LLC

Tuesday, May 25, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Item Level—Part 2

Page copy protected against web site content infringement by Copyscape

SHOULD YOU USE THE ITEM MEAN OR MEDIAN? As mentioned previously, when the anchor distribution is negatively skewed, the mean will always be lower than the median, as it is for all three items displayed in the previous blog. The reason is that the mean is drawn toward the few extremely low ratings of SD. The mean is sensitive to extreme scores. (Statisticians’ Concern: For as long as statisticians can remember, the mean has always had an attraction for extreme scores. Granted, this affinity for outliers is not normal. However, statisticians have tolerated this abnormal relationship for years, but also felt compelled to create another index that is not so easily swayed: the median. You may now resume this paragraph already in progress.) Depending on the degree of skew, the bias in interpreting the mean can be significant or insignificant.

DIRECTION AND DEGREE OF ITEM MEAN BIAS: The problem is that the mean misrepresents the actual ratings in a negatively skewed distribution by portraying lower class ratings than actually occurred. This bias MAKES THE INSTRUCTOR APPEAR WORSE in teaching performance, on all of the items, than the students’ ratings indicate. This is not very desirable, especially if these results are used for summative decisions by your department chair or associate dean.

Although means are reported on most commercially published scales, it is strongly recommended that MEDIANS SHOULD BE REPORTED ALONG WITH THE MEANS. Although the median is less discriminating as an index, it is more accurate, more representative, and less biased than the mean for markedly skewed distributions. The lower the degree of skew, the more similar both measures will be. In a perfectly normal distribution, the mean and median are identical. However, keep in mind that ratings of faculty, administrators, courses, programs, and fast food are typically skewed. Therein lays the importance of picking the right index.

BOTTOM LINE RECOMMENDATION: Use both mean and median.

INTERPRETATION: PROFILE OF STRENGTHS AND WEAKNESSES: Since the item means/medians are based on the total N for the class, they can be compared. They display a profile of strengths and weaknesses related to the different teaching behaviors and course characteristics. On a 0−3 scale, means/medians above 1.5 indicate strengths; those below 1.5 denote weaknesses. The means/medians in conjunction with the anchor percentages provide meaningful diagnostic information on areas that might need attention. Again, this report is intended for the instructor's use primarily, although the results on course characteristics may have curricular implications.

Next, the meaning and uses of subscale and total scale scores will be discussed. They are basically summaries of the item scores. Hope you’re finding this stuff helpful. If not, let me know.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, May 23, 2010

A BerksNotes® GUIDE TO INTERPRETING STUDENT RATING RESULTS: Item Level—Part 1

Page copy protected against web site content infringement by Copyscape

WHAT ARE ITEM SCORES?
The next level is the item, where a statistic such as a mean or median is reported. Since most anchor distributions are usually negatively skewed and answers are on a ranked, or ordinal, scale, the median is the most appropriate measure of central tendency. However, given the range of distributions that can occur, you may see both the mean and median on your report form.

“WAIT!! Back up. How did you get from responses of SD, D, etc. to means and medians?” Great question! Glad you’re on the ball. First, you have to convert the “verbal” anchors into “numbers.”

(MEASUREMENT ALERT: Keep in mind that we started with a “qualitative scale” of verbal expressions of how students feel about each behavior and now we’re converting the words into a “quantitative scale” for the convenience of performing analysis of those feelings. Actually, this conversion involves an arbitrary numerical coding scheme.)

CREATE A ZERO-BASED NUMERICAL SCORE SCALE: For simplicity and interpretability, a zero-based scale is recommended, so that the most negative anchor, such as SD, would be coded as “0.” Zero-based scoring was originally recommended by Likert (1932), who created this scaling method. Then the other anchors would be coded in 1-point increments above 0.

Higher values weight more desirable or positive ratings higher than negative ones. SA or Strongly Agreeing with a desirable teaching behavior or course characteristic is weighted with the highest value of 3. An example of this coding for a 4-point, agree–disagree scale is shown below:

SD    D    A    SA
 0     1    2     3

The score range for this single item is 0 to 3. (Note: These score points will vary with the number of anchors and the base number on different scales. Yours may be one of these. Sometimes the number 1 is used as the base instead of 0. Although the number scale may be different, the final interpretation will be the similar.)

COMPUTATION OF ITEM MEANS AND MEDIANS: If you hate stat, this section may make you hurl. Skip it. (SIDEBAR: Over 30 years of teaching stat, I had lots of student hurlers.) For you interested nonhurlers, here are the simple computational definitions:

MEAN = the sum of all students’ scores to each item, divided by the number of students or N. This is the average score for an item, within the range of 0–3 for this example.

MEDIAN = the middle score, after all students’ scores are ranked from high to low.

An example report, based on the anchor data shown in the previous blog, is shown below:

                      SD       D         A        SA       N    Mean   Median
Statement 1   1.0%   3.1%   37.5%   58.6%   96    2.52    3.00
Statement 2   1.0     3.1     24.0     71.9     96    2.65    3.00
Statement 3   1.1     1.1     28.9     68.9     90    2.57    3.00

The median score of 3 means the typical student in the middle of the distribution rated those behaviors as SA. The means were slightly lower with ratings between A and SA. Those are very respectable scores. Of course, they are consistent with the anchor percentage distribution, where the highest percentages are concentrated on the A and SA anchors.

So which index should you use? Mean? Median? Or both? Ah ha! The statistical plot thickens. Stay tuned…

COPYRIGHT © 2010 Ronald A. Berk, LLC

Thursday, May 20, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Anchor Level—Part 2

Page copy protected against web site content infringement by Copyscape

How do you interpret the percentage distribution across the anchors? Is a skewed distribution a good or bad sign for teaching performance? How do any of these results relate to your teaching? Keep perusing, dear colleague.

SKEWED DISTRIBUTION: The overall response pattern shown in the previous blog indicates a negatively skewed distribution of responses, which is the most common outcome and typically a desirable one as well. It occurs when the majority of the responses are A and SA, but there is also a sprinkling of a few Ds and SDs. The extreme SD responses, or outliers, create the skew. Realistically, a few students might mark SD to every statement to express their desire to see you whacked, while the majority of satisfied customers will choose the two “Agree” anchors. These distributions occur in more than 90% of the courses I’ve reviewed at different institutions. It’s rare to receive ratings without any Ds or SDs. Other anchors may yield a different pattern of responses.

DIAGNOSTIC PROFILE: This anchor distribution information provides you with the most detailed profile of responses to a single item. The percentages reveal the degrees of agreement and disagreement with each statement. It is diagnostic of how the class felt about each behavior or characteristic. Examine each distribution carefully to pinpoint your strength behaviors (high percentage of SAs and As) and your weakness behaviors (relatively high percentages of SDs and Ds). You can then consider specific changes in your teaching, evaluation, or course behaviors to shift the distribution farther to the right, in the A–SA zone, the next time the class is taught.

The next blog will examine responses at the item level. Is that info of any value after reviewing the anchor distributions? I will solve that mystery!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Monday, May 17, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Overview

Page copy protected against web site content infringement by Copyscape

APPLICATION TO DIFFERENT FORMS
Although each of you is using a different rating form with different numbers of items and scores, those differences do not matter in score interpretation. Whether you’re using a commercial package, such as IDEA, SIR II, PICES, or CIEQ DU SOLEIL, or a “homegrown scale,” there are only so many score reporting possibilities for any form in Likert-type format. So my suggestions are generic and should be applicable to your form. I encourage you to consult the guidelines or manual for your reporting system for more specific information.

FIVE BASIC CATEGORIES OF RESULTS
There are 5 possible categories of results reported for most student rating forms:

1. anchor distribution of percentages
2. item statistics (mean and/or median)
3. subscale statistics (mean and/or median)
4. total scale statistics (mean and/or median)
5. summary of comments to open-ended questions

Your report form may not provide all of the above, but it should certainly give you at least 2 and 4.

WHAT DO FACULTY NEED?
That's a lot of information. You could use all of those results, however, 1 and 2, in particular, provide the most valuable diagnostic info to revise teaching or course materials that will benefit your next course-load of students. These are called formative decisions about teaching. Category 5 can explain the reasons for the ratings to 1 and 2.

WHAT DO ADMINISTRATORS NEED?
Summative decisions about annual contract renewal, merit pay, or promotion and tenure review by department chairs, associate deans, etc. can be based on 3 and 4 and possibly the global item scores.

This blog series will focus primarily on the faculty needs. My next blog will examine the 1st level of interpretation: ANCHOR-WORLD!

COPYRIGHT © 2010 Ronald A. Berk, LLC

Sunday, May 16, 2010

A BerksNotes® GUIDE TO STUDENT RATING SCORE INTERPRETATION: Anchor Level—Part 1

Page copy protected against web site content infringement by Copyscape

RESPONSE RATE WARNING: As you begin to analyze the results from your class evaluations, please consider the response rate in your interpretations:

1. For class sizes of 30 to Super Bowl attendance, it is desirable to have at least 70% response, preferably 80 or 90%, to assure reasonably representative ratings. Anything less may be biased (aka "evil") in some unknown direction. Since all responses are anonymous, there is no way to assess the degree and direction of response bias. Just be cautious in your inferences about your teaching behaviors.
2. For classes less than 30, especially seminars of 5 to 10 students, be particularly careful in your interpretations based on both a less than desirable response rate and inadequate number of responses.

BOTTOM LINE: When your results are used to guide teaching improvement, view your ratings as suggestive of possible areas for change rather than as conclusive.

WHAT’S THE 1ST LEVEL OF INTERPRETATION?

IT’S ANCHOR-WORLD! The first level of score reporting is anchor results for each item. The anchors are the response options on the scale, such as STRONGLY DISAGREE or DISAGREE. Usually the percentage of students picking each anchor is reported, item by item. An example is shown below for student rating scale items with agree–disagree anchors:

                     SD       D         A          SA        N
Statement 1  1.0%   3.1%   37.5%    58.6%    96
Statement 2  1.0     3.1     24.0       71.9      96
Statement 3  1.1     1.1     28.9       68.9      90

ANCHOR SCALE: These results can be reported for any word, phrase, or statement stimuli and any response anchors. The anchor abbreviations for "Strongly Disagree" (SD), "Disagree" (D), "Agree" (A), and "Strongly Agree" (SA) are listed horizontally, left to right, from unfavorable to favorable ratings. This is the same order as the original scale. These anchors measure the degree or intensity of your feeling toward each statement. The N is the number of students that responded to the statement.

Other anchors may ask you evaluate the quality of a behavior, how frequently a behavior occurs, the quantity or the extent to which a behavior occurs, or how a behavior in one course compares to a behavior in another course. There are a variety of possible anchors and number of anchors on the scales now in use.

PERCENTAGE RESPONSES: The percentages for the four anchors indicate the percentage distribution based on the actual N. When the statements are positive teaching behaviors or course characteristics, you would expect low percentages for the first two “Disagree” anchors and high percentages for the second two “Agree” anchors, with the highest for SA. The percentages taper off drastically from right to left, with tiny percentages for D and SD for all three items. (NOTE: The percentages are slightly different for statement 3 compared to statements 1 and 2, particularly the 1.1% to SD and D. This was due, in part, to the six fewer students who responded to that item [N = 90].)

The next blog will examine the meaning of these distributions in terms of a diagnostic profile of your teaching strengths and weaknesses. Stay on board.

COPYRIGHT © 2010 Ronald A. Berk, LLC

Friday, May 14, 2010

WHAT ARE YOU DOING WITH YOUR STUDENT RATING FORM RESULTS?

Page copy protected against web site content infringement by Copyscape

ADMINISTRATION OF STUDENT RATING FORMS
Yup! It’s that time of the year. The Cherry Blossoms are gone and many professors on the east coast are sneezing and wheezing their brains out from the sky-high pollen counts. They’re medicating themselves with megadoses of antihistamines like Benadryl, Zyrtec, Claritin, Allegra, Clarinex, Flonase, Nazonex, and Sneezwhizz.

This is the perfect time to administer those end-of-course student rating forms and interpret the scores. The results usually appear better when you’re drowsy and punchy from those medications. This month has the highest number of forms administered world-wide, except for December when the same professors are sneezing and wheezing from the common cold.

The forms may be administered in class or online, but the results have to be reported in some format. You may receive the results in a couple of days to several months, depending on your processing system. Of course, all of this happens in between commencement exercises and end-of-year parties. Hopefully, you can glance at the form results before your next course begins in the summer or fall.

YOUR "RATING ANGEL"
I am your Rating Angel for this next week. If you’re not sure how to interpret the scores or use the results, this blog series is for YOU! I want you to milk those ratings for all their worth, to squeeze every drip of information that can guide your teaching improvement.

If you already know how to interpret these scores, STOP reading this blog immediately. Disregard it and get back to work, class, or lunch. Stop fooling around and wasting time on this blog. You should be ashamed of yourself.

SOURCE ALERT
My Thirteen Strategies book goes into considerable detail on the how, why, computing, reporting, formatting, and who cares about those scores. This series will simply focus on the scores you can use to guide decisions about your teaching improvement.

GOAL OF BLOG SERIES
This blog series is intended to present the BerksNotes® version of student rating scale interpretation. Hopefully, by cutting to the chase, whatever that is, you’ll be able to get the most out of your scores in lickety-split time or faster. Hold on to your keyboards. My next blog will begin with an overview of the different types of scores.

COPYRIGHT © 2010 Ronald A. Berk, LLC