My coworker K and I attended one day of this conference, being held in Austin this year, on Tuesday. We went to three presentations in the morning, had a lunch of good-tasting food in fairly small portions (including the thinnest sliver of Italian cream cake I have ever had the good fortune of encountering) at the hotel while being serenaded by a cowboy singer, and took an afternoon field trip to a couple locations in Wimberley (where I had never been).
Unbeatable Visual Highlight:
Floyd fell into the water at Jacob's Well when the unstable planks making a bridge across the water slipped off the rock they were rather half-assedly braced against. He took this in much better spirits than one might expect, especially given that his cell phone was ruined. Fortunately, he fell into the waist-deep water on the near side of the planks and did not fall into the actual cave. A couple young swimmers, including one with full body tattoos that a fellow visitor aptly described as giving him the appearance of a scaled fish (and not, in my view, in a bad way), held the boards steady so I was able to get back across with dry feet, though I was prepared to walk back through the water with my backpack held over my head.
Best Nature Sighting:
The locally rare Chatterbox Orchid growing along the stream bank. Pretty.
Coincidence #1:
I let K pick the seminars we went to, so not having paid much attention at all to what they were about, I was surprised about 90 seconds into the very first one to discover that it was about the development of trails and parks by Jefferson County Open Space in Colorado, where Tam lives. It kept me awake (who had been half-snoozing in a comfy chair in the hotel lobby before the session began) hearing about and looking at photos of places I had actually been - Crown Hill Park (which Jeffco Open Space did not originally support, by the way), Wheat Ridge Recreation Center, the trail along the creek in Golden, Golden Gate State Park (not operated by Jeffco, but in the area), and more. Otherwise, the most interesting thing about this session was seeing how fucking much money these people have to play with. (Presentations at these kinds of conferences are generally given by groups with the kind of money that my agency - and many to most others - can only dream of. I wish there were more "Public Input Processes on a Shoestring" and "High Quality, Fast, Cheap, Pick 2, Maybe: Planning in the Real World" type sessions available, but somehow, nobody has really figured out how to do a good job of things in those situations.) The other was a comment by another attendee that surveys of users of their park/trail system have found a higher percentage of people bringing dogs than bringing children. This pointed up to me the strange disconnect, common to a great many park systems, between actual visitors and the interpretive programming that is made available; too often interpreters provide "family" or "child" oriented programming but nothing for adults.
Coincidence #2:
Our musical entertainment, Jane Leche, is a member of the US Forest Service group "The Fiddlin' Foresters." One of the items available at the Silent Auction was their CD, which featured a cover of one of my favorite songs, "Cold Missouri Waters," which tells the story of the famous Mann Gulch fire in Montana. The leader of the fire crew independently came up with the idea of setting a fire to an area of grass and lying down in the burned out area; because the fire was deprived of fuel in that spot, the flames went around him, saving his life. I was disappointed that Jane did not sing "Cold Missouri Waters" for us at our lunch concert, but we can all listen to her version here. Not as great as Richard Shindell's amazing, goose-bump-inducing version, but it's a good song nonetheless.
Silliest Question:
After I attested to the difficulty of shooting a traditional English longbow given a draw weight of 100 pounds or more that you have to hold to aim, unlike a compound bow that has a much lower hold weight, Floyd of the wet pants asked me, "Sally, are you an Englishman?" I said no. (I'm not much of an archer either. Or an expert on medieval weapons.) This was in the context of discussing the story (story, not reality) that the origins of the raised middle finger and phrase "fuck you" arose from the English bowmen at the Battle of Agincourt raising their middle fingers to the French (who had threatened to cut the middle finger off any captured Englishmen, thereby making them incapable of shooting bows which required three fingers to pull) and saying that they could still "pluck the yew" (the tree from which the bows were made). Although this story is not true, it does make for an amusing bon mot. May "pluck yew" join "Chuck you, Farley" in my personal lexicon of "fuck you" alternatives.
Most Surprising Overheard Comment:
"I spent $5,000 on it used and got another 320,000 miles out of that [American - can't remember which brand] truck before I rebuilt the engine."
Unplanned Side Stop for the Wimberley Field Trip:
Sally's apartment complex! Robert had taken me to the hotel downtown in the morning and was planning to pick me up after the field trip was over. But since we were driving by the apartment on the way back from Wimberley, I asked if they would mind dropping me off on the way back to the hotel. Either they really didn't mind or they admired my pluck/ballsiness in asking enough that they went along with it, though a couple of people said that a spot of single malt scotch would go down quite well before they returned to the hotel. (Little did they know that I actually do have several quite nice bottles in my kitchen.) This saved me a good 45 minutes of commute time, possibly more. I had fun calling Robert on my cell phone and, barely able to hear him and uncertain how well he could hear me, saying "Go home. I have a ride. OK? Just go home!" He got it. Somebody asked me how much my apartment cost, and I told him, and several people from other states seemed surprised - welcome to Austin real estate! (To their credit, they could not know that it's a luxuriously comfortable 1350 square foot 2 bedroom/2 bathroom apartment with two people and one rabbit. They did not actually get the tour.)
Wednesday, May 9, 2007
Lesson in Electrical Engineering
Upon inspection of our freezer, the repairman told Robert: "It's a good thing your rabbit chewed the white wire and not the black one."
The voltage in the white wire is basically zero. The black wire can give you a nasty shock (and electrocute a rabbit).
So remember this, little bunnies:
Chewing wire
Can be dire:
If you must bite,
Nip the white.
If you chomp black,
It will fight back;
Don't make this chew
The last thing you do.
The voltage in the white wire is basically zero. The black wire can give you a nasty shock (and electrocute a rabbit).
So remember this, little bunnies:
Chewing wire
Can be dire:
If you must bite,
Nip the white.
If you chomp black,
It will fight back;
Don't make this chew
The last thing you do.
Sunday, May 6, 2007
A Series of Annoying Events
Thursday evening, our air conditioner broke. The apartment complex repair guy came to look at it on Friday and determined that they needed to replace the motor. A couple other guys, who work for a company the complex contracts out to, came out on Saturday morning and replaced it. But the A/C didn't seem to really be working right even after that, so they came back to check it out and realized that the person at the parts supply store gave them the wrong motor. So they will have to get another one on Monday when the store is open again and switch them out.
Today, I was doing some sewing in Leo's room (where I have to block him away from the sewing table area because he loves to pick up the sewing pedal and toss it out from under my foot) and when I finished, for some reason I can't quite explain (but I assume must have been the result of a sub-conscious appreciation that the room wasn't as noisy as it should be), I decided to open our one-week-old freezer (which contains 5 dozen muffins, 8 pieces of lasagna, and 4 bowls of soup) that is also in Leo's room. And everything was unfrozen and just kind of chilled feeling. Robert came to look at it and discovered that Leo had chewed a hole in the cord. This was quite mysterious because Robert had specifically positioned the freezer unit and the cord against the wall such that Leo's lovely round body could not possibly fit behind it. Yet this was the only explanation consistent with the bitten cord. I was starting to wonder whether rabbits have some amazing ability to squeeze into impossibly small places, but as Robert pointed out, Leo would have had to dislocate his hip to get back there.
But on further examination of the freezer, Robert noticed that it wasn't actually as flush to the wall as it had been, which he found confusing until I realized: on Saturday, the contractor needed to access the breaker box in Leo's room, which is next to the freezer. He must have moved the freezer to get to it easier and then pushed it back, only not as far as it had been, thus leaving Leo with just a large enough opening to get in.
So we currently have a half-assed A/C running and no freezer. The repair phone line the store referred Robert to is not open on the weekend, so we won't know until tomorrow (or later) what will have to be done to get the freezer back in action. Presumably they can simply replace the cord at who knows what expense. (The possibility that we have a $500, ~12 cubic foot paperweight is too depressing to contemplate.)
This, on top of the fact that Robert can't get his new desktop computer that he purchased to run his huge SAS programs to talk to his laptop, has made for a weekend of all kinds of stuff just not working right. I even ran into some difficulty with my sewing machine that finally resolved itself with quite a bit of frustration but no serious mistakes. Our Internet has been iffy all weekend as well, only connecting about half of the time and thus frequently requiring Robert to perform an increasingly complicated series of rituals to get it to work.
At this point, I am almost afraid to try starting my car in the morning. (Actually, for the past couple weeks, it has been giving me a little bit of a hassle when first start it to come home from work; not enough to worry me, but just enough to notice.)
Next: Robert and I make the stereo burst into flames just by looking at it funny.
Today, I was doing some sewing in Leo's room (where I have to block him away from the sewing table area because he loves to pick up the sewing pedal and toss it out from under my foot) and when I finished, for some reason I can't quite explain (but I assume must have been the result of a sub-conscious appreciation that the room wasn't as noisy as it should be), I decided to open our one-week-old freezer (which contains 5 dozen muffins, 8 pieces of lasagna, and 4 bowls of soup) that is also in Leo's room. And everything was unfrozen and just kind of chilled feeling. Robert came to look at it and discovered that Leo had chewed a hole in the cord. This was quite mysterious because Robert had specifically positioned the freezer unit and the cord against the wall such that Leo's lovely round body could not possibly fit behind it. Yet this was the only explanation consistent with the bitten cord. I was starting to wonder whether rabbits have some amazing ability to squeeze into impossibly small places, but as Robert pointed out, Leo would have had to dislocate his hip to get back there.
But on further examination of the freezer, Robert noticed that it wasn't actually as flush to the wall as it had been, which he found confusing until I realized: on Saturday, the contractor needed to access the breaker box in Leo's room, which is next to the freezer. He must have moved the freezer to get to it easier and then pushed it back, only not as far as it had been, thus leaving Leo with just a large enough opening to get in.
So we currently have a half-assed A/C running and no freezer. The repair phone line the store referred Robert to is not open on the weekend, so we won't know until tomorrow (or later) what will have to be done to get the freezer back in action. Presumably they can simply replace the cord at who knows what expense. (The possibility that we have a $500, ~12 cubic foot paperweight is too depressing to contemplate.)
This, on top of the fact that Robert can't get his new desktop computer that he purchased to run his huge SAS programs to talk to his laptop, has made for a weekend of all kinds of stuff just not working right. I even ran into some difficulty with my sewing machine that finally resolved itself with quite a bit of frustration but no serious mistakes. Our Internet has been iffy all weekend as well, only connecting about half of the time and thus frequently requiring Robert to perform an increasingly complicated series of rituals to get it to work.
At this point, I am almost afraid to try starting my car in the morning. (Actually, for the past couple weeks, it has been giving me a little bit of a hassle when first start it to come home from work; not enough to worry me, but just enough to notice.)
Next: Robert and I make the stereo burst into flames just by looking at it funny.
Saturday, May 5, 2007
Student Ratings of Teaching, Part 1
Note: If this long post starts to wear you down, buck up - it is followed by a video of the funniest thing it has been my pleasure to see in a long time - 1:00 of hilarious bunny-filled brilliance. (Thanks, Mom, for telling me about this commercial and "thank you, science.")
In the comments to the post on people’s ability to evaluate an instructor’s personality based on viewing 6 seconds of silent video, Tam asks:
"I also wonder what the results are of studies comparing these factors (the ones that influence student evaluations, or just the results of student evaluations themselves) to actual effectiveness in teaching - i.e., how much students learn."
The easy answer is, it depends on whom you ask. The validity of student evaluation of teachers is a matter of great controversy among researchers (to say nothing of the larger general academic public). At the most extreme, this divides into two camps: those who have made research predicated on the necessity that student evaluations are reasonably valid measures a prominent aspect of their career and are tempted to believe this despite the contradictory evidence that may appear and those who believe that student evaluations are of obvious bogosity and are tempted to hold researchers in this area to a standard that perhaps is unfairly rigorous (and that they are unlikely to match in their own areas of interest) such that their opponents can never make a good enough case to satisfy them.
Of course, my own tendencies toward critical assessment make me naturally inclined to be skeptical of student evaluations and a half-assed cursory examination of the literature does not make me any less a tentative ally of those in the second camp.
Here I will discuss at length one particular journal article that I liked a lot (written by a critic of student evaluations of teaching named Olivares), briefly another that had a useful run-down of some empirical findings (written by Ahmadi and Cotton), and my general thoughts on the subject. (Sources at the end of the post as usual.)
The Olivares article begins with a review of the as-near-to-universally-accepted-as-I-can-imagine definitions of validity. It poses the general question, what would it mean to say that student ratings of teaching (SRTs) are valid? At its most basic level, a valid measure is one that measures what it is supposed to measure. There are several types of validity that psychologists talk about:
- Content validity: Does the measure (SRTs) represent all aspects of teacher effectiveness?
- Criterion validity: Is there a meaningful relationship between the measure (SRTs) and some measure of the relevant behavior (such as “student learning”)? Note that this is related to the question that Tam asks.
- Construct validity: Do SRTs measure a trait or characteristic of interest? Does “teacher effectiveness” exist?
Olivares argues that SRTs do not hold up well to an examination of their validity. One major problem is that without a good definition of “teacher effectiveness,” it becomes next to impossible to judge whether SRTs do a good job of measuring it. In the absence of such a definition, the ratings themselves become the de facto operational definition of teaching effectiveness. He quotes another critic who has issues with the most obvious definition of teacher effectiveness – how much students learn: “The best teaching is not that which produces the most learning, since what is learned may be worthless.”
I am inclined to agree that the lack of a well-formulated definition of teacher effectiveness is problematic, and the pervasive use of SRTs as the de facto operational definition can put the field in the uncomfortable position of being caught in a circular reference: What is teacher effectiveness? What this test of teacher effectiveness measures. I feel sure that most researchers who use SRTs as a measure in their work appreciate the fact that teacher effectiveness is a multi-faceted concept, but I know that there is a strong tendency to privilege in your mind whatever aspect of some complex thing you can measure and do something with. In my opinion, psychologists, who generally work in an experimental mode and hence have a bit better control over their datasets, can be less prone to this than other social scientists* (and even doctors perhaps**), but it is a danger in all of these fields. I think it’s entirely appropriate for curmudgeonly critics (and I am obviously a student member of the Curmudgeonly Critics of America) to occasionally remind researchers of this fact.
* As Robert has said, economists are forced to use whatever data they can find and thus use very strange measures indeed, like tractors-per-capita, as proxies for their variables of interest, and when asked to explain what one of these measures means, are inclined to respond, “I don’t know, but it explains 73% of the variance.”
** My mom recently commented to me that she was starting to wonder if her doctor’s insistent focus on her cholesterol level was a true reflection of the importance of that level to her health or was simply an artifact of cholesterol being something that she could measure.
It is interesting to think about the ways that even “how much students learn” fails as a universally acceptable definition for teacher effectiveness. Even assuming there was some way to get a very good measure of this (using some kind of pre-/post- measure of knowledge and controlling for student variables like intelligence, motivation, study habits, etc., that could impact learning), maximizing the sheer amount of learning is not always the sole (or in some cases, primary) goal of teaching. One aspect that I think is important is a teacher’s ability to stimulate interest in and future study/thinking about a subject in students. It is easy to imagine the instructor who by blunt force crams a significant amount of knowledge into students’ heads long enough to take the exam, but whose students come away hating the subject and eager to forget this boring crap as soon as possible. Depending on the situation, the ideal balance between knowledge and interest may shift, but in most cases, I believe you want to do both.
It’s certainly true that learning a great amount of trivial information is less desirable than mastery of the fundamental concepts of a subject. (For instance, an American history student may be able to regurgitate a large number of names, dates, and places without understanding how any of it ties together.)
Also, it’s possible that in specialized situations, teachers would focus on very different things. For example, a teacher working with students disadvantaged by some combination of circumstances (e.g. socio-economic status) and innate abilities (e.g. learning disability), and with a history of low achievement, may emphasize increasing the students’ motivation to learn and feelings of self-efficacy toward learning at the expense of short-term mastery of the subject material. (By this I mean students who have done poorly in classes that progressed at a normal pace might be placed in a more slowly paced course that allows them to realize, hey, with effort, I can learn something; it just takes me longer to do it.) Most to all junior colleges and some universities teach what at Rice and (to my knowledge) most other schools is a two-semester course in calculus over three semesters, presumably because they recognize that their students are not prepared to take on this material at such a fast pace. On the flip side, as Robert pointed out, other institutions use a “weed out” process to separate out those who can advance through difficult material very quickly from the masses who cannot. And even I can see the value of this to an elite program for astronauts, specialized doctors, or that kind of thing.
Another huge issue to me, that is more methodological than theoretical, is that effective teaching results in learning that lasts beyond the final exam.
The article lays out four assumptions about student ratings of teachers that Olivares believes are not sufficiently met:
"- rating forms adequately capture the domain of teacher effectiveness across instructional settings, academic disciplines, instructors and course levels and types;
- students know what effective teaching is, hold a common view of teacher effectiveness, and are objective and reliable sources of teacher effectiveness data;
- relatedly, ratings are, for all intents and purposes, unaffected by potential biasing variables; and, collaterally;
- teacher effectiveness is being measured as opposed to, for example, course difficulty or differences in disciplines, student characteristics, grading leniency, teacher expressiveness, teacher popularity or any number of other variables."
Olivares states (in a sentence I enjoy very much), “To think that students, who have no training in evaluation, are not content experts, and possess myriad idiosyncratic tendencies, would not be susceptible to errors in judgment is specious.” I agree that to the degree that the validity of SRTs is dependent on believing otherwise, the SRT project is doomed.
Ahmadi and Cotton, who conclude in their article that “in general, student ratings tend to be statistically reliable, valid, and relatively free from bias or the need for control,” give a run-down of some findings that may be useful and, more significantly, on point to answering Tam’s question.
First, they report that studies have found correlations between exam grades and SRTs such that classes that gave higher ratings tend to be the ones where students learned more (i.e. scored higher on an [I believe standardized] exam). However, they acknowledge that many variables related to student learning are themselves related to student ability rather than teacher performance. This might imply that the answer to Tam’s question is “Yes, but.…”
They list the following factors that have been found to not be related to student ratings: instructor research productivity, age, teaching experience, race, and gender (though there can be interaction effects between student race or gender and teacher race or gender, with students giving higher ratings to instructors with similar characteristics); student age, level (e.g. freshman, grad student), GPA, and personality; class size and time of day.
They also list factors that are related: faculty rank and teacher expressiveness; students’ expected grades and motivation; work load and difficulty (perhaps surprisingly and reassuringly, those correlate positively, with classes perceived as more difficult getting higher ratings), level of course (higher level are rated more positively), and academic field (humanities and arts>social sciences>math and science).
Getting back to our fellow curmudgeonly critic Olivares. He talks about how, when pressed, many supporters of SRTs will fall back on an argument for their utility; he quotes one who wrote, “Student ratings almost certainly contain useful information that is independent of their correlation with student achievement. That is, student ratings provide information on how well students like a course.”
Of course, this is where yet another of my buttons get pushed – customer satisfaction. So join me later for the continuation of this discussion, focusing on the theory of customer satisfaction and practice of its measurement, SRTs as a customer satisfaction measure, comparisons of c-sat in teaching and a field I know quite a bit about, viewing students as the single relevant consumer group, the implications of SRTs for teacher behavior, and the use of SRTs. I do not mean use as in, are Likert scales, which are considered ordinal level data following Stevens' arguably invalid definitions in measurement theory, appropriately reported using parametric statistics (which is an interesting if highly geeky debate with implications for the calculation of student GPAs as well), but rather: the use and misuse of SRTs (understood as a c-sat measure) as an element in making instructor personnel decisions, such as granting tenure.
Sources:
A Conceptual and Analytic Critique of Student Ratings of Teachers in the USA with Implications for Teacher Effectiveness and Student Learning, Teaching in Higher Education, Vol. 8, No. 2, 2003, pp. 233–245, ORLANDO J. OLIVARES
Assessing Students’ Ratings of Faculty, Assessment Update, September–October 1998, Volume 10, Number 5, Reza T. Ahmadi, Samuel E. Cotton
In the comments to the post on people’s ability to evaluate an instructor’s personality based on viewing 6 seconds of silent video, Tam asks:
"I also wonder what the results are of studies comparing these factors (the ones that influence student evaluations, or just the results of student evaluations themselves) to actual effectiveness in teaching - i.e., how much students learn."
The easy answer is, it depends on whom you ask. The validity of student evaluation of teachers is a matter of great controversy among researchers (to say nothing of the larger general academic public). At the most extreme, this divides into two camps: those who have made research predicated on the necessity that student evaluations are reasonably valid measures a prominent aspect of their career and are tempted to believe this despite the contradictory evidence that may appear and those who believe that student evaluations are of obvious bogosity and are tempted to hold researchers in this area to a standard that perhaps is unfairly rigorous (and that they are unlikely to match in their own areas of interest) such that their opponents can never make a good enough case to satisfy them.
Of course, my own tendencies toward critical assessment make me naturally inclined to be skeptical of student evaluations and a half-assed cursory examination of the literature does not make me any less a tentative ally of those in the second camp.
Here I will discuss at length one particular journal article that I liked a lot (written by a critic of student evaluations of teaching named Olivares), briefly another that had a useful run-down of some empirical findings (written by Ahmadi and Cotton), and my general thoughts on the subject. (Sources at the end of the post as usual.)
The Olivares article begins with a review of the as-near-to-universally-accepted-as-I-can-imagine definitions of validity. It poses the general question, what would it mean to say that student ratings of teaching (SRTs) are valid? At its most basic level, a valid measure is one that measures what it is supposed to measure. There are several types of validity that psychologists talk about:
- Content validity: Does the measure (SRTs) represent all aspects of teacher effectiveness?
- Criterion validity: Is there a meaningful relationship between the measure (SRTs) and some measure of the relevant behavior (such as “student learning”)? Note that this is related to the question that Tam asks.
- Construct validity: Do SRTs measure a trait or characteristic of interest? Does “teacher effectiveness” exist?
Olivares argues that SRTs do not hold up well to an examination of their validity. One major problem is that without a good definition of “teacher effectiveness,” it becomes next to impossible to judge whether SRTs do a good job of measuring it. In the absence of such a definition, the ratings themselves become the de facto operational definition of teaching effectiveness. He quotes another critic who has issues with the most obvious definition of teacher effectiveness – how much students learn: “The best teaching is not that which produces the most learning, since what is learned may be worthless.”
I am inclined to agree that the lack of a well-formulated definition of teacher effectiveness is problematic, and the pervasive use of SRTs as the de facto operational definition can put the field in the uncomfortable position of being caught in a circular reference: What is teacher effectiveness? What this test of teacher effectiveness measures. I feel sure that most researchers who use SRTs as a measure in their work appreciate the fact that teacher effectiveness is a multi-faceted concept, but I know that there is a strong tendency to privilege in your mind whatever aspect of some complex thing you can measure and do something with. In my opinion, psychologists, who generally work in an experimental mode and hence have a bit better control over their datasets, can be less prone to this than other social scientists* (and even doctors perhaps**), but it is a danger in all of these fields. I think it’s entirely appropriate for curmudgeonly critics (and I am obviously a student member of the Curmudgeonly Critics of America) to occasionally remind researchers of this fact.
* As Robert has said, economists are forced to use whatever data they can find and thus use very strange measures indeed, like tractors-per-capita, as proxies for their variables of interest, and when asked to explain what one of these measures means, are inclined to respond, “I don’t know, but it explains 73% of the variance.”
** My mom recently commented to me that she was starting to wonder if her doctor’s insistent focus on her cholesterol level was a true reflection of the importance of that level to her health or was simply an artifact of cholesterol being something that she could measure.
It is interesting to think about the ways that even “how much students learn” fails as a universally acceptable definition for teacher effectiveness. Even assuming there was some way to get a very good measure of this (using some kind of pre-/post- measure of knowledge and controlling for student variables like intelligence, motivation, study habits, etc., that could impact learning), maximizing the sheer amount of learning is not always the sole (or in some cases, primary) goal of teaching. One aspect that I think is important is a teacher’s ability to stimulate interest in and future study/thinking about a subject in students. It is easy to imagine the instructor who by blunt force crams a significant amount of knowledge into students’ heads long enough to take the exam, but whose students come away hating the subject and eager to forget this boring crap as soon as possible. Depending on the situation, the ideal balance between knowledge and interest may shift, but in most cases, I believe you want to do both.
It’s certainly true that learning a great amount of trivial information is less desirable than mastery of the fundamental concepts of a subject. (For instance, an American history student may be able to regurgitate a large number of names, dates, and places without understanding how any of it ties together.)
Also, it’s possible that in specialized situations, teachers would focus on very different things. For example, a teacher working with students disadvantaged by some combination of circumstances (e.g. socio-economic status) and innate abilities (e.g. learning disability), and with a history of low achievement, may emphasize increasing the students’ motivation to learn and feelings of self-efficacy toward learning at the expense of short-term mastery of the subject material. (By this I mean students who have done poorly in classes that progressed at a normal pace might be placed in a more slowly paced course that allows them to realize, hey, with effort, I can learn something; it just takes me longer to do it.) Most to all junior colleges and some universities teach what at Rice and (to my knowledge) most other schools is a two-semester course in calculus over three semesters, presumably because they recognize that their students are not prepared to take on this material at such a fast pace. On the flip side, as Robert pointed out, other institutions use a “weed out” process to separate out those who can advance through difficult material very quickly from the masses who cannot. And even I can see the value of this to an elite program for astronauts, specialized doctors, or that kind of thing.
Another huge issue to me, that is more methodological than theoretical, is that effective teaching results in learning that lasts beyond the final exam.
The article lays out four assumptions about student ratings of teachers that Olivares believes are not sufficiently met:
"- rating forms adequately capture the domain of teacher effectiveness across instructional settings, academic disciplines, instructors and course levels and types;
- students know what effective teaching is, hold a common view of teacher effectiveness, and are objective and reliable sources of teacher effectiveness data;
- relatedly, ratings are, for all intents and purposes, unaffected by potential biasing variables; and, collaterally;
- teacher effectiveness is being measured as opposed to, for example, course difficulty or differences in disciplines, student characteristics, grading leniency, teacher expressiveness, teacher popularity or any number of other variables."
Olivares states (in a sentence I enjoy very much), “To think that students, who have no training in evaluation, are not content experts, and possess myriad idiosyncratic tendencies, would not be susceptible to errors in judgment is specious.” I agree that to the degree that the validity of SRTs is dependent on believing otherwise, the SRT project is doomed.
Ahmadi and Cotton, who conclude in their article that “in general, student ratings tend to be statistically reliable, valid, and relatively free from bias or the need for control,” give a run-down of some findings that may be useful and, more significantly, on point to answering Tam’s question.
First, they report that studies have found correlations between exam grades and SRTs such that classes that gave higher ratings tend to be the ones where students learned more (i.e. scored higher on an [I believe standardized] exam). However, they acknowledge that many variables related to student learning are themselves related to student ability rather than teacher performance. This might imply that the answer to Tam’s question is “Yes, but.…”
They list the following factors that have been found to not be related to student ratings: instructor research productivity, age, teaching experience, race, and gender (though there can be interaction effects between student race or gender and teacher race or gender, with students giving higher ratings to instructors with similar characteristics); student age, level (e.g. freshman, grad student), GPA, and personality; class size and time of day.
They also list factors that are related: faculty rank and teacher expressiveness; students’ expected grades and motivation; work load and difficulty (perhaps surprisingly and reassuringly, those correlate positively, with classes perceived as more difficult getting higher ratings), level of course (higher level are rated more positively), and academic field (humanities and arts>social sciences>math and science).
Getting back to our fellow curmudgeonly critic Olivares. He talks about how, when pressed, many supporters of SRTs will fall back on an argument for their utility; he quotes one who wrote, “Student ratings almost certainly contain useful information that is independent of their correlation with student achievement. That is, student ratings provide information on how well students like a course.”
Of course, this is where yet another of my buttons get pushed – customer satisfaction. So join me later for the continuation of this discussion, focusing on the theory of customer satisfaction and practice of its measurement, SRTs as a customer satisfaction measure, comparisons of c-sat in teaching and a field I know quite a bit about, viewing students as the single relevant consumer group, the implications of SRTs for teacher behavior, and the use of SRTs. I do not mean use as in, are Likert scales, which are considered ordinal level data following Stevens' arguably invalid definitions in measurement theory, appropriately reported using parametric statistics (which is an interesting if highly geeky debate with implications for the calculation of student GPAs as well), but rather: the use and misuse of SRTs (understood as a c-sat measure) as an element in making instructor personnel decisions, such as granting tenure.
Sources:
A Conceptual and Analytic Critique of Student Ratings of Teachers in the USA with Implications for Teacher Effectiveness and Student Learning, Teaching in Higher Education, Vol. 8, No. 2, 2003, pp. 233–245, ORLANDO J. OLIVARES
Assessing Students’ Ratings of Faculty, Assessment Update, September–October 1998, Volume 10, Number 5, Reza T. Ahmadi, Samuel E. Cotton
Thursday, May 3, 2007
Judging Teacher Effectiveness in 6 Seconds
A couple of highlights from a 1993 article describing the outcome of a series of experiments in which female students were shown brief, video only (no sound) clips of a university instructor in the classroom and then rated the instructors on a series of personality variables. The scores on these personality variables were compared to the ratings instructors received in the university's normal end of semester student evaluation program.
"Teachers who were rated higher by their students [on the end of semester evaluations] were judged to be significantly more optimistic, confident, dominant, active, enthusiastic, likable, warm, competent, and supportive on the basis of their nonverbal behavior."
A different group of female students also rated the physical attractiveness of the instructors based on still photographs. Correlations suggested that student evaluation ratings are somewhat influenced by physical attractiveness (r=.32), but that when controlling for the scores on the personality variables, the relationship between physical attractiveness and teacher evaluation ratings dropped to r=.14.
They performed the same kind of study of high school teachers, using principals' ratings of teacher effectiveness, and found similar results - except the correlation between physical attractiveness and ratings from the principal was negative. However, scores on the personality variables given by strangers watching a video clip still predicted the principals' ratings.
I was surprised that ratings on the personality variables were not more reliable when people watched 3 10-second clips than watched 3 2-second clips. People are able to make these kinds of judgments extremely quickly and with a great deal of consensus.
An examination of nonverbal behaviors showed that nodding and laughing are related to higher student evaluation scores while fidgeting with hands or an object is related to lower ratings.
(Source: Half a minute: Predicting teacher evaluations from thin slices of nonverbal behavior and physical attractiveness., By: Ambady, Nalini, Rosenthal, Robert, Journal of Personality and Social Psychology, 00223514, 19930301, Vol. 64, Issue 3)
Article linked from Marginal Revolution. Like Tyler, and as I have talked about before in the context of credibility, I think giving the appearance of confidence is very helpful in the public speaking arena.
"Teachers who were rated higher by their students [on the end of semester evaluations] were judged to be significantly more optimistic, confident, dominant, active, enthusiastic, likable, warm, competent, and supportive on the basis of their nonverbal behavior."
A different group of female students also rated the physical attractiveness of the instructors based on still photographs. Correlations suggested that student evaluation ratings are somewhat influenced by physical attractiveness (r=.32), but that when controlling for the scores on the personality variables, the relationship between physical attractiveness and teacher evaluation ratings dropped to r=.14.
They performed the same kind of study of high school teachers, using principals' ratings of teacher effectiveness, and found similar results - except the correlation between physical attractiveness and ratings from the principal was negative. However, scores on the personality variables given by strangers watching a video clip still predicted the principals' ratings.
I was surprised that ratings on the personality variables were not more reliable when people watched 3 10-second clips than watched 3 2-second clips. People are able to make these kinds of judgments extremely quickly and with a great deal of consensus.
An examination of nonverbal behaviors showed that nodding and laughing are related to higher student evaluation scores while fidgeting with hands or an object is related to lower ratings.
(Source: Half a minute: Predicting teacher evaluations from thin slices of nonverbal behavior and physical attractiveness., By: Ambady, Nalini, Rosenthal, Robert, Journal of Personality and Social Psychology, 00223514, 19930301, Vol. 64, Issue 3)
Article linked from Marginal Revolution. Like Tyler, and as I have talked about before in the context of credibility, I think giving the appearance of confidence is very helpful in the public speaking arena.
Wednesday, May 2, 2007
Orwellian Traffic Signs
As I've discussed before, there is a stretch of the southbound I-35 access road between the Slaughter and Slaughter Creek Overpass exits that has two strange merge signs that my fellow drivers do not understand. Or, I should say, had these signs. About a month ago, the confusing straight line/broken line signs that indicated that the left lane ends and the right lane continues were replaced by signs saying "Lane Ends Merge Left." Um, yes, that's correct - they changed it so that the right lane ends and the left lane continues (which both Robert and RB independently decided the road appeared to suggest by design, though it looks completely ambiguous to me). Once I got over boggling at the turn of events, I thought it was more logical - the right lane is usually bogged down by cars turning into the Wal-Mart shopping center or onto Old San Antonio Road, so it made more sense for the left lane to be the through lane. And traffic seemed to me to be slightly more sane along that part of the road after we all got used to the change and/or the rules of the road changed to conform to what we were already doing.
But then - in a bizarre switcharoo that had me talking in my car at loud volume "What?! Now they're just fucking with me!" - on Tuesday night, the signs had been changed again; they now read "Lane Ends Merge Right," thus re-establishing the original protocol, only with words instead of those pictograms that are equally incomprehensible in every language. Well, I say "now" but really, who knows? Since 5:15 this afternoon, they could have been changed to the semi-mysterious drawing showing that we should merge left or replaced by a series of Burma Shave-esque signs:
You must suffer
From aphasia
We were always at war
with Eurasia
Merge now
But then - in a bizarre switcharoo that had me talking in my car at loud volume "What?! Now they're just fucking with me!" - on Tuesday night, the signs had been changed again; they now read "Lane Ends Merge Right," thus re-establishing the original protocol, only with words instead of those pictograms that are equally incomprehensible in every language. Well, I say "now" but really, who knows? Since 5:15 this afternoon, they could have been changed to the semi-mysterious drawing showing that we should merge left or replaced by a series of Burma Shave-esque signs:
You must suffer
From aphasia
We were always at war
with Eurasia
Merge now
Subscribe to:
Posts (Atom)
