Chronicle of Higher Education, Teaching Evaluations Are Broken. Can They Be Fixed?:
Superficial assessments hurt professors and students, but reform is hard.
When Phoebe Young began working at the University of Colorado at Boulder as an assistant professor of history in 2009, her annual teaching reviews were fairly perfunctory. Everyone knew, she says, that student course evaluations were essentially popularity contests. But they were also the only measure used to determine merit raises. You got more money if you were above the average, and less money if you were below it. Professors would often strategize on how to increase their scores — say, bringing in donuts while students filled out the forms.
Several of the questions were problematic, she recalls, including a “super weird” one that asked students to rate the intellectual challenge of the course. The most impenetrable professors might score the highest because that’s how some students interpreted the question. Young scored low on that question, she believes, because she always tried to make her courses accessible.
When a faculty member came up for tenure, the process was marginally better. The department would scramble to find a colleague or two to sit in on a class. The letters they wrote were all roughly the same, she recalls. A long exegesis on the content of the class. A summary of the person’s teaching style that might boil down to, “They’re great!” (Nobody wanted to get a colleague in trouble.) And a comment or two along the lines of, “Maybe reconsider the font on those PowerPoint presentations.”
“Everybody’s like, ‘Well, whatever. You can’t really measure teaching anyway,’” Young says of the fatalistic view she shared with her colleagues. “And so we just do what’s required by the system.”
Today Young is a full professor, and her early classroom experience is not unusual. A core part of a professor’s job — and arguably the most central role of higher education — is teaching. Yet no matter how challenging the subject, how invested the professor, or how varied the students, teaching capability is often reduced to a number on a point scale. Congratulations: You are a 6.3 out of 7. Or, try harder, you’re a point below average. The additions of peer review and a written reflection on your own teaching may appear to give the process more depth, but on many campuses professors say that these tools still come with little guidance or forethought.
Some colleges and universities, including Boulder, are trying to change that equation. They are investing in more thoughtfully designed course evaluations, preparing faculty members to substantively critique their colleagues, and fostering discussions of teaching in departmental meetings, which is often where change happens. There is even a national effort underway, led by several higher-education groups and research universities, to transform the way teaching is evaluated.
But the movement is still young, and the work, according to many reformers, is difficult. They face the familiar obstacles of entrenched norms, disagreement about what it means to be a good teacher, and limited time. This methodical, collaborative approach is also a relatively new way of looking at teaching, which has traditionally been considered the purview of the faculty member, particularly at research universities, where stellar teaching has often operated in the shadow of high-profile research. …
[T]he ways in which many departments and colleges across the country assess teaching skills remain more ad hoc than deliberative, more superficial than substantive. Considering that most instructors enter their first classroom having received little guidance in graduate school on how to teach, getting minimal — and often useless — feedback only compounds the problem. …
A system that fails to evaluate teaching effectively, reformers say, shortchanges students. A growing body of research shows that effective teaching is hugely influential in determining whether students succeed in college, and that it is a key lever in helping support students who may have come into college with fewer educational advantages than their classmates. A slapdash evaluation of teaching, in other words, undermines higher education’s ability to deliver on its promise.
It wasn’t that long ago that teaching was seen as more of an art than a science. Examining one’s own teaching or the teaching strategies of others seemed to serve little purpose. The excellent teacher was the charismatic or engaging professor. Questions on course evaluations still reflect that. Two common, and highly subjective, questions ask students to rate their course and rate their instructor over all. Sometimes answers to those two questions have been the only ones used to determine merit raises.
A great deal of scholarship challenges this narrative and offers alternative scripts. The effective instructor, teaching experts tell us, is one who is well organized, whose course material and class discussions align with what students are tested on, who sparks students’ curiosity and fosters their confidence as learners, who is willing to adapt as the student body changes, and who stays on top of the latest teaching innovations.
Course evaluations, too, have come under the microscope: Dozens of studies have shown they are subject to racial and gender bias. A meta-analysis found little to no correlation between how highly students rate their instructor and how well they have learned the subject. In 2019 more than a dozen scholarly organizations endorsed a statement that describes the current use of student evaluations as “problematic” and recommends a holistic approach to merit and promotion reviews.
So why haven’t evaluations of teaching kept pace with these developments? Weaver says it’s a chicken and egg problem. Professors don’t want to put in the work of developing and undergoing a more rigorous review of teaching until that work is more highly valued. But many administrators won’t understand that good teaching is a complex endeavor worthy of more attention until the old measurements are scrapped and replaced with something better.
Another dilemma, particularly at research-intensive universities, is that administrators still tout scholarly contributions in ways that make them appear more valuable than teaching, such as announcing how much money their researchers attract. …
Since 2017, Transforming Higher Education — Multidimensional Evaluation of Teaching, or TEval, has involved hundreds of faculty members and administrators at Boulder, UMass-Amherst, and the University of Kansas in the work of creating new evaluation processes on their campuses. It is funded by the National Science Foundation and has the backing of the Association of American Universities and other higher-education organizations. The goal of the project is to advance the use of evidence-based teaching methods by changing the way teaching is evaluated. …
Conditions have made it more likely that colleges will consider scrapping their old evaluation systems in favor of a process that is more thoughtful, coherent, and based in research. The question now is whether that will be enough to propel them through the skepticism and uncertainty that has kept a flawed system in place this long.



