Showing posts with label teacher evaluation. Show all posts
Showing posts with label teacher evaluation. Show all posts

Wednesday, March 2, 2011

Last In, First Out by Matt Bromme


Matt Bromme is a former NYC teacher, assistant principal, principal, district superintendent and high level official at Tweed. He has seen this issue from many angles, and his words are to be heeded:

Much has been stated and argued over the policy used to lay-off teachers by seniority (Last In First Out). Many new teachers to the school systems across America are enthusiastic and willing to bring new methodologies to the classroom. However, many senior teachers are extremely competent and their body of work makes them deserving of our respect, not deserving of being tainted as incompetent and unsatisfactory.

In 1975 I was terminated as a New York City teacher during that fiscal crisis. I had three years of experience as a teacher, and I was extremely upset about my situation. Were there senior teachers who possibly could have been let go if there was an objective criteria? Someone please define what an objective criteria is? Were some of these teachers not terminated deserving of an unsatisfactory rating? The answer is yes. That was management’s responsibility to rid the system of anyone who was incompetent.

However, I supported then and I support now the concept of a due process procedure to ensure that all staff have their rights protected and that good teachers have their seniority rights protected.

Just like senior staff in the police department, fire department and military, have a strong knowledge base, so do senior teachers and school leaders. There is a need for a professional to develop a reputation in the community that they serve. This takes time. Time is not an idea embraced by the new corporate mentality swarming over our schools.

In 1976-1977 I returned to the classroom. In 1984 I became a middle school assistant principal, in 1988 I became an elementary school principal and in 1991 I became a middle school principal. My last position was at the rank of school district superintendent in New York City, with the last six months of my career assigned to the Tweed Courthouse.

The Tweed Court House environment was fascinating from my point of view as a civil servant. It seemed that everyone hired by the Department of Education came from an Ivy League College, had not yet reached their twenty fifth birthday and looked at all of us “old” educators as failures because we stayed in our position for more than five years. What really disturbed me was that these young perky preppy staff members had no clue as to what needed to be done in the communities we served.

Fast forward to today’s argument that newer teachers are automatically better then senior teachers. This corporate mentality (Bloomberg, Black, Gates, etc.), has taken over from the the philosophy that teaching is an art and not a science. In their desire to use data (which too many times is incorrect), they are missing the point that teachers are like ministers preaching to a very difficult group and trying to convert them to accept a better life. Before the system went data crazy, in many of our most challenging schools, we used music, art, and drama to motivate our students to do better. Today, too many schools have had to give up their assembly programs and art programs because there is not enough time, since a child’s educational experience today is dominated by test prep.

The corporate mentality also does not understand that it takes time for school teachers and school leaders to develop and to gain the respect of the community they serve. In my experience, especially at the middle school and high school level, it takes at least three years for a reputation to develop and ergo for the educator to be respected. It is also my experience that for competent teachers and principals each year they improve in their skills, just as police fire and military personnel do. As educators mature they learn numerous different techniques to enhance their classroom performance.

Tenure has become an obscene word among politicians and the new educational corporate mentality. Tenure at the public school level (as opposed to the world of the university) only guarantees a due process procedure for those accused of egregious behavior. I rated teachers and school leaders unsatisfactory. I had numerous grievances on all levels, including arbitration hearings and court cases.

Where I and my staff did our homework, we won our cases. In some cases, staff was terminated, fined or chose to retire. Where supervisors, either at the school or district level, failed to meet their contractual obligations, we did not win. I would not have it any other way.

Due process protects teachers who speak out for their students. Due process protects school leaders who administer their buildings and often are subject to political pressures. It also protects school leaders who are brave enough to challenge their central district corporate leaders who know nothing about schools and the community the schools serve.

While Mayor Bloomberg and his “people” should have been focused on class size issues to enhance school performance, he caused this issue by creating numerous schools within schools. Many of them will be overcrowded when they reach their full maturity.

Mayor Bloomberg created this issue by allowing principals to automatically refuse to hire teachers of the schools being replaced by the new small schools. Mayor Bloomberg also created these problems by readjusting the budget process, so that “average” salary was replaced at the school level with “exact” salary. Therefore this motivated new principals not to hire from the ranks of teachers that were available, but to go out and hire “inexpensive” teachers.

There is a crisis today regarding LIFO that was not caused by unions and or senior staff members. If you look deep enough into the corporate mentality,you will find that this is more of a budget issue than an educational issue. If the principals had hired the senior staff that was assigned to ATR status, none of their schools would be looking at draconian cuts. However, under a false sense of security the Mayor rolled the dice and decided to ignore these career ATR teachers, many of them competent and capable. Instead they went for the young and the restless, most of whom will leave the system within five years.

-- Matt

Tuesday, December 28, 2010

Fact-checking "Waiting for Superman": False data and fraudulent claims

In response to critical comments, I have added clarifications and corrections in bold italics to my original post. I apologize for some sloppy math in not annualizing what was a six year rate in the apparent source article for the film. After further analysis, the movie’s claims remain clearly inaccurate as well as misleading; though the source article was only partially erroneous, at least as far as I can tell. Sorry for the mistake -- and thanks to all my readers, and especially those of you who checked my figures so assiduously.

In the movie Waiting for Superman, nominated for an Oscar as the best Documentary of 2010, the following statement is made:

" ...in Illinois, 1 in 57 doctors loses his or her medical license, and 1 in 97 attorneys loses his or her law license, but only 1 teacher in 2500 has ever lost his or her credentials."

Since the movie was released, these figures have been repeated frequently. They take up five pages in the Google search engine, were cited in the NY Times review of the film, the British newspaper the Independent, as well as Brian Williams of NBC in the television program Education Nation.

But apparently not a single one of these news outlets, or the makers of Waiting for Superman, ever bothered to check them.

While looking for the source of this claim, which is repeated without citation in the movie and its companion book, I came upon a 2007 newspaper article by Scott Reeder of the Small Newspaper Group:

During the past six years, 1 in 2,500 Illinois educators have lost their teaching credentials through suspension, revocation or surrender. By comparison, during the same period 1 in 57 doctors practicing in Illinois lost their medical licenses and 1 in 97 Illinois attorneys lost their law licenses.

"Either Illinois teachers are 43 times better behaved than doctors or they are being held to a considerably lower professional standard than other professions,'' said Jeff Mays, executive director of the Illinois Business Roundtable and an advocate for educator accountability standards. ``Just like doctors and lawyers, teachers are members of an important and demanding profession. It's time that they be held to the same professional standards."

One should note that the data cited in the source article is substantially different from the claim made in the film. In the movie, the period of six years is omitted for the disbarment of physicians and/or attorneys– making indefinite the time span over which the data was collected. The film also says that only 1 in 2500 Illinois teachers have “ever” lost his or her credentials, rather than over six years.

In an effort to verify these claims, I first consulted the annual summary put out by the Federation of State Medical Boards. In reality, 121 doctors lost their licenses in Illinois in 2009, out of 43,670 physicians. That means an average of 0.3% of doctors per year lost their licenses; or 3 out 1,000 per year. Over six years, this would equal 1.8% -- substantially the same as the 1 in 57 figure cited in the source material.

I also checked the claim that 1 in 97 attorneys in Illinois lose their licenses over six years. According to data reported by the American Bar Association, 26 lawyers in Illinois were disbarred in 2009, out of a total of 58,457 - in some cases, by mutual consent.

Since 2001, the average rate of Illinois attorneys disbarred is 32 per year – with more than half of them leaving their professions “voluntarily.” This is an annual rate of about 0.05%, for a six year rate of 0 .3% -- 3 out of 1,000 – not one out of 97, as the source material claimed. As mentioned above, the movie did not specify the time frame over which this disbarment is supposed to have occurred.

The total number of lawyers disbarred in the entire country, either involuntarily or by mutual consent, is 800 per year out of 1,180,386; which is about 0.07% per year, or 7 out of 10,000. The number of those involuntarily disbarred is 441- about 0 .04% or 4 out of 10,000 per year. The six year rate for disbarment nationally would be 0.42% -- about ten times the figure cited in the film of one in 2500 Illinois teachers who “ever” lost their credentials.

I could not find any independent data verifying the number of Illinois of teachers who lose their credentials each year. According to the NY Daily News, over the past three years, 88 out of about 80,000 New York City schoolteachers have lost their jobs for "poor performance." This represents an annual rate of about 30 out of 80,000, or 0.03%, which is about the same rate as attorneys who are involuntarily disbarred each year nationally.

According to the Houston Chronicle, over the last five years, 364 Houston teachers have been fired, out of about 12,000: "Of those, 140 were ousted for performance reasons, a broad category that generally covers teachers not fulfilling their job duties."

So the rate of Houston teachers who lose their jobs due to poor performance is about 0.2% per year - higher than the rate of either doctors or attorneys in the state of Texas removed from their profession annually. For example, only 32 Texan attorneys were disbarred in 2009 out of 75,087; for an annual rate of 0.04% -- at one fifth the rate. 64 doctors per year on average lost their licenses in Texas between 2005 and 2009; out of about 60,000 physicians, at an annual rate of about 0.1 % -- about half the percentage.

Moreover, many more teachers who are untenured and/or uncertified are removed from their jobs for poor performance. Roughly 3.7% of New York City teachers were denied tenure this year, according to the NY Times.

The overall attrition rate of teachers is much higher - many of whom would probably otherwise be cited for poor performance, but who leave the profession either willingly, or "counseled" out. In New York City, the four year attrition rate is more than 40% -- a mind-boggling figure.

In reality, one of the most serious problems plaguing our urban schools, along with excessive class sizes, overcrowding, and poor support for teachers and students, is the fact that we have far too many inexperienced educators revolving through our high-needs schools each year.

Can you imagine if 40% of physicians or attorneys left their jobs after four years? A national emergency would be declared, with a commission appointed to find out how their working conditions could be improved.

Yet instead of examining this critical issue objectively, the movie Waiting for Superman cites false statistics in their effort to scapegoat teachers, unfairly blaming them for all the failures of our urban schools. The film features the views of Eric Hanushek of the Hoover Institute, a well-known conservative critic of equitable educational funding, claiming that the best way to improve our schools would be to fire 5-10% of teachers each year.

To the contrary, eliminating teacher tenure and seniority protections would likely produce an even less experienced and less effective teaching force - especially in our urban public schools, which already suffer from excessively high rates of turnover.

As a parent, I support a higher standard for teacher tenure and more rigorous teacher evaluation systems. I have seen my own children benefit from excellent teachers over the years, but also occasionally suffer as a result of poor teaching, though the latter has occurred as often in schools without union protections as those that were unionized. An improved evaluation system would take into account not only test score data, but also feedback from other teachers, administrators, students and parents.

But at this point, we simply cannot trust the corporate oligarchy currently making policies for our schools to create a fair evaluation system, including those who backed Waiting for Superman, given their proclivity to misuse and distort data, as shown by the egregiously inaccurate figures cited in the film.

Rather than a documentary, perhaps the movie should be re-categorized, with an appropriate disclaimer, as an urban myth.

Friday, December 3, 2010

Silver Linings

It's probably no surprise to regular readers of this blog that I'm a Democrat. That said, in my professional life, I have worked for non-partisan and non-profit organizations committed to working with public officials of all political persuasions. And I think that both political parties -- as well as political independents -- are potential partners in improving pubic education. Nonetheless, I can't honestly say that November 2010 was an uplifting month from my (electoral) perspective.

But this is the Education OPTIMISTS blog, right? I guess I'm a little self-aware in that I recognize that a majority of my posts grouse about something or other and too few are offered in a truly optimistic vein. So here is my attempt at weaving a silk purse from an elephant's(?) ear.

Newly elected Republican governors and state legislators have opportunities to improve education in ways that some of their Democratic counterparts may not. They may be more politically willing and able to take on certain vestiges of the education status quo that may not be research-based, may not be working, and may be not be in service of student outcomes. I'm talking about things like the traditional steps-and-lanes teacher salary schedule, the length of the school day and year, and largely purposeless teacher evaluation systems. They may be able to construct new human capital systems and reform outdated school practices and processes. Certainly, these opportunities will vary depending upon numerous contextual factors, including state systems of educational governance, the intellectual and policy foundations for such work, the existence of non-governmental advocates and thought partners, existing vehicles for collaboration (such as p-16 councils), etc.

A number of past Republican governors emerged as leaders, or so-called "education governors." Some that immediately come to mind are former Massachusetts Gov. Bill Weld, former Ohio Gov. Bob Taft, outgoing Alabama Gov. Bob Riley, and former Tennessee Gov. (and current U.S. Senator) Lamar Alexander (who also served as U.S Education Secretary under President George H.W. Bush).

That said, here's where my pessimism takes over. There are certainly members of the current Republican gubernatorial circle that are certainly not in the running for such an earned label, such as New Jersey Governor Chris Christie. With less than a year in office, he has moved in a decidedly confrontational direction. His actions and rhetoric makes him better suited for talk radio than as a collaborative education governor who needs to work with a Democratic-controlled state legislature and, yes, with teachers. Christie is casting himself as a modern-day Archie Bunker and is relishing the attention and headlines he is getting. The sad fact is that such rhetoric is what garners attention in today's media rather than the dogged policy work that actually changes the equation for students and teachers.

Christie seems to think that leadership consists of rhetoric rather than results. And that's dangerous if our collective goal is the creation of meaningful reform as opposed to simply talk or threats of it. Christie has attacked his own state's relatively good educational performance in order to further his demonizing of the state teachers' union, which appears to be his primary priority. He fired his first education commissioner who dared to strike a compromise with teachers around New Jersey's aborted Race to the Top application. That former commissioner said that the Governor "placed fighting with the state teachers unions and his persona on talk radio above education reform." Good luck to those who seek to thrust him upon America as a 2012 candidate for President.

My fear is that there those within the ranks of new Republican governors that are more likely to fashion themselves in the style of Christie than as "old school" education governors. Texas Governor Rick Perry's ascension to head the RGA probably doesn't help either.

With the exception of the 12 Race to the Top winners, one challenge that these new office holders of both parties will face is the distinct lack of new resources to inject into the educational system, either from state or federal sources. They won't be able to buy reforms by increasing education funding or refashion teacher pay with a major infusion of cash.

My hope is that governors of both parties will seek to work with educators to accomplish their policy objectives rather than force their desires onto them. There is good evidence to suggest that collaboration leads to more resilient and relevant policies that are likely to trickle down to actually change practices and processes within classrooms and schools. That's where the real work gets done.

I am hopeful that some quiet leaders will emerge.

Tuesday, August 31, 2010

Adding Value to the Value-Added Debate

Seeing as I am not paid to blog as part of my daily job, it's basically impossible for me to be even close to first out of the box on the issues of the day. Add to that being a parent of two small children (my most important job – right up there with being a husband) and that only adds to my sometimes frustration of not being able to weigh in on some of these issues quickly.

That said, here is my attempt to distill some key points and share my opinions -- add value, if you will -- to the debate that is raging as a result of the Los Angeles Times's decision to publish the value-added scores of individual teachers in the L.A. Unified School District.

First of all, let me address the issue at hand. I believe that the LA Times's decision to publish the value-added scores was irresponsible. Given what we know about the unreliability and variability in such scores and the likelihood that consumers of said scores will use them at face value without fully understanding all of the caveats, this was a dish that should have been sent back to the kitchen.

Although the LA Times is not a government or public entity, it does operate in the public sphere. And it has a responsibility as such an actor. Its decision to label LA teachers as 'effective' and 'ineffective' based on suspect value-added data alone is akin to an auditor secretly investigating a firm or agency without an engagement letter and publishing findings that may or may not hold water.

Frankly, I don't care what positive benefits this decision by the LA Times might have engendered.
Yes, the district and the teachers union have agreed to begin negotiations on a new evaluation system. Top district officials have said they want at least 30% of a teacher's review to be based on value-added and have wisely said that the majority of the evaluations should depend on classroom observations. Such a development exonerates the LA Times, as some have argued. In my mind, any such benefits are purloined and come at the expense of sticking it -- rightly in some cases, certainly wrongly in others -- to individual teachers who mostly are trying their best.

Oh, I know, I know. It's not about the teachers anymore. Their day has come and gone. "It's about the kids" now, right? But you know what? The decisions we make about how we license, compensate, evaluate and dismiss teachers affects them as individual people, as husbands and wives, as mothers and fathers. It effects who may or may not choose to enter the profession in the coming years. If we mistakenly catch a bunch of teachers in a wrong-headed, value-added dragnet based upon a missionary zeal and 'head in the sand' conviction that numbers don't lie, we will be doing a disservice both to teachers and to the kids. And, if we start slicing and dicing teachers left and right, who exactly will replace them?

(1) Value-added test scores should not be used as the primary means of informing high-stakes decisions, such as tenure and dismissal.
One primary piece of evidence was released just this week from the well-respected, nonpartisan
Economic Policy Institute. The EPI report, co-authored by numerous academic experts, said:

  • Student test scores are not reliable indicators of teacher effectiveness, even with the addition of value-added modeling (VAM).
  • Though VAM methods have allowed for more sophisticated comparisons of teachers than were possible in the past, they are still inaccurate, so test scores should not dominate the information used by school officials in making high-stakes decisions about the evaluation, discipline and compensation of teachers.
  • Neither parents nor anyone else should believe that the Los Angeles Times analysis actually identifies which teachers are effective or ineffective in teaching children because the methods are incapable of doing so fairly and accurately.
  • Analyses of VAM results show that they are often unstable across time, classes and tests; thus, test scores, even with the addition of VAM, are not accurate indicators of teacher effectiveness. Student test scores, even with VAM, cannot fully account for the wide range of factors that influence student learning, particularly the backgrounds of students, school supports and the effects of summer learning loss. As a result, teachers who teach students with the greatest educational needs appear to be less effective than they are.
Other experts, such as Mathematica Policy Research, Rick Hess, and Dan Goldhaber have offered important cautions as well.

The findings of the IES-funded Mathematica report were “largely driven by findings from the literature and new analyses that more than 90 percent of the variation in student gain scores is due to the variation in student-level factors that are not under the control of the teacher. Thus, multiple years of performance data are required to reliably detect a teacher’s true long-run performance signal from the student-level noise…. Type I and II error rates [‘false positives’ and ‘false negatives’] for teacher-level analyses will be about 26 percent if three years of data are used for estimation.
In a typical performance measurement system, more than 1 in 4 teachers who are truly average in performance will be erroneously identified for special treatment, and more than 1 in 4 teachers who differ from average performance by 3 months of student learning in math or 4 more in reading will be overlooked. In addition, Type I and II error rates will likely decrease by only about one half (from 26 to 12 percent) using 10 years of data.”

Hess has “three serious problems with what the LAT did. First … I'm increasingly nervous at how casually reading and math value-added calculations are being treated as de facto determinants of "good" teaching…. Second, beyond these kinds of technical considerations, there are structural problems. For instance, in those cases where students receive substantial pull-out instruction or work with a designated reading instructor, LAT-style value-added calculations are going to conflate the impact of the teacher and this other instruction…. Third, there's a profound failure to recognize the difference between responsible management and public transparency.”

Goldhaber, in a Seattle Times op-ed, says that he “support[s] the idea of using value-added methods as one means of judging teacher performance, but strongly oppose[s] making the performance estimates of individual teachers public in this way. First, there are reasons to be concerned that individual value-added estimates may be misleading indicators of true teacher performance. Second, performance estimates that look different from one another on paper may not truly be distinct in a statistically significant sense. Finally, and perhaps most important, I cannot think of a profession in either the public or private sector where individual employee performance estimates are made public in a newspaper.”

Multiple measures to inform teacher evaluation seems like the right approach, including the use of multiple years of value-added student data (one thing the LA Times DID get right). That said, the available research would seem to suggest that states (particularly in Race to the Top) that have proposed basing 50% or more of an individual educators evaluation on a value-added score may have gone too far down the path. LA Unified officials have said (LA Times, 8/30/2010) they want at least 30% of a teacher's review to be based on value-added and that the majority of the evaluations should depend on observations. That might be a more appropriate stance.

(2) Embracing the status quo is unacceptable.
As reports such at The New Teacher Project's
Widget Effect have chronicled, current approaches to teacher evaluation are broken. They don’t work for anyone involved. Critics of VAM cannot simply draw a line in the sand and state that, "This will not stand!" If not this, then what? Certainly not the current system! Fortunately, efforts led by organizations such as the American Federation of Teachers and the Hope Street Group are developing or have offered thoughtful solutions to this issue. [Disclosure: I participated in Hope Street's effort and my New Teacher Center colleague Eric Hirsch serve on AFT’s evaluation committee.] Sadly, LA Unified and the LA Teachers Union both are culpable –along with the LA Times – in bringing this upon the city's teachers by refusing to act to analyze or utilize available value-added data. An adherence to the status quo created a void that the LA Times sought to fill in order to sell more newspapers in a wrong-headed attempt to inform the public.

(3) The ‘lesser of two evils’ axiom should not be invoked.
Even if you agree that all the factors we currently use to select and sort teachers is worse than a value added only alternative,
as argued by Education Sector's Chad Aldeman, our current arsenal does not meaningfully inform high-stakes decisions (apart from entry tests with largely low passing scores and the aforementioned impossible-to-fail evaluations). That's, of course, both a condemnation of the current system's inability and/or unwillingness to differentiate between teachers, but it's also a recognition that we haven't struck the right balance or developed the value-added systems to inform high-stakes decisions in this regard in all but a few promising places.

(4) Don't lose sight of the utility of value-added data to inform formative assessment of teaching practice.
If one of the takeaways from research is that value-added data shouldn't be used to drive high-stakes decisions, it is helpful to think about the use of this data to inform teacher development. Analysis of student work, including relevant test scores, is an important professional development opportunity that all teachers, especially new ones, should have regular opportunities to engage in. Systems such as the NTC’s
Formative Assessment System provide such a tool in states and districts with whom it works on teacher induction. Sadly, this is not the norm in American schools, but is built into high-quality professional development approaches, as Sara Mead wisely discusses in her recent Ed Week blog post. As I noted under #2, LA Unified missed an opportunity to embrace such data to inform its educators in such a way. In the LA Times value added series, several teachers bemoaned the fact that they had never had the opportunity to see such data until it was published in the newspaper.

(5) Valid and reliable classroom observation conducted by trained evaluators is critical.
Other elements of an evaluation system are even more important than value-added methodology if for no other reason that the majority of teachers do not teach tested subjects. Unless we, God forbid, develop multiple-choice assessments of more and more subjects and grade levels, we're going to need valid and reliable ways of assessing the practice of educators who cannot be assessed by value-added student achievement scores. Despite some of the criticisms lobbed at the District of Columbia's new
IMPACT evaluation system, this is an element at the heart of DC’s approach to teacher evaluation. Further, the Gates Foundation’s on-going teacher effectiveness study holds great promise.

(6) We've got to get beyond this focus on the 'best' and 'worst' teachers.
How about we focus on strengthening the effectiveness of the 80-90% of teachers in the middle? We know how to do that through
comprehensive new teacher induction and high-quality professional development, but we're just lacking the collective will to pull it off and invest in what makes a difference. These are similar roadblocks to what has prevented the use of student outcomes from being considered in teacher evaluations. It raises discomfort, requires a change in prevailing (often mediocre) practices, demands greater accountability, and necessitates viewing teaching not as a private activity but as a collective endeavor. But I keep making this point over and over again about the importance of a teacher development focus within the teacher effectiveness conversation because I see too few reform advocates taking it seriously. Take off the blinders, folks. It is not primarily about firing teachers.

(7) Teacher effectiveness is contextual.
Teaching and learning conditions impact an individual educator’s ability to succeed. It is entirely possible that an individual teacher's value-added score is significantly determined by the teaching and learning conditions (supportive leadership, opportunities to collaborate, classroom resources) present at their school site than about their individual knowledge, skills and practices. In Seinfeldian terms, teachers are not 'masters of their domain' necessarily. The EPI report makes this point. So do my New Teacher Center colleagues through statewide teaching and learning conditions surveys. So does Duke University economist Helen Ladd (also a co-signed on the EPI report) and the University of Toronto’s Kenneth Leithwood.

Saturday, August 28, 2010

Problems with the use of Student Test Scores to Evaluate Teachers

originally posted at Daily Kos

If new laws or policies specifically require that teachers be fired if their students’ test scores do not rise by a certain amount, then more teachers might well be terminated than is now the case. But there is not strong evidence to indicate either that the departing teachers would actually be the weakest teachers, or that the departing teachers would be replaced by more effective ones. There is also little or no evidence for the claim that teachers will be more motivated to improve student learning if teachers are evaluated or monetarily rewarded for student test score gains.


That is a quote from the Executive Summary of one of the most important policy briefs about education in recent years. At a time when the Dept. of Education is pushing to tie teacher evaluation and compensation to student test scores, this Economic Policy Institute Briefing Paper (whose title is the same as this diary, and which is a pdf), pulls together the extensive relevant research that demonstrates the dangers of pursuing such a path. Please continue reading as I explore this important document, released at 12:01 AM today, August 29.

First, let me clarify several things.

This is a very long diary. That is because I am trying to reasonably thoroughly cover the contents of an extremely important document. My purpose in doing so is to convince people of the document's importance. Thus I will be perfectly happy should you decide you do not need to further read what I have written below. You can follow the link for the brief (which I have provided you again), download the pdf, and begin reading. The executive summary is only four pages. The brief itself, without the critical apparatus of footnotes and sources, another 17. So if you want, one more time follow this link.


This document has been in the works for several months, and was NOT hurriedly put together as a response to the recent series by the Los Angeles Times which used value-added assessment to label teachers in the Los Angeles Unified School District. Second, the ten scholars whose names are on the document are some of the most eminent in educational circles, including among their midst former Presidents of the American Educational Research Association and the National Council on Measurement in Education, two of the three professional organizations most involved with psychological measurement, of which school-related testing is a subset. One of the scholars, Robert Linn, has not only presided over both of those organizations, he has also serve as chair of the National Research Council's Board on Testing and Assessment. The group also includes the immediate past president of the National Academy of Education, Lorrie Shepard, Dean of the School of Education at Colorado. A brief and applicable curricula vitae of each of the ten authors can be found at the end of the document, and briefer descriptions at the beginning, where each author is listed, along with the following statement:
Authors, each of whom is responsible for this brief as a whole, are listed alphabetically.
An email address is provided for further contact.

The ten authors, alphabetically, are as follows:
Eva L. Baker
Paul E. Barton
Linda Darling-Hammond
Edward Haertel
Helen F. Ladd
Robert E. Linn
Diane Ravitch
Richard Rothstein
Richard J. Shavelson
Lorrie A. Shepard

Let me be blunt. I do not know how anyone who knows the work of these scholars and who reads this brief can accept the idea of placing any stakes as to firing or awarding of merit pay based on the current status of Value-Added Assessment methodologies. The document is thorough. It reviews all the relevant studies, including one not yet in print. Those includes studies by Mathematica for the US Department of Education: by Rand: by the Educational Testing Service; done for the National Center for Education Statistics of the Institute of Education Sciences of the U. S. Dept. of Education; issued by the Board of Testing and Assessment of the Division of Behavioral and Social Sciences and Education of the National Academy of Sciences, and so on. There are citations from books, from peer reviewed journals.

I am not a scholar. I am a high school social studies teacher. During now abandoned doctoral studies in educational policy I got interested in value-added assessment and devoured what studies there were in the educational literature. I also talked extensively with the technical person for one organization that offered a value-added methodology who cautioned me that the approach was not stable enough for it to be used as the basis for decisions with any kind of meaningful stakes. That was about a decade ago. What I had read since, and what I have absorbed from this study convinces me that the situation is not significantly better now.

But you do not have to take my word for it. Let me offer a few key examples from the study. Those who follow me on Daily Kos already have seen in the study by Mathematica the high rate of error in determining superior and inferior teachers beyond the broad middle. In this diary, written on August 27, I noted that the error rate with 2 years of data was 36%, with 3 years 26%, and even with 10 years of data still 12%.

But that is just the tip of the iceberg of the technical problems with using such an approach.

Without recapitulating the entire brief, let me offer a couple of other key points.

1. Results for individual teachers are not stable:
One study found that across five large urban districts, among teachers who were ranked in the top 20% of effectiveness in the first year, fewer than a third were in that top group the next year, and another third moved all the way down to the bottom 40%. Another found that teachers’ effectiveness ratings in one year could only predict from 4% to 16% of the variation in such ratings in the following year.


2. One key question is whether one is really accounting for teacher effects and excluding other influences in the results one gets from value-added assessment. Jesse Rothstein reported something interesting, about which I quote from the Executive Summary:
A study designed to test this question used VAM methods to assign effects to teachers after controlling for other factors, but applied the model backwards to see if credible results were obtained. Surprisingly, it found that students’ fifth grade teachers were good predictors of their fourth grade test scores. Inasmuch as a student’s later fifth grade teacher cannot possibly have influenced that student’s fourth grade performance, this curious result can only mean that VAM results are based on factors other than teachers’ actual effectiveness.


3. The brief notes that arguments that the private sector evaluates professional employees using quantitative measures that are parallel. The authors of the brief point out that rarely are such quantitative measures the sole or even the primary factor, noting that management experts warning against using such measures for making salary or bonus decisions. They remind us that some of the distortion on Wall Street was the result of emphasizing short term gains that could be easily measured. They also touch on medicine:
In both the United States and Great Britain, governments have attempted to rank cardiac surgeons by their patients’ survival rates, only to find that they had created incentives for surgeons to turn away the sickest patients.


4. Students are not randomly assigned to teachers. While some control for school effects is possible, scholars are reluctant to place any weight on comparisons for teachers in different schools even within the same system. And even within a school, teachers may have varying numbers of students who are learning English or have learning disabilities or are homeless or who move multiple times, each of which is a factor that can affect learning.

5. Sample sizes are often too small. Even if the class makeup stays stable during the year, and all the students show up regularly, the N=30 of a large elementary class is too small a sample to provide a result that can allow strong inferences to be drawn. Often the makeup of the class changes during the year. If you exclude students who were not there all year, or whose absences exceed some designated level, the N decreases, providing a result of even less reliability.

6. Some argue that statewide data banks can address the question of student mobility. But if you derive results on a year or two years of data where the student has moved, how much of the improvement can properly be assigned to any one teacher? Even in elementary school, do we account for pull-out instruction, or possible tutoring (that could in some cases be counterproductive) as a possible influence on the test results upon which we base our analysis?

7. Even with value-added analysis, to date scholars have not been able to isolate the impact of outside learning experiences, home and school supports, and differences in student characteristics and starting points when trying to measure their growth.

8. A proper system of value-added assessment would have vertically scaled tests. Most states do not currently have such tests, for example, neither New York nor California does. That is, the tests in one grade are not necessarily congruent with those of the next along a continuum from year to year - we are not testing the same thing each year. As testing expert Dan Koretz of Harvard is quoted as noting,
"because of the need for vertically scaled tests, value-added systems may be even more incomplete than some status or cohort-to-cohort systems"
Here it is worth noting that cohort to cohort is comparing this year's fourth graders to last years, which is how Adequate Yearly Progress under No Child Left Behind has been calculated.

9. If measuring end of year to end of year, even if there are vertically scaled tests, there is still the well-documented issue of summer learning loss, which falls disproportionally upon those of lesser economic means, which also means it falls disproportionally upon those of color, who are more heavily represented at the lower end of the economic scale. IF we do not control for summer learning loss, our results are skewed. Allow me to quote a relevant portion of the study:
researchers have found that three-fourths of schools identified as being in the bottom 20% of all schools, based on the scores of students during the school year, would not be so identified if differences in learning outside of school were taken into account. Similar conclusions apply to the bottom 5% of all schools.
The authors also cite a study that shows "two-thirds of the difference between the ninth grade test scores of high and low socioeconomic status students can be traced to summer learning differences over the elementary years."

There is more, but this should give a real sense of how much there is in this paper, how thoroughly the authors examine relevant material to demonstrate that value-added assessment, the supposed magic bullet to allow us to tie student learning back to the effectiveness of teachers, cannot properly fulfill the task some wish to give to it.

The authors acknowledge that value-added approaches are superior to some of the alternatives methods of using test scores to evaluate teachers. These are

status test-score comparisons - compare average scores of students of one teacher to those of another

over change measures - compare the average test results of a single teacher from one year to the next - remember, these are different students

over growth measures - a comparison of the scores of the students of the teacher this year to the scores of those same students the previous year when they had different teachers.

Each of these approaches has serious problems with it. One can read the detailed explanation on p. 9. Value-added assessments may be an improvement, but
the claim that they can “level the playing field” and provide reliable, valid, and fair comparisons of individual teachers is overstated. Even when student demographic characteristics are taken into account, the value-added measures are too unstable (i.e., vary widely) across time, across the classes that teachers teach, and across tests that are used to evaluate instruction, to be used for the high-stakes purposes of evaluating teachers.



Let me offer a few of the quotes about value-added assessment that the authors of the brief offer from scholars who have examined the approach over the years, and then I will offer a few observations of my own.

in 2003, a research team at Rand concluded
The research base is currently insufficient to support the use of VAM for high-stakes decisions about individual teachers or schools.


In 2004, Donald Rubin opined
We do not think that their analyses are estimating causal quantities, except under extreme and unrealistic assumptions.


Henry Braun, then at ETS, offered this in 2005:
VAM results should not serve as the sole or principal basis for making consequential decisions about teachers. There are many pitfalls to making causal attributions of teacher effectiveness on the basis of the kinds of data available from typical school districts. We still lack sufficient understanding of how seriously the different technical problems threaten the validity of such interpretations.


Last year the Board on Testing and Assessment of the National Research Council of the National Academy of Sciences wrote to the Department of Education saying
...VAM estimates of teacher effectiveness should not be used to make operational decisions because such estimates are far too unstable to be considered fair or reliable.


Finally, this year, a report of a workshop run jointly by The National Research Council and the National Academy of Education offered this:
Value-added methods involve complex statistical models applied to test data of varying quality. Accordingly, there are many technical challenges to ascertaining the degree to which the output of these models provides the desired estimates. Despite a substantial amount of research over the last decade and a half, overcoming these challenges has proven to be very difficult, and many questions remain unanswered...


Let me repeat that last sentence, written this year: Despite a substantial amount of research over the last decade and a half, overcoming these challenges has proven to be very difficult, and many questions remain unanswered...

And yet this administration wants to move ahead with using student test scores, perhaps analyzed through value-added assessment methodologies, as a significant component of teacher evaluation. It is including this as part of the criteria to win Race to the Top Funds. In fairness, the Department does not specify using value-added (although anything else is far worse) nor does it specify what percentage of the evaluation is to depend upon the test scores - both of these decisions are still left to the states, some of which have left themselves wiggle room in their applications, using terms like "significant" to indicate the proportion of the evaluation that will depend upon student test scores.

The original Bush proposal for No Child Left Behind, as it went up on the White House website shortly after the inauguration of the 43rd president, proposed giving a 1% bonus of Title I money to schools that would give parents the value-added scores of the teachers of their students. That, fortunately, did not make it into the final legislation. Now we have the Los Angeles Times action, about which the Secretary of Education has offered a somewhat mixed and confusing response, even as he seems to support the idea of using such evaluations in assessing of teachers. Since the Times story broke we have seen some who write or advocate about education who have praised what the paper did, while others have condemned it. While mine might not be a major voice on education, I find myself very much in the latter camp.

One problem is that too many who write about education are close to ignorant about the limits of the information one can get from various kinds of assessment. We tend to what hard numbers as a society, we are obsessed with comparisons and rankings. In the process we often give far more credence to quantitative measures than they warrant.

I do not dispute that tests, including tests external to the school, have some utility. I also recognize that value-added assessment is beginning to offer some useful additional information. By itself that information is not sufficiently reliable that people's livelihoods should be either solely or heavily determined by the information they provide. They MAY indicate a teacher outside the norm - either well above or well below - but as the various studies you will encounter in this brief demonstrate, that is not necessarily the case, the results are not yet stable for individual teachers from year to year, we do not yet know how to properly control for non-instructional factors that can influence the scores upon which the analysis is based, nor can we properly distribute responsibility for student learning among the different adults who interact with a child at school.

I am a high school teacher. Let me offer a hypothetical - if I do more work in a social studies class on a particular kind of writing and that is what is assessed on the English exam, does the English teacher properly deserve the credit or blame for how students do on that part of the test? Those of us who teach in high school are aware that students often learn about our content either in other classes or from interactions outside of our classroom. Sometimes what they learn is correct and increases their performance in our class, sometimes it is incorrect and undercuts what we are instructing. To date, even value-added assessment is insufficient to control for such influences and allow proper inferences to be drawn about the actual impact of the teacher upon the learning of the students.

I have only explored a small portion of the material in the brief. You can download it without paying. If you are worried about whether you will be able to understand the contents, don't. You can start with the executive summary, in which you will find most of the key takeaways, written in language and presented in a style that is easily accessible. It is a bit less than four pages. The brief itself runs from pages 5-21, followed by three columns (over a page and a half) of footnotes, and 5 columns (over three and half pages) of sources. You can read through the brief without having to check the footnotes, or you can if you want glance at the back to see who is being cited if that is not clear in the text.

Let me clear. The authors are not opposed to value-added assessment. They are not even opposed to it being included in the process of teacher evaluation, although they offer some serious cautions that policy makers would be well advised to consider.

The title is accurate - there are still serious problems with using test scores to evaluate teachers. These problems are not solved by resorting to a value-added methodology.

We need to be careful not to denigrate nor discourage our teaching corps. We will not improve education if the end result of our efforts is to drive away the very teachers who most connect with students, who are able to inspire those students to persist when they are struggling, who are willing to take on the harder to teach. We have other methods of ascertaining whether teachers are in fact effective. We should not be abandoning them in favor of quantitative measures that cannot, as yet, fully carry the load.

The authors of this study have enough prestige that one can hope our media will give some attention to it. Those responsible for educational policy at local, state and national levels are not doing their jobs if they are unwilling to read and be sure they understand the implications of this brief.

That said, and adding that I will try to bring to the attention of as many policy makers as I can, I do not have high hopes that our wrongheaded headlong pursuit of quantitative measures of teacher effectiveness can even be slowed. I will add what voice I have to the efforts of these scholars. Perhaps after you read the brief, you will add yours?

Thanks.

Problems with the use of Student Test Scores to Evaluate Teachers

originally posted at Daily Kos

If new laws or policies specifically require that teachers be fired if their students’ test scores do not rise by a certain amount, then more teachers might well be terminated than is now the case. But there is not strong evidence to indicate either that the departing teachers would actually be the weakest teachers, or that the departing teachers would be replaced by more effective ones. There is also little or no evidence for the claim that teachers will be more motivated to improve student learning if teachers are evaluated or monetarily rewarded for student test score gains.


That is a quote from the Executive Summary of one of the most important policy briefs about education in recent years. At a time when the Dept. of Education is pushing to tie teacher evaluation and compensation to student test scores, this Economic Policy Institute Briefing Paper (whose title is the same as this diary, and which is a pdf), pulls together the extensive relevant research that demonstrates the dangers of pursuing such a path. Please continue reading as I explore this important document, released at 12:01 AM today, August 29.

First, let me clarify several things.

This is a very long diary. That is because I am trying to reasonably thoroughly cover the contents of an extremely important document. My purpose in doing so is to convince people of the document's importance. Thus I will be perfectly happy should you decide you do not need to further read what I have written below. You can follow the link for the brief (which I have provided you again), download the pdf, and begin reading. The executive summary is only four pages. The brief itself, without the critical apparatus of footnotes and sources, another 17. So if you want, one more time follow this link.


This document has been in the works for several months, and was NOT hurriedly put together as a response to the recent series by the Los Angeles Times which used value-added assessment to label teachers in the Los Angeles Unified School District. Second, the ten scholars whose names are on the document are some of the most eminent in educational circles, including among their midst former Presidents of the American Educational Research Association and the National Council on Measurement in Education, two of the three professional organizations most involved with psychological measurement, of which school-related testing is a subset. One of the scholars, Robert Linn, has not only presided over both of those organizations, he has also serve as chair of the National Research Council's Board on Testing and Assessment. The group also includes the immediate past president of the National Academy of Education, Lorrie Shepard, Dean of the School of Education at Colorado. A brief and applicable curricula vitae of each of the ten authors can be found at the end of the document, and briefer descriptions at the beginning, where each author is listed, along with the following statement:
Authors, each of whom is responsible for this brief as a whole, are listed alphabetically.
An email address is provided for further contact.

The ten authors, alphabetically, are as follows:
Eva L. Baker
Paul E. Barton
Linda Darling-Hammond
Edward Haertel
Helen F. Ladd
Robert E. Linn
Diane Ravitch
Richard Rothstein
Richard J. Shavelson
Lorrie A. Shepard

Let me be blunt. I do not know how anyone who knows the work of these scholars and who reads this brief can accept the idea of placing any stakes as to firing or awarding of merit pay based on the current status of Value-Added Assessment methodologies. The document is thorough. It reviews all the relevant studies, including one not yet in print. Those includes studies by Mathematica for the US Department of Education: by Rand: by the Educational Testing Service; done for the National Center for Education Statistics of the Institute of Education Sciences of the U. S. Dept. of Education; issued by the Board of Testing and Assessment of the Division of Behavioral and Social Sciences and Education of the National Academy of Sciences, and so on. There are citations from books, from peer reviewed journals.

I am not a scholar. I am a high school social studies teacher. During now abandoned doctoral studies in educational policy I got interested in value-added assessment and devoured what studies there were in the educational literature. I also talked extensively with the technical person for one organization that offered a value-added methodology who cautioned me that the approach was not stable enough for it to be used as the basis for decisions with any kind of meaningful stakes. That was about a decade ago. What I had read since, and what I have absorbed from this study convinces me that the situation is not significantly better now.

But you do not have to take my word for it. Let me offer a few key examples from the study. Those who follow me on Daily Kos already have seen in the study by Mathematica the high rate of error in determining superior and inferior teachers beyond the broad middle. In this diary, written on August 27, I noted that the error rate with 2 years of data was 36%, with 3 years 26%, and even with 10 years of data still 12%.

But that is just the tip of the iceberg of the technical problems with using such an approach.

Without recapitulating the entire brief, let me offer a couple of other key points.

1. Results for individual teachers are not stable:
One study found that across five large urban districts, among teachers who were ranked in the top 20% of effectiveness in the first year, fewer than a third were in that top group the next year, and another third moved all the way down to the bottom 40%. Another found that teachers’ effectiveness ratings in one year could only predict from 4% to 16% of the variation in such ratings in the following year.


2. One key question is whether one is really accounting for teacher effects and excluding other influences in the results one gets from value-added assessment. Jesse Rothstein reported something interesting, about which I quote from the Executive Summary:
A study designed to test this question used VAM methods to assign effects to teachers after controlling for other factors, but applied the model backwards to see if credible results were obtained. Surprisingly, it found that students’ fifth grade teachers were good predictors of their fourth grade test scores. Inasmuch as a student’s later fifth grade teacher cannot possibly have influenced that student’s fourth grade performance, this curious result can only mean that VAM results are based on factors other than teachers’ actual effectiveness.


3. The brief notes that arguments that the private sector evaluates professional employees using quantitative measures that are parallel. The authors of the brief point out that rarely are such quantitative measures the sole or even the primary factor, noting that management experts warning against using such measures for making salary or bonus decisions. They remind us that some of the distortion on Wall Street was the result of emphasizing short term gains that could be easily measured. They also touch on medicine:
In both the United States and Great Britain, governments have attempted to rank cardiac surgeons by their patients’ survival rates, only to find that they had created incentives for surgeons to turn away the sickest patients.


4. Students are not randomly assigned to teachers. While some control for school effects is possible, scholars are reluctant to place any weight on comparisons for teachers in different schools even within the same system. And even within a school, teachers may have varying numbers of students who are learning English or have learning disabilities or are homeless or who move multiple times, each of which is a factor that can affect learning.

5. Sample sizes are often too small. Even if the class makeup stays stable during the year, and all the students show up regularly, the N=30 of a large elementary class is too small a sample to provide a result that can allow strong inferences to be drawn. Often the makeup of the class changes during the year. If you exclude students who were not there all year, or whose absences exceed some designated level, the N decreases, providing a result of even less reliability.

6. Some argue that statewide data banks can address the question of student mobility. But if you derive results on a year or two years of data where the student has moved, how much of the improvement can properly be assigned to any one teacher? Even in elementary school, do we account for pull-out instruction, or possible tutoring (that could in some cases be counterproductive) as a possible influence on the test results upon which we base our analysis?

7. Even with value-added analysis, to date scholars have not been able to isolate the impact of outside learning experiences, home and school supports, and differences in student characteristics and starting points when trying to measure their growth.

8. A proper system of value-added assessment would have vertically scaled tests. Most states do not currently have such tests, for example, neither New York nor California does. That is, the tests in one grade are not necessarily congruent with those of the next along a continuum from year to year - we are not testing the same thing each year. As testing expert Dan Koretz of Harvard is quoted as noting,
"because of the need for vertically scaled tests, value-added systems may be even more incomplete than some status or cohort-to-cohort systems"
Here it is worth noting that cohort to cohort is comparing this year's fourth graders to last years, which is how Adequate Yearly Progress under No Child Left Behind has been calculated.

9. If measuring end of year to end of year, even if there are vertically scaled tests, there is still the well-documented issue of summer learning loss, which falls disproportionally upon those of lesser economic means, which also means it falls disproportionally upon those of color, who are more heavily represented at the lower end of the economic scale. IF we do not control for summer learning loss, our results are skewed. Allow me to quote a relevant portion of the study:
researchers have found that three-fourths of schools identified as being in the bottom 20% of all schools, based on the scores of students during the school year, would not be so identified if differences in learning outside of school were taken into account. Similar conclusions apply to the bottom 5% of all schools.
The authors also cite a study that shows "two-thirds of the difference between the ninth grade test scores of high and low socioeconomic status students can be traced to summer learning differences over the elementary years."

There is more, but this should give a real sense of how much there is in this paper, how thoroughly the authors examine relevant material to demonstrate that value-added assessment, the supposed magic bullet to allow us to tie student learning back to the effectiveness of teachers, cannot properly fulfill the task some wish to give to it.

The authors acknowledge that value-added approaches are superior to some of the alternatives methods of using test scores to evaluate teachers. These are

status test-score comparisons - compare average scores of students of one teacher to those of another

over change measures - compare the average test results of a single teacher from one year to the next - remember, these are different students

over growth measures - a comparison of the scores of the students of the teacher this year to the scores of those same students the previous year when they had different teachers.

Each of these approaches has serious problems with it. One can read the detailed explanation on p. 9. Value-added assessments may be an improvement, but
the claim that they can “level the playing field” and provide reliable, valid, and fair comparisons of individual teachers is overstated. Even when student demographic characteristics are taken into account, the value-added measures are too unstable (i.e., vary widely) across time, across the classes that teachers teach, and across tests that are used to evaluate instruction, to be used for the high-stakes purposes of evaluating teachers.



Let me offer a few of the quotes about value-added assessment that the authors of the brief offer from scholars who have examined the approach over the years, and then I will offer a few observations of my own.

in 2003, a research team at Rand concluded
The research base is currently insufficient to support the use of VAM for high-stakes decisions about individual teachers or schools.


In 2004, Donald Rubin opined
We do not think that their analyses are estimating causal quantities, except under extreme and unrealistic assumptions.


Henry Braun, then at ETS, offered this in 2005:
VAM results should not serve as the sole or principal basis for making consequential decisions about teachers. There are many pitfalls to making causal attributions of teacher effectiveness on the basis of the kinds of data available from typical school districts. We still lack sufficient understanding of how seriously the different technical problems threaten the validity of such interpretations.


Last year the Board on Testing and Assessment of the National Research Council of the National Academy of Sciences wrote to the Department of Education saying
...VAM estimates of teacher effectiveness should not be used to make operational decisions because such estimates are far too unstable to be considered fair or reliable.


Finally, this year, a report of a workshop run jointly by The National Research Council and the National Academy of Education offered this:
Value-added methods involve complex statistical models applied to test data of varying quality. Accordingly, there are many technical challenges to ascertaining the degree to which the output of these models provides the desired estimates. Despite a substantial amount of research over the last decade and a half, overcoming these challenges has proven to be very difficult, and many questions remain unanswered...


Let me repeat that last sentence, written this year: Despite a substantial amount of research over the last decade and a half, overcoming these challenges has proven to be very difficult, and many questions remain unanswered...

And yet this administration wants to move ahead with using student test scores, perhaps analyzed through value-added assessment methodologies, as a significant component of teacher evaluation. It is including this as part of the criteria to win Race to the Top Funds. In fairness, the Department does not specify using value-added (although anything else is far worse) nor does it specify what percentage of the evaluation is to depend upon the test scores - both of these decisions are still left to the states, some of which have left themselves wiggle room in their applications, using terms like "significant" to indicate the proportion of the evaluation that will depend upon student test scores.

The original Bush proposal for No Child Left Behind, as it went up on the White House website shortly after the inauguration of the 43rd president, proposed giving a 1% bonus of Title I money to schools that would give parents the value-added scores of the teachers of their students. That, fortunately, did not make it into the final legislation. Now we have the Los Angeles Times action, about which the Secretary of Education has offered a somewhat mixed and confusing response, even as he seems to support the idea of using such evaluations in assessing of teachers. Since the Times story broke we have seen some who write or advocate about education who have praised what the paper did, while others have condemned it. While mine might not be a major voice on education, I find myself very much in the latter camp.

One problem is that too many who write about education are close to ignorant about the limits of the information one can get from various kinds of assessment. We tend to what hard numbers as a society, we are obsessed with comparisons and rankings. In the process we often give far more credence to quantitative measures than they warrant.

I do not dispute that tests, including tests external to the school, have some utility. I also recognize that value-added assessment is beginning to offer some useful additional information. By itself that information is not sufficiently reliable that people's livelihoods should be either solely or heavily determined by the information they provide. They MAY indicate a teacher outside the norm - either well above or well below - but as the various studies you will encounter in this brief demonstrate, that is not necessarily the case, the results are not yet stable for individual teachers from year to year, we do not yet know how to properly control for non-instructional factors that can influence the scores upon which the analysis is based, nor can we properly distribute responsibility for student learning among the different adults who interact with a child at school.

I am a high school teacher. Let me offer a hypothetical - if I do more work in a social studies class on a particular kind of writing and that is what is assessed on the English exam, does the English teacher properly deserve the credit or blame for how students do on that part of the test? Those of us who teach in high school are aware that students often learn about our content either in other classes or from interactions outside of our classroom. Sometimes what they learn is correct and increases their performance in our class, sometimes it is incorrect and undercuts what we are instructing. To date, even value-added assessment is insufficient to control for such influences and allow proper inferences to be drawn about the actual impact of the teacher upon the learning of the students.

I have only explored a small portion of the material in the brief. You can download it without paying. If you are worried about whether you will be able to understand the contents, don't. You can start with the executive summary, in which you will find most of the key takeaways, written in language and presented in a style that is easily accessible. It is a bit less than four pages. The brief itself runs from pages 5-21, followed by three columns (over a page and a half) of footnotes, and 5 columns (over three and half pages) of sources. You can read through the brief without having to check the footnotes, or you can if you want glance at the back to see who is being cited if that is not clear in the text.

Let me clear. The authors are not opposed to value-added assessment. They are not even opposed to it being included in the process of teacher evaluation, although they offer some serious cautions that policy makers would be well advised to consider.

The title is accurate - there are still serious problems with using test scores to evaluate teachers. These problems are not solved by resorting to a value-added methodology.

We need to be careful not to denigrate nor discourage our teaching corps. We will not improve education if the end result of our efforts is to drive away the very teachers who most connect with students, who are able to inspire those students to persist when they are struggling, who are willing to take on the harder to teach. We have other methods of ascertaining whether teachers are in fact effective. We should not be abandoning them in favor of quantitative measures that cannot, as yet, fully carry the load.

The authors of this study have enough prestige that one can hope our media will give some attention to it. Those responsible for educational policy at local, state and national levels are not doing their jobs if they are unwilling to read and be sure they understand the implications of this brief.

That said, and adding that I will try to bring to the attention of as many policy makers as I can, I do not have high hopes that our wrongheaded headlong pursuit of quantitative measures of teacher effectiveness can even be slowed. I will add what voice I have to the efforts of these scholars. Perhaps after you read the brief, you will add yours?

Thanks.