Showing posts with label Aaron Pallas. Show all posts
Showing posts with label Aaron Pallas. Show all posts

Thursday, September 30, 2010

Why the school grading system, and Joel Klein, still deserve a big "F"

Amidst all the hype and furor of the release of today’s NYC school "progress reports", everyone should remember how the grades are not to be trusted. By their inherent design, the grades are statistically invalid, and the DOE must be fully aware of this fact. Why?

See this Daily News oped I wrote in 2007, in which all the criticisms still hold true, “Why parents and teachers should reject the new grades”.
In part, this is because 85% of each school’s grade depends on one year’s test scores alone – which according to experts, is highly unreliable. Researchers have found that 32 to 80% of the annual fluctuations in a typical school’s scores are random or due to one time factors alone, unrelated to the amount of learning taking place. Thus, given the formula used by the Department of Education, a school’s grade may be based more on chance than anything else.
(source: Thomas Kane, Douglas O. Staiger, “The Promise and Pitfalls of Using Imprecise School Accountability Measures, The Journal of Economic Perspectives, Autumn, 2002.)

Now Jim Liebman admitted this fact, that one year’s test score data was inherently unreliable, in testimony to the City Council, and to numerous parent groups, including to CEC D2, as recounted on p. 121 of Beth Fertig’s book, Why can’t U teach me 2 read.” In responding to Michael Markowitz’s observations that the grading system was designed to provide essentially random results, he admitted:

“There’s a lot I actually agree with, he said in a concession to his opponent…He then proceeded to explain how the system would eventually include three years’ worth of data on every school so the risk of big fluctuations from one year to the next wouldn’t be such a problem.”

Nevertheless, the DOE and Liebman have refused to comply with this promise, which reveals a basic intellectual dishonesty. This is what Suransky emailed me about the issue, a couple of weeks ago, when I asked him about it before our NY Law school “debate.”

“We use one year of data because it is critical to focus schools’ attention on making progress with their students every year. While we have made gains as a system over the last 9 years, we still have a long way to reach our goal of ensuring that all students who come out of a New York City school are prepared for post-secondary opportunities. Measuring multiple years’ results on the Progress Report could allow some schools to “ride the coattails” of prior years’ success or unduly punish schools that rebound quickly from a difficult year.”

Of course, this is nonsense. No educators would “coast” on a prior year’s “success”, but they would be far more confident in a system that didn’t give them an inherently inaccurate rating.

Given the fact that that school grades bounce up and down each year, most teachers, administrators and even parents have long figured out how they should be discounted, and justifiably believe that any administration that would punish or reward a school based on such invalid measures is not to be trusted.

That DOE has changed the school grading formula in other ways every year for the last three years also doesn’t give one any confidence….though they refuse to change the most fundamental flaw. Yet another major problem is while the teacher data reports take class size into account as a significant limiting factor in how much schools can get student test scores to improve, the progress reports do not.

There are lots more problems with the school grading system, including the fact that they are primarily based upon state exams that we know are themselves completely unreliable. As MIT professor Doug Ariely recently wrote about the damaging nature of value-added teacher pay, because of the way they are based on highly unreliable measurements:

…What if, after you finished kicking [a ball] somebody comes and moves the ball either 20 feet right or 20 feet left? How good would you be under those conditions? It turns out you would be terrible. Because human beings can learn very well in deterministic systems, but in a probabilistic system—what we call a stochastic system, with some random error—people very quickly become very bad at it.

So now imagine a schoolteacher. A schoolteacher is doing what [he or she] thinks is best for the class, who then gets feedback. Feedback, for example, from a standardized test. How much random error is in the feedback of the teacher? How much is somebody moving the ball right and left? A ton. Teachers actually control a very small part of the variance. Parents control some of it. Neighborhoods control some of it. What people decide to put on the test controls some of it. And the weather, and whether a kid is sick, and lots of other things determine the final score.

So when we create these score-based systems, we not only tend to focus teachers on a very small subset of [what we want schools to accomplish], but we also reward them largely on things that are outside of their control. And that's a very, very bad system.”

Indeed. The invalid nature of the school grades are just one more indication of the fundamentally dishonest nature of the Bloomberg/Klein administration, and yet another reason for the cynicism, frustration and justifiable anger of teachers and parents.

Also be sure to check out this Aaron Pallas classic: Could a Monkey Do a Better Job of Predicting Which Schools Show Student Progress in English Skills than the New York City Department of Education?

Wednesday, December 2, 2009

Tying tenure to test scores: not ready for prime time


Lots of interesting letters to the Times today, deploring the Mayor's proposal to base tenure decisions on test scores. [see “Mayor to Link Teacher Tenure to Test Scores” ]

In the same vein, Aaron Pallas has a column in Gotham schools, Teacher Education in New York State: A skoolboy’s-Eye View, in which he lucidly explains how the evaluation of teachers based on value-added student test scores is not ready for prime time. Pallas recently appeared on a panel at Teachers College with David Steiner, new NY Commissioner of State Education, (photo to the right), and Merryl Tisch, head of the Board of Regents. (You can see a webcast of this event here.)

In his column, Pallas urges Steiner and Tisch to start working on improving the state exams, which have gotten radically easier over time, before beginning to consider a system that would base decision-making on their results. He also points out how the long-standing practice of having high schools score their own Regents exams is a system ripe for abuse.

As part of the state's "Race to the Top" proposal, Commissioner Steiner recently also proposed that they expand the awarding of teaching degrees -- allowing providers other than institutions of higher learning to offer teacher preparation programs, with the Board of Regents granting Master’s degrees to candidates who "graduate" from these programs.

There is so much lacking in terms of the state's current oversight -- of district spending practices, of cheating, of "credit recovery", of the proper reporting of graduation rates, of whether schools are even providing the minimal services to kids that they are entitled to by law.

Given the awful mess at State Ed which Steiner has not yet begun to clean up, I would hate to see him allow further abuses to occur by deregulating the awarding of teaching degrees -- which could easily make a teaching certificate as meaningless as passing the Regents exam is now.

Wednesday, September 30, 2009

Live By The Sword, Die By The Sword?

The problem with Jay Mathews' defense ("Measuring Progress At Shaw With More Than Numbers") of a Washington, DC school principal who did not demonstrate student learning gains at his school after one year is that the principal operates within an accountability system that demands such a result. In this case, both Mathews -- and DC Schools Chancellor Michelle Rhee, as described in Mathews' WP column -- are right not to have lowered the boom on Brian Betts, principal of the DC's Shaw Middle School at Garnet-Patterson, based on a single year's worth of test scores.
The state superintendent of education's Web site says Shaw dropped from 38.6 to 30.5 in the percentage of students scoring at least proficient in reading, and from 32.7 to 29.2 in math.

But those were not the numbers Rhee read to Betts over the phone.

Only 17 percent of Shaw's 2009 students had attended the school in 2008, distorting the official test score comparisons. Rhee instead recited the 2008 and 2009 scores of the 44 students who had been there both years. It didn't help much.

The students' decline in reading was somewhat smaller; it went from 34.5 to 29.7. Their math proficiency increased a bit, from 26.2 to 29.5. But Shaw is still short of the 30 percent mark, far below where Rhee and Betts want to be....

Despite the sniping at Rhee, the best teachers I know think that what happened at Shaw is a standard part of the upgrading process. I have watched Betts, his staff, students and parents for a year. The improvement of poor-performing schools has been the focus of my reporting for nearly three decades. The Shaw people are doing nearly everything that the most successful school turnaround artists have done.

They have raised expectations for students. They have recruited energetic teachers who believe in the potential of impoverished students. They have organized themselves into a team that compares notes on youngsters. They regularly review what has been learned, what some critics dismiss as "teaching to the test." They consider it an important part of their jobs.

That's how it's done, usually with a strong and engaging principal like Betts.

Mathews' take -- including consideration of contextual factors, such as the fact that only 17% of the school's students had attended the prior year and the contention that school turnaround requires more than a single year -- is how the education world should work. Embrace the complexity of learning and trying to measure it! To do so would disallow the use of single-year changes in test scores for making high-stakes decisions about schools and individual school personnel. It would also remove the unrealistic pressure on school turnarounds to bear fruit in a single year. Test scores would be used responsibly in combination with other data and evidence to paint a fuller picture about individual school contexts and inform judgments about school leadership and student success.

But Michelle Rhee and other education reform advocates have publicly argued that student performance as measured by test scores is basically the be all and end all. According to this Washington Post story ("Testing Tactics Helped Fuel D.C. School Gains"), Rhee supports strengthening No Child Left Behind to "emphasize year-to-year academic growth." Such a stance creates a problem for such reformers when they are leading a district and staking their leadership on uncomplicated test score gains. Others will assess their leadership and judge their success by this measure -- an ill-advised one in its simplest form.

I would argue that, in addition to doing the right thing (as happened in this instance), reform advocates and school leaders like Rhee also have a responsibility to say and advocate for the right thing. They have a responsibility to be honest about the complexity of student learning and the inability of student assessments to accurate capture all of the nuance going on within schools and classrooms. While the reformers' challenge of the adult-focused policies of the educational status quo is often warranted, some reforms -- accountability, chief among them -- have been taken too far. Student learning, school leadership and teaching cannot be measured and judged good or bad based on a single set of test scores. Test scores must be part of the consideration -- and supporting systems such as accountability, compensation and evaluation must be informed by such data -- but they should not single-handedly define success or failure.

The complexity as presented by Mathews in his article -- and, more importantly, by existing research (such as by Robert Linn, Aaron Pallas, Tim Sass, and embedded within Sunny Ladd's RttT comments) about year-to-year comparisons of both overall test scores and test score gains -- strongly suggests that educational accountability systems should be designed more thoughtfully than they have been to date, but unfortunately that does not seem to be the direction that policymaking is headed at either the federal or state levels. Part of being more thoughtful is moving away from NCLB-style adequate yearly progress and toward a value-added approach, but thoughtfulness also requires not making high-stakes decisions based exclusively on volatile student data. Do I hear "multiple measures"? Sure, but Sherman Dorn offers some provocative thoughts on this subject in a 2007 blog post.

With regard to educational accountability, policymakers first should do their homework -- and then they clearly have more work to do in creating a better system and undoing parts of the existing system that aren't evidence-based and accomplish only in simplifying a truly complex art: learning.

-------------------

For those of you that have gotten this far, there's a related post on the New America Foundation's Ed Money Watch blog discussing a new GAO report that analyzes state spending on student assessment tests -- $640 million in 2007-08.
The increasing cost of developing and scoring assessments has also led many states to implement simpler and more cost-effective multiple choice tests instead of open response tests. In fact, although five states have changed their assessments to include more open response items in both reading and math since 2002, 11 and 13 states have removed open items from their reading and math tests, respectively over the same time period.... This reliance on multiple choice tests has forced states to limit the content and complexity of what they test. In fact, some states develop academic standards for testing separately from standards for instruction, which are often un-testable in a multiple choice system. As a result, state NCLB assessments tend to test and measure memorization of facts and basic skills rather than complex cognitive abilities.
------------

And here's a new story hot off the presses from Education Week. It discusses serious questions raised about New York City's school grading system.

Eighty-four percent of the city’s 1,058 public elementary and middle schools received an A on the city’s report cards this year, compared with 38 percent in 2008, while 13 percent received a B, city officials announced this month.

“It tells us virtually nothing about the actual performance of schools,” Aaron M. Pallas, a professor of sociology and education at Teachers College, Columbia University, said of the city’s grades.

Diane Ravitch, an education historian at New York University, was even sharper: She declared the school grades “bogus” in a Sept. 9 opinion piece for the Daily News of New York, saying the city’s report card system “makes a mockery of accountability.”

But Andrew J. Jacob, a spokesman for the New York City Department of Education, defended the ratings, even as he said the district’s demands on schools would continue to rise next year....

The city employs a complex methodology to devise its overall letter grades, with the primary driver being results from statewide assessments in reading and mathematics, which have also encountered considerable skepticism lately.

The city’s grades are based on three categories: student progress on state tests from one year to the next, which accounts for 60 percent; student performance for the most recent school year, which accounts for 25 percent; and school environment, which makes up 15 percent.

Mr. Pallas of Teachers College argues that one key flaw with the city’s rating system is that it depends heavily on a what he deems a “wholly unreliable” measure of student growth on test scores from year to year that fails to account adequately for statistical error.


Monday, May 4, 2009

Cheerleading for NCLB

I guess my reaction to today's Washington Post shout out to No Child Left Behind ("'No Child' in Action") from former U.S. Education Secretary Margaret Spellings is a question: "If NCLB's accountability alone is such a silver bullet, then how come test scores at the high school level didn't improve?"

Although Spellings mentions that NCLB requires math and reading tests in grades 3-8, it is quite disingenuous of her not to mention that such tests were also required in high school. If the achievement gains aren't sustained through high school, what real difference does it make?

The wise Aaron Pallas offers his take on this issue ("Wishful Thinking"), calling into question Spellings's claims:
But what portion of those trends can be attributed to NCLB? Margaret Spellings refers to changes since 1999, which is convenient for her story, because there were sharp increases in grade 4 reading between 2000 and 2002, and in grade 4 and grade 8 math between 2000 and 2003. But NCLB was signed into law in January, 2002; the first final regulations dealing with assessment were issued in December, 2002; and initial state accountability plans were approved by the U.S. Department of Education no later than June, 2003. The 2003 main NAEP was administered between January and March of 2003. Is it realistic to claim that NCLB affected scores before the 2003 NAEP administration? I, and a great many other analysts, think not.

Only in Margaret Spellings’ world can NCLB affect NAEP scores for the four years before the law was passed and implemented. Now that’s wishful thinking.

UPDATE -- Diane Ravitch comes to similar conclusions in her blog post.
Thus, when one looks at the patterns, it suggests the following: First, our students are making gains, though not among 17-year-olds. Second, the gains they have made since NCLB are smaller than the gains they made in the years preceding NCLB. Third, even when they are significant, the gains are small. Fourth, the Long Term Trend data are not a resounding endorsement of NCLB. If anything, the slowing of the rate of progress suggests that NCLB is not a powerful instrument to improve student performance.
Caveat emptor.

Friday, March 13, 2009

Stupid Stuff from Skoolboy

Kudos to Dr. Aaron Pallas (AKA skoolboy) for his terrific post ("It's The Stupid System") today on the Gotham Schools blog.

He takes New York City schools chancellor Joel Klein and the Reverend Al Sharpton to task for their Huffinington Post piece which implies that there is an easy achievement-gap fix -- namely value-added assessment and merit pay alone.
As usual, skoolboy’s main concern is that Klein and Sharpton are talking about effective teachers without ever once discussing what it is that they do. Reward the good ones, get rid of the bad ones, it’s all about sorting teachers–and never about actually improving instruction. Let’s suppose that Klein, Sharpton and others are right–that it is difficult to tell which teachers are going to be highly successful when they start teaching, because the instruction teachers receive prior to taking over a classroom can’t fully prepare them for the challenges of an urban classroom. Why not focus on professional development, and assisting novice teachers in learning effective practices on the job? How does giving effective teachers merit pay and dismissing poor performers actually improve anyone’s practice?
I wholeheartedly agree with Pallas's take on this. I said as much in my post on Monday ("Measurement Is Not Destiny"). The human capital challenge can't just be about rewarding the best and dismissing the worst. It must also be about a focused effort to make the vast majority of educators more effective. That will require a comprehensive effort, including high-quality, job-embedded, sustained professional development and robust induction support.

UPDATE: Corey Bunje Bower at Ed Policy Thoughts has some thoughts on the Klein/Sharpton piece as well.

Thursday, August 28, 2008

EduProfs in the 21st Century

Edu-Academia is taking giant leaps forward (in my opinion) with the willingness of faculty (and future faculty) to get out there in the blogosphere and talk with the commoners. If we only hear the thoughts of our colleagues when they appear in peer-reviewed journals, then we're often engaging in an out-of-date conversation.

So, props for the day go to my Facebook pal Aaron Pallas; Columbia faculty and as it turns out skoolboy extraordinaire. It's a brave, brave new world, and the bigger the population the more powerful it is....

And my oh my, I had NO idea that I was *already* friends with skoolboy and eduwonkette. My known circle of friends already included Kevin Carey. Now things are SO interesting....