Rubric Was Built Twice
A record of building Rubric twice: we failed trying to move assessment out of gut feeling into a tool, then changed course toward supporting teachers' judgment.
One Essay, Five Scores
Last fall, I passed a single essay by one student around to five teachers at Epic Prep. The scores came back between 72 and 94. A 22-point spread. Nobody was wrong. Each of them was simply looking at something different. One looked at the structure of the writing, another at vocabulary, another at how far the student had come since the month before.
The problem is that those scores end up in parent meetings. When a parent asks, “Is my child’s writing actually getting better?” and the answer changes every time the homeroom teacher does, the academy loses trust. But if headquarters sends down a scoring sheet and unifies the numbers, you lose the very eye that was noticing “how far the student had come since the month before.” The more you standardize assessment, the thinner the judgment gets; the more you leave it to judgment, the less it can be reproduced. That is what makes this problem hard.
I thought software could solve it. I was half right and half wrong.
The First Version, 18 Percent
The first version of Rubric, which shipped early this year, was an entry form with 24 assessment items. After class, teachers filled out the items for each student, and as long as the form got filled, per-class statistics and growth graphs came out automatically. Opening the form took three clicks; finishing one student took more than twenty entries. The entry rate in the first week was 81 percent. By week four it had dropped to 18. The teachers had stopped using it.
I still remember what Jamie told me when I went to the campus to ask why. “It takes seven minutes per student. With a class of eight, that’s an hour. The kids get more out of it if I spend that hour preparing the next lesson.” Fair enough. The remark that stung more came next. “And honestly, this doesn’t feel like a screen that helps me assess. It feels like being checked on whether I did my job.” I had tried to turn teachers’ judgment into data, and the teachers read that, accurately, as surveillance. Statistics coming out automatically was the builder’s point of pride, not the user’s gain.
Changing Course
Deciding to write off two months of development did not take long. The numbers were unambiguous. What took long was deciding what to replace it with. I spent a few days sitting in on classes at the campus Rachel directs, and what I saw was that teachers were already doing assessment. In the margins of textbooks, in group chat announcements, in memories hastily dredged up right before a parent meeting. There was no shortage of judgment. The problem was that none of it accumulated.
So the second version went the opposite way. We cut the 24 items down to six, and instead of score entry we made the starting point the class comments teachers were writing anyway. We settled on what Rubric’s job actually was: not to make the judgment for the teacher, but to let scattered judgments pile up along one student’s timeline. Before a parent meeting, a teacher can reread their own notes and set them next to the previous teacher’s. The tool would go that far and no further.
The entry rate for the revised version currently sits in the 60s. Not 100. But I no longer treat that number as the target. Every time part of the other 40 percent didn’t use it, there was a reason, and listening to those reasons is becoming the backlog for the next version.
The Math That Remains
Here is the arithmetic. Scrapping the first version cost two months of development. What we gained: teachers telling us that prep for parent meetings has dropped from an average of 40 minutes to around 15, and student records that no longer break when the homeroom teacher changes. That 15-minute figure is, of course, the teachers’ felt sense, not something I timed with a stopwatch. Even so, the trade comes out ahead.
One doubt I am keeping, though. Whether the meetings got better because of Rubric, or because living through that failure taught us to count teachers’ time more seriously than before, I still cannot tell them apart. The latter may deserve the larger share of the credit. So I am leaving this post as a record rather than a promotion. The next time I feel the urge to push a feature through, I intend to open this page again.