Spelling Bee 2026 Post Mortem
What BEELO got right and wrong about this year's batch of spellers (and the words they faced).
In my last post, I gave a brief explainer of my new model, BEELO, which went live for this year’s Scripps National Spelling Bee. Outside of a few bugs (due mostly to the Bee site itself not updating correctly), the live updater held up really well. In fact, I feel really good about prospects for doing this again next year.
Of course, the real question is: How did the model actually perform? The answer? Well, there’s good news and there’s bad news.
And the winner is…?
Let’s start with the good: Pre-Bee speller ratings. For the most part, the speller priors held up to how the Bee played out.
It’s hard to look at that table and not feel really good about how the ratings worked. Of the final 12 spellers (including the 9 finalists), only two were higher than 22nd: Aiden Meng (37th) and Logan Bailey (72nd), the latter of which is a 12-year-old home school student who I expect will make a return next year. Every single one of the top five was among the top 12 spellers.
The biggest disappointment, personally, was that the pre-Bee favorite, Sarv Dharavane, didn’t end up winning. In fact, he ended up right back where he finished last year (3rd). The winner, on the other hand, Shrey Parikh, was 3rd in the pre-Bee ratings.
I think there’s a reasonable explanation for this, though. Shrey had competed before, but that was back in 2024, when he, like Sarv, finished 3rd. He missed 2025 when he lost in his school bee due, at least in part, to an illness. The model is trained to regress spellers who have gap years, meaning it’s no surprise that his score was knocked a little relative to Sarv. Had he been at the 2025 Bee, there’s a good chance he would have gone in as a favorite (or may not have even been eligible, if he’d won last year).
As for the rest of the spellers, BEELO had a solid showing for the top end of the list, but that was expected. We’ve seen them before, and the model has a better idea about how strong they are. Further down the list, however, the results are a little murkier.
I divided the spellers into quartiles and measured their average round reached. As expected, the spellers in the top quartile performed the best, reaching, on average, somewhere between the 6th and 7th round. The spellers in the third quartile reached round 3.3 on average. Meanwhile the spellers in the bottom two quartiles were basically indistinguishable, averaging rounds 2.6 and 2.7.
If I’m being honest, I’m not all that shocked. The Bee is made up of 247 spellers, and most of them have no prior history. The model bases their rating on a combination of their age and grade. There may be some way to improve this in the future,1 but for the most part the naivety is a feature, not a bug. The fact is that 70% of the spellers in 2026 had no prior, so the model lumps them mostly into a pretty narrow band. Non-returning spellers’ starting ELOs ranged from 1514 to 1542, a difference of less than 30 points.
In other words, the model expects returners to finish higher, and this year bore that out. If you’re looking to find a Bee winner, you generally don’t look at the rookies, you look at those with Bees under their belt.
What’s the word?
The bad news? The word model was not exactly as sharp as I’d hoped it’d be, particularly on words that the model was seeing for the first time and guessing at with features.
To be fair, that calibration doesn’t look particularly rosy for either side. Originally I thought that just calibrating based on the round number would improve things, but that was slightly off as well. The model ended up timing the difficulty ramp-up incorrectly in the end.
I think part of that may have been due to some adjustments I made in the run-up. Rather than take raw round numbers, I adjusted to remove the vocab and test rounds. This might have skewed the results, causing the spike to happen sooner. I’ve already gone back and tweaked that part of the code, so hopefully that’s a boon next year.
Otherwise, I’m not sure how else to improve the word model. There are some spelling bee training resources that look promising. Outside of that, though, I think the only methods are finding a way to incorporate more sound features (schwa sounds, anyone?) or refining the language of origin piece of this.
Ultimately, the word model feels like the best way to make BEELO better next year. Better word ratings means more accurate priors, better accuracy at the round-by-round results, and quicker convergence on who the real favorites are. That way we don’t end up with so many big misses, like these:
I’ve toyed with including hometown population as a possible feature. It’d be great if I could get regional spelling bee data, but I’m not holding my breath.






