Three Wordle grids side by side in one room, each at a different stage, all on the same word

The 2/6 that broke the group chat

Someone in my group posted a 2/6 last winter. Two guesses. The chat went off for a solid ten minutes. Screenshots, disbelief, one accusation of cheating.

Here's the thing I knew and didn't say, because nobody wants that guy in the chat: a 2/6 is mostly luck. Run a perfect solver over the whole five-letter answer pool and it lands on the word in two guesses 8.5% of the time. Not because it got clever on those days. Because the second guess happened to be the answer rather than one of the four other words that fit equally well.

That got me curious about the rest of it. If a 2/6 says almost nothing, what does the daily grid say? I build a multiplayer word game, so I had every incentive to find that it says nothing at all. The numbers came back messier than that, and the last one is genuinely bad for me. It's at the bottom.

First, the part the sceptics get wrong

Wordle serves the same word to everyone on the same day. You, your sister, a stranger in Auckland: identical puzzle, identical difficulty. So the comparison is not rigged. Nobody drew an easier word than you.

That surprises people who assume the grids are meaningless. They aren't meaningless. They're just an extremely small sample of something extremely noisy.

One word a day. That's the whole problem.

Same word, different door

Take twelve reasonable opening words. Not gimmicks: the twelve highest-information five-letter openers in my English dictionary, the sort of thing anyone who has read one strategy article might land on.

CRATE, TRACE, RAISE, SLATE, IRATE, ARISE, CRANE, SNARE, ALERT, LEAST, TRADE, REACT.

Now solve all 1,632 possible answers twelve times over, once per opener, using the identical method after that first move. Same player, same brain, same discipline. The only thing that changes is which door they walked in through.

On a given answer, the twelve openers…Share of words
all take the same number of guesses2.9 %
land one guess apart49.1 %
land two or more guesses apart47.9 %

Fewer than three words in a hundred are indifferent to your opener. Pick any two of those twelve openers and they disagree on the score 48.1% of the time, with an average spread of 1.55 guesses between the best and worst of the twelve.

So when your friend posts 3/6 and you post 4/6, the most likely explanation is not that they out-thought you. It's that their habitual opener happened to suit today's word. Tomorrow it'll be yours.

How long until the group chat knows who's best

You can push this all the way, and it gets absurd.

CRATE really is a better opener than CANOE. That's a genuine, measurable edge across the whole dictionary:

OpenerAverage guesses
CRATE3.407
CANOE3.442

Three and a half hundredths of a guess. Here's what that edge looks like on a single day:

CRATE wins
25.6 %
Tie
51.7 %
CANOE wins
22.7 %

Half the time they finish on the same number and the shared grid has nothing to say at all. The rest of the time the better opener wins slightly more often than it loses.

How many words would you have to play before that real advantage becomes statistically visible at 95% confidence?

2,241 words. At one puzzle a day, that is six years and two months of daily Wordle to establish that CRATE beats CANOE. And that's only the gap between two opening words. Separating two actual human players would take considerably longer.

Your group chat is not a leaderboard. It's a ritual, which is a fine thing to be, but it is not measuring anyone.

So how do you actually play together

Three families of answer, and they solve different problems.

Paste your grid. What everyone already does. No friction, no setup, play whenever you like. In exchange: asynchronous, one word a day, nobody watches anybody play, and nothing ever concludes.

Agree to play the same puzzle at the same time. Slightly better. You know you're facing identical difficulty simultaneously. But you're coordinating by hand and still trading screenshots at the end.

A shared room in real time. One code, everyone joins, same word, same six guesses, and you watch the other grids fill in while you think. This is the only one where something actually happens: rounds end, somebody wins, you play again.

The multiplayer games, honestly

I build one of these. All the more reason to say plainly what the others do well, because several of them do things I don't.

Squabble is the battle royale version. Blitz mode for a handful of players, Royale for dozens at once, join as a guest with no account. It's the most spectacular format in the category and nobody else has really matched it.

WordleOff does the straightforward thing well: create a session, share the ID, everyone races the same word.

Duordle is French, up to four players in a room, no account and no data collection. It has a cooperative mode I don't offer, where you hunt the word together instead of against each other. If that's what you want for a family evening, go there.

Tusmo runs duels on the French Motus rules, with longer words and the first letter revealed. That ruleset changes the difficulty more than people expect, which I measured in a separate article.

GlyphDuel, mine, covers ground the others don't: four languages (English, French, Spanish, German) and three word lengths (4, 5 or 6 letters) in the same room, plus a spectator mode for watching a match you're not in. There are bots too, for when you want to play right now and nobody's around.

Starting a match, concretely

On GlyphDuel:

  1. Pick a nickname, the number of players (2, 3 or 4), the word length, and optionally a 60, 90 or 120 second clock per round if you want the pressure.
  2. You get a five-character code. Drop it in the chat.
  3. Everyone types it in and lands in the room. When the last person arrives, the host starts the match.

No account, no email, no app. The room's language is whatever the interface was set to when you created it, and it stays fixed for the whole match.

What real time actually changes

Scoring is simple. Whoever solves the word first takes 7 − guesses used. Solve in three and you bank four points; solve in five and you bank two. First to ten points takes the match.

Two consequences, and the second one is the interesting one.

It's short. Running the same simulation on two comparable players, someone reaches ten points in 4.15 rounds on average. A match is four words, roughly ten minutes. You get through a lunch break what the daily puzzle would ration out over four days.

And speed becomes the tiebreak. Remember that number from earlier: more than half the time, two comparable players finish on exactly the same guess count. In the group chat those days are silent draws that vanish. In a live room they're decided by who typed the word first, which is admittedly a strange skill, but at least it names somebody.

And no, it still proves nothing

Here's the paragraph I could have quietly left out.

I simulated two hundred thousand complete matches between a CRATE player and a CANOE player, using the real scoring and the real ten-point target. CRATE, genuinely the stronger of the two, wins 52.9% of matches. Establishing that at the same confidence as before would take roughly 1,165 matches.

Eleven hundred matches of four rounds each is close to five thousand words. That's more than the 2,241 you'd need by simply comparing daily scores. The knockout-round format throws information away: it records who won each round and forgets by how much.

So if you came here looking for an instrument that settles who's better, neither format is it. No consumer format is. There's too much luck in the draw and not enough rounds in a lifetime.

Which, thinking about it, settles the question anyway. Since neither one seriously ranks anybody, you may as well pick whichever is more fun to do together. A room where you watch the other grids fill in, where someone actually wins, and where you immediately start another one, isn't a better leaderboard. It's just a better half hour.

Where the numbers come from

The English dictionary in GlyphDuel, specifically the 1,632 five-letter words eligible as answers. The solver picks whichever remaining candidate maximises Shannon entropy over the surviving field, which makes it deterministic: given the same opener it plays the same way every time. So all the variation you see comes from the word or the opener, never from the player having a good day.

Sample sizes are ordinary 95% confidence calculations from the standard deviation of the per-word difference. The match simulation draws a random answer each round, gives the round to whoever solved it in fewer guesses, and flips a coin when they tie.

Three caveats. A solver is not a person: real players scatter much more widely than this, which makes every duration above an underestimate. My answer list is not the New York Times list, so exact values would shift with a different dictionary, though the orders of magnitude wouldn't. And modelling ties as coin flips is crude, since in practice they come down to typing speed and connection as much as thinking.

Details on Squabble, WordleOff, Duordle and Tusmo come from their own sites, checked in July 2026.

One last thing

The funny part of the whole exercise was finding out that the format I'm selling measures people worse than the one I set out to debunk. I kept the number in.

A daily grid is a postcard. You send it, someone reads it later, nothing passes between you. A shared room is a phone call. Both have their place, and it turns out neither one was ever going to tell you who's best.

If you want the mechanics underneath: the best opening word, the trap words, and why short words are harder.

Ready to practice? Apply what you've learned on GlyphDuel, free multiplayer Wordle in 4 languages.