Four rows of tiles, hardest to easiest: CASAS in Spanish, VACHE in French, LIGHT in English, MEINE in German

The first answer I got was Jerry

I wanted to settle a simple question. Of the four languages GlyphDuel runs in, which one is hardest?

So I wrote the script, ran it, and asked it for the largest group of Spanish words differing by a single letter. It came back with:

jerry, terry, perry, gerry, ferry, kerry, berry

Seven words. Six of them American first names.

It got worse the further I looked. My Spanish dictionary had spock, kensi, deeks, sookie and stewie in it (television characters), along with a healthy serving of plain English: world, think, where, black. German had the identical problem, with texas, harry and peter sitting near the top.

The cause was embarrassing and simple. Both dictionaries are built from a frequency corpus scraped out of film subtitles. In subtitles, character names are some of the most frequent words in the entire corpus. My script took the most frequent words first and then truncated the list to size, so I was keeping precisely the slice where all the proper nouns live, and throwing away the actual vocabulary underneath.

Two evenings of cleanup later I could run it again. Everything below is post-repair. Spanish still wins, this time for real reasons.

The whole answer is one number

Spanish has 27 letters. In Wordle it behaves like it has 16.9.

That's the finding. The rest of this article is just unpacking it.

Letters don't pull equal weight. English E fills 10.4 % of all positions; Q fills 0.22 %. A letter you never meet isn't really contributing: it adds no variety to the words you're trying to tell apart.

You can measure this properly. Shannon entropy gives the average information carried by a single letter of the dictionary, and converting it back gives the effective alphabet: how many equally likely letters would produce the same variety.

LanguageReal alphabetEffective alphabet
🇩🇪 German26 (+3 umlauts)19.8
🇬🇧 English2619.8
🇫🇷 French2617.2
🇪🇸 Spanish26 (+ñ)16.9

A Spanish player is drawing from a bag of seventeen letters. A German player draws from twenty. Three letters sounds like nothing. Spread over five positions it isn't nothing at all: less variety means more words that resemble each other, which means more moments where the colours stop telling you anything.

Where Spanish loses its variety

Vowels. Spanish has a consonant-vowel structure of almost punishing regularity:

Language4 letters5 letters6 letters
🇪🇸 Spanish50.8 %45.7 %45.6 %
🇫🇷 French45.6 %46.2 %45.5 %
🇬🇧 English38.3 %37.2 %37.8 %
🇩🇪 German36.8 %36.2 %34.9 %

A four-letter Spanish word is more than half vowels. There are five useful vowels. When half your positions can only hold five values, collisions stop being a risk and become a guarantee.

Spanish A alone takes 17.1 % of every position at four letters, nearly one letter in six. No letter in any of the four languages comes close to that.

Which is how you get CA_A: casa, cada, cara, cama, caja, caza, capa, caña, caía, cava, cala, cata. Twelve completely ordinary words, one letter apart.

Repeated letters, where English gets lucky

Second mechanism, and this is the one English handles best.

A word with a repeat gives you less information. BELLE tests three distinct letters instead of five: a whole guess spent on three facts.

Language4 letters5 letters6 letters
🇪🇸 Spanish30.4 %39.8 %58.6 %
🇫🇷 French19.2 %36.5 %56.2 %
🇩🇪 German23.9 %38.4 %54.8 %
🇬🇧 English18.4 %30.5 %49.1 %

English is cleaner at every length, and its consonant range is why: it can build long words without doubling up.

But look at that right-hand column again before feeling smug. Even in English, roughly half of all six-letter words contain a repeat. If you're mentally ruling out doubles, you're ruling out half the dictionary.

Very concrete consequence: never assume all the letters are distinct, especially at six letters. You'd be wrong more than half the time, and the assumption makes you discard the right answer without ever looking at it.

Dead letters

Every language carries letters that barely exist in ordinary words. Playing one throws away a tile.

LanguageEffectively unusable (5 letters)
🇫🇷 FrenchW (0.07 %), Q (0.27 %)
🇩🇪 GermanQ (0.10 %), X (0.20 %)
🇪🇸 SpanishX (0.15 %), Ñ (0.25 %), Q (0.26 %)
🇬🇧 EnglishQ (0.22 %), J (0.25 %)

French W is the extreme: seven positions in ten thousand. Putting it in an opener means playing with four tiles instead of five.

I'm fond of Spanish Ñ at 0.25 %. It's the language's signature letter, the one that goes on the posters, and it's statistically invisible.

The ranking

All of it funnels into one measure: the share of words with five or more neighbours, the words that reasoning cannot get you out of.

RankLanguage4 letters5 letters6 letters
1🇪🇸 Spanish65.9 %26.6 %6.4 %
2🇫🇷 French47.6 %19.0 %3.3 %
3🇬🇧 English47.4 %15.2 %5.2 %
4🇩🇪 German31.5 %13.8 %9.2 %

Two Spanish words in three at four letters. That's a lot.

One nice exception: at six letters the order flips and German becomes the hardest. Its verb-inflection clusters (_EINEN, _IESEN, _ASSEN, eight or nine words each) sail straight through the extra length, while the French and Spanish ones dissolve.

The German paradox

A few months ago I wrote why German is hard to learn: three genders, four cases, compounds that run off the edge of the page. And here it is, the easiest language in the game at four and five letters.

Not a contradiction. One cause, seen from two sides. Everything that makes German punishing for a learner (heavy morphology, long words, few vowels) is exactly what makes its words easy to tell apart. A German word carries a lot of information. Burden for the student, gift for the guesser.

Worth saying plainly, because the intuition really does mislead here: learning difficulty and guessing difficulty are unrelated. One is grammar. The other is nothing but the shape of the dictionary.

Where the numbers come from

GlyphDuel's validation dictionaries: around 25,000 words, split across French (1,048 / 1,808 / 2,631 words at four, five and six letters), English, Spanish and German.

Effective alphabet is 2^H, where H is the Shannon entropy of the letter distribution across all words of a given length. Neighbour counts use Hamming-1 distance: two words are neighbours if they differ by exactly one letter at one position.

Three caveats while I'm here. The measurements run over validation lists rather than secret-word pools; English and French draw secrets from a smaller curated set. The Spanish cleanup is solid, because I could cross the corpus against a real Spanish lexicon that carries the conjugations but no proper nouns. The German cleanup is weaker: the available German lexicon contained texas and harry itself, so I had to list proper nouns by hand, and there are certainly English words I failed to catch. If German moves in this ranking, that's where it'll come from.

One last thing

Spanish is the hardest Wordle language and the reason fits in a sentence: it uses too few distinct letters. Effective alphabet of 16.9, over half vowels in short words, an A holding one position in six. Its words simply don't have enough ways to differ from one another.

German is the easiest for the mirror-image reason, which happens to be the same thing that makes it miserable to learn.

Maximum difficulty, if you want it, is identified: Spanish, four letters. I'm not going back.

Same series: the trap words and how to escape them and why short words are harder than long ones. For letter-by-letter frequencies, start here.

Ready to practice? Apply what you've learned on GlyphDuel, free multiplayer Wordle in 4 languages.