You can count them in a second. A language model never even sees the letters. Find out what it receives instead, using the real tokenizer behind GPT-4o.
No sign-up, nothing to install. Your input stays in your browser.
st302
raw1618
berry19772
1Count the r’s
Before a model gets the word, try it yourself. How many times does the letter r appear in this word?
strawberrystrawberry
This exact question became a well-known stumbling block: language models often got it wrong [1]. To see the first part of the reason, look at what a model receives.
2What the model receives
A language model never receives your letters. A tokenizer first cuts the text into pieces from a fixed list, its vocabulary, and replaces every piece with a number, its token ID. Below is the question, split by o200k_base: the tokenizer that OpenAI’s library tiktoken assigns to GPT-4o and later models [2][8].
How5299
␣many1991
␣r428
's885
␣are553
␣in306
␣strawberry101830
?30
Look at “strawberry”: together with the space in front of it, the whole word is one single vocabulary entry, number 101830. The model receives that one number, not ten letters.
The number is only a label, like a locker number: 101830 is not “bigger” than 428 in any way that matters. The model uses it to look up a list of learned values. Which letters hide inside 101830 is not part of the number. Whatever a model knows about the spelling of this token, it had to pick up during training.
3Same word, other numbers
The vocabulary stores pieces with their exact characters. So a space, a capital letter or a different tokenizer leads to different pieces.
“ strawberry” in mid-sentence (with a space)
␣strawberry101830
“strawberry” at the very start
st302
raw1618
berry19772
“Strawberry”
Str3504
aw1134
berry19772
“STRAWBERRY”
ST1117
RAW46176
B33
ERRY132354
“strawberry” in GPT-4’s older tokenizer, cl100k_base
str496
aw675
berry15717
The piece “berry” keeps the same number, 19772, in “strawberry” and “Strawberry”: the tokenizer reuses pieces. And a new model generation can come with a new tokenizer. GPT-4’s cl100k_base cuts the same word into “str”, “aw”, “berry”, with entirely different IDs [2]. An ID only means something together with its tokenizer.
4Your turn
Type anything: your name, a long word, a phone number, an emoji. The real o200k_base tokenizer runs right here in your browser, and nothing you type is sent anywhere.
Or try:
The tokenizer loads when you start typing.
st302
raw1618
berry19772
10 characters → 3 tokens
Pieces shown as hex numbers (like F0 9F 8D) are raw bytes: only part of a character.
5Three more surprises
Numbers are cut into threes
1237633
45619354
722
“1234567” becomes “123”, “456”, “7”: the tokenizer groups digits in threes from the left [3]. Written the usual way, the same number is grouped from the right: 1,234,567. The tokenizer’s pieces do not match those groups: the final “7” is the ones digit, while “123” mixes millions, hundred-thousands and ten-thousands. In a study, GPT-3.5 and GPT-4 calculated considerably better when numbers were grouped from the right instead, by writing them with thousands separators [4].
This emoji is not even one token
F0 9F 8D102415
93241
In UTF-8, the usual way computers store text, 🍓 takes four bytes: F0 9F 8D 93. o200k_base cuts them into two tokens, “F0 9F 8D” and “93”. Neither is a whole character; only together do they make the strawberry. Because the tokenizer can always fall back to bytes, it handles any text, even text it never saw while its vocabulary was built [5].
Some languages need more tokens
Language
Sentence
Unicode characters
o200k_base
cl100k_base
English
I would like a cup of coffee, please.
37
10
10
German
Ich hätte gern eine Tasse Kaffee, bitte.
40
10
13
Greek
Θα ήθελα ένα φλιτζάνι καφέ, παρακαλώ.
37
18
37
Hindi
मुझे एक कप कॉफ़ी चाहिए, कृपया।
30
13
34
The same request, translated. English and German need 10 tokens each, Greek needs 18 and Hindi 13. With GPT-4’s older tokenizer, Greek needed 37 and Hindi 34. More tokens use more of the model’s limited input and, where pricing is per token, cost more; across languages, a 2023 study of tokenizers found differences of up to 15 times [6]; newer tokenizers like o200k_base narrow the gap, as the table shows, but do not close it.
6So is the tokenizer to blame?
Partly. It explains why the letters are not directly in the input: the model receives 101830, not s-t-r-a-w-b-e-r-r-y. But research suggests that is not the whole story.
In a 2024 benchmark, most models tested seemed to know how their tokens are spelled, yet failed to use that knowledge when asked to change a word letter by letter [7].
A 2024 study tested eight models, including GPT-4o, on counting letters. The models mostly recognised which letters a word contains, but often miscounted them, especially letters that occur more than twice. How common a word or its tokens were made no significant difference. The authors conclude that tokenization is not the fundamental issue [1].
So counting repeated letters is a weakness of its own, on top of the tokenizer. Here is what spelling it out does, and what it does not do:
“strawberry” at the start of a text: 3 tokens (mid-sentence it is the single token 101830)
st302
raw1618
berry19772
“s t r a w b e r r y”: 10 tokens
s82
␣t260
␣r428
␣a261
␣w286
␣b287
␣e319
␣r428
␣r428
␣y342
Now every letter has a position of its own, and the r arrives as the same number, 428, three times. The letters are in the input itself; the model no longer has to recall them from inside a token. That removes the first obstacle. The second remains: the model still has to keep track of how many r’s it has passed, and that is exactly where the 2024 study found the errors. Whether a particular chatbot answers the strawberry question correctly today depends on the model and its version.
Go deeper
This page is part of KI einfach verstehen, a free, bilingual (German and English) course that explains step by step how AI works, without hype and without formulas up front.
This explainer and the whole website are open source. Found a mistake or have a better example? Open an issue. If you want to find the project again later, a star on GitHub works as a bookmark.
Checked on 6 October 2026. Token pieces and IDs are reproducible with tiktoken or the library listed last.
Fu, Ferrando, Conde, Arriaga, Reviriego (2024): Why Do Large Language Models (LLMs) Struggle to Count Letters? arXiv:2412.18626Counting letters such as the r’s in “strawberry” is hard for LLMs. In tests of eight models including GPT-4o, the models recognised which letters a word contains but miscounted them, most of all letters that occur more than twice; word and token frequency had no significant effect, and the authors conclude that tokenization is not the fundamental issue. GPT-4o’s tokenizer splits “strawberry” into “st”, “raw”, “berry”.