Interactive explainer · about 5 minutes

How many r’s are in “strawberry”?

You can count them in a second. A language model never even sees the letters. Find out what it receives instead, using the real tokenizer behind GPT-4o.

Start exploring

No sign-up, nothing to install. Your input stays in your browser.

Count the r’s

Before a model gets the word, try it yourself. How many times does the letter r appear in this word?

strawberry

Your answer

This exact question became a well-known stumbling block: language models often got it wrong [1]. To see the first part of the reason, look at what a model receives.

What the model receives

A language model never receives your letters. A tokenizer first cuts the text into pieces from a fixed list, its vocabulary, and replaces every piece with a number, its token ID. Below is the question, split by o200k_base: the tokenizer that OpenAI’s library tiktoken assigns to GPT-4o and later models [2][8].

  1. How5299
  2. ␣many1991
  3. ␣r428
  4. 's885
  5. ␣are553
  6. ␣in306
  7. ␣strawberry101830
  8. ?30

Look at “strawberry”: together with the space in front of it, the whole word is one single vocabulary entry, number 101830. The model receives that one number, not ten letters.

The number is only a label, like a locker number: 101830 is not “bigger” than 428 in any way that matters. The model uses it to look up a list of learned values. Which letters hide inside 101830 is not part of the number. Whatever a model knows about the spelling of this token, it had to pick up during training.

Same word, other numbers

The vocabulary stores pieces with their exact characters. So a space, a capital letter or a different tokenizer leads to different pieces.

“ strawberry” in mid-sentence (with a space)

  1. ␣strawberry101830

“strawberry” at the very start

  1. st302
  2. raw1618
  3. berry19772

“Strawberry”

  1. Str3504
  2. aw1134
  3. berry19772

“STRAWBERRY”

  1. ST1117
  2. RAW46176
  3. B33
  4. ERRY132354

“strawberry” in GPT-4’s older tokenizer, cl100k_base

  1. str496
  2. aw675
  3. berry15717

The piece “berry” keeps the same number, 19772, in “strawberry” and “Strawberry”: the tokenizer reuses pieces. And a new model generation can come with a new tokenizer. GPT-4’s cl100k_base cuts the same word into “str”, “aw”, “berry”, with entirely different IDs [2]. An ID only means something together with its tokenizer.

Your turn

Type anything: your name, a long word, a phone number, an emoji. The real o200k_base tokenizer runs right here in your browser, and nothing you type is sent anywhere.

Or try:

The tokenizer loads when you start typing.

  1. st302
  2. raw1618
  3. berry19772

10 characters → 3 tokens

Three more surprises

Numbers are cut into threes

  1. 1237633
  2. 45619354
  3. 722

“1234567” becomes “123”, “456”, “7”: the tokenizer groups digits in threes from the left [3]. Written the usual way, the same number is grouped from the right: 1,234,567. The tokenizer’s pieces do not match those groups: the final “7” is the ones digit, while “123” mixes millions, hundred-thousands and ten-thousands. In a study, GPT-3.5 and GPT-4 calculated considerably better when numbers were grouped from the right instead, by writing them with thousands separators [4].

This emoji is not even one token

  1. F0 9F 8D102415
  2. 93241

In UTF-8, the usual way computers store text, 🍓 takes four bytes: F0 9F 8D 93. o200k_base cuts them into two tokens, “F0 9F 8D” and “93”. Neither is a whole character; only together do they make the strawberry. Because the tokenizer can always fall back to bytes, it handles any text, even text it never saw while its vocabulary was built [5].

Some languages need more tokens

LanguageSentenceUnicode characterso200k_basecl100k_base
EnglishI would like a cup of coffee, please.371010
GermanIch hätte gern eine Tasse Kaffee, bitte.401013
GreekΘα ήθελα ένα φλιτζάνι καφέ, παρακαλώ.371837
Hindiमुझे एक कप कॉफ़ी चाहिए, कृपया।301334

The same request, translated. English and German need 10 tokens each, Greek needs 18 and Hindi 13. With GPT-4’s older tokenizer, Greek needed 37 and Hindi 34. More tokens use more of the model’s limited input and, where pricing is per token, cost more; across languages, a 2023 study of tokenizers found differences of up to 15 times [6]; newer tokenizers like o200k_base narrow the gap, as the table shows, but do not close it.

So is the tokenizer to blame?

Partly. It explains why the letters are not directly in the input: the model receives 101830, not s-t-r-a-w-b-e-r-r-y. But research suggests that is not the whole story.

  • In a 2024 benchmark, most models tested seemed to know how their tokens are spelled, yet failed to use that knowledge when asked to change a word letter by letter [7].
  • A 2024 study tested eight models, including GPT-4o, on counting letters. The models mostly recognised which letters a word contains, but often miscounted them, especially letters that occur more than twice. How common a word or its tokens were made no significant difference. The authors conclude that tokenization is not the fundamental issue [1].

So counting repeated letters is a weakness of its own, on top of the tokenizer. Here is what spelling it out does, and what it does not do:

“strawberry” at the start of a text: 3 tokens (mid-sentence it is the single token 101830)

  1. st302
  2. raw1618
  3. berry19772

“s t r a w b e r r y”: 10 tokens

  1. s82
  2. ␣t260
  3. ␣r428
  4. ␣a261
  5. ␣w286
  6. ␣b287
  7. ␣e319
  8. ␣r428
  9. ␣r428
  10. ␣y342

Now every letter has a position of its own, and the r arrives as the same number, 428, three times. The letters are in the input itself; the model no longer has to recall them from inside a token. That removes the first obstacle. The second remains: the model still has to keep track of how many r’s it has passed, and that is exactly where the 2024 study found the errors. Whether a particular chatbot answers the strawberry question correctly today depends on the model and its version.

Go deeper

This page is part of KI einfach verstehen, a free, bilingual (German and English) course that explains step by step how AI works, without hype and without formulas up front.

Continue with the lessons

Built in the open

This explainer and the whole website are open source. Found a mistake or have a better example? Open an issue. If you want to find the project again later, a star on GitHub works as a bookmark.

View the code on GitHub

Questions?

Ask in the community, where other learners and the project discuss the lessons.

Go to the community

Sources

Checked on 6 October 2026. Token pieces and IDs are reproducible with tiktoken or the library listed last.

  1. Fu, Ferrando, Conde, Arriaga, Reviriego (2024): Why Do Large Language Models (LLMs) Struggle to Count Letters? arXiv:2412.18626Counting letters such as the r’s in “strawberry” is hard for LLMs. In tests of eight models including GPT-4o, the models recognised which letters a word contains but miscounted them, most of all letters that occur more than twice; word and token frequency had no significant effect, and the authors conclude that tokenization is not the fundamental issue. GPT-4o’s tokenizer splits “strawberry” into “st”, “raw”, “berry”.
  2. OpenAI tiktoken: tiktoken/model.py (model-to-encoding mapping)tiktoken assigns the o200k_base encoding to GPT-4o, GPT-4.1, o1, o3 and GPT-5, and cl100k_base to GPT-4.
  3. OpenAI tiktoken: tiktoken_ext/openai_public.py (split patterns)Before merging, cl100k_base and o200k_base cut runs of digits into groups of one to three digits (pattern \p{N}{1,3}).
  4. Singh, Strouse (2024): Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs. arXiv:2402.14903GPT-3.5 and GPT-4 use separate tokens for 1-, 2- and 3-digit numbers; enforcing right-to-left grouping with comma separators improved their arithmetic substantially.
  5. OpenAI tiktoken: READMEBPE as used in tiktoken is reversible and lossless and works on arbitrary text, even text not in the tokenizer’s training data.
  6. Petrov, La Malfa, Torr, Bibi (2023): Language Model Tokenizers Introduce Unfairness Between Languages. NeurIPS 2023, arXiv:2305.15425The same text translated into different languages can have very different tokenization lengths, up to 15 times in some cases, with consequences for cost, latency and how much context fits.
  7. Edman, Schmid, Fraser (2024): CUTE: Measuring LLMs’ Understanding of Their Tokens. EMNLP 2024, arXiv:2409.15452Most LLMs tested seem to know the spelling of their tokens, yet fail to use this information effectively to manipulate text.
  8. gpt-tokenizer 4.0.0 (MIT), a JavaScript port of tiktoken’s encodingsAll token pieces and IDs on this page were computed with this library’s o200k_base and cl100k_base encodings; the free input runs it in your browser.