GPT Token Counter

Paste text to see its token count under both OpenAI tokenizers. Counting happens in your browser, so the text never leaves your machine.

GPT-4o, GPT-4.1, o-series
o200k_base tokenizer
0

tokens

GPT-4, GPT-3.5 Turbo
cl100k_base tokenizer
0

tokens

Characters
Total character count
0

characters

Words
Word count
0

words

Char Estimate
~4 chars per token
0

estimated tokens

Word Estimate
~0.75 words per token
0

estimated tokens

How to use the token counter

  1. Type or paste your text into the box. The counts update about 300 ms after you stop typing.
  2. Read the first card for GPT-4o, GPT-4.1 and the o-series models, which use the o200k_base tokenizer.
  3. Read the second card for GPT-4 and GPT-3.5 Turbo, which use cl100k_base.
  4. Use the character and word estimates only as a ballpark. They apply rules of thumb (4 characters or 0.75 words per token), not a tokenizer.

Example: the same sentence in English and Indonesian

These two sentences are almost the same length, 61 and 62 characters:

Good morning, how are you? I am learning how to count tokens.
Selamat pagi, apa kabar? Saya sedang belajar menghitung token.

The English sentence is 15 tokens under both tokenizers. The Indonesian one is 15 tokens under o200k_base and 20 under cl100k_base.

o200k_base has about twice the vocabulary of cl100k_base (roughly 200,000 entries against 100,000), so it splits non-English words into fewer pieces. Sent to GPT-4, the Indonesian sentence costs a third more tokens than it does on GPT-4o.

Frequently asked questions

Which tokenizer does my model use?

GPT-4o, GPT-4o mini, GPT-4.1 and the o1, o3 and o4-mini reasoning models use o200k_base. GPT-4, GPT-4 Turbo and GPT-3.5 Turbo use cl100k_base.

Does it count Claude or Gemini tokens?

No. Claude and Gemini use their own tokenizers, so these numbers are only a rough guide for them. For exact counts, Anthropic's API has a count_tokens endpoint and the Gemini API has countTokens.

Why is my API bill higher than this count?

Chat requests add a few tokens per message for roles and formatting, and the system prompt, tool definitions and the model's reply are billed too. This tool counts only the text you paste.

How many tokens is 1,000 words?

For English prose, roughly 1,300 to 1,400 tokens. Code, numbers and non-English text usually take more. Paste a sample of your own text to get a real number.

Is my text sent anywhere?

No. The tokenizer runs in your browser with the gpt-tokenizer library. Nothing you type is sent to a server.