← Dragon WisprTermsPrivacyRefundsLegal noticeSecurity

How we measure token savings

Last updated: 2 October 2026

  • We compare Simplified Chinese with your own language, never with English. In our test, Chinese used 40% fewer tokens than Greek on the GPT-4o / GPT-5 tokenizer.
  • For English, Chinese costs more: 13% more tokens on the same tokenizer. If you write to AI in English, Dragon mode will not save you tokens.
  • Chinese prompts can also lower answer quality on some tasks. That is why Dragon mode will be a toggle.
  • The test messages, the script and the results are published below, so you can check every number.

1. Method

  1. We wrote 12 realistic messages to an AI assistant or coding agent in English, 23 to 33 words each. The mix: 4 coding requests with English identifiers, 2 writing help, 2 analysis, 2 everyday, 1 CI/devops, 1 personal finance explanation.
  2. Each message was translated faithfully into 16 other languages and into Simplified Chinese (zh-Hans), the way Dragon mode is designed to write it. Code identifiers, file paths and product names stay in Latin script in every language (as Dragon mode will keep them). Each language uses its own natural conventions for numerals and common technical loanwords.
  3. Every message was tokenized separately with tiktoken 0.14.0 on 2 October 2026, using two open tokenizers:
    • GPT-4o / GPT-5: OpenAI o200k_base (GPT-4o, GPT-4.1, GPT-5 family)
    • GPT-4: OpenAI cl100k_base (GPT-4, GPT-3.5-turbo)
  4. For each language we add up the tokens of all 12 messages and compare: change = (tokens in Chinese − tokens in your language) / tokens in your language. A negative number means Chinese uses fewer tokens. We also report the range across the individual messages.

2. Results

Total input tokens for the same 12 messages. “Range” is the smallest and largest change for a single message on the GPT-4o / GPT-5 tokenizer.

Your languageGPT-4o / GPT-5: yours → ChineseGPT-4: yours → ChineseRange per message
Greek755 → 456 −40%1,707 → 599 −65%−50% to −30%
German523 → 456 −13%604 → 599 −1%−27% to +14%
French (no reliable saving)496 → 456 −8%575 → 599 +4%−17% to +5%
Spanish (no reliable saving)479 → 456 −5%545 → 599 +10%−18% to +11%
Italian (no reliable saving)539 → 456 −15%590 → 599 +2%−31% to +3%
Portuguese (Brazilian) (no reliable saving)476 → 456 −4%543 → 599 +10%−18% to +21%
Russian521 → 456 −13%851 → 599 −30%−26% to +3%
Ukrainian644 → 456 −29%1,069 → 599 −44%−40% to −19%
Polish633 → 456 −28%724 → 599 −17%−41% to −16%
Turkish547 → 456 −17%709 → 599 −16%−29% to −4%
Arabic493 → 456 −8%1,038 → 599 −42%−19% to +5%
Hebrew574 → 456 −21%1,355 → 599 −56%−28% to −8%
Hindi609 → 456 −25%1,734 → 599 −66%−37% to −15%
Japanese642 → 456 −29%857 → 599 −30%−40% to −18%
Korean581 → 456 −22%879 → 599 −32%−31% to −5%
Vietnamese533 → 456 −14%859 → 599 −30%−30% to +3%

“No reliable saving” means Chinese did not use fewer tokens on both tokenizers. For French, Spanish, Italian and Portuguese the saving is small or turns into a cost depending on the tokenizer, so we don’t present them as savers.

Why we never compare with English

The same messages in English took 405 tokens on the GPT-4o / GPT-5 tokenizer; in Chinese they took 456 (+13%). On GPT-4 the difference was +47%. Chinese costs more tokens than English. Independent research finds the same on Claude, GPT and Gemini tokenizers (arXiv:2604.14210). Dragon mode helps people who write to AI in a language that tokenizes poorly, like Greek; it is not for people who already write in English.

3. Caveats

  • Small corpus (12 messages per language); treat results as indicative, not as a benchmark.
  • Translations were written by hand for this measurement; a different translator or Dragon mode's live model output may be longer or shorter, especially for zh-Hans.
  • Results depend on the domain mix: messages full of code identifiers save less, because identifiers stay in Latin script and cost the same in every language.
  • Only OpenAI's open tokenizers are measured; other models (Claude, Gemini, Llama, Qwen) use different tokenizers and will give different numbers.
  • Only input tokens are measured. Model replies are not part of this comparison.
  • For English, Chinese costs MORE tokens; Dragon mode is not a token saver for English speakers.
  • Claude (Anthropic) uses its own tokenizer and is not part of these numbers yet. The script adds it when run with an Anthropic API key, using Anthropic’s free token-counting endpoint.
  • Fewer tokens is not the same as better answers. Published coding benchmarks show models solving somewhat fewer tasks when the prompt is written in Chinese rather than English, and the effect compared with your own language depends on the model and the task. Use Dragon mode where token cost matters, and keep it off where precision matters most.
  • Dragon mode is not available yet. When it launches, its live translations may be longer or shorter than the hand-written ones measured here; we will re-measure with its real output and update this page.

4. Check it yourself

  • token-corpus.json: the parallel test messages in every language.
  • measure-tokens.py: the measurement script (Python, tiktoken).
  • token-data.json: the results this page and the calculator use.

With Python 3, put the first two files in one folder and run:

python -m pip install tiktoken
python measure-tokens.py --corpus token-corpus.json --out token-data.json

Apart from the date, the output should match the published results exactly. You can paste any of the test messages into OpenAI’s tokenizer page or Anthropic’s token-counting API to spot-check a single message.

5. Updates

Measured on 2 October 2026. We update this page whenever we re-measure, and every token number on the site, in the Dragon mode calculator included, comes from the same results file.