See how an AI model like ChatGPT chops your prompt into tokens, which lines cost the most, and which wordy phrases to cut, with the savings measured, not guessed.
For anyone who writes prompts and pays per token: developers, prompt engineers, and anyone curious why a sentence "costs" what it does. A free LLM token counter for the terminal, using OpenAI's real tokenizers (GPT-4, GPT-4o, GPT-3.5) offline.
Download for Windows (.exe) · macOS and Linux · Project site · Docs
| Platform | How |
|---|---|
| Windows | Download token-visualizer-windows-x86_64.exe, double-click it, paste your prompt. (Unsigned, so SmartScreen may ask: More info, then Run anyway.) |
| Windows (Scoop) | scoop bucket add mattbusel https://github.com/Mattbusel/scoop-bucket; scoop install mattbusel/token-visualizer |
| macOS / Linux (Homebrew) | brew install mattbusel/tap/token-visualizer |
| macOS / Linux (script) | curl -fsSL https://raw.githubusercontent.com/Mattbusel/Token-Visualizer/main/install.sh | sh |
| Any OS with Python 3.8+ | pipx install git+https://github.com/Mattbusel/Token-Visualizer |
The downloads have the GPT tokenizers built in and work offline. More options (PowerShell one-liner, tarballs, Hugging Face tokenizers): docs/REFERENCE.md.
- Cut. The model's own tokenizer (
tiktoken) splits each line into tokens, the numbered pieces a model reads and bills for. - Count. Every line gets a token count, colored green, yellow or red, and the whole prompt gets a total and a rough cost.
- Measure. It swaps wordy phrases for short ones, squeezes extra whitespace, runs the new text through the same tokenizer, and reports the real difference.
All real output from token-visualizer 0.3.1 today.
1. A 5-line support prompt (examples/support-prompt.txt), the end of the report:
$ token-visualizer examples/support-prompt.txt -m gpt-4o
TOKEN ANALYSIS - GPT-4O tiktoken o200k_base
Total tokens: 70
Total characters: 357
...
COMPRESSION SUGGESTIONS
Verbose phrases found:
'in order to' → 'to'
'due to the fact that' → 'because'
'in the event that' → 'if'
MEASURED SAVINGS
Applying the phrase and whitespace fixes: 70 → 61 tokens (-9, 13%)
Re-tokenized with the same tokenizer; $0.0003 less per request at $0.03 per 1K.2. One wordy sentence, piped in:
$ echo "In order to help you, due to the fact that you asked." | token-visualizer
Total tokens: 14
...
Applying the phrase and whitespace fixes: 14 → 8 tokens (-6, 43%)3. Exact token boundaries. In a terminal each token is a colored chip (the GIF above); with colors off you get an indexed grid:
$ echo "In order to help you, due to the fact that you asked." | token-visualizer --no-color
TOKEN BREAKDOWN:
[0:In] [1: order] [2: to] [3: help] [4: you] [5:,] [6: due] [7: to] [8: the]
[9: fact] [10: that] [11: you] [12: asked] [13:.\n]Most tokens carry the space in front of the word, which is why order and order are different tokens.
- Get it: download the .exe above, or
brew/scoop/pipx. - Run it on your prompt:
token-visualizer prompt.txt -m gpt-4o(or double-click the .exe and paste, then Ctrl+Z and Enter). - Cut what it flags and run it again to see the new count.
token-visualizer --help lists every option with examples.
| Doc | What is in it |
|---|---|
| Reference | All options, what it checks, install details, Hugging Face tokenizers, use from Python, limitations |
| Token-Visualizer or tokenviz? | Which of the two sibling tools to use for what |
| Changelog | What changed in each release |
Anthropic publishes no Claude tokenizer, so -m claude-3-sonnet falls back to whitespace splitting and the header says so. For Llama and other open models, pass a Hugging Face model ID (from source, with transformers). Details.
MIT, see LICENSE.
