by Haytham ElFadeel - hfadeelm@gmail.com
published 2024, updated 2025
2025 Update:
I added a section on H-Net / Dynamic Chunking (2025), a newer end-to-end approach that learns segmentation jointly with the model.
Tokenization is useful because it significantly decreases training and inference cost by shortening the effective sequence length seen by the main model. However, tokenization also comes with well-known drawbacks: sensitivity to noise, weaker character/number handling, representational biases, and extra system complexity.
This article surveys why people want to remove tokenization, and summarizes a few research directions toward tokenizer-free (or tokenizer-lite) language modeling.
Motivation
What
Modern LLM pipelines typically look like:
- String → bytes (e.g., UTF-8)
- Bytes → tokens (usually subwords via BPE / Unigram LM / WordPiece-style schemes)
- Tokens → model (embedding + “global” sequence model)
- Model → tokens → bytes → string (detokenization / decoding)
Tokenization is primarily a compression layer for sequence modeling: fewer “symbols” means fewer steps for the main network, which matters a lot when your backbone has quadratic cost (e.g. attention) or large per-token FFN cost.
Most modern tokenizers (BPE / unigram variants) create tokens based on substring frequency (e.g., car will most likely be one token, but frequency could be “freq”, “ue”, and “ncy”)
Why we want to remove it
The right unit of computation and “concepts”
Subword tokenization is a heuristic compromise: it compresses text well, but it hard-codes a particular segmentation that may not align with semantics. Humans don’t think in subwords; humans think in concepts (e.g., Beyoncé — not “Bey”, “once”, “é”). Concepts are language- and modality-agnostic and often represent higher-level structure.
If a model could learn its own units of computation (and do so in a way that’s stable and efficient), it could allocate compute where it matters and potentially learn abstractions that generalize better across domains and languages.
Numbers
For numbers, substring-based tokenization often produces mixed / irregular chunking: adjacent integers can tokenize differently (e.g., 480 as one token but 481 split), which breaks digit-level regularities and makes “+1” and carry behavior harder to internalize. Algorithmic operations (addition, comparison, rounding) are much easier when token boundaries align with the algorithm’s structure (digits, place values, carries). When boundaries are inconsistent, the model is nudged toward memorizing frequent token sequences rather than learning systematic rules. Pretraining frequency compounds this: rare numbers and formats receive fewer gradient updates, yielding noisier, less reliable representations and weaker generalization.
Character-level understanding
Tokenization can make character-level tasks awkward because the model is operating on a non-uniform, tokenizer-induced alphabet. Two superficially similar strings may map to very different token sequences; conversely, very different strings might share token fragments. This can hurt tasks that require manipulating orthography (spelling) or exact numerics.
For example, LLaMA 3 reportedly scores ~27% on CUTE (simple character manipulation tasks), while scoring ~79% on HellaSwag (commonsense NLI) [BLT paper, GPT-3 example].
Robustness and adversarial attacks
Tokenization adds a brittle boundary between “raw text” and “what the model actually sees”. Small perturbations (whitespace, Unicode variants, rare merges) can produce large changes in tokenization, creating an attack surface and also a source of evaluation/train-test mismatch.
Examples:
- SolidGoldMagikarp: some tokens are much more common during tokenizer training than in LM training, leading to weird activation regimes at test time [LessWrong blog].
- Trailing whitespaces: seemingly harmless formatting can create surprising distribution shifts [Scottlogic blog].
- Adding noise to eval data like HellaSwag reportedly causes large drops for some tokenized models (BLT discusses noise robustness results) [BLT paper].
End-to-end training
Tokenization means the system is not fully end-to-end: there is a separate learned (or partially learned) component optimized for a different objective (compression / likelihood under a tokenizer training corpus), plus a hard non-differentiable boundary (string → ids). Ideally, we want the model to learn segmentation and representation jointly with the language modeling objective.
Research
A lot of tokenizer-free work follows a common pattern:
- Operate on bytes / characters (no fixed vocab).
- Use a local model to encode fine-grained inputs.
- Apply a chunking / grouping mechanism to produce a shorter sequence.
- Run a global model on the compressed representation.
- Use a local model to decode back to byte-level outputs.
MegaByte (paper)
MegaByte proposes modeling bytes instead of tokens using a multiscale Transformer. It groups bytes into fixed-size patches of size P (often 4 or 8). Each patch is embedded and fed into a global Transformer operating over the patch sequence. The global output is then consumed by a smaller local decoder that autoregressively produces byte-level logits.
A useful way to view MegaByte is: fixed-rate downsampling of a byte stream, where the global model sees length ~L/P instead of L.
MambaByte (paper)
MambaByte improves compute efficiency by replacing the local Transformer blocks with Mamba / SSM blocks, avoiding quadratic attention cost in the local path and improving scalability for long byte sequences.
SpaceByte (paper)
Early work focused heavily on compute, but used arbitrary chunking (e.g., fixed-size). SpaceByte takes a step toward linguistically meaningful segmentation by grouping bytes based on spacelike characters rather than fixed size.
From the paper (paraphrased): a byte is “spacelike” if it does not encode a letter, number, or UTF-8 continuation byte; apply global blocks after spacelike boundaries (and BOS). Intuitively, this makes the global model operate closer to “word starts”.
SpaceByte was one of the first tokenizer-less LMs to approach tokenized baselines while staying practical.
Byte Latent Transformer / BLT (paper)
Space-based segmentation is a good first step, but not all regions of text carry the same information density. BLT uses entropy-based grouping to dynamically segment bytes into patches, allocating more compute where next-byte prediction is uncertain.
BLT computes next-byte entropy under a small auxiliary byte LM:
and uses that entropy signal to decide patch boundaries (e.g., above a threshold, or “high relative to previous entropy”, which encourages roughly monotone entropy decrease within a patch).
BLT shows strong results on character-level tasks (e.g., CUTE), noisy inputs, and scaling trends suggesting patching can be a better compute-allocation primitive than fixed tokenization.
Large Concept Model / LCM (paper)
Another direction is to bypass token-level modeling and operate directly on higher-level units (e.g., sentences / “concepts”). LCM introduces:
- a sentence encoder mapping text → fixed-size embedding,
- a concept-language model predicting the next embedding,
- and a decoder mapping embeddings → surface text.
Advantages:
- Direct modeling at a higher level of abstraction.
- Efficient handling of long contexts (fewer steps at the concept level).
- Potentially better generalization and multimodality.
Open questions:
- Granularity: sentences vary a lot in complexity; a single fixed unit may be too coarse.
- Specificity: if the concept model doesn’t observe exact characters/tokens, will it struggle with precise reasoning, exact numbers, or faithful copying?
H-Net / Dynamic Chunking (2025) (paper)
A limitation of SpaceByte/BLT-style methods is that segmentation is still driven by external heuristics or auxiliary predictors (whitespace rules, entropy from a separate LM). H-Net pushes further toward a true end-to-end system by learning content- and context-dependent segmentation jointly with the model.
High-level idea: replace the implicit “tokenize → LM → detokenize” stack with a single hierarchical architecture that (1) reads raw bytes, (2) dynamically chunks them into a shorter latent sequence, (3) runs a powerful backbone on the latent chunks, and (4) “dechunks” back to byte resolution—while keeping the whole process differentiable enough to train stably.
Key points:
- Explicit hierarchy (U-Net-like): small encoder → chunking (downsample) → main network on compressed sequence → dechunking (upsample) → decoder.
- Dynamic Chunking (DC): a learned mechanism that predicts boundaries between adjacent elements (a “router”) plus a smoothing / interpolation step that reduces instability from hard discrete boundary decisions.
- Recursive / multi-stage hierarchies: the “main network” can itself be another H-Net stage, enabling multiple abstraction levels (bytes → subword-like chunks → higher-order chunks).
- Reported results: compute- and data-matched byte-level H-Nets outperform strong BPE-tokenized Transformer baselines; deeper hierarchies improve scaling and robustness, with particularly strong gains on domains where tokenization heuristics are weaker (e.g., Chinese, code, even DNA sequences).
Why this matters in the broader trend: H-Net is a concrete attempt at learning the segmentation operator itself without hand-designed rules, while still matching tokenized systems on efficiency.
Future: promise and open questions
We’re getting closer to removing tokenization entirely. My (speculative) predictions:
- By 2026 / 2027: ~ 25% of models may abandon classical tokenization.
- By 2028 / 2029: > 50% of models may abandon classical tokenization.
Open research questions:
- Learned grouping strategies: entropy and whitespace are good baselines, but end-to-end learned chunking (e.g., H-Net) raises new questions about stability, inductive bias, and what “good” boundaries look like.
- Scaling laws: do tokenizer-free models scale as reliably as tokenized ones across data regimes and architectures?
- Inference mechanics: how do we do fast streaming generation when chunking is dynamic (and potentially depends on context in non-trivial ways)?
- Multimodal unification: can the same chunking/latent interface work cleanly for text, audio, vision tokens, and structured data?