Tachiwin multilingual tokenizer

A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.

Corpus weighting

Component Target
Modern + old exotic-language data 70%
English 10%
Spanish 10%
Code 10%

The complete available exotic-language corpus is used as the 70% anchor. Its existing modern/old composition is preserved.

Tokenizer

  • Model: BPE
  • Vocabulary target: 256,000
  • Initial alphabet: complete ByteLevel alphabet
  • ByteLevel GPT-2 regex: disabled
  • Unicode normalizer: none
  • Special tokens: 94
  • Human-language tags: 60

Corpus size

Total materialized corpus:

50,027,186 bytes (0.047 GiB)

Important training note

The Hugging Face BPE trainer does not expose an internal resumable merge-state checkpoint. The recipe therefore treats the completed tokenizer.json as the training checkpoint:

  • corpus preparation is resumable;
  • recipe/statistics/checksums are stored in recipe/;
  • if tokenizer.json already exists, subsequent runs skip BPE training;
  • an interrupted BPE computation itself must be restarted.

This avoids changing the training procedure merely to obtain artificial checkpointing.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support