GPT-2 Turkish
A cased GPT-2 pretrained on Turkish, published on Hugging Face as a base for other people to fine-tune. Anyone who wants a Turkish starting point can take the weights instead of running the pretraining themselves.
why turkish needs its own vocabulary
Turkish is agglutinative and its suffixes stack productively, so one word does what English spreads across a phrase. Evlerimizden means "from our houses". Nothing fixes where the stacking stops, and the set of words a corpus produces is effectively open-ended.
A vocabulary built for English handles that badly. It shreds ordinary Turkish words into fragments and buries the morpheme boundaries. So I built the vocabulary from the Turkish corpus itself. It came to 52K byte-level BPE tokens, trained with Hugging Face's Tokenizers library on the same text the model was later pretrained on.
how it was trained
The corpus was Turkish text from OSCAR. Training ran for five epochs on two RTX 2080 Ti cards.
I published the weights for PyTorch, TensorFlow and JAX, with safetensors alongside them, so whichever of the three someone already has installed will load it.
what it is for
Downloads have run at a few hundred a month since 2022. I never announced it anywhere and never wrote it up, so whoever pulls it down found it without help from me.
It is a base model and was never instruction-tuned. A question put to it just draws more text in the same vein, sentence after sentence, and none of it answers. Fine-tuning on something narrower is what it is for, and that is why it went up.
- Python
- PyTorch
- TensorFlow
- Transformers
- Tokenizers