HB hakanbogan.com
hakanbogan/gpt2-turkish-cased

GPT-2 Turkish

A cased GPT-2 pretrained on Turkish, published on Hugging Face as a base for other people to fine-tune. Anyone who wants a Turkish starting point can take the weights instead of running the pretraining themselves.

Turkish is agglutinative and its suffixes stack productively, so one word does what English spreads across a phrase. Evlerimizden means "from our houses". Nothing fixes where the stacking stops, and the set of words a corpus produces is effectively open-ended.

A vocabulary built for English handles that badly. It shreds ordinary Turkish words into fragments and buries the morpheme boundaries. So I built the vocabulary from the Turkish corpus itself. It came to 52K byte-level BPE tokens, trained with Hugging Face's Tokenizers library on the same text the model was later pretrained on.

The corpus was Turkish text from OSCAR. Training ran for five epochs on two RTX 2080 Ti cards.

I published the weights for PyTorch, TensorFlow and JAX, with safetensors alongside them, so whichever of the three someone already has installed will load it.

Downloads have run at a few hundred a month since 2022. I never announced it anywhere and never wrote it up, so whoever pulls it down found it without help from me.

It is a base model and was never instruction-tuned. A question put to it just draws more text in the same vein, sentence after sentence, and none of it answers. Fine-tuning on something narrower is what it is for, and that is why it went up.

## stack
  • Python
  • PyTorch
  • TensorFlow
  • Transformers
  • Tokenizers
→ huggingface.co/hakanbogan/gpt2-turkish-cased