Models / Datasets / Open Source

Open-source models and datasets.

Selected speech recognition, language model, and dataset releases. The full archive is on Hugging Face.

View all on Hugging Face

Automatic Speech Recognition

Open on Hugging Face

Qwen3-ASR 0.6B — Amharic Mixed

Amharic ASR checkpoint trained for broader generalization across speech domains.

Automatic Speech Recognition

Open on Hugging Face

Qwen3-ASR 0.6B — Amharic Gold + Silver

Amharic ASR checkpoint trained on the curated gold and silver dataset mix.

Language Model

Open on Hugging Face

Gemma 2 27B — Amharic CPT

Gemma 2 27B adapted for Amharic through tokenizer expansion and continued pretraining.

Language Model

Open on Hugging Face

Gemma 2 27B — Amharic SFT

Instruction-tuned Amharic model built on the continued-pretraining checkpoint.

The personal site shows selected releases. Hugging Face remains the complete archive.

View all on Hugging Face