Skip to content
AIBuilding

Hayaan — Somali Multimodal AI

A Somali-first multimodal AI system I'm building to better understand and work with Somali language, text, images, video, and long-form documents.

PythonPyTorchTransformers

Why I'm building Hayaan#

Hayaan is an experiment to explore what a Somali-first AI system could look like when Somali is treated as a first-class language rather than an afterthought.

The goal is to build a system that can work across different forms of information, including Somali text, images, video, and long-form documents such as PDFs.

Where it is now#

Hayaan is still being built.

The core idea is larger than the current implementation, and there is still significant engineering and research work ahead.

The current focus is on understanding what is required to make the system reliable, useful, and genuinely capable with Somali.

The context challenge#

Building a capable Somali AI system is not only a model problem.

It requires:

  • high-quality Somali data
  • appropriate evaluation datasets
  • research into Somali language understanding
  • multimodal data
  • reliable training and inference infrastructure
  • significant compute
  • careful evaluation

The availability and quality of these resources will heavily influence what Hayaan can eventually become.

Long-context and multimodal exploration#

One of the directions I'm exploring is long-context understanding, with a target context window of around 1M tokens.

The goal is to eventually investigate how Hayaan could reason over large amounts of information rather than treating every interaction as a short prompt.

I'm also exploring multimodal capabilities across:

  • text
  • images
  • video
  • PDFs and other long-form documents

These capabilities are still under development and should be treated as research goals rather than finished capabilities.

What I need to solve#

There are several major challenges ahead.

Compute#

Training and evaluating capable multimodal models requires substantial compute.

Hayaan will need significantly more computational resources as the project moves beyond experimentation.

Infrastructure#

At larger scales, infrastructure becomes a major part of the problem.

I'm interested in eventually exploring dedicated AI infrastructure and data-center capacity that could support training, evaluation, and inference.

Data#

A Somali-first system needs better data.

This includes finding, cleaning, structuring, and evaluating high-quality Somali datasets while being careful about data quality, licensing, privacy, and representation.

Evaluation#

A model appearing to understand Somali is not enough.

Hayaan needs meaningful evaluation.

I want to investigate benchmarks that measure things such as:

  • Somali language understanding
  • reasoning
  • translation
  • long-context comprehension
  • document understanding
  • multimodal understanding
  • factuality
  • robustness

The evaluation work is as important as the model itself.

Research questions#

Some of the questions I'm currently interested in:

  • How well can modern multimodal architectures handle Somali?
  • What kinds of Somali data are most valuable for training?
  • How should Somali-specific evaluation datasets be designed?
  • How much does additional context actually improve Somali reasoning?
  • What happens when long documents contain Somali text, tables, images, and other information?
  • What infrastructure would be required to operate a system like this at meaningful scale?

I don't have all the answers yet.

That's the point of the Lab.

Current status#

Building

Hayaan is an active experiment.

There is still substantial work to do before I can make strong claims about its capabilities.

What comes next#

The next stages will focus on research, data, evaluation, infrastructure, and compute.

I'll document meaningful progress here as the system develops.

For now, Hayaan is an experiment — not a finished product.

#AI#Somali AI#Multimodal AI#Language Models#Research