Language intelligence for Odia

We are building the AI layer for 38 million Odia speakers — tokenizer optimization, language models, speech, and translation for India’s underserved languages.

Get in touch
Odisha, IndiaFounded 2026Apache 2.0
ଓଡ଼ିଆODIAଲେଖନୀLEKHANIଶ୍ରୁତିSHRUTIଅନୁବାଦANUVADAଖୋଜKHOJAପତ୍ରPATRAଅର୍ଥARTHA

ଲେଖନୀ

Lekhani — the foundation

One tokenizer. Everything stands on it.

Odia costs other models three or four tokens for a single word. That is not a rounding error — it multiplies across every request, every context window, every rupee of inference. We rebuilt the tokenizer for Odia and made it roughly 3x more efficient than generic alternatives, with Unicode awareness for Brahmic script.

Everything Maelis builds — the models, the speech, the translation, the search and the document understanding — is built on top of that one foundation. It is the only reason the rest is affordable.

ଲେଖନୀLekhani

Model family

Odia-exclusive LLMs fine-tuned on open-weight base models. Tiered from lightweight edge deployment to flagship reasoning. Apache 2.0, and not yet shipped.

ଶ୍ରୁତିShruti

Speech recognition & synthesis

Odia speech recognition and synthesis, built for government, media, and enterprise use.

ଅନୁବାଦAnuvada

Translation

Odia translation in both directions, carrying the cultural context a literal translation loses.

The foundation

Tokenizer

Odia-optimized tokenization, evaluated across the configurations we tested. Our champion tokenizer was the most efficient configuration in that sweep. Lower cost and better performance on every API call, because the arithmetic underneath everything else finally adds up.

The same vocabulary, segmented

Each rule is a token edge

The Odia language gap in AI

38 million speakers

Where every other language has an answer

Dutch25M
Greek13M
Norwegian5M
Odia38M

Dutch, Greek and Norwegian all have robust AI support. Odia has no robust, general-purpose equivalent.

The gap

  • 01Global AI models are measurably weaker at Odia than at English
  • 02No commercial Odia-specific AI API currently exists
  • 03Existing academic projects use non-commercial licenses
  • 04~38 million speakers (35M native, 2011 Census) remain underserved

How Maelis answers it

  • 01Odia-exclusive tokenizer with 3x better efficiency
  • 02Full commercial license (Apache 2.0) for enterprise use
  • 03SLA-backed API with dedicated support
  • 04Built specifically for Odisha government and business needs

We are Odia specialists, not 22-language generalists. Every part of the stack is designed for one language done exceptionally well, not many languages done adequately.

Open-source by default under Apache 2.0, with commercial licensing available for enterprises that need it. Models ship either self-hosted or through our managed API. Founded 2026, in Odisha.

Research & publications

Open by default · Apache 2.0

Our research spans tokenizer optimization and dataset construction for low-resource Indian languages. What we can show you, we show in public.

Published dataset

Screened for benchmark contamination

odia-eval-benchmark

A unified evaluation suite for Odia: 121,947 evaluation rows consolidated from 34 public Odia datasets into one comparable leaderboard across 7 task families, published under the MaelisResearch organisation so anyone can reproduce our numbers instead of taking our word for them.

Hugging Face · MaelisResearch/odia-eval-benchmark

121,947

Evaluation rows

7

Task families

34

Source datasets

CC-BY-4.0

Licence, MIT where held

  • 01Question answering
  • 02Reading comprehension
  • 03Reasoning
  • 04Summarisation
  • 05Classification
  • 06Generation
  • 07Instruction following

Research · not yet public

Efficient Odia Tokenization for Large Language Models

Empirical evaluation of tokenizer configurations for Odia, focused on Brahmic-script efficiency. Findings remain internal until compliance clearance.

Internal

Licence, model weights and paper detail are released as each clears compliance. If you need something specific for a diligence process, ask and we will tell you plainly whether we can share it yet.

One language, the whole stack

Six named surfaces, one foundation. The model family is named after the Odia literary tradition — Pada, Chhanda, Kavya, Mahakavya.

Nothing here is a finished product yet. This is the architecture we are building, in the order we are building it.

ଲେଖନୀLekhani · The writer

Model family

Odia-exclusive LLMs fine-tuned on open-weight base models, tiered from lightweight edge deployment to flagship reasoning. Apache 2.0.

ଶ୍ରୁତିShruti · The listener

Speech recognition & synthesis

Odia speech recognition and synthesis, built for government, media, and enterprise use.

ଅନୁବାଦAnuvada · The translator

Translation

Translation between Odia and other languages, with the cultural context a literal rendering loses.

ଖୋଜKhoja · The searcher

Semantic search

Semantic search and information retrieval trained specifically for Odia content.

ପତ୍ରPatra · The reader

Document understanding

Optical character recognition and document understanding for Odia text.

ଅର୍ଥArtha · The reasoner

Chain-of-thought reasoning

Advanced reasoning for complex Odia language tasks and analysis.

Two people, one mother tongue

Maelis is native Odia speakers doing the infrastructure nobody else would take on. Neither of us needed to be convinced this language deserved a real tokenizer — we had lived the absence of one.

The corpus pipeline, the tokenizer research and the evaluation benchmark were built here, by these two, and published in the open where we are able to.

01

Sai Dutta Abhishek Dash

Engineer Founder

Building language AI for his mother tongue, Odia. Background in NLP research, tokenizer optimization, and full-stack engineering. AWS Certified Cloud Practitioner. Previously built OffSage, a digital agency, and multiple open-source NLP projects.

Owns — Corpus pipeline, tokenizer research, evaluation benchmark

02

Saurav Mahalik

CTO, Kalinga Series Owner

CS engineer with hands-on experience across mobile development, machine learning, and full-stack web. Built Android applications, ML prediction models, and did deep-learning research on ECG signal analysis. Combines technical breadth with a user-first mindset from production support work.

Owns — Evaluation harness, experimentation, Kalinga series publication

Talk to us.

Whether you are reviewing a grant application, running a procurement process, or evaluating the tokenizer, we would rather answer you directly than put you through a form.

contact@sdad.proRead the full investor brief

The short record

Founded
2026
Entity
Indian Pvt Ltd, planned
Recognition
DPIIT startup, targeted
Licence
Apache 2.0
Base
Odisha, India
Named
MAY-liss

Maelis is a language infrastructure company building the AI layer for Odia and low-resource Indian languages. Founded in 2026, in Odisha, India.

Maelis Research © 2026 · Odisha, India

Open-source by default · Apache 2.0