Building language AI for his mother tongue, Odia. Background in NLP research, tokenizer optimization, and full-stack engineering. AWS Certified Cloud Practitioner. Previously built OffSage, a digital agency, and multiple open-source NLP projects.
Language intelligence for Odia
We are building the AI layer for 38 million Odia speakers — tokenizer optimization, language models, speech, and translation for India’s underserved languages.
ଲେଖନୀ
One tokenizer. Everything stands on it.
Odia costs other models three or four tokens for a single word. That is not a rounding error — it multiplies across every request, every context window, every rupee of inference. We rebuilt the tokenizer for Odia and made it roughly 3x more efficient than generic alternatives, with Unicode awareness for Brahmic script.
Everything Maelis builds — the models, the speech, the translation, the search and the document understanding — is built on top of that one foundation. It is the only reason the rest is affordable.
Odia-exclusive LLMs fine-tuned on open-weight base models. Tiered from lightweight edge deployment to flagship reasoning. Apache 2.0, and not yet shipped.
Odia speech recognition and synthesis, built for government, media, and enterprise use.
Odia translation in both directions, carrying the cultural context a literal translation loses.
Tokenizer
Odia-optimized tokenization, evaluated across the configurations we tested. Our champion tokenizer was the most efficient configuration in that sweep. Lower cost and better performance on every API call, because the arithmetic underneath everything else finally adds up.
The Odia language gap in AI
38 million speakers
- Global AI models are measurably weaker at Odia than at English
- No commercial Odia-specific AI API currently exists
- Existing academic projects use non-commercial licenses
- ~38 million speakers (35M native, 2011 Census) remain underserved
- Odia-exclusive tokenizer with 3x better efficiency
- Full commercial license (Apache 2.0) for enterprise use
- SLA-backed API with dedicated support
- Built specifically for Odisha government and business needs
We are Odia specialists, not 22-language generalists. Every part of the stack is designed for one language done exceptionally well, not many languages done adequately.
Open-source by default under Apache 2.0, with commercial licensing available for enterprises that need it. Models ship either self-hosted or through our managed API. Founded 2026, in Odisha.
Research & publications
Our research spans tokenizer optimization and dataset construction for low-resource Indian languages. What we can show you, we show in public.
odia-eval-benchmark
A unified evaluation suite for Odia: 121,947 evaluation rows consolidated from 34 public Odia datasets into one comparable leaderboard across 7 task families, published under the MaelisResearch organisation so anyone can reproduce our numbers instead of taking our word for them.
Hugging Face · MaelisResearch/odia-eval-benchmark121,947
7
34
CC-BY-4.0
- Question answering
- Reading comprehension
- Reasoning
- Summarisation
- Classification
- Generation
- Instruction following
Efficient Odia Tokenization for Large Language Models
Empirical evaluation of tokenizer configurations for Odia, focused on Brahmic-script efficiency. Findings remain internal until compliance clearance.
Licence, model weights and paper detail are released as each clears compliance. If you need something specific for a diligence process, ask and we will tell you plainly whether we can share it yet.
One language, the whole stack
Six named surfaces, one foundation. The model family is named after the Odia literary tradition — Pada, Chhanda, Kavya, Mahakavya.
Nothing here is a finished product yet. This is the architecture we are building, in the order we are building it.
Odia-exclusive LLMs fine-tuned on open-weight base models, tiered from lightweight edge deployment to flagship reasoning. Apache 2.0.
Odia speech recognition and synthesis, built for government, media, and enterprise use.
Translation between Odia and other languages, with the cultural context a literal rendering loses.
Semantic search and information retrieval trained specifically for Odia content.
Optical character recognition and document understanding for Odia text.
Advanced reasoning for complex Odia language tasks and analysis.
Two people, one mother tongue
Maelis is native Odia speakers doing the infrastructure nobody else would take on. Neither of us needed to be convinced this language deserved a real tokenizer — we had lived the absence of one.
The corpus pipeline, the tokenizer research and the evaluation benchmark were built here, by these two, and published in the open where we are able to.
CS engineer with hands-on experience across mobile development, machine learning, and full-stack web. Built Android applications, ML prediction models, and did deep-learning research on ECG signal analysis. Combines technical breadth with a user-first mindset from production support work.
Talk to us.
Whether you are reviewing a grant application, running a procurement process, or evaluating the tokenizer, we would rather answer you directly than put you through a form.
contact@sdad.proRead the full investor brief- 2026
- Indian Pvt Ltd, planned
- DPIIT startup, targeted
- Apache 2.0
- Odisha, India
- MAY-liss
Maelis is a language infrastructure company building the AI layer for Odia and low-resource Indian languages. Founded in 2026, in Odisha, India.