A full-stack AI research project built from first principles — corpus curation, custom tokenizer, model training from scratch, autonomous agent, and deployed enterprise product. Built by someone who has spent years closing enterprise deals across Southeast Asia and decided that selling technology was no longer enough.
Procurement cycles in SEA run 12-24 months. Buying decisions involve stakeholders who rarely appear in org charts and relationships that take years to build. Western AI models have no context for this. They were not trained on it.
Every major foundation model was trained on English-dominant web data. The enterprise intelligence that matters in SEA — conglomerate governance structures, ASEAN regulatory frameworks, Bahasa-inflected business communication, regional procurement customs — is systematically absent from every model in production today.
Enterprise buyers in SEA can tell when AI doesn't understand their world. Generic responses to specific questions destroy trust faster than any competitor can. Domain-specific intelligence is not a nice-to-have. It is the product.
22.7 million records curated from academic, financial, and enterprise sources. Every source chosen deliberately. Every record filtered for quality and domain relevance.
32,000 vocabulary items. Trained so that enterprise terminology — conglomerate governance, regional procurement frameworks, enterprise technology adoption patterns — is represented at the token level, not fragmented into meaningless subwords.
355 million parameters. GPT architecture. Trained from random weights — not fine-tuned, not a wrapper. Completed at step 100,000 with final loss 0.2952.
1.4B parameters. GPT-NeoX architecture. Planned as the next phase: training on H100 80GB HBM3 in bfloat16 over the 57GB corpus (33 sources) with the custom BPE tokenizer.
Applied proof-of-concept. An enterprise pitch simulator that generates a 5-person executive buying committee grounded in actual SEA enterprise buying behaviour, cultural context, and decision-making patterns.
try it →Autonomous agent running live. Posts original analysis on enterprise AI adoption in SEA. Builds a timestamped public track record. Not a chatbot — a reasoning system with memory, goals, and a perspective.
follow the agent →355M parameters. 100,000 steps. Loss: 0.2952. Done.
Supervised fine-tuning on enterprise intelligence. The base model learns to reason about this domain specifically.
H100 80GB HBM3. 1.4B params. bfloat16. Planned next phase.
The model deployed and queryable. Demonstrably better on SEA enterprise context than anything currently available.
The models that will matter in Southeast Asia are the ones that understand how enterprise decisions actually get made here.
That understanding cannot be scraped from the web. It cannot be approximated from English-language benchmarks. It has to be earned in the field — in the meetings, the negotiations, and the deals that never make it into any dataset.
maverickai-sea is an attempt to encode that understanding at the weight level.
Not a product. Not a startup. A proof that the gap between enterprise domain knowledge and model architecture can be closed by a single person who has lived on both sides of it.