Token Factory
Layer 05 · Compute & Inference — Net-new this year
Serve every open-weight and custom model behind one governed, OpenAI-compatible API, metered by the token and running on Navon metal across NVIDIA or AMD GPUs, so no prompt or output ever leaves the country and no single GPU vendor can gate us.
Capabilities
Section titled “Capabilities”- One OpenAI-compatible API for chat, embeddings, and rerank
- Open-weight model catalog, including African-language models hosted nowhere else
- Serverless shared endpoints and dedicated endpoints with an SLA
- Portable inference across NVIDIA and AMD GPUs, so silicon is a choice not a lock-in
- Optimized serving: speculative decoding, paged attention, continuous batching, quantization
- Batch inference at roughly half the price of live requests
- Fine-tuning, LoRA, and distillation, deployed in one step
- Retrieval that resolves through the Data Lake ontology
- Per-token billing, zero retention by default, usage fully observable
This product exposes an API. The reference is generated from its OpenAPI spec (openapi/token-factory.yaml).