Local-first AI · Certified SLM marketplace (in development)
Every open-source AI model. One local runtime. Your data never leaves your device.
NIKITRIA is a local-first AI platform and certified marketplace of specialist Small Language Models — orchestrated on-device, with full provenance, no cloud required.
Working prototype, running today on an 8GB Apple M1 Mac, self-capped to 4GB of RAM — no cloud dependency, no server, no data leaving the machine.
Mission & vision
Sovereign AI, by default
Mission
To unify the world's open-source small language models into a single, local-first runtime — so anyone, anywhere, can run specialized AI privately on the hardware they already own.
Vision
A world where advanced AI is sovereign by default: models come to your data, not the other way around. NIKITRIA aims to be the trusted, certified marketplace and orchestration layer where every open-source SLM, skill, tool, and MCP converges — auditable, adaptable, and owned by the user.
-
Local-first
Runs entirely on hardware you own. No cloud round-trips, no third-party inference.
-
Sovereign by design
Data residency is absolute: models come to your data, never the reverse.
-
Certified & provenance-backed
Marketplace models will be adversarially tested before listing, and every answer traces back to the expert that produced it.
-
Adaptable
Planned: an AI-assisted pipeline that continuously generates new SLMs, skills, tools, and MCPs.
Prototype demo
Watch the prototype run
Orchestrator v0, running entirely on an 8GB Apple M1 Mac: a question routed across a panel of specialist experts and synthesized into one answer, with per-expert provenance, relevance scores, and real, unedited timings — no cloud in the loop.
Short public preview shown here; a full walkthrough is available to investors and design partners on request.
If the video does not load
The 45-second silent screen recording shows Orchestrator v0 receiving a question, routing it across specialist experts, and returning a synthesized answer with per-expert provenance and unedited stage timings — all running locally on the machine. A downloadable copy is available on request via the contact page.
How it works
A panel of experts, on your machine
Instead of one giant generalist model in someone else's datacenter, NIKITRIA runs a panel of small specialist models on your own hardware and combines their answers — with every step visible and attributable.
-
Ask
Your question is processed entirely on your device. Nothing is sent anywhere.
-
Route
An embedding-based router matches the question to the best-suited specialist expert — semantic matching, not keywords, and still no network calls.
-
Panel of experts
One or more specialist Small Language Models produce candidate answers within a strict on-device memory budget.
-
Synthesize with provenance
The answers are combined into a single response, and every part of it traces back to the expert that produced it.
Architecture shown at the conceptual level.
Benchmarks
Honest numbers, or no numbers
We publish only measured values, with full methodology — the way MLPerf-style benchmarks are reported: disclosed hardware, software versions, workload, and a method anyone can re-run. Until our internal benchmark rulings are finalized, every figure below is an explicit placeholder. Nothing on this page is an estimate.
Generation (decode)
[TO BE MEASURED]
on a [TO BE MEASURED] 8GB Apple M1 Mac, self-capped to 4GB of RAM
Reported separately from prompt evaluation.
Prompt eval (prefill)
[TO BE MEASURED]
on a [TO BE MEASURED] 8GB Apple M1 Mac, self-capped to 4GB of RAM
Long prompts add time-to-first-token; we report both directions.
Configuration
- Model
- Qwen2.5-1.5B-Instruct · GGUF Q4_K_M · llama-cpp-python
- Context length
- [TO BE MEASURED]
- Threads
- [TO BE MEASURED]
- Acceleration
- [TO BE MEASURED]
- Model size on disk
- ~1 GB
- Peak RAM
- [TO BE MEASURED]
Methodology
- Runs
- [TO BE MEASURED]
- Averaging
- [TO BE MEASURED]
- llama.cpp build
- [TO BE MEASURED]
- Measured on
- [TO BE MEASURED]
Decode and prefill are reported separately, and run-to-run variance will be disclosed alongside the published figures. The hardware named above is the same machine shown in the demo video.
Why placeholders instead of numbers?
Fabricated or context-free benchmark figures destroy trust with exactly the technical audience this page is written for. These values will be filled in from our benchmark rulings — hardware, build, context, thread count, run count, and date included — and not before.
A separate Track B pilot explores 1-bit models (BitNet via bitnet.cpp). Any Track B result will be published separately, with its own hardware context — never blended with the figures above.
Why this is the future
One runtime for all open-source SLMs
The open-model ecosystem is exploding, regulation is pushing AI on-premises, and the economics of small models keep improving. NIKITRIA's bet is that the missing layer is a trusted local runtime that unifies all of it. The evidence below is external research, labeled with its source — separated from anything NIKITRIA claims about itself.
-
2M+
Open models on Hugging Face
Hugging Face hosts more than two million open models — the second million arrived roughly 335 days after the first. Qwen2.5-1.5B-Instruct, the base model in our prototype, is among the most-downloaded text models. Open families now include Qwen, Phi, Gemma, Llama, SmolLM, and BitNet.
Hugging Face, 2025
-
10–30×
SLMs are the future of agentic AI
An NVIDIA Research position paper argues small language models are sufficiently powerful, inherently more suitable, and necessarily more economical for many agentic tasks — and that heterogeneous, multi-model systems are the natural architecture. It reports that serving a ~7B SLM is 10–30× cheaper in latency, energy, and FLOPs than a 70–175B LLM. That is the panel-of-experts thesis.
Belcak et al., NVIDIA Research, arXiv:2506.02153, 2025
-
On-device
The platform shift is already underway
Apple Intelligence (on-device models plus Private Cloud Compute), Microsoft's Copilot+ PCs and Phi family, and Google's Gemma and Gemini Nano all point the same direction: inference is moving onto the device.
Public product announcements, 2024–2025
-
$15–$30
Per million output tokens, frontier cloud
Frontier flagship models charge on the order of $15–$30 per million output tokens. A capable free on-device floor pushes routine inference toward zero marginal cost.
Published provider price lists, 2025
-
€35M / 7%
Regulation makes sovereignty a requirement
India's DPDP Act 2023, the EU AI Act, and GDPR are turning data locality from a preference into an obligation. The EU AI Act's Article 99(3) sets fines for prohibited practices of up to €35M or 7% of total worldwide annual turnover, whichever is higher — above GDPR's €20M / 4% tier.
EU AI Act Art. 99(3); DPDP Act 2023; GDPR
-
6× / −82%
The 1-bit efficiency frontier
Microsoft's BitNet b1.58 and bitnet.cpp run 1-bit LLMs on CPUs, with Microsoft reporting up to 6× faster inference and 82% less energy than conventional approaches — reinforcing the modest-hardware thesis and the direction of our exploratory Track B pilot.
Microsoft, BitNet b1.58 / bitnet.cpp
All figures in this section are external research, quoted with source and date. None are NIKITRIA measurements.
Business model
Five pillars, one adaptable platform
The model we are building toward. Every pillar below is a plan under active development, not a shipped product — the working prototype today is the orchestration runtime they all sit on.
-
Certified SLM marketplace
A marketplace of specialist models, each adversarially tested before certification, with creator-friendly revenue sharing to build the supply side.
-
MCP, skills, tools & apps marketplace
The same certification and distribution rails extended to MCPs, skills, tools, and SLM-powered apps.
-
Hosted inference tier
Usage-metered inference in a customer's VPC or on managed servers, for teams that want the platform without operating it.
-
BYO-server enterprise tier
Software licensing on the customer's own infrastructure — the sovereign-AI wedge. Data never leaves their perimeter.
-
Integration services
Hands-on integration as the cold-start go-to-market, productized over time into SDKs.
Adaptability as moat
The strategic thesis behind all five pillars: an AI-assisted pipeline (planned) that continuously generates new SLMs, skills, MCPs, and tools — so the platform's catalog compounds faster than any static product could.
BYO-server · on-premise
Your servers. Your models. Your data.
The planned enterprise tier licenses the entire NIKITRIA platform to run on infrastructure you already control — your servers, your VPC, your air-gapped rack. No cloud dependency, no vendor lock-in, and data that never crosses your perimeter.
-
Data sovereignty, literally
Models, routing, and synthesis all execute inside your network boundary. There is no telemetry path to us.
-
No lock-in
Built on open-source model formats (GGUF) and open protocols — your models remain yours, portable by construction.
-
Regulatory alignment
Designed for DPDP, GDPR, and EU AI Act environments where off-premises inference is a compliance risk, not a convenience.
Gartner's 2025 Hype Cycle for Government Services predicts that by 2028, 65% of governments worldwide will introduce technological-sovereignty requirements to improve independence and protect against extraterritorial regulatory interference. — Gartner, 2025
Deloitte's 8th State of AI in the Enterprise report (3,235 leaders across 24 countries, fielded Aug–Sep 2025) found 77% of companies now factor an AI solution's country of origin into vendor selection, and 58% build their AI stacks primarily with local vendors. — Deloitte AI Institute, Jan 2026
On-premises deployments dominate sovereign-AI infrastructure for the most sensitive workloads. — NextMSC
In-company fine-tuning
Your data, distilled into your own private experts
A planned NIKITRIA service: fine-tune a company's or industry's proprietary data into domain-specific Small Language Models that run on the customer's own hardware — private by construction, because the data and the model never leave the building.
-
Specialists beat generalists at the job that matters
A model trained on your processes, vocabulary, and edge cases — retaining the context generic pilots lose.
-
Fast, cheap iteration
NVIDIA's SLM research notes small models can be fine-tuned overnight rather than over weeks, and are 10–30× cheaper to serve than frontier-scale LLMs.
-
Private by construction
Training data, adapters, and inference all stay on infrastructure you control.
MIT Project NANDA's "The GenAI Divide: State of AI in Business 2025" (July 2025) found just 5% of integrated AI pilots are extracting millions in value, while the vast majority show no measurable P&L impact — attributing the gap to a "learning gap" that specialized, context-retaining models are positioned to close. (150 executive interviews, a 350-employee survey, 300 public deployments analyzed; lead author Aditya Challapally.) — MIT Project NANDA, Jul 2025
SLMs can be fine-tuned in hours and served at a fraction of frontier-LLM cost — 10–30× cheaper in latency, energy, and FLOPs. — Belcak et al., NVIDIA Research, arXiv:2506.02153, 2025
Offered initially as an integration-services engagement (pillar 5), productizing over time.
Education vertical
Curriculum-specific AI for India's classrooms — offline and affordable
Indian institutions are being asked to teach AI at national scale on constrained budgets. NIKITRIA's planned education offering pairs curriculum-specific SLMs with hardware schools already own: no GPU cluster, no per-student cloud bill, and it works with no internet at all.
-
Curriculum-specific experts
Planned: specialist models aligned to specific syllabi and courses, for student practice and faculty support.
-
Runs on existing computer labs
The same local-first runtime shown in our prototype — designed for modest, consumer-grade hardware rather than GPU clusters.
-
Offline by design
Connectivity gaps stop being a blocker: everything runs on the machine in the room.
NEP 2020 emphasizes AI-integrated, personalized learning. CBSE offers a 15-hour AI module from Class VI and AI as an elective in Classes IX–XII; NCERT has embedded AI content in Class XI CS/IP textbooks; AICTE published a model AI & Data Science curriculum (2021) and runs faculty development programs. — NEP 2020; CBSE; NCERT; AICTE
The IndiaAI Mission (₹10,371.92 crore approved March 2024, seven pillars including FutureSkills) aims to expand AI courses across UG, PG, and PhD programs. — IndiaAI Mission, Mar 2024
The honest read on the tailwind
As of Feb 2026, only ~₹400 crore of the IndiaAI Mission's ₹10,372 crore outlay had been released, and the FY2026-27 allocation was cut roughly in half. Institutional GPU and compute budgets are genuinely constrained — which strengthens, not weakens, the case for low-cost, offline SLM deployments. — Budget reporting, Feb 2026
Market opportunity
A market measured in ranges, not cherry-picked numbers
Market-size estimates for small language models diverge 3–5× across research firms, so we present the range with each firm named and dated — and we separate external projections from anything NIKITRIA claims about itself.
TAM — global SLM market
-
$0.93B (2025) → $5.45B (2032)
28.7% CAGR; 2024 base $0.74B
MarketsandMarkets, Mar 2025
-
$7.8B (2023) → $20.7B (2030)
15.1% CAGR
Grand View Research
-
$6.5B (2024)
25.7% CAGR to 2034
GM Insights
-
$6.98B (2024)
23.6% CAGR
Polaris
Estimates differ because firms scope 'SLM' differently. We cite the range rather than selecting one number.
Adjacent — edge AI
-
$1.95B (2024) → $8.91B (2030)
Edge AI software, 29.2% CAGR
Grand View Research
-
Tens of billions by 2030
Broader edge-AI market, estimates vary across firms
Multiple firms
SAM — sovereign/on-prem + India
-
$10–12B
India AI services revenue today
NASSCOM
-
~$17B by 2027
India AI market projection, 25–35% CAGR
NASSCOM–BCG
-
₹10,371.92 crore
IndiaAI Mission outlay — a policy tailwind signal
IndiaAI Mission, Mar 2024
SOM — our wedge
Bottom-up, not a top-down percentage: a targeted outreach list of ~242 privacy-sensitive and sovereign-AI organizations for design-partner LOIs, converting into the certified-marketplace beta and early BYO-server deployments.
All market figures are external vendor research — projections, not audited facts — cited with firm and date.
Competitive differentiation
Everyone wraps the same engine. Nobody integrates the stack.
Ollama, LM Studio, Jan, and GPT4All all wrap the same llama.cpp engine — raw token throughput between them is close to a wash. They differ on packaging, GUI, and API surface, not fundamental capability. None offers a certified marketplace, panel-of-experts orchestration, full provenance, and a BYO-server sovereign tier as one integrated product. That combination is what NIKITRIA is building: orchestration and provenance run in the v0 prototype today; the certified marketplace, certification, and BYO-server tier are planned.
| Product | Core engine | Marketplace | Orchestration | Provenance | BYO-server / on-prem | Certification |
|---|---|---|---|---|---|---|
| Ollama | llama.cpp | Model registry | No | No | Self-host (DIY) | No |
| LM Studio | llama.cpp | HF browse | No | No | Desktop/local + server | No |
| Jan | llama.cpp | Hub | No | No | Local (offline-first) | No |
| GPT4All | llama.cpp | Curated | No (LocalDocs RAG) | No | Local | No |
| llamafile | llama.cpp | N/A (single-file) | No | No | Local | No |
| Hugging Face | Hub | Yes (2M+ models) | No | Model cards | Hosted/inference | No |
| Cloud LLM APIs | Proprietary | No | No | No | No (cloud) | No |
| NIKITRIA | llama.cpp base + hot-swap LoRA experts | Certified SLM + MCP/skills/tools/apps (planned) | Panel-of-experts routing + synthesis | Full provenance | Yes — sovereign wedge (planned) | Adversarial testing (planned) |
What the integrated stack adds
- Certification through adversarial testing before a model is listed (planned)
- Embedding-routed orchestration and synthesis across specialist experts
- Full provenance: every answer attributable to the expert that produced it
- Creator revenue share to build the marketplace supply side (planned)
- Enterprise on-prem licensing — the sovereign wedge (planned)
- The adaptability pipeline: continuously generated SLMs, skills, MCPs, and tools (planned)
Roadmap
Where this goes — labeled honestly
One thing on this roadmap exists today; everything else is a plan with a target, and is labeled that way.
-
Working today
Orchestrator v0 prototype
Panel-of-experts routing, synthesis with provenance, fully local — running on an 8GB Apple M1 Mac, self-capped to 4GB of RAM.
-
Target: 10–12 months
Certified-marketplace beta
Certified specialist SLMs with adversarial testing and creator revenue share — the core milestone of the pre-seed plan.
-
Planned
Hosted inference tier
Usage-metered VPC/server deployments for teams that want the platform managed.
-
Planned
Enterprise BYO-server GA
General availability of the sovereign licensing tier on customer infrastructure.
-
Exploratory
Track B: 1-bit models
A pilot exploring BitNet 1-bit models via bitnet.cpp for even lower-footprint inference. Results, if any, will be reported separately with their own hardware context.
Founder
Who is behind NIKITRIA
Pranith Avula
Founder & CEO
Hyderabad, India
nikitria.admin@gmail.comOr reach out through the investor and design-partner forms on this site.
Pranith Avula is the solo founder of NIKITRIA and designed and built the Orchestrator v0 prototype end-to-end — the routing, synthesis, provenance, and packaging that this site demonstrates — before raising outside capital.
NIKITRIA is built from Hyderabad, one of India's fastest-growing deep-tech startup ecosystems, with the certification framework and marketplace design as the current focus.
Why a solo founder?
Deliberately, at this stage: v0 was specified, built, and tested by one person, which kept the architecture coherent and the burn near zero. The pre-seed round exists precisely to change that — the first hires are engineering and model-certification roles.
For investors
Request the deck
NIKITRIA is raising a pre-seed round. The deck covers the architecture, the certified-marketplace plan, the market thesis with sourced ranges, and the terms of the round.
Prefer a call? Say so in the note and we'll reply with a booking link.
Request received
Thank you — your request is in. The founder reads every submission and will reply with the deck personally.
Submissions go only to the founder, are used only to respond to you, and are never added to a mailing list. Privacy policy
For privacy-sensitive & sovereign-AI organizations
Become a design partner
We are selecting a small group of design partners — organizations whose data cannot leave their infrastructure — to shape the certified marketplace and the BYO-server tier. Design partnership starts with a non-binding letter of intent.
Interest received
Thank you — we read every submission and will follow up personally about design partnership.
Submissions go only to the founder and are used solely to follow up about design partnership. Privacy policy