Local-first AI · Certified SLM marketplace (in development)

Every open-source AI model. One local runtime. Your data never leaves your device.

NIKITRIA is a local-first AI platform and certified marketplace of specialist Small Language Models — orchestrated on-device, with full provenance, no cloud required.

Working prototype, running today on an 8GB Apple M1 Mac, self-capped to 4GB of RAM — no cloud dependency, no server, no data leaving the machine.

Mission & vision

Sovereign AI, by default

Mission

To unify the world's open-source small language models into a single, local-first runtime — so anyone, anywhere, can run specialized AI privately on the hardware they already own.

Vision

A world where advanced AI is sovereign by default: models come to your data, not the other way around. NIKITRIA aims to be the trusted, certified marketplace and orchestration layer where every open-source SLM, skill, tool, and MCP converges — auditable, adaptable, and owned by the user.

  • Local-first

    Runs entirely on hardware you own. No cloud round-trips, no third-party inference.

  • Sovereign by design

    Data residency is absolute: models come to your data, never the reverse.

  • Certified & provenance-backed

    Marketplace models will be adversarially tested before listing, and every answer traces back to the expert that produced it.

  • Adaptable

    Planned: an AI-assisted pipeline that continuously generates new SLMs, skills, tools, and MCPs.

Prototype demo

Watch the prototype run

Orchestrator v0, running entirely on an 8GB Apple M1 Mac: a question routed across a panel of specialist experts and synthesized into one answer, with per-expert provenance, relevance scores, and real, unedited timings — no cloud in the loop.

Short public preview shown here; a full walkthrough is available to investors and design partners on request.

If the video does not load

The 45-second silent screen recording shows Orchestrator v0 receiving a question, routing it across specialist experts, and returning a synthesized answer with per-expert provenance and unedited stage timings — all running locally on the machine. A downloadable copy is available on request via the contact page.

How it works

A panel of experts, on your machine

Instead of one giant generalist model in someone else's datacenter, NIKITRIA runs a panel of small specialist models on your own hardware and combines their answers — with every step visible and attributable.

  1. Ask

    Your question is processed entirely on your device. Nothing is sent anywhere.

  2. Route

    An embedding-based router matches the question to the best-suited specialist expert — semantic matching, not keywords, and still no network calls.

  3. Panel of experts

    One or more specialist Small Language Models produce candidate answers within a strict on-device memory budget.

  4. Synthesize with provenance

    The answers are combined into a single response, and every part of it traces back to the expert that produced it.

Panel-of-experts flow diagram Diagram: a query flows into an on-device router, which dispatches it to specialist experts; their answers merge in a synthesis step that returns one response with provenance. A boundary line marks that everything happens on the device. Everything inside this line runs on your device — no cloud Query Router embedding match Expert A specialist SLMs Expert B specialist SLMs Expert C specialist SLMs Synthesis one answer Provenance who said what

Architecture shown at the conceptual level.

Benchmarks

Honest numbers, or no numbers

We publish only measured values, with full methodology — the way MLPerf-style benchmarks are reported: disclosed hardware, software versions, workload, and a method anyone can re-run. Until our internal benchmark rulings are finalized, every figure below is an explicit placeholder. Nothing on this page is an estimate.

Generation (decode)

[TO BE MEASURED]

on a [TO BE MEASURED] 8GB Apple M1 Mac, self-capped to 4GB of RAM

Reported separately from prompt evaluation.

Prompt eval (prefill)

[TO BE MEASURED]

on a [TO BE MEASURED] 8GB Apple M1 Mac, self-capped to 4GB of RAM

Long prompts add time-to-first-token; we report both directions.

Configuration

Model
Qwen2.5-1.5B-Instruct · GGUF Q4_K_M · llama-cpp-python
Context length
[TO BE MEASURED]
Threads
[TO BE MEASURED]
Acceleration
[TO BE MEASURED]
Model size on disk
~1 GB
Peak RAM
[TO BE MEASURED]

Methodology

Runs
[TO BE MEASURED]
Averaging
[TO BE MEASURED]
llama.cpp build
[TO BE MEASURED]
Measured on
[TO BE MEASURED]

Decode and prefill are reported separately, and run-to-run variance will be disclosed alongside the published figures. The hardware named above is the same machine shown in the demo video.

Why placeholders instead of numbers?

Fabricated or context-free benchmark figures destroy trust with exactly the technical audience this page is written for. These values will be filled in from our benchmark rulings — hardware, build, context, thread count, run count, and date included — and not before.

A separate Track B pilot explores 1-bit models (BitNet via bitnet.cpp). Any Track B result will be published separately, with its own hardware context — never blended with the figures above.

Why this is the future

One runtime for all open-source SLMs

The open-model ecosystem is exploding, regulation is pushing AI on-premises, and the economics of small models keep improving. NIKITRIA's bet is that the missing layer is a trusted local runtime that unifies all of it. The evidence below is external research, labeled with its source — separated from anything NIKITRIA claims about itself.

  • 2M+

    Open models on Hugging Face

    Hugging Face hosts more than two million open models — the second million arrived roughly 335 days after the first. Qwen2.5-1.5B-Instruct, the base model in our prototype, is among the most-downloaded text models. Open families now include Qwen, Phi, Gemma, Llama, SmolLM, and BitNet.

    Hugging Face, 2025

  • 10–30×

    SLMs are the future of agentic AI

    An NVIDIA Research position paper argues small language models are sufficiently powerful, inherently more suitable, and necessarily more economical for many agentic tasks — and that heterogeneous, multi-model systems are the natural architecture. It reports that serving a ~7B SLM is 10–30× cheaper in latency, energy, and FLOPs than a 70–175B LLM. That is the panel-of-experts thesis.

    Belcak et al., NVIDIA Research, arXiv:2506.02153, 2025

  • On-device

    The platform shift is already underway

    Apple Intelligence (on-device models plus Private Cloud Compute), Microsoft's Copilot+ PCs and Phi family, and Google's Gemma and Gemini Nano all point the same direction: inference is moving onto the device.

    Public product announcements, 2024–2025

  • $15–$30

    Per million output tokens, frontier cloud

    Frontier flagship models charge on the order of $15–$30 per million output tokens. A capable free on-device floor pushes routine inference toward zero marginal cost.

    Published provider price lists, 2025

  • €35M / 7%

    Regulation makes sovereignty a requirement

    India's DPDP Act 2023, the EU AI Act, and GDPR are turning data locality from a preference into an obligation. The EU AI Act's Article 99(3) sets fines for prohibited practices of up to €35M or 7% of total worldwide annual turnover, whichever is higher — above GDPR's €20M / 4% tier.

    EU AI Act Art. 99(3); DPDP Act 2023; GDPR

  • 6× / −82%

    The 1-bit efficiency frontier

    Microsoft's BitNet b1.58 and bitnet.cpp run 1-bit LLMs on CPUs, with Microsoft reporting up to 6× faster inference and 82% less energy than conventional approaches — reinforcing the modest-hardware thesis and the direction of our exploratory Track B pilot.

    Microsoft, BitNet b1.58 / bitnet.cpp

All figures in this section are external research, quoted with source and date. None are NIKITRIA measurements.

Business model

Five pillars, one adaptable platform

The model we are building toward. Every pillar below is a plan under active development, not a shipped product — the working prototype today is the orchestration runtime they all sit on.

  1. Certified SLM marketplace

    A marketplace of specialist models, each adversarially tested before certification, with creator-friendly revenue sharing to build the supply side.

  2. MCP, skills, tools & apps marketplace

    The same certification and distribution rails extended to MCPs, skills, tools, and SLM-powered apps.

  3. Hosted inference tier

    Usage-metered inference in a customer's VPC or on managed servers, for teams that want the platform without operating it.

  4. BYO-server enterprise tier

    Software licensing on the customer's own infrastructure — the sovereign-AI wedge. Data never leaves their perimeter.

  5. Integration services

    Hands-on integration as the cold-start go-to-market, productized over time into SDKs.

Adaptability as moat

The strategic thesis behind all five pillars: an AI-assisted pipeline (planned) that continuously generates new SLMs, skills, MCPs, and tools — so the platform's catalog compounds faster than any static product could.

BYO-server · on-premise

Your servers. Your models. Your data.

The planned enterprise tier licenses the entire NIKITRIA platform to run on infrastructure you already control — your servers, your VPC, your air-gapped rack. No cloud dependency, no vendor lock-in, and data that never crosses your perimeter.

  • Data sovereignty, literally

    Models, routing, and synthesis all execute inside your network boundary. There is no telemetry path to us.

  • No lock-in

    Built on open-source model formats (GGUF) and open protocols — your models remain yours, portable by construction.

  • Regulatory alignment

    Designed for DPDP, GDPR, and EU AI Act environments where off-premises inference is a compliance risk, not a convenience.

Gartner's 2025 Hype Cycle for Government Services predicts that by 2028, 65% of governments worldwide will introduce technological-sovereignty requirements to improve independence and protect against extraterritorial regulatory interference. — Gartner, 2025

Deloitte's 8th State of AI in the Enterprise report (3,235 leaders across 24 countries, fielded Aug–Sep 2025) found 77% of companies now factor an AI solution's country of origin into vendor selection, and 58% build their AI stacks primarily with local vendors. — Deloitte AI Institute, Jan 2026

On-premises deployments dominate sovereign-AI infrastructure for the most sensitive workloads. — NextMSC

In-company fine-tuning

Your data, distilled into your own private experts

A planned NIKITRIA service: fine-tune a company's or industry's proprietary data into domain-specific Small Language Models that run on the customer's own hardware — private by construction, because the data and the model never leave the building.

  • Specialists beat generalists at the job that matters

    A model trained on your processes, vocabulary, and edge cases — retaining the context generic pilots lose.

  • Fast, cheap iteration

    NVIDIA's SLM research notes small models can be fine-tuned overnight rather than over weeks, and are 10–30× cheaper to serve than frontier-scale LLMs.

  • Private by construction

    Training data, adapters, and inference all stay on infrastructure you control.

MIT Project NANDA's "The GenAI Divide: State of AI in Business 2025" (July 2025) found just 5% of integrated AI pilots are extracting millions in value, while the vast majority show no measurable P&L impact — attributing the gap to a "learning gap" that specialized, context-retaining models are positioned to close. (150 executive interviews, a 350-employee survey, 300 public deployments analyzed; lead author Aditya Challapally.) — MIT Project NANDA, Jul 2025

SLMs can be fine-tuned in hours and served at a fraction of frontier-LLM cost — 10–30× cheaper in latency, energy, and FLOPs. — Belcak et al., NVIDIA Research, arXiv:2506.02153, 2025

Offered initially as an integration-services engagement (pillar 5), productizing over time.

Education vertical

Curriculum-specific AI for India's classrooms — offline and affordable

Indian institutions are being asked to teach AI at national scale on constrained budgets. NIKITRIA's planned education offering pairs curriculum-specific SLMs with hardware schools already own: no GPU cluster, no per-student cloud bill, and it works with no internet at all.

  • Curriculum-specific experts

    Planned: specialist models aligned to specific syllabi and courses, for student practice and faculty support.

  • Runs on existing computer labs

    The same local-first runtime shown in our prototype — designed for modest, consumer-grade hardware rather than GPU clusters.

  • Offline by design

    Connectivity gaps stop being a blocker: everything runs on the machine in the room.

NEP 2020 emphasizes AI-integrated, personalized learning. CBSE offers a 15-hour AI module from Class VI and AI as an elective in Classes IX–XII; NCERT has embedded AI content in Class XI CS/IP textbooks; AICTE published a model AI & Data Science curriculum (2021) and runs faculty development programs. — NEP 2020; CBSE; NCERT; AICTE

The IndiaAI Mission (₹10,371.92 crore approved March 2024, seven pillars including FutureSkills) aims to expand AI courses across UG, PG, and PhD programs. — IndiaAI Mission, Mar 2024

The honest read on the tailwind

As of Feb 2026, only ~₹400 crore of the IndiaAI Mission's ₹10,372 crore outlay had been released, and the FY2026-27 allocation was cut roughly in half. Institutional GPU and compute budgets are genuinely constrained — which strengthens, not weakens, the case for low-cost, offline SLM deployments. — Budget reporting, Feb 2026

Market opportunity

A market measured in ranges, not cherry-picked numbers

Market-size estimates for small language models diverge 3–5× across research firms, so we present the range with each firm named and dated — and we separate external projections from anything NIKITRIA claims about itself.

TAM — global SLM market

  • $0.93B (2025) → $5.45B (2032)

    28.7% CAGR; 2024 base $0.74B

    MarketsandMarkets, Mar 2025

  • $7.8B (2023) → $20.7B (2030)

    15.1% CAGR

    Grand View Research

  • $6.5B (2024)

    25.7% CAGR to 2034

    GM Insights

  • $6.98B (2024)

    23.6% CAGR

    Polaris

Estimates differ because firms scope 'SLM' differently. We cite the range rather than selecting one number.

Adjacent — edge AI

  • $1.95B (2024) → $8.91B (2030)

    Edge AI software, 29.2% CAGR

    Grand View Research

  • Tens of billions by 2030

    Broader edge-AI market, estimates vary across firms

    Multiple firms

SAM — sovereign/on-prem + India

  • $10–12B

    India AI services revenue today

    NASSCOM

  • ~$17B by 2027

    India AI market projection, 25–35% CAGR

    NASSCOM–BCG

  • ₹10,371.92 crore

    IndiaAI Mission outlay — a policy tailwind signal

    IndiaAI Mission, Mar 2024

SOM — our wedge

Bottom-up, not a top-down percentage: a targeted outreach list of ~242 privacy-sensitive and sovereign-AI organizations for design-partner LOIs, converting into the certified-marketplace beta and early BYO-server deployments.

All market figures are external vendor research — projections, not audited facts — cited with firm and date.

Competitive differentiation

Everyone wraps the same engine. Nobody integrates the stack.

Ollama, LM Studio, Jan, and GPT4All all wrap the same llama.cpp engine — raw token throughput between them is close to a wash. They differ on packaging, GUI, and API surface, not fundamental capability. None offers a certified marketplace, panel-of-experts orchestration, full provenance, and a BYO-server sovereign tier as one integrated product. That combination is what NIKITRIA is building: orchestration and provenance run in the v0 prototype today; the certified marketplace, certification, and BYO-server tier are planned.

Category positioning as of Aug 2026 — not an exhaustive feature audit; competitors iterate quickly. NIKITRIA's marketplace, certification, and BYO-server entries describe the planned product; its orchestration, provenance, and local-runtime entries describe the working Orchestrator v0 prototype.
ProductCore engineMarketplaceOrchestrationProvenanceBYO-server / on-premCertification
Ollama llama.cppModel registryNoNoSelf-host (DIY)No
LM Studio llama.cppHF browseNoNoDesktop/local + serverNo
Jan llama.cppHubNoNoLocal (offline-first)No
GPT4All llama.cppCuratedNo (LocalDocs RAG)NoLocalNo
llamafile llama.cppN/A (single-file)NoNoLocalNo
Hugging Face HubYes (2M+ models)NoModel cardsHosted/inferenceNo
Cloud LLM APIs ProprietaryNoNoNoNo (cloud)No
NIKITRIA llama.cpp base + hot-swap LoRA expertsCertified SLM + MCP/skills/tools/apps (planned)Panel-of-experts routing + synthesisFull provenanceYes — sovereign wedge (planned)Adversarial testing (planned)

What the integrated stack adds

  • Certification through adversarial testing before a model is listed (planned)
  • Embedding-routed orchestration and synthesis across specialist experts
  • Full provenance: every answer attributable to the expert that produced it
  • Creator revenue share to build the marketplace supply side (planned)
  • Enterprise on-prem licensing — the sovereign wedge (planned)
  • The adaptability pipeline: continuously generated SLMs, skills, MCPs, and tools (planned)

Roadmap

Where this goes — labeled honestly

One thing on this roadmap exists today; everything else is a plan with a target, and is labeled that way.

  1. Working today

    Orchestrator v0 prototype

    Panel-of-experts routing, synthesis with provenance, fully local — running on an 8GB Apple M1 Mac, self-capped to 4GB of RAM.

  2. Target: 10–12 months

    Certified-marketplace beta

    Certified specialist SLMs with adversarial testing and creator revenue share — the core milestone of the pre-seed plan.

  3. Planned

    Hosted inference tier

    Usage-metered VPC/server deployments for teams that want the platform managed.

  4. Planned

    Enterprise BYO-server GA

    General availability of the sovereign licensing tier on customer infrastructure.

  5. Exploratory

    Track B: 1-bit models

    A pilot exploring BitNet 1-bit models via bitnet.cpp for even lower-footprint inference. Results, if any, will be reported separately with their own hardware context.

Founder

Who is behind NIKITRIA

Pranith Avula

Founder & CEO

Hyderabad, India

nikitria.admin@gmail.com

Or reach out through the investor and design-partner forms on this site.

Pranith Avula is the solo founder of NIKITRIA and designed and built the Orchestrator v0 prototype end-to-end — the routing, synthesis, provenance, and packaging that this site demonstrates — before raising outside capital.

NIKITRIA is built from Hyderabad, one of India's fastest-growing deep-tech startup ecosystems, with the certification framework and marketplace design as the current focus.

Why a solo founder?

Deliberately, at this stage: v0 was specified, built, and tested by one person, which kept the architecture coherent and the burn near zero. The pre-seed round exists precisely to change that — the first hires are engineering and model-certification roles.

For investors

Request the deck

NIKITRIA is raising a pre-seed round. The deck covers the architecture, the certified-marketplace plan, the market thesis with sourced ranges, and the terms of the round.

Prefer a call? Say so in the note and we'll reply with a booking link.

Submissions go only to the founder, are used only to respond to you, and are never added to a mailing list. Privacy policy

For privacy-sensitive & sovereign-AI organizations

Become a design partner

We are selecting a small group of design partners — organizations whose data cannot leave their infrastructure — to shape the certified marketplace and the BYO-server tier. Design partnership starts with a non-binding letter of intent.

Submissions go only to the founder and are used solely to follow up about design partnership. Privacy policy