Feasibility, evidence and advantages of the local compute tier — Plan A1 / A2 / A3 and the Edge AI Box.
Riotouch Cloud · Deployment series, part 1 of 4. Series: 1 On-Premise Compute · 2 Cloud Compute · 3 Cloud Services · 4 Applications.
Quick answer. Riotouch's on-premise tier puts an AI-capable server inside the school — a 42U rack unit sized for 200 students (Plan A1), 200–1,000 students (Plan A2), or a multi-campus fleet (Plan A3) — optionally extended by one Jetson-class Edge AI Box per few classrooms. Every model runs on campus: LLM inference, RAG over school-owned content, OCR, speech-to-text and text-to-speech, image generation, and a model registry with canary release and rollback. Nothing here is exotic or vendor-specific: it is the same class of on-device and on-premises AI that Microsoft ships on Windows PCs and Google ships as an on-premises, air-gapped cloud rack. This article sets out the evidence, with sources, and the specific advantages it buys a school. Compare it with the alternative modes in the Riotouch Cloud Solution overview.
The Riotouch platform is seven layers deep (parts 3 and 4 cover the service and application layers). In on-premise mode, layers 1–3 and layer 4's Local AI Engine all run on the school's own hardware; the AI Gateway sits in front of every request so that routing, redaction, auditing, degradation and metering are decided before anything leaves the building.
| Tier | Scale | Compute (as published) | Model capacity |
|---|---|---|---|
| Plan A1 — Light | ≤ 200 students | Intel i9-14900K / AMD Ryzen 9 7950X (24C/32T), DDR5 64 GB, 1× RTX 4090 24 GB, 2 TB NVMe Gen4 + 8 TB NAS RAID5, 2.5 GbE, 1 kVA UPS | ~14B parameters (Q4) |
| Plan A2 — Standard | 200–1,000 students | AMD Threadripper PRO 7965WX / Intel Xeon w5-3435X (24C), DDR5 ECC 128 GB, 2× RTX 4090 or 1× L40S 48 GB, 4 TB NVMe Gen4 + 24 TB NAS RAID6, 10 GbE, 3 kVA online UPS | 32B (Q4) / 70B (Q4) |
| Plan A3 — Flagship | 1,000+ students, multi-campus | 2× AMD EPYC 9454 (96C total) / 2× Intel Xeon Gold 6448Y, DDR5 ECC 256–512 GB, 4× A6000 48 GB or 2× H100 80 GB, 8 TB NVMe RAID10 + 100 TB+ Ceph, 25 GbE + SD-WAN, 6 kVA + generator | 72B (FP16) / 405B (Q4), multi-model |
| Edge AI Box | one per N classrooms | NVIDIA Jetson Orin Nano / NX class, PoE-powered | small-model inference (handwriting, board-face tracking, quick responses) |
What runs locally in the Local AI Engine: LLM inference (Ollama / vLLM, Qwen2.5-class models), RAG search (LangChain + pgvector over the school's own knowledge base), an agent framework (LangGraph, ReAct / tool calling), ASR and TTS (Whisper, VITS2), OCR and computer vision (PaddleOCR, board-face recognition), bilingual CN/EN embeddings, optional image generation (SD / SDXL), and a model registry with versioning, canary release and rollback.
Room requirements are part of the deal — precision HVAC, anti-static flooring, physical access control, smoke and water-leak sensors, KVM and a 42U rack. These are the line items most often left out of an integrator's BOM, and the ones that decide whether a deployment survives its third summer.
Microsoft ships local model execution as a first-class Windows feature. Foundry Local lets applications run large language models directly on a Windows device as an alternative to cloud calls, and Microsoft documents Foundry on Windows as its route to integrating local AI capabilities into Windows apps [1][2]. The capability Riotouch uses on campus — local LLM execution behind an application API — is a supported, documented platform capability.
Google ships on-premises and air-gapped cloud racks. Google Distributed Cloud extends Google Cloud infrastructure, services and Kubernetes to customer data centers and edge locations, explicitly to meet regulatory, latency and sovereignty requirements; the air-gapped appliance is an integrated hardware and software platform for environments outside a data center that creates an isolated environment [3][4]. If Google's answer to data-residency rules is a rack in your building, a school running classroom AI on its own server is not a compromise — it is the same architecture.
The latency advantage is measured, not asserted. A measurement study of 8,456 end-users against 6,341 edge servers and 69 cloud locations found that 58% of end-users can reach a nearby edge server in under 10 ms, while only 29% obtain comparable latency from a nearby cloud location [5]. For a classroom, that is the difference between handwriting recognition that feels instantaneous and one that visibly stutters whenever the network is busy.
Privacy-preserving distributed learning is a solved research problem. Nature Communications work on decentralized federated learning shows that institutions in highly regulated domains need models trained without centralizing sensitive data, and demonstrates proxy-model-sharing methods that preserve privacy while keeping performance [6]. That is the argument a school data-protection officer needs: better models without shipping student data anywhere.
Cloud-native gravity is real, but it is not a reason to give up local control. Even the cloud-infrastructure research community frames the next phase as AI-native computation spanning cloud and edge [9]; the choice schools face is not cloud or nothing, but which layer runs where — the question the three Riotouch deployment modes answer.
| Advantage | Why it matters to a school | Supporting evidence |
|---|---|---|
| Student data never leaves the campus | Names, faces, voices, handwriting, exam answers, grades and live board content stay inside the building; only redacted derivatives may leave, and only under the AI Gateway policy | Microsoft's local-AI model keeps inference on-device [1][2]; Google's air-gapped GDC exists precisely for isolated environments [4]; federated learning methods serve regulated institutions that cannot centralize data [6] |
| Sub-10 ms interaction on the LAN | Handwriting recognition, board-face tracking and voice feedback feel instant, even on a saturated school network | 58% vs 29% sub-10 ms reachability, edge vs cloud [5] |
| Works when the internet does not | Playback and local inference continue from cache during an outage — a documented property of hybrid and degraded operation | Edge and on-premises AI are designed for environments with unreliable connectivity [3][4] |
| Predictable, non-metered cost | One capital purchase instead of per-token cloud billing; the AI Gateway token metering shows what the same workload would have cost in the cloud | Metering is a documented gateway function; LLM operational cost drivers are studied in the peer-reviewed literature [10] |
| Model choice and lifecycle control | Run Qwen2.5-class open models at 14B–405B, pin versions, canary-release and roll back without vendor coordination | Model registry with versioning, canary release and rollback is a standard MLOps capability; local open-weight execution is documented by Microsoft [1] |
| Classroom-level offload | An Edge AI Box handles handwriting, board-face tracking and small-model responses with no server round-trip, so one busy room cannot degrade the whole site | Edge tiers are promoted exactly for latency-critical, bandwidth-hungry tasks [5] |
| Situation | Recommendation |
|---|---|
| Data-protection rules prohibit student data leaving campus | On-premise (Mode 1) — Plan A1–A3 by headcount |
| One or two buildings, under 200 students | Plan A1 |
| 200–1,000 students, several buildings, IT staff on site | Plan A2 (or hybrid with an A2 baseline) |
| Multi-campus / district, 1,000+ users, one dashboard | Plan A3 + SD-WAN (typically in hybrid mode) |
| Unreliable or expensive internet | Plan A1/A2 — local inference plus local cache keeps teaching running |
Pair the server tier with the display layer it serves: the interactive flat panel for classrooms, an RK3588 smart board or OPS module for panel-side compute, and the Aether Orb Edge AI Box for classroom-level offload. Full interface and power numbers are published on the product specifications page.
Next in the series: part 2 covers the cloud compute tier — elastic GPU capacity for peaks, redacted-only cloud inference, and what the hyperscaler evidence says about scale.
Yes. Riotouch ships a local compute tier — Plan A1 (a 42U rack unit sized for about 200 students), Plan A2 for 200–1,000 students and Plan A3 for multi-campus fleets — that runs LLM inference, RAG search, OCR, speech-to-text and text-to-speech inside the school. This is the same class of on-premises execution that Microsoft ships as a Windows feature and that Google offers as an on-premises, air-gapped cloud rack, so it is a mainstream engineering pattern rather than a vendor-specific experiment.
Plan A2 — Standard. Published configuration: AMD Threadripper PRO 7965WX or Intel Xeon w5-3435X (24C), DDR5 ECC 128 GB, 2× RTX 4090 or 1× L40S 48 GB, 4 TB NVMe Gen4 plus 24 TB NAS RAID6, 10 GbE networking and a 3 kVA online UPS, supporting 32B and 70B models at Q4 quantization. Add one Jetson-class Edge AI Box per few classrooms for handwriting and board-tracking workloads.
Precision HVAC, anti-static flooring, physical access control, smoke and water-leak sensors, a KVM switch and a 42U rack. These line items are the ones most often left out of an integrator BOM — and the ones that decide whether a deployment survives its third summer.
Not for classroom workloads. A measurement study of 8,456 end-users, 6,341 edge servers and 69 cloud locations found that 58% of users can reach a nearby edge server in under 10 ms, while only 29% obtain comparable latency from a nearby cloud location. Model capacity is a sizing decision: Plan A1 targets ~14B parameters, Plan A2 reaches 70B and Plan A3 reaches 405B at Q4 or 72B at FP16 — with a model registry that supports version pinning, canary release and rollback.
Teaching keeps running. Playback and local inference continue from local cache and on-campus hardware, because nothing in the inference path depends on an external API call in on-premise mode.
All links verified live on 2026-10-09. Deployment specifications in section 1 are Riotouch's published configurations and are subject to change with hardware generations.
Riotouch Cloud Solution — one platform, three deployment modes for interactive flat panels. Hardware: RK3588 Smart Board · OPS modules · Aether Orb Edge AI Box.
Tell us what you're looking for and we'll tailor a quote to your requirements — with pricing, lead time, and expert advice.