DaisyOS ▼
Products ▼
Solutions ▼
Support ▼
About ▼
Blogs
Get a Quote
✍️ By Suki Wang | sukiwang@riotouch.com 📅 October 09, 2026 📂 Classroom AI ⏱ 10 min read

Riotouch On-Premise Compute: Running Classroom AI on the School's Own Server

On-premise compute for classroom AI: school server room with monitoring wall and rack

Feasibility, evidence and advantages of the local compute tier — Plan A1 / A2 / A3 and the Edge AI Box.

Riotouch Cloud · Deployment series, part 1 of 4. Series: 1 On-Premise Compute · 2 Cloud Compute · 3 Cloud Services · 4 Applications.

Quick answer. Riotouch's on-premise tier puts an AI-capable server inside the school — a 42U rack unit sized for 200 students (Plan A1), 200–1,000 students (Plan A2), or a multi-campus fleet (Plan A3) — optionally extended by one Jetson-class Edge AI Box per few classrooms. Every model runs on campus: LLM inference, RAG over school-owned content, OCR, speech-to-text and text-to-speech, image generation, and a model registry with canary release and rollback. Nothing here is exotic or vendor-specific: it is the same class of on-device and on-premises AI that Microsoft ships on Windows PCs and Google ships as an on-premises, air-gapped cloud rack. This article sets out the evidence, with sources, and the specific advantages it buys a school. Compare it with the alternative modes in the Riotouch Cloud Solution overview.

1. What on-premise compute means in the Riotouch stack

The Riotouch platform is seven layers deep (parts 3 and 4 cover the service and application layers). In on-premise mode, layers 1–3 and layer 4's Local AI Engine all run on the school's own hardware; the AI Gateway sits in front of every request so that routing, redaction, auditing, degradation and metering are decided before anything leaves the building.

TierScaleCompute (as published)Model capacity
Plan A1 — Light≤ 200 studentsIntel i9-14900K / AMD Ryzen 9 7950X (24C/32T), DDR5 64 GB, 1× RTX 4090 24 GB, 2 TB NVMe Gen4 + 8 TB NAS RAID5, 2.5 GbE, 1 kVA UPS~14B parameters (Q4)
Plan A2 — Standard200–1,000 studentsAMD Threadripper PRO 7965WX / Intel Xeon w5-3435X (24C), DDR5 ECC 128 GB, 2× RTX 4090 or 1× L40S 48 GB, 4 TB NVMe Gen4 + 24 TB NAS RAID6, 10 GbE, 3 kVA online UPS32B (Q4) / 70B (Q4)
Plan A3 — Flagship1,000+ students, multi-campus2× AMD EPYC 9454 (96C total) / 2× Intel Xeon Gold 6448Y, DDR5 ECC 256–512 GB, 4× A6000 48 GB or 2× H100 80 GB, 8 TB NVMe RAID10 + 100 TB+ Ceph, 25 GbE + SD-WAN, 6 kVA + generator72B (FP16) / 405B (Q4), multi-model
Edge AI Boxone per N classroomsNVIDIA Jetson Orin Nano / NX class, PoE-poweredsmall-model inference (handwriting, board-face tracking, quick responses)

What runs locally in the Local AI Engine: LLM inference (Ollama / vLLM, Qwen2.5-class models), RAG search (LangChain + pgvector over the school's own knowledge base), an agent framework (LangGraph, ReAct / tool calling), ASR and TTS (Whisper, VITS2), OCR and computer vision (PaddleOCR, board-face recognition), bilingual CN/EN embeddings, optional image generation (SD / SDXL), and a model registry with versioning, canary release and rollback.

Room requirements are part of the deal — precision HVAC, anti-static flooring, physical access control, smoke and water-leak sensors, KVM and a 42U rack. These are the line items most often left out of an integrator's BOM, and the ones that decide whether a deployment survives its third summer.

2. Feasibility: a mainstream engineering pattern, not a Riotouch invention

Microsoft ships local model execution as a first-class Windows feature. Foundry Local lets applications run large language models directly on a Windows device as an alternative to cloud calls, and Microsoft documents Foundry on Windows as its route to integrating local AI capabilities into Windows apps [1][2]. The capability Riotouch uses on campus — local LLM execution behind an application API — is a supported, documented platform capability.

Google ships on-premises and air-gapped cloud racks. Google Distributed Cloud extends Google Cloud infrastructure, services and Kubernetes to customer data centers and edge locations, explicitly to meet regulatory, latency and sovereignty requirements; the air-gapped appliance is an integrated hardware and software platform for environments outside a data center that creates an isolated environment [3][4]. If Google's answer to data-residency rules is a rack in your building, a school running classroom AI on its own server is not a compromise — it is the same architecture.

The latency advantage is measured, not asserted. A measurement study of 8,456 end-users against 6,341 edge servers and 69 cloud locations found that 58% of end-users can reach a nearby edge server in under 10 ms, while only 29% obtain comparable latency from a nearby cloud location [5]. For a classroom, that is the difference between handwriting recognition that feels instantaneous and one that visibly stutters whenever the network is busy.

Privacy-preserving distributed learning is a solved research problem. Nature Communications work on decentralized federated learning shows that institutions in highly regulated domains need models trained without centralizing sensitive data, and demonstrates proxy-model-sharing methods that preserve privacy while keeping performance [6]. That is the argument a school data-protection officer needs: better models without shipping student data anywhere.

Cloud-native gravity is real, but it is not a reason to give up local control. Even the cloud-infrastructure research community frames the next phase as AI-native computation spanning cloud and edge [9]; the choice schools face is not cloud or nothing, but which layer runs where — the question the three Riotouch deployment modes answer.

3. Advantages of the on-premise tier — each tied to a source

AdvantageWhy it matters to a schoolSupporting evidence
Student data never leaves the campusNames, faces, voices, handwriting, exam answers, grades and live board content stay inside the building; only redacted derivatives may leave, and only under the AI Gateway policyMicrosoft's local-AI model keeps inference on-device [1][2]; Google's air-gapped GDC exists precisely for isolated environments [4]; federated learning methods serve regulated institutions that cannot centralize data [6]
Sub-10 ms interaction on the LANHandwriting recognition, board-face tracking and voice feedback feel instant, even on a saturated school network58% vs 29% sub-10 ms reachability, edge vs cloud [5]
Works when the internet does notPlayback and local inference continue from cache during an outage — a documented property of hybrid and degraded operationEdge and on-premises AI are designed for environments with unreliable connectivity [3][4]
Predictable, non-metered costOne capital purchase instead of per-token cloud billing; the AI Gateway token metering shows what the same workload would have cost in the cloudMetering is a documented gateway function; LLM operational cost drivers are studied in the peer-reviewed literature [10]
Model choice and lifecycle controlRun Qwen2.5-class open models at 14B–405B, pin versions, canary-release and roll back without vendor coordinationModel registry with versioning, canary release and rollback is a standard MLOps capability; local open-weight execution is documented by Microsoft [1]
Classroom-level offloadAn Edge AI Box handles handwriting, board-face tracking and small-model responses with no server round-trip, so one busy room cannot degrade the whole siteEdge tiers are promoted exactly for latency-critical, bandwidth-hungry tasks [5]
On-premise compute architecture: local AI engine and Plan A1, A2 and A3 servers with the Edge AI Box
The local compute tier: the Local AI Engine runs on the school's own hardware, in the light (A1), standard (A2) or flagship (A3) configuration, optionally offloaded classroom by classroom by an Edge AI Box.

4. When to choose it

SituationRecommendation
Data-protection rules prohibit student data leaving campusOn-premise (Mode 1) — Plan A1–A3 by headcount
One or two buildings, under 200 studentsPlan A1
200–1,000 students, several buildings, IT staff on sitePlan A2 (or hybrid with an A2 baseline)
Multi-campus / district, 1,000+ users, one dashboardPlan A3 + SD-WAN (typically in hybrid mode)
Unreliable or expensive internetPlan A1/A2 — local inference plus local cache keeps teaching running

Pair the server tier with the display layer it serves: the interactive flat panel for classrooms, an RK3588 smart board or OPS module for panel-side compute, and the Aether Orb Edge AI Box for classroom-level offload. Full interface and power numbers are published on the product specifications page.

Next in the series: part 2 covers the cloud compute tier — elastic GPU capacity for peaks, redacted-only cloud inference, and what the hyperscaler evidence says about scale.

Frequently asked questions

Can classroom AI really run on a school server instead of the cloud?

Yes. Riotouch ships a local compute tier — Plan A1 (a 42U rack unit sized for about 200 students), Plan A2 for 200–1,000 students and Plan A3 for multi-campus fleets — that runs LLM inference, RAG search, OCR, speech-to-text and text-to-speech inside the school. This is the same class of on-premises execution that Microsoft ships as a Windows feature and that Google offers as an on-premises, air-gapped cloud rack, so it is a mainstream engineering pattern rather than a vendor-specific experiment.

What server do we need for 500 students?

Plan A2 — Standard. Published configuration: AMD Threadripper PRO 7965WX or Intel Xeon w5-3435X (24C), DDR5 ECC 128 GB, 2× RTX 4090 or 1× L40S 48 GB, 4 TB NVMe Gen4 plus 24 TB NAS RAID6, 10 GbE networking and a 3 kVA online UPS, supporting 32B and 70B models at Q4 quantization. Add one Jetson-class Edge AI Box per few classrooms for handwriting and board-tracking workloads.

What room infrastructure does the on-premise tier need?

Precision HVAC, anti-static flooring, physical access control, smoke and water-leak sensors, a KVM switch and a 42U rack. These line items are the ones most often left out of an integrator BOM — and the ones that decide whether a deployment survives its third summer.

Does local inference give up model quality compared with cloud AI?

Not for classroom workloads. A measurement study of 8,456 end-users, 6,341 edge servers and 69 cloud locations found that 58% of users can reach a nearby edge server in under 10 ms, while only 29% obtain comparable latency from a nearby cloud location. Model capacity is a sizing decision: Plan A1 targets ~14B parameters, Plan A2 reaches 70B and Plan A3 reaches 405B at Q4 or 72B at FP16 — with a model registry that supports version pinning, canary release and rollback.

What happens if the school internet goes down?

Teaching keeps running. Playback and local inference continue from local cache and on-campus hardware, because nothing in the inference path depends on an external API call in on-premise mode.

References

  1. Microsoft Learn — Get started with Foundry Local (local LLM execution on Windows devices). learn.microsoft.com
  2. Microsoft Learn — Use local AI with Microsoft Foundry on Windows. learn.microsoft.com
  3. Google Cloud — Google Distributed Cloud. cloud.google.com
  4. Google Cloud documentation — About Google Distributed Cloud air-gapped appliance. docs.cloud.google.com
  5. B. Charyyev, E. Arslan, M. H. Gunes — Latency Comparison of Cloud Datacenters and Edge Servers (8,456 users; 6,341 edge servers; 69 cloud locations). par.nsf.gov
  6. Decentralized federated learning through proxy model sharing, Nature Communications (2023). nature.com
  7. Differentially private knowledge transfer for federated learning, Nature Communications (2023). nature.com
  8. Microsoft Learn — What are Windows AI APIs? learn.microsoft.com
  9. Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native (arXiv:2401.12230). arxiv.org
  10. Reconciling the contrasting narratives on the environmental impact of large language models, Scientific Reports (2024). nature.com

All links verified live on 2026-10-09. Deployment specifications in section 1 are Riotouch's published configurations and are subject to change with hardware generations.

Related Product

Riotouch Cloud Solution — one platform, three deployment modes for interactive flat panels. Hardware: RK3588 Smart Board · OPS modules · Aether Orb Edge AI Box.

Related Articles

Suki Wang
International Sales, Riotouch Technology
Tel: +86 769 8258 3996  |  Mobile/WhatsApp: +86 137 1299 3879
Email: sukiwang@riotouch.com

📩 Request a Quote

Need a Solution That
Fits Your Project?

Tell us what you're looking for and we'll tailor a quote to your requirements — with pricing, lead time, and expert advice.

Headquarters5th Floor, Huihang High-tech Industrial Park, No. 9 Zhongnan Middle Road, Chang An Town, Dongguan City, Guangdong Province
Get Your Custom Quote
Fill out the form and we'll get back to you within 12 hours.

By submitting this form you agree to our Privacy Policy. We use your details only to respond to your enquiry.

Quick InquiryGet a quote
✓ Sent! We'll reply within 4 hours.
Chat on WhatsAppReply in minutes