Position paper · v1 · Apache 2.0

Democratizing distributed AI inference.

OpenPrism Network is an open, UMA-first architecture for running large language model inference across clusters of unified-memory machines — so universities, labs, hospitals, and independent operators can participate as infrastructure providers, not just as customers of closed APIs.

UMA-first
Apple Silicon · Snapdragon X
Open by default
Apache 2.0 · CC-BY-4.0
Workload scope
Batch · throughput-tolerant
The problem

Inference is concentrating into a few GPU-rich providers.

Hyperscalers and a handful of frontier labs control most large-model serving capacity. For universities, hospitals, public-sector organizations, and independent developers, the result is the same: high capital cost, scarce accelerators, dependency on closed APIs, and data that must leave the building to be useful.

Capital barrier

Comparable GPU deployments routinely run into six- to seven-figure budgets before a single token is served.

Access barrier

Accelerator supply is allocated to the largest buyers first. Smaller institutions wait, or rent at a premium.

Sovereignty barrier

Sensitive data — clinical, governmental, regulated research — often legally or practically cannot leave the institution.

How it works

A network shaped around unified memory.

UMA hardware changes the economics of inference: large models fit in one box, but one box isn't enough. OpenPrism Network coordinates many of them into a verifiable inference fabric.

  1. 01

    Layers stay put

    Transformer layers are statically assigned to UMA nodes. Model weights stay resident in unified memory — no shuffling, no reloading.

  2. 02

    Only activations move

    Small activation tensors transit within low-latency metro clusters between sequential layer owners. The wire never carries weights.

  3. 03

    Outputs are verified

    Each request runs on multiple independent nodes. Output fingerprints are compared to detect faulty or dishonest execution.

  4. 04

    Settlement, not compute

    A blockchain handles payment and reputation only. It stays off the critical path — never gating an inference call.

Scope, honestly: the network targets batch and throughput-tolerant workloads — offline document processing, summarization, synthetic generation, model evaluation. It is not designed for real-time interactive chat.

Two ways to deploy

Mesh for reach. Micro data center for control.

OpenPrism Network is designed as a single architecture with two complementary deployment shapes. Most institutions will mix both.

Model 1

Distributed Mesh

Harvests idle institutional UMA hardware that's already powered for another purpose — lab workstations, teaching Macs, developer machines. Near-zero marginal capital: no new hardware, no new data center.

  • Brings new participants in at the lowest possible cost
  • Uses energy that's already being spent
  • Independent operator entry tier from a single Mac Studio M3 Ultra (under $15k)
Model 2

UMA Micro Data Center

A purpose-built, locally owned, sovereign inference facility — on the order of a single rack of Apple Silicon nodes. For organizations that legally or practically cannot send data to hyperscalers.

  • Healthcare, government, regulated biotech, air-gapped sites
  • Data residency and on-premises operation by design
  • The pitch is control and predictability — not undercutting cloud
Built to be open

Open participation, open licensing, open metrics.

OpenPrism Network is not a product owned by a company. It is a coordination layer that only works if many independent parties can build on, audit, and challenge it.

Node operators

Contribute UMA capacity — a single workstation or a full rack — and earn settlement for verified inference work.

Runtime implementers

Build and improve reference runtimes against publicly versioned interfaces. No single vendor owns the implementation.

Benchmark maintainers

Join the open Benchmark Working Group. Results and methodology are published under CC-BY-4.0.

Application integrators

Wire OpenPrism into pipelines for offline document work, evaluation, summarization, and synthetic data generation.

Open commitments
Reference code
Apache 2.0
Benchmarks
CC-BY-4.0
Interfaces
Publicly versioned
Metrics
No single vendor controls them

Energy efficiency is reported as a first-class metric — not buried, not omitted.

Open research

Help us solve these.

OpenPrism Network is a position and architecture paper, not a finished product. Two problems in particular are open invitations to the research community.

01

Consensus under floating-point non-determinism

Independent nodes running the same model on the same input may produce subtly different outputs because of floating-point ordering and hardware-specific math. We need verification protocols that accept honest divergence while still catching faulty or dishonest execution.

02

Dynamic layer reassignment without full redistribution

When a node joins, leaves, or fails, the network needs to reassign layer ownership without re-broadcasting full model weights. The open question is how to do this cheaply, safely, and with predictable latency.

Contribute

The network is built by the people who run it.

Contributions are open on routing, node certification, benchmarks, the reference runtime, and governance. If you build, operate, or study inference systems — we'd like your help.