Atoms
Join the Praxis beta

Sector · Energy and utilities

Physical AI for energy and utilities

Grids and plants run on control systems, maintenance records and field crews that were never designed to share one picture of risk. These are the questions engineers ask first, answered from 3 published CodeNinja Atoms reference architectures.

  1. What does a physical AI system for energy and utilities look like?
  2. Which AI models can an energy and utilities operator run on its own hardware?
  3. How much compute and hardware does AI in energy and utilities need?
  4. Is it cheaper to own AI hardware or rent cloud GPUs in energy and utilities?
  5. What ontology or object model does an energy and utilities AI system need?
  6. Who approves the decisions an AI system makes in energy and utilities?

What does a physical AI system for energy and utilities look like?

A complete physical AI design for energy and utilities names what to sense, which existing systems to join, the object model that joins them, the models and hardware, the three-year cost and the person who approves every action. CodeNinja Atoms has published 3 such reference architectures for energy and utilities, each free to reuse under CC BY 4.0.

DesignCountryWhat it does
Grid Context WatchUnited Statesone system of context that sees every device on a utility's control networks, flags abnormal behaviour and assembles the evidence a NERC CIP audit asks for, on the operator's own hardware.
Reliability AtlasSaudi Arabiaa records-based reliability and availability assessment that joins each plant's work orders, trip history and design documents into one reliability object model the operator owns, so every availability figure, criticality rank and root cause finding traces to its source.
Feeder FirewatchUnited Statesa live ignition and outage risk score for every distribution feeder segment, built on the cooperative's own hardware, with every de-energisation and fast-trip change approved by a named operator.

Which AI models can an energy and utilities operator run on its own hardware?

Each published energy and utilities design names its models and why, and every model is open-weight or no model is used at all, so the operator can run it on hardware it owns.

DesignModels
Grid Context WatchGLM 5.2 (MIT) for reasoning and Qwen3-Embedding-0.6B (Apache-2.0) for retrieval; detection stays in the procured sensors
Reliability AtlasNone: the register is empty by design because no inference workload is contracted
Feeder FirewatchSelf-hosted open-weight models: GLM 5.2 under MIT (reasoning) on site, RF-DETR (vision) at the edge
DesignChoiceWhat was pickedWhy
Grid Context WatchReasoning modelGLM 5.2, 753B parameters, frontier class, FP8, 1M contextPlain MIT license with no revenue trigger or security-review clause; one node of eight 141 GB HBM GPUs holds it; the work surface's role demands long-context agentic reasoning
Grid Context WatchEmbedding modelQwen3-Embedding-0.6B, Apache-2.0, 32K windowWhole playbooks and CIP procedures retrieve as single passages; runs beside the frontier model for about 1.2 GB at BF16
Grid Context WatchSizing rule1.2 planning factor over FP8 weights753 GB of weights times 1.2 is 904 GB against 1,128 GB of node memory, leaving governed headroom for KV cache and activations
Grid Context WatchFrontier computeOne node of eight 141 GB HBM GPUsThe smallest class that holds the FP8 weights; sized from the model's filed parameters
Grid Context WatchEdge computeFanless, extended-temperature, cabinet-mount classCompute stays out of electronic security perimeters wherever topology allows
Grid Context WatchEnclosure and powerNEMA 3R/4X class cabinets with uninterruptible power supplies under Network UPS ToolsClean shutdown ordering protects capture integrity environment
Grid Context WatchTime synchronizationOCP Time Card GNSS grandmaster with holdover, Linuxptp and chronyPcap, sensor and log clocks must agree for evidence to correlate
Grid Context WatchOne-way transferHardware data diodes with the one-way transfer appliance at highest-impact perimetersDemonstrable one-way flow where policy demands it; constrained DMZ conduit elsewhere
Grid Context WatchSite networkingHardened, fanless, dual-DC industrial switches with SPAN/TAP portsSegmentation remains the operator's policy decision; the design reads a copy, never the control path
Grid Context WatchSensingPassive network traffic capture via the procured OT monitoring sensors at SPAN/TAP sourcesThe scope watches OT traffic, not premises; no cameras or video models
Grid Context WatchPrimary patternSystem of Context over one Hyper-style object model of 14 objectsThe lasting asset is the living model of the OT estate, not any single alert
Grid Context WatchGroundThe operator's own on-premises OT hardware and data centerOwnership of weights, ontology, decision record and boundary stays with the operator; identity and update paths run inside the air gap
Reliability AtlasModelsNone: the register is empty by designNo inference workload is contracted; the deliverable is the study, and an empty register recorded plainly outranks a speculative entry the scope cannot justify
Reliability AtlasHardware classes and sizing rulesNone procured; the study runs on the operator's own in-country servers, storage and data center capacityNo field hardware, no edge compute and no serving runtime is installed under this scope
Reliability AtlasSensingNo cameras, no sensors and no positioning; site verification is by engineer walkdownsThe world is read through records, so nothing is watched and nothing needs country legality stated
Reliability AtlasThe pattern it stands onSystem of context as the primary pattern, with the reliability object model as a projection over the three systems of recordUnavailability ranking, sparing confirmation and RCA validation are cross-silo joins that no single source system can produce alone
Reliability AtlasThe ground it runs onIn-country hosting in Saudi Arabia, with identity through the operator's own directory and single sign-on, procured by classThe operator's requirement keeps data inside the country's boundary and under the operator's identity, which the operator already operates
Feeder FirewatchLanguage model parameters filed, FP8, MIT licenseGLM 5.2, mixture-of-experts, 753 billion weights the cooperative owns; one 8-GPUAgentic reasoning over the ontology with FP8 node holds it
Feeder FirewatchDetection model to 68 MB and 2XL at about 254 MB, BF16, Apache-2.0RF-DETR, Nano to Large checkpoints at 61 detection on the edge box class, fine-tuned on the cooperative's own fire-season framesSmoke, pole, vegetation and equipment
Feeder FirewatchEdge compute class and sealed, NEMA 3R/4 enclosuresIndustrial edge accelerator modules, fanless streams, rated for dust, heat, hail and coldSized decode-first from the actual camera
Feeder FirewatchSite inference server center, one node of 8 GPUs of the 141 GB HBM classRuggedised server class at the operations sized from filed parameters at FP81,128 GB against a 904 GB requirement,
Feeder FirewatchServing runtimes Runtime or OpenVINO class at the edgevLLM on site; Triton Inference Server, ONNX a bench measurement of the real streamsThe edge runtime is pinned at design against
Feeder FirewatchSizing rules thermal beats visible; reuse existing cameras or not; size GPUs from filed parametersSize edge compute from streams; when measured duty rather than a vendor defaultEvery hardware choice derives from a
Feeder FirewatchSensing camera feeds, truck and drone cameras, mesonet wind and humidity, 310 recloser fault indicators, AMI last gasp from 61,000 meters, crew AVL on 20 trucks42 substation PTZ cameras, 12 wildfire thermal only where a coverage gap is provenReused through the reuse gates; new fixed
Feeder FirewatchPatterns read-only systems of record; human-approved write-back; one-way diode for SCADASystem of context; adapters-only ingestion; record authoritative while the platform joins themKeeps the grid protected and the systems of
Feeder FirewatchGround as primary; sovereign cloud for overflow and recovery onlyCooperative servers at the operations center during a storm, and the risk picture cannot live outside the boundaryThe model must survive an internet outage

How much compute and hardware does AI in energy and utilities need?

The compute follows from the models: the published energy and utilities designs size it as follows, from no new hardware to a full GPU node.

DesignPartThe design
Grid Context WatchComputeOne node of eight 141 GB HBM-class GPUs: 753 GB of FP8 weights, 904 GB with the 1.2 planning factor, against 1,128 GB
Grid Context WatchBoundaryRead only and one way out of the control networks; nothing writes into an electronic security perimeter, and no data leaves the operator
Reliability AtlasComputeNo hardware bought; the study runs on the operator's own in-country servers
Feeder FirewatchFrontier computeOne node of eight 141 GB HBM-class GPUs holds GLM 5.2 at FP8 (753 GB of weights, 904 GB with headroom)
Feeder FirewatchEdgeSealed industrial boxes at substations and on patrol trucks detect smoke and damaged equipment when cellular coverage drops

Is it cheaper to own AI hardware or rent cloud GPUs in energy and utilities?

Each energy and utilities design prices three years of ownership in its Appendix A, with every price cited, against renting the same capacity from a cloud region at its deepest three-year commitment where hardware is bought.

DesignLineThree years
Grid Context WatchThree-year costAbout 510,000 dollars to own the reasoning node, about four fifths the deepest three-year AWS commitment (Appendix A)
Reliability AtlasThree-year costNo hardware line to price: the cost is the assessment work itself (Appendix A)
Feeder FirewatchThree-year cost, ownedAbout 841,000 US dollars with support and power at the Texas industrial power price
Feeder FirewatchThree-year cost, rented0.88 million to 1.92 million US dollars for the same GPUs around the clock; ownership costs about the same as the deepest three-year commitment (version 2)
Feeder FirewatchClosed model break-evenThe cheapest closed model matches the owned stack at about 41 users; above that, ownership is cheaper and the gap grows with every user

What ontology or object model does an energy and utilities AI system need?

The object model is the part that makes the system an ontology rather than a pipeline: typed objects for the things in the energy and utilities operation, with properties, status values and typed links. Every published design ships its object model as hyper-ontology/1 JSON that loads into Hyper Ontology.

DesignSizeObjectsDownload
Grid Context Watch14 objects, 14 linksSubstation, OT Asset, Network Sensor, Anomaly Alert, Investigation Case, Vulnerability Finding, Detection Rule, Threat Intelligence Item, Pcap Evidence Record, CIP Evidence Artefact, OT Security Analyst, SIEM, Energy management system, Substation data platformobjects.json
Reliability Atlas13 objects, 14 linksPlant, Production train, Equipment, Failure mode, Work order, Downtime event, Availability prediction, Availability target, Criticality index, RCA report, Major event, Spare part, Reliability programobjects.json
Feeder Firewatch14 objects, 12 linksSubstation, Feeder, Feeder segment, Pole, Recloser, Meter, Pole inspection record, Outage event, Feeder segment ignition risk score, Red flag warning, Wildfire camera station, Field crew, Work order, PSPS decision recordobjects.json

Who approves the decisions an AI system makes in energy and utilities?

In every published energy and utilities design a named person makes the decision that changes the physical world; the system prepares it.

DesignHuman control
Grid Context WatchEvery triage, case and risk acceptance carries a named OT security analyst; agents draft, the analyst decides
Reliability AtlasPredictions are accepted only by a named reviewer and RCA reports validated only by a named engineer
Feeder FirewatchEvery public safety power shutoff (PSPS) recommendation becomes a decision record a named operator approves or declines; the design never opens or closes a recloser

Every answer on this page is drawn from the papers linked in it: each paper's At a glance table, model register and cost appendix. Full text for agents: llms-full.txt. Designed on Praxis; object models load into Hyper Ontology.