Sector · Energy and utilities
Physical AI for energy and utilities
Grids and plants run on control systems, maintenance records and field crews that were never designed to share one picture of risk. These are the questions engineers ask first, answered from 3 published CodeNinja Atoms reference architectures.
- What does a physical AI system for energy and utilities look like?
- Which AI models can an energy and utilities operator run on its own hardware?
- How much compute and hardware does AI in energy and utilities need?
- Is it cheaper to own AI hardware or rent cloud GPUs in energy and utilities?
- What ontology or object model does an energy and utilities AI system need?
- Who approves the decisions an AI system makes in energy and utilities?
What does a physical AI system for energy and utilities look like?
A complete physical AI design for energy and utilities names what to sense, which existing systems to join, the object model that joins them, the models and hardware, the three-year cost and the person who approves every action. CodeNinja Atoms has published 3 such reference architectures for energy and utilities, each free to reuse under CC BY 4.0.
| Design | Country | What it does |
|---|---|---|
| Grid Context Watch | United States | one system of context that sees every device on a utility's control networks, flags abnormal behaviour and assembles the evidence a NERC CIP audit asks for, on the operator's own hardware. |
| Reliability Atlas | Saudi Arabia | a records-based reliability and availability assessment that joins each plant's work orders, trip history and design documents into one reliability object model the operator owns, so every availability figure, criticality rank and root cause finding traces to its source. |
| Feeder Firewatch | United States | a live ignition and outage risk score for every distribution feeder segment, built on the cooperative's own hardware, with every de-energisation and fast-trip change approved by a named operator. |
Which AI models can an energy and utilities operator run on its own hardware?
Each published energy and utilities design names its models and why, and every model is open-weight or no model is used at all, so the operator can run it on hardware it owns.
| Design | Models |
|---|---|
| Grid Context Watch | GLM 5.2 (MIT) for reasoning and Qwen3-Embedding-0.6B (Apache-2.0) for retrieval; detection stays in the procured sensors |
| Reliability Atlas | None: the register is empty by design because no inference workload is contracted |
| Feeder Firewatch | Self-hosted open-weight models: GLM 5.2 under MIT (reasoning) on site, RF-DETR (vision) at the edge |
| Design | Choice | What was picked | Why |
|---|---|---|---|
| Grid Context Watch | Reasoning model | GLM 5.2, 753B parameters, frontier class, FP8, 1M context | Plain MIT license with no revenue trigger or security-review clause; one node of eight 141 GB HBM GPUs holds it; the work surface's role demands long-context agentic reasoning |
| Grid Context Watch | Embedding model | Qwen3-Embedding-0.6B, Apache-2.0, 32K window | Whole playbooks and CIP procedures retrieve as single passages; runs beside the frontier model for about 1.2 GB at BF16 |
| Grid Context Watch | Sizing rule | 1.2 planning factor over FP8 weights | 753 GB of weights times 1.2 is 904 GB against 1,128 GB of node memory, leaving governed headroom for KV cache and activations |
| Grid Context Watch | Frontier compute | One node of eight 141 GB HBM GPUs | The smallest class that holds the FP8 weights; sized from the model's filed parameters |
| Grid Context Watch | Edge compute | Fanless, extended-temperature, cabinet-mount class | Compute stays out of electronic security perimeters wherever topology allows |
| Grid Context Watch | Enclosure and power | NEMA 3R/4X class cabinets with uninterruptible power supplies under Network UPS Tools | Clean shutdown ordering protects capture integrity environment |
| Grid Context Watch | Time synchronization | OCP Time Card GNSS grandmaster with holdover, Linuxptp and chrony | Pcap, sensor and log clocks must agree for evidence to correlate |
| Grid Context Watch | One-way transfer | Hardware data diodes with the one-way transfer appliance at highest-impact perimeters | Demonstrable one-way flow where policy demands it; constrained DMZ conduit elsewhere |
| Grid Context Watch | Site networking | Hardened, fanless, dual-DC industrial switches with SPAN/TAP ports | Segmentation remains the operator's policy decision; the design reads a copy, never the control path |
| Grid Context Watch | Sensing | Passive network traffic capture via the procured OT monitoring sensors at SPAN/TAP sources | The scope watches OT traffic, not premises; no cameras or video models |
| Grid Context Watch | Primary pattern | System of Context over one Hyper-style object model of 14 objects | The lasting asset is the living model of the OT estate, not any single alert |
| Grid Context Watch | Ground | The operator's own on-premises OT hardware and data center | Ownership of weights, ontology, decision record and boundary stays with the operator; identity and update paths run inside the air gap |
| Reliability Atlas | Models | None: the register is empty by design | No inference workload is contracted; the deliverable is the study, and an empty register recorded plainly outranks a speculative entry the scope cannot justify |
| Reliability Atlas | Hardware classes and sizing rules | None procured; the study runs on the operator's own in-country servers, storage and data center capacity | No field hardware, no edge compute and no serving runtime is installed under this scope |
| Reliability Atlas | Sensing | No cameras, no sensors and no positioning; site verification is by engineer walkdowns | The world is read through records, so nothing is watched and nothing needs country legality stated |
| Reliability Atlas | The pattern it stands on | System of context as the primary pattern, with the reliability object model as a projection over the three systems of record | Unavailability ranking, sparing confirmation and RCA validation are cross-silo joins that no single source system can produce alone |
| Reliability Atlas | The ground it runs on | In-country hosting in Saudi Arabia, with identity through the operator's own directory and single sign-on, procured by class | The operator's requirement keeps data inside the country's boundary and under the operator's identity, which the operator already operates |
| Feeder Firewatch | Language model parameters filed, FP8, MIT license | GLM 5.2, mixture-of-experts, 753 billion weights the cooperative owns; one 8-GPU | Agentic reasoning over the ontology with FP8 node holds it |
| Feeder Firewatch | Detection model to 68 MB and 2XL at about 254 MB, BF16, Apache-2.0 | RF-DETR, Nano to Large checkpoints at 61 detection on the edge box class, fine-tuned on the cooperative's own fire-season frames | Smoke, pole, vegetation and equipment |
| Feeder Firewatch | Edge compute class and sealed, NEMA 3R/4 enclosures | Industrial edge accelerator modules, fanless streams, rated for dust, heat, hail and cold | Sized decode-first from the actual camera |
| Feeder Firewatch | Site inference server center, one node of 8 GPUs of the 141 GB HBM class | Ruggedised server class at the operations sized from filed parameters at FP8 | 1,128 GB against a 904 GB requirement, |
| Feeder Firewatch | Serving runtimes Runtime or OpenVINO class at the edge | vLLM on site; Triton Inference Server, ONNX a bench measurement of the real streams | The edge runtime is pinned at design against |
| Feeder Firewatch | Sizing rules thermal beats visible; reuse existing cameras or not; size GPUs from filed parameters | Size edge compute from streams; when measured duty rather than a vendor default | Every hardware choice derives from a |
| Feeder Firewatch | Sensing camera feeds, truck and drone cameras, mesonet wind and humidity, 310 recloser fault indicators, AMI last gasp from 61,000 meters, crew AVL on 20 trucks | 42 substation PTZ cameras, 12 wildfire thermal only where a coverage gap is proven | Reused through the reuse gates; new fixed |
| Feeder Firewatch | Patterns read-only systems of record; human-approved write-back; one-way diode for SCADA | System of context; adapters-only ingestion; record authoritative while the platform joins them | Keeps the grid protected and the systems of |
| Feeder Firewatch | Ground as primary; sovereign cloud for overflow and recovery only | Cooperative servers at the operations center during a storm, and the risk picture cannot live outside the boundary | The model must survive an internet outage |
How much compute and hardware does AI in energy and utilities need?
The compute follows from the models: the published energy and utilities designs size it as follows, from no new hardware to a full GPU node.
| Design | Part | The design |
|---|---|---|
| Grid Context Watch | Compute | One node of eight 141 GB HBM-class GPUs: 753 GB of FP8 weights, 904 GB with the 1.2 planning factor, against 1,128 GB |
| Grid Context Watch | Boundary | Read only and one way out of the control networks; nothing writes into an electronic security perimeter, and no data leaves the operator |
| Reliability Atlas | Compute | No hardware bought; the study runs on the operator's own in-country servers |
| Feeder Firewatch | Frontier compute | One node of eight 141 GB HBM-class GPUs holds GLM 5.2 at FP8 (753 GB of weights, 904 GB with headroom) |
| Feeder Firewatch | Edge | Sealed industrial boxes at substations and on patrol trucks detect smoke and damaged equipment when cellular coverage drops |
Is it cheaper to own AI hardware or rent cloud GPUs in energy and utilities?
Each energy and utilities design prices three years of ownership in its Appendix A, with every price cited, against renting the same capacity from a cloud region at its deepest three-year commitment where hardware is bought.
| Design | Line | Three years |
|---|---|---|
| Grid Context Watch | Three-year cost | About 510,000 dollars to own the reasoning node, about four fifths the deepest three-year AWS commitment (Appendix A) |
| Reliability Atlas | Three-year cost | No hardware line to price: the cost is the assessment work itself (Appendix A) |
| Feeder Firewatch | Three-year cost, owned | About 841,000 US dollars with support and power at the Texas industrial power price |
| Feeder Firewatch | Three-year cost, rented | 0.88 million to 1.92 million US dollars for the same GPUs around the clock; ownership costs about the same as the deepest three-year commitment (version 2) |
| Feeder Firewatch | Closed model break-even | The cheapest closed model matches the owned stack at about 41 users; above that, ownership is cheaper and the gap grows with every user |
What ontology or object model does an energy and utilities AI system need?
The object model is the part that makes the system an ontology rather than a pipeline: typed objects for the things in the energy and utilities operation, with properties, status values and typed links. Every published design ships its object model as hyper-ontology/1 JSON that loads into Hyper Ontology.
| Design | Size | Objects | Download |
|---|---|---|---|
| Grid Context Watch | 14 objects, 14 links | Substation, OT Asset, Network Sensor, Anomaly Alert, Investigation Case, Vulnerability Finding, Detection Rule, Threat Intelligence Item, Pcap Evidence Record, CIP Evidence Artefact, OT Security Analyst, SIEM, Energy management system, Substation data platform | objects.json |
| Reliability Atlas | 13 objects, 14 links | Plant, Production train, Equipment, Failure mode, Work order, Downtime event, Availability prediction, Availability target, Criticality index, RCA report, Major event, Spare part, Reliability program | objects.json |
| Feeder Firewatch | 14 objects, 12 links | Substation, Feeder, Feeder segment, Pole, Recloser, Meter, Pole inspection record, Outage event, Feeder segment ignition risk score, Red flag warning, Wildfire camera station, Field crew, Work order, PSPS decision record | objects.json |
Who approves the decisions an AI system makes in energy and utilities?
In every published energy and utilities design a named person makes the decision that changes the physical world; the system prepares it.
| Design | Human control |
|---|---|
| Grid Context Watch | Every triage, case and risk acceptance carries a named OT security analyst; agents draft, the analyst decides |
| Reliability Atlas | Predictions are accepted only by a named reviewer and RCA reports validated only by a named engineer |
| Feeder Firewatch | Every public safety power shutoff (PSPS) recommendation becomes a decision record a named operator approves or declines; the design never opens or closes a recloser |
Every answer on this page is drawn from the papers linked in it: each paper's At a glance table, model register and cost appendix. Full text for agents: llms-full.txt. Designed on Praxis; object models load into Hyper Ontology.