Benchmark · PADI
Physical AI Design Index
One question about system design: given the same short requirement for physical AI in a real operation, does Praxis produce a better system design than frontier AI on its own? PADI scores both on fresh tasks across 13 industries in three countries.
RSI Loop 16: Praxis platform vs frontier AI, checks passed
Defense and Intelligence: Praxis platform vs frontier AI, all RSI loops
Industries measured across three countries
PADI is the benchmark CodeNinja uses to improve Praxis, one RSI loop (recursive self improvement loop) at a time. Scores are far from saturated: the goal is every check on every task in all 13 industries.
Industries
Checks passed in each industry, all RSI loops to date, for the Praxis platform and frontier AI on its own, on the same tasks. Three industries are in the bank and not yet measured.
- Agriculture and Earth ObservationRSI Loop 14 · 20 checks
- Aviation Manufacturing3 tasks in the bank
- Defense and IntelligenceRSI Loops 13, 15 · 37 checks
- Discrete Manufacturing and AutomotiveRSI Loops 13, 15 · 37 checks
- Energy and UtilitiesRSI Loop 13 · 76 checks
- Heavy Industry and ConstructionRSI Loop 12 · 56 checks
- Maritime and PortsRSI Loops 12, 16 · 75 checks
- MiningRSI Loops 14, 15, 16 · 91 checks
- Oil and GasRSI Loop 14 · 39 checks
- Rail TransportRSI Loops 14, 15, 16 · 77 checks
- Semiconductors3 tasks in the bank
- Supply Chain and LogisticsRSI Loops 15, 16 · 38 checks
- Warehousing and Intralogistics3 tasks in the bank
Trajectory
Every RSI loop on a scale to 100%, so the distance to the goal stays in view.
Different tasks each RSI loop, five or six tasks per RSI loop, so RSI loops are not like for like.
Results
Every figure is checks passed as a percentage of checks judged, for each arm.
By RSI loop
| RSI Loop | Date | Tasks | Checks per arm | Praxis platform | Frontier AI | Tasks: platform ahead · tied · behind | Rubric |
|---|---|---|---|---|---|---|---|
| RSI Loop 12 | 3 Oct 2026 | 6 | 112 | 37.5% | 38.4% | 3 · 0 · 3 | 0.1 |
| RSI Loop 13 | 4 Oct 2026 | 6 | 112 | 29.5% | 32.1% | 3 · 0 · 3 | 0.2 |
| RSI Loop 14 | 5 Oct 2026 | 6 | 111 | 43.2% | 39.6% | 2 · 1 · 3 | 0.2 |
| RSI Loop 15 | 5 Oct 2026 | 6 | 115 | 41.7% | 35.7% | 5 · 0 · 1 | 0.2 |
| RSI Loop 16 | 7 Oct 2026 | 5 | 96 | 36.5% | 26.0% | 5 · 0 · 0 | 0.2 |
| All RSI loops | 29 | 546 | 37.7% | 34.6% | 18 · 1 · 10 |
By industry
| Industry | RSI Loop | Checks per arm | Praxis platform | Frontier AI |
|---|---|---|---|---|
| Agriculture and Earth Observation | RSI Loop 14 | 20 | 35.0% | 40.0% |
| Defense and Intelligence | RSI Loop 13 | 18 | 38.9% | 22.2% |
| Defense and Intelligence | RSI Loop 15 | 19 | 57.9% | 47.4% |
| Discrete Manufacturing and Automotive | RSI Loop 13 | 18 | 27.8% | 33.3% |
| Discrete Manufacturing and Automotive | RSI Loop 15 | 19 | 36.8% | 31.6% |
| Energy and Utilities | RSI Loop 13 | 76 | 27.6% | 34.2% |
| Heavy Industry and Construction | RSI Loop 12 | 56 | 44.6% | 46.4% |
| Maritime and Ports | RSI Loop 12 | 56 | 30.4% | 30.4% |
| Maritime and Ports | RSI Loop 16 | 19 | 42.1% | 15.8% |
| Mining | RSI Loop 14 | 33 | 66.7% | 54.5% |
| Mining | RSI Loop 15 | 39 | 35.9% | 35.9% |
| Mining | RSI Loop 16 | 19 | 31.6% | 26.3% |
| Oil and Gas | RSI Loop 14 | 39 | 28.2% | 25.6% |
| Rail Transport | RSI Loop 14 | 19 | 42.1% | 42.1% |
| Rail Transport | RSI Loop 15 | 20 | 55.0% | 40.0% |
| Rail Transport | RSI Loop 16 | 38 | 34.2% | 28.9% |
| Supply Chain and Logistics | RSI Loop 15 | 18 | 27.8% | 22.2% |
| Supply Chain and Logistics | RSI Loop 16 | 20 | 40.0% | 30.0% |
By check family
| RSI Loop | Must name: Praxis vs frontier AI | Must flag: Praxis vs frontier AI | Must never: Praxis vs frontier AI |
|---|---|---|---|
| RSI Loop 12 | 36.8% vs 33.3% | 15.2% vs 15.2% | 72.7% vs 86.4% |
| RSI Loop 13 | 22.8% vs 17.5% | 17.6% vs 20.6% | 66.7% vs 90.5% |
| RSI Loop 14 | 41.1% vs 28.6% | 21.2% vs 18.2% | 81.8% vs 100.0% |
| RSI Loop 15 | 35.7% vs 23.2% | 20.0% vs 17.1% | 87.5% vs 91.7% |
| RSI Loop 16 | 26.0% vs 14.0% | 11.5% vs 0.0% | 95.0% vs 90.0% |
By country
| RSI Loop | United States: Praxis vs frontier AI | Saudi Arabia: Praxis vs frontier AI | Pakistan: Praxis vs frontier AI |
|---|---|---|---|
| RSI Loop 12 | 34.2% vs 44.7% | 38.9% vs 44.4% | 39.5% vs 26.3% |
| RSI Loop 13 | 24.3% vs 37.8% | 30.8% vs 30.8% | 33.3% vs 27.8% |
| RSI Loop 14 | 35.0% vs 30.0% | 66.7% vs 54.5% | 31.6% vs 36.8% |
| RSI Loop 15 | 56.4% vs 43.6% | 33.3% vs 28.2% | 35.1% vs 35.1% |
| RSI Loop 16 | 35.9% vs 28.2% | 39.5% vs 23.7% | 31.6% vs 26.3% |
Coverage
10 of 13 industries measured. 60 tasks in the bank, at least one per industry in each country, because operations differ with the place.
| Industry | United States | Saudi Arabia | Pakistan |
|---|---|---|---|
| Agriculture and Earth Observation | RSI Loop 14 | Held out | Held out |
| Aviation Manufacturing | Not run | Held out | Held out |
| Defense and Intelligence | RSI Loop 15 | Held out | RSI Loop 13 |
| Discrete Manufacturing and Automotive | Not runNot runNot run | Not runRSI Loop 15Held out | RSI Loop 13 |
| Energy and Utilities | RSI Loop 13Held outRSI Loop 13 | RSI Loop 13RSI Loop 13Held out | Held out |
| Heavy Industry and Construction | RSI Loop 12 | RSI Loop 12 | RSI Loop 12 |
| Maritime and Ports | RSI Loop 12Not runNot run | RSI Loop 12RSI Loop 16Held out | RSI Loop 12 |
| Mining | Not runNot runHeld out | RSI Loop 14RSI Loop 14RSI Loop 15 | RSI Loop 15RSI Loop 16 |
| Oil and Gas | RSI Loop 14 | Held out | RSI Loop 14 |
| Rail Transport | RSI Loop 15RSI Loop 16Held out | RSI Loop 16Held outNot run | RSI Loop 14 |
| Semiconductors | Not run | Held out | Not run |
| Supply Chain and Logistics | RSI Loop 16 | Not run | RSI Loop 15 |
| Warehousing and Intralogistics | Held out | Held out | Not run |
A task
What one task looks like, described generically. Task text, packs and checks are never published, so the held out set stays clean.
1 · The requirement
A short pack
What a customer would send for one operation: at most two pages, anonymised, for a site in the United States, Saudi Arabia or Pakistan. Packs leave out facts a good design should ask for.
2 · The design
A system design
Both arms get the same pack and the same instruction: produce the first system design a customer will argue with, under fixed headings, among them:
- What it does
- What it reads
- Where things run
- Sensing classes
- Models and provenance
- Human decisions
- Phase one and its exit test
- Open questions
- Risks
3 · The checks
Yes or no
15 to 20 checks per task in three families. Illustrative examples, not taken from any task:
Must nameThe design states which decisions a person confirms before anything acts, and what phase one must demonstrate before it scales.
Must flagThe design asks whether the existing cameras have been tested for accuracy in this site's own lighting before any reported figure depends on them.
Must neverThe design never names a camera, sensor or server by part number or model code. An open weight model named for self hosting is not a part number.
Methodology
Both arms run the same model on the same pack, so the difference between them is the platform, not the model.
Arms
Same model, twice
Frontier AI · glm-5.3-flash
The model alone, given the same instruction and the same requirement pack. Nothing else.
Praxis platform · glm-5.3-flash
The same model inside the Praxis planner, with its sourced corpus of reference designs and its own quality gate.
Judge
Three lenses
glm-4.6 reads each design and answers every check yes or no through three lenses: strict reviewer, plant engineer and technical assessor.
A check passes when the majority of the lenses say yes.
Families
Name, flag, never
Must name · 50% of the bank
What the design has to state: the systems it reads, where things run, which decisions a person confirms, what phase one must prove.
Must flag · 30% of the bank
What the design has to raise: a missing fact asked as a question, thin evidence said out loud.
Must never · 20% of the bank
What disqualifies a design: a part number, an invented saving, a customer name, raw data leaving the site when residency forbids it.
Fresh tasks
Never run twice
A third of the bank (17 of 60 tasks) is held out for a final blind evaluation and never run in an RSI loop.
Each RSI loop picks up to six unused tasks from the rest, two per country where the pool allows. No task or check ever enters the platform's corpus or briefs.
Score
Checks passed
Checks passed divided by checks total, for each arm. No partial credit and no weighting.
Industry, family and country figures are the same count over their subset. RSI loops use different tasks, so compare percentages, not counts.
Rubric
Version 0.2
Version 0.1 judged RSI Loop 12. Version 0.2 (4 October 2026) judges RSI Loops 13 onward: an open weight model named for self hosting, or a product named with its licence, is not a part number.
Limitations
What PADI is good for: seeing whether Praxis improves on design judgement from one release to the next, on tasks it has never seen, and where it still misses. What it is not:
- Design judgement, not deployment. PADI scores designs on paper. No plant, sensor or model was run.
- One judge model. A single model reads every design through three lenses. Judging the same text again can move a few checks. No human has scored the outputs yet.
- Small samples. Each RSI loop is five or six tasks and roughly a hundred checks per arm. A swing of a few points is within noise, so no single RSI loop is a trend.
- Different tasks every RSI loop. Fresh tasks keep the measure honest, and they also mean RSI loops are not like for like.
- Written by CodeNinja. Tasks and checks were written by CodeNinja with model assistance and reviewed adversarially. Independent expert grading is planned and not yet done.
- Two arms. Only frontier AI on its own and the same model inside Praxis are scored. No other system or model is in the index yet.
- Partial coverage. Ten of thirteen industries are measured. Aviation manufacturing, semiconductors and warehousing have not been run.
- One rubric change. RSI Loop 12 used rubric 0.1. Every later RSI loop uses 0.2.
Questions
What is the Physical AI Design Index?
PADI is CodeNinja's benchmark for system design in physical AI. Each task is a short anonymised requirement for an AI system in a real kind of operation, such as a mine, a port, a rail line or a substation, plus yes or no checks. A judge model scores the system design from frontier AI on its own and from the same model inside the Praxis platform.
What does PADI measure?
Whether a system design is complete and sound: what it must state (the systems it reads, where things run, which decisions a person confirms, what phase one must prove), what it must raise (a missing fact asked as a question, thin evidence said out loud) and what it must never do (a part number, an invented saving, a customer name, data leaving the site when residency forbids it).
How is PADI scored?
Checks passed divided by checks total, for each arm, with no partial credit and no weighting. The judge is glm-4.6 through three lenses (a strict reviewer, a plant engineer and a technical assessor); a check passes when the majority say yes. Both arms use glm-5.3-flash.
How often is PADI updated?
Once per RSI loop, on tasks never run before. A third of the task bank is held out for a final blind evaluation and never run in an RSI loop.
Cadence
Updated each RSI loop. Last published: RSI Loop 16, 7 October 2026. RSI Loop 17 is under way.
Contact
Questions about PADI, or a task from your own operation you would like measured: hello@codeninjaconsulting.com