Atoms
Join the Praxis beta

Benchmark · PADI

Physical AI Design Index

One question about system design: given the same short requirement for physical AI in a real operation, does Praxis produce a better system design than frontier AI on its own? PADI scores both on fresh tasks across 13 industries in three countries.

36.5%vs 26.0%

RSI Loop 16: Praxis platform vs frontier AI, checks passed

48.6%vs 35.1%

Defense and Intelligence: Praxis platform vs frontier AI, all RSI loops

10of 13

Industries measured across three countries

PADI is the benchmark CodeNinja uses to improve Praxis, one RSI loop (recursive self improvement loop) at a time. Scores are far from saturated: the goal is every check on every task in all 13 industries.

Industries

Checks passed in each industry, all RSI loops to date, for the Praxis platform and frontier AI on its own, on the same tasks. Three industries are in the bank and not yet measured.

Praxis platformFrontier AI
  • Agriculture and Earth ObservationRSI Loop 14 · 20 checks
    Praxis35.0%
    Frontier40.0%
  • Aviation Manufacturing3 tasks in the bank
    Not yet measured
  • Defense and IntelligenceRSI Loops 13, 15 · 37 checks
    Praxis48.6%
    Frontier35.1%
  • Discrete Manufacturing and AutomotiveRSI Loops 13, 15 · 37 checks
    Praxis32.4%
    Frontier32.4%
  • Energy and UtilitiesRSI Loop 13 · 76 checks
    Praxis27.6%
    Frontier34.2%
  • Heavy Industry and ConstructionRSI Loop 12 · 56 checks
    Praxis44.6%
    Frontier46.4%
  • Maritime and PortsRSI Loops 12, 16 · 75 checks
    Praxis33.3%
    Frontier26.7%
  • MiningRSI Loops 14, 15, 16 · 91 checks
    Praxis46.2%
    Frontier40.7%
  • Oil and GasRSI Loop 14 · 39 checks
    Praxis28.2%
    Frontier25.6%
  • Rail TransportRSI Loops 14, 15, 16 · 77 checks
    Praxis41.6%
    Frontier35.1%
  • Semiconductors3 tasks in the bank
    Not yet measured
  • Supply Chain and LogisticsRSI Loops 15, 16 · 38 checks
    Praxis34.2%
    Frontier26.3%
  • Warehousing and Intralogistics3 tasks in the bank
    Not yet measured

Trajectory

Every RSI loop on a scale to 100%, so the distance to the goal stays in view.

Praxis platformFrontier AI
0%25%50%75%100%Goal: every checkRSI Loop 12RSI Loop 13RSI Loop 14RSI Loop 15RSI Loop 1638.4%32.1%39.6%35.7%26.0%37.5%29.5%43.2%41.7%36.5%

Different tasks each RSI loop, five or six tasks per RSI loop, so RSI loops are not like for like.

Results

Every figure is checks passed as a percentage of checks judged, for each arm.

By RSI loop

RSI LoopDateTasksChecks per armPraxis platformFrontier AITasks: platform ahead · tied · behindRubric
RSI Loop 123 Oct 2026611237.5%38.4%3 · 0 · 30.1
RSI Loop 134 Oct 2026611229.5%32.1%3 · 0 · 30.2
RSI Loop 145 Oct 2026611143.2%39.6%2 · 1 · 30.2
RSI Loop 155 Oct 2026611541.7%35.7%5 · 0 · 10.2
RSI Loop 167 Oct 202659636.5%26.0%5 · 0 · 00.2
All RSI loops2954637.7%34.6%18 · 1 · 10

By industry

IndustryRSI LoopChecks per armPraxis platformFrontier AI
Agriculture and Earth ObservationRSI Loop 142035.0%40.0%
Defense and IntelligenceRSI Loop 131838.9%22.2%
Defense and IntelligenceRSI Loop 151957.9%47.4%
Discrete Manufacturing and AutomotiveRSI Loop 131827.8%33.3%
Discrete Manufacturing and AutomotiveRSI Loop 151936.8%31.6%
Energy and UtilitiesRSI Loop 137627.6%34.2%
Heavy Industry and ConstructionRSI Loop 125644.6%46.4%
Maritime and PortsRSI Loop 125630.4%30.4%
Maritime and PortsRSI Loop 161942.1%15.8%
MiningRSI Loop 143366.7%54.5%
MiningRSI Loop 153935.9%35.9%
MiningRSI Loop 161931.6%26.3%
Oil and GasRSI Loop 143928.2%25.6%
Rail TransportRSI Loop 141942.1%42.1%
Rail TransportRSI Loop 152055.0%40.0%
Rail TransportRSI Loop 163834.2%28.9%
Supply Chain and LogisticsRSI Loop 151827.8%22.2%
Supply Chain and LogisticsRSI Loop 162040.0%30.0%

By check family

RSI LoopMust name: Praxis vs frontier AIMust flag: Praxis vs frontier AIMust never: Praxis vs frontier AI
RSI Loop 1236.8% vs 33.3%15.2% vs 15.2%72.7% vs 86.4%
RSI Loop 1322.8% vs 17.5%17.6% vs 20.6%66.7% vs 90.5%
RSI Loop 1441.1% vs 28.6%21.2% vs 18.2%81.8% vs 100.0%
RSI Loop 1535.7% vs 23.2%20.0% vs 17.1%87.5% vs 91.7%
RSI Loop 1626.0% vs 14.0%11.5% vs 0.0%95.0% vs 90.0%

By country

RSI LoopUnited States: Praxis vs frontier AISaudi Arabia: Praxis vs frontier AIPakistan: Praxis vs frontier AI
RSI Loop 1234.2% vs 44.7%38.9% vs 44.4%39.5% vs 26.3%
RSI Loop 1324.3% vs 37.8%30.8% vs 30.8%33.3% vs 27.8%
RSI Loop 1435.0% vs 30.0%66.7% vs 54.5%31.6% vs 36.8%
RSI Loop 1556.4% vs 43.6%33.3% vs 28.2%35.1% vs 35.1%
RSI Loop 1635.9% vs 28.2%39.5% vs 23.7%31.6% vs 26.3%

Coverage

10 of 13 industries measured. 60 tasks in the bank, at least one per industry in each country, because operations differ with the place.

RSI Loop 16
measured in that RSI loop
Held out
kept for a final blind evaluation
Not run
in the bank, not yet measured
IndustryUnited StatesSaudi ArabiaPakistan
Agriculture and Earth Observation
RSI Loop 14
Held out
Held out
Aviation Manufacturing
Not run
Held out
Held out
Defense and Intelligence
RSI Loop 15
Held out
RSI Loop 13
Discrete Manufacturing and Automotive
Not runNot runNot run
Not runRSI Loop 15Held out
RSI Loop 13
Energy and Utilities
RSI Loop 13Held outRSI Loop 13
RSI Loop 13RSI Loop 13Held out
Held out
Heavy Industry and Construction
RSI Loop 12
RSI Loop 12
RSI Loop 12
Maritime and Ports
RSI Loop 12Not runNot run
RSI Loop 12RSI Loop 16Held out
RSI Loop 12
Mining
Not runNot runHeld out
RSI Loop 14RSI Loop 14RSI Loop 15
RSI Loop 15RSI Loop 16
Oil and Gas
RSI Loop 14
Held out
RSI Loop 14
Rail Transport
RSI Loop 15RSI Loop 16Held out
RSI Loop 16Held outNot run
RSI Loop 14
Semiconductors
Not run
Held out
Not run
Supply Chain and Logistics
RSI Loop 16
Not run
RSI Loop 15
Warehousing and Intralogistics
Held out
Held out
Not run

A task

What one task looks like, described generically. Task text, packs and checks are never published, so the held out set stays clean.

1 · The requirement

A short pack

What a customer would send for one operation: at most two pages, anonymised, for a site in the United States, Saudi Arabia or Pakistan. Packs leave out facts a good design should ask for.

2 · The design

A system design

Both arms get the same pack and the same instruction: produce the first system design a customer will argue with, under fixed headings, among them:

  1. What it does
  2. What it reads
  3. Where things run
  4. Sensing classes
  5. Models and provenance
  6. Human decisions
  7. Phase one and its exit test
  8. Open questions
  9. Risks

3 · The checks

Yes or no

15 to 20 checks per task in three families. Illustrative examples, not taken from any task:

Must nameThe design states which decisions a person confirms before anything acts, and what phase one must demonstrate before it scales.

Must flagThe design asks whether the existing cameras have been tested for accuracy in this site's own lighting before any reported figure depends on them.

Must neverThe design never names a camera, sensor or server by part number or model code. An open weight model named for self hosting is not a part number.

Methodology

Both arms run the same model on the same pack, so the difference between them is the platform, not the model.

Arms

Same model, twice

Frontier AI · glm-5.3-flash
The model alone, given the same instruction and the same requirement pack. Nothing else.

Praxis platform · glm-5.3-flash
The same model inside the Praxis planner, with its sourced corpus of reference designs and its own quality gate.

Judge

Three lenses

glm-4.6 reads each design and answers every check yes or no through three lenses: strict reviewer, plant engineer and technical assessor.

A check passes when the majority of the lenses say yes.

Families

Name, flag, never

Must name · 50% of the bank
What the design has to state: the systems it reads, where things run, which decisions a person confirms, what phase one must prove.

Must flag · 30% of the bank
What the design has to raise: a missing fact asked as a question, thin evidence said out loud.

Must never · 20% of the bank
What disqualifies a design: a part number, an invented saving, a customer name, raw data leaving the site when residency forbids it.

Fresh tasks

Never run twice

A third of the bank (17 of 60 tasks) is held out for a final blind evaluation and never run in an RSI loop.

Each RSI loop picks up to six unused tasks from the rest, two per country where the pool allows. No task or check ever enters the platform's corpus or briefs.

Score

Checks passed

Checks passed divided by checks total, for each arm. No partial credit and no weighting.

Industry, family and country figures are the same count over their subset. RSI loops use different tasks, so compare percentages, not counts.

Rubric

Version 0.2

Version 0.1 judged RSI Loop 12. Version 0.2 (4 October 2026) judges RSI Loops 13 onward: an open weight model named for self hosting, or a product named with its licence, is not a part number.

Limitations

What PADI is good for: seeing whether Praxis improves on design judgement from one release to the next, on tasks it has never seen, and where it still misses. What it is not:

  • Design judgement, not deployment. PADI scores designs on paper. No plant, sensor or model was run.
  • One judge model. A single model reads every design through three lenses. Judging the same text again can move a few checks. No human has scored the outputs yet.
  • Small samples. Each RSI loop is five or six tasks and roughly a hundred checks per arm. A swing of a few points is within noise, so no single RSI loop is a trend.
  • Different tasks every RSI loop. Fresh tasks keep the measure honest, and they also mean RSI loops are not like for like.
  • Written by CodeNinja. Tasks and checks were written by CodeNinja with model assistance and reviewed adversarially. Independent expert grading is planned and not yet done.
  • Two arms. Only frontier AI on its own and the same model inside Praxis are scored. No other system or model is in the index yet.
  • Partial coverage. Ten of thirteen industries are measured. Aviation manufacturing, semiconductors and warehousing have not been run.
  • One rubric change. RSI Loop 12 used rubric 0.1. Every later RSI loop uses 0.2.

Questions

What is the Physical AI Design Index?

PADI is CodeNinja's benchmark for system design in physical AI. Each task is a short anonymised requirement for an AI system in a real kind of operation, such as a mine, a port, a rail line or a substation, plus yes or no checks. A judge model scores the system design from frontier AI on its own and from the same model inside the Praxis platform.

What does PADI measure?

Whether a system design is complete and sound: what it must state (the systems it reads, where things run, which decisions a person confirms, what phase one must prove), what it must raise (a missing fact asked as a question, thin evidence said out loud) and what it must never do (a part number, an invented saving, a customer name, data leaving the site when residency forbids it).

How is PADI scored?

Checks passed divided by checks total, for each arm, with no partial credit and no weighting. The judge is glm-4.6 through three lenses (a strict reviewer, a plant engineer and a technical assessor); a check passes when the majority say yes. Both arms use glm-5.3-flash.

How often is PADI updated?

Once per RSI loop, on tasks never run before. A third of the task bank is held out for a final blind evaluation and never run in an RSI loop.

Cadence

Updated each RSI loop. Last published: RSI Loop 16, 7 October 2026. RSI Loop 17 is under way.

Contact

Questions about PADI, or a task from your own operation you would like measured: hello@codeninjaconsulting.com