Scoring, version 1.0
100 points: 40 from a static review of the package, 60 from a live benchmark in NX. Nothing here is hidden from entrants except the referee itself.
Computed when the zip arrives, by reading it. Nothing in the package is executed. Every score is returned with the evidence it was awarded on and, where points were lost, what would earn them.
The static review measures how well a skill is written, by pattern. It cannot tell whether the skill works. That is why it is the smaller half.
A skill is a folder an agent can load. It has to arrive as one.
The description is all an agent reads before deciding to load the skill.
How much of a ship the skill can actually address, counted by area.
A recipe that was never run is a guess with good formatting.
The traps are the part of a skill an agent cannot work out for itself.
An agent should be able to act from the first screen.
After a reviewer approves the run, an agent is given your skill, the NX tools and the steps below. It has no other skill to fall back on. When it finishes, a referee script reads the saved model and reports plane positions, sheet areas, stored thicknesses and solid properties. The score is computed from those measurements, never from what the agent said it did.
A run that could not happen, because the workstation or the referee failed, is reported as could-not-run. It is never recorded as a zero.
One assembly, built step by step by an agent that has only the uploaded skill.
A hull whose every hydrostatic value has a closed form, so the answer is not negotiable.
The agent reports what it did. The referee checks whether that is true.
Published in full, so nobody is guessing. Dimensions are millimetres.
Skills rank by total score. Ties go to the skill with more referee checks passed, then to the earlier upload. One row per author and skill shows the best upload. A skill without a live result is ranked on its static score and labelled.