A rule that only moves things
· 6 min · rules · testing
There is a working reconstruction of this on the home page. Flip R5 from orders to books and look at the silver piece.
Every design we make needs a routing: the ordered list of steps a piece takes through the factory. Wax, casting, filing, setting, polishing, plating, inspection. The ERP holds a certified routing for every item it already makes. What it cannot give you is the routing for a design that does not exist yet, and new designs arrive every day.
So the routing is generated. It starts from a template, and rules adjust it. The rules live in tables the planners edit themselves. They know the plant and I do not, and a rule that needs me and a deployment before it can change will be out of date by the time it ships.
That part is well understood. It took me longer to learn that a rule editor holds two kinds of rule that look almost exactly alike, and that confusing them gets past every check I had.
Templates are mined, not written
The first decision was not to hand-write templates at all. The ERP's certified routings are the most honest record of how the plant actually works, so the templates come from them. Each routed item is reduced to its chain of steps, items are grouped by product scope, and the template is the chain the group shares.
One unglamorous step comes first. Work centres get retired and renamed over the years, and old routings still carry the old codes. Retired centres have to be folded onto their live replacements before anything is compared. Skip that, and two routings that describe the same work look different, so every difference the miner finds is noise.
The cost of mining is that the templates are only as fresh as the last build. Mine once went twelve days stale, in both directions, and nobody noticed until a template disagreed with a routing everyone knew was right.
Two rules that read the same
A rule has a condition and an effect. "When the piece has stones, book SETTING." "When the metal is gold, book HALLMARK." Each of those adds a step.
Then a planner asked for something that sounds like the same kind of rule: the hallmark has to go on before the final polish, because polishing cleans up the mark. In an editor built around steps, the natural way to write that is a rule on HALLMARK saying what it must be preceded by or come before.
The trouble was what "preceded by" meant in our engine: it booked the step it named. A rule written to say where the hallmark goes would also have said every piece gets a hallmark. That includes every silver piece, which an earlier rule says has none. Applied to the whole catalogue, it would have added a step to hundreds of templates that correctly had none.
Nothing flagged it. The rule was valid, the expression parsed, the build passed and the types checked. Every generated routing was a well-formed routing. It was simply wrong, and that kind of wrong only shows up if you already know what the routing should have been.
Every kind of rule should say what it may change
What caught it was a check I now run on every rule change, and which should have been there from the start. An ordering rule has a precise contract: it may change the order of steps and nothing else. So before and after an ordering change, each item's routing must contain exactly the same steps, as a multiset: the same steps, each the same number of times. Only the sequence may move.
A machine can run that check over every item in seconds, and it has no opinion on whether the new order is better. It only asks whether the rule stayed in its lane. The ordering change that finally shipped reported what it should: no steps gained, no steps lost, a few dozen routings reordered, and no ordering violations left.
The engine now keeps the two kinds apart in the data rather than in anyone's head. A booking rule adds a step. An ordering rule constrains where a step goes, and if either end of its constraint is missing from a routing, it has nothing to say. They are different kinds of row with different effects, and the editor shows which is which.
Graded against the ERP
A rule set you can only demonstrate is a rule set you have to take on trust. So the second half of the engine is a simulator. It runs every rule over the real catalogue and compares each result with the certified routing the ERP already holds, step by step. The two are aligned by longest common subsequence, so one missing step does not make everything after it look wrong.
It is deliberately not a second implementation. The comparison calls the same simulation the editor uses, so what gets graded is what the planners actually run.
The first honest run matched whole routings exactly zero percent of the time. Recall was 67% and precision 48%. It looked like a disaster, and it was the most useful number the project produced, because the alignment said where the error was. Most of the lost precision came from three rules. One of them had been set to fire on everything, so it booked its step onto every item in the run, where the ERP runs that step on a few dozen. Fixing three rows changed the picture, and nobody had to rethink the engine.
The templates graded well on their own, the best reaching 99.5% recall against the ERP. That split is the real result. The mined part was sound. The error lived in the hand-written part, which is exactly where you would expect it and exactly where nobody had looked.
Undecided is not false
The simulator forced one more distinction. Some rules depend on facts the ERP does not hold. For those items the honest answer is not "the rule does not fire" but "I cannot tell", and scoring that as false would have made the rule look wrong on every one of them.
So those items come back as undecided and are counted separately. For one rule, the ERP turned out to run the step on every item that came back undecided. That does not prove the rule is right, but it is strong evidence. If the engine had treated a missing fact as a no, the same rule would have been scored with hundreds of failures.
Handing over the editor
Letting the business edit its own rules is the right design, and I would build it again. But handing over the editor hands over its failure modes too. The dangerous rules are not the invalid ones, because the parser catches those. They are the valid rules that do more than their author meant. The defence is to give each kind of rule a contract for what it is allowed to change, and to check that contract mechanically, on every item, every time.