MySite

From Distribution Lists to a Rule Evaluation Engine

Imagine a school with five separate locations and one of them asks: “Give us a distribution list containing all the employees excluding volunteers and interns”. Not a big problem. We have our IAM system (HelloID) and we can make a bussiness rule with the following scope:

  • location = “school A”
  • role != “volunteer”
  • role != “intern”
  • type = “employee”

This business rule is applied to all users and if it matches the criteria we run a script that adds the user to the distribution list. Case closed? Nah, not really.

School B comes along and says “We want the same distribution list but we also want to include the volunteers”. So we add a new rule with the following scope:

  • location = “school B”
  • role != “intern”
  • type = “employee”

Again, this rule is applied to all users and if it matches the criteria we run a script that adds the user to the distribution list.

Showing some issues.

So now we have two rules that are applied to all users. And the problem begins to emerge:

  • each rule has to be applied to all users, so the more rules we have, the more processing time it takes to evaluate all users against all rules.
  • each rule has simular logic to the other rules, so we have a lot of duplicate code.
  • and with each new rule we have to add more code, and the code is getting more complex and harder to maintain.

Wouldn’t it be nice to let the code do the scoping for us and get rid of the duplicate bussiness rules and code?

Making spaghetti.

At first that was exactly what I did to solve the problem in the code, scope the rule to every employee and let the code do the further scoping. Something like this:

if (user.type == "employee") {
    if (user.location.name == "school A" && user.role != "volunteer" && user.role != "intern") {
        addToDistributionList("List A", user);
    } else if (user.location.nae == "school B" && user.role != "intern") {
        addToDistributionList("List B", user);
    }
}

Should work right?

Well, I don’t know about you, but I think it’s rather ugly. And imagine the other schools wanting to have disctribution list, teams for certain groups of people, etc. The code will get more and more complex and harder to maintain. Before you know it you’ll be adding people to the wrong distribution list or team. That’s a collection of special cases waiting to break.

Analyzing the problem.

So I investigated further to find a better solution.

I realized that regardless of the desired result, I was doing exactly the same thing :

  • Check if a condition is met.
  • If it is, evaluate the next condition.
  • If all conditions match, produce a result.

Looking at the conditions themselves, they always boiled down to the same three things:

  • Identify the property to check.
  • Choose an operator to use.
  • Choose a value to compare against.

For example, a condition such as: “location.name equals school A”

contains exactly those three elements:

  • Property: location.name
  • Operator: Equals
  • Value: "school A"

A real-world business rule is simply a collection of such conditions that all need to be satisfied before a result can be produced.

This is when I realized I was no longer solving a distribution list problem.

I was solving a rule evaluation problem.

A declarative way of doing things.

Looking back at the analysis, the next question became obvious:

Why express conditions in code at all?

If a condition can always be described by a property, an operator and a value, perhaps I could describe it as data and let the code evaluate it.

A condition such as:

location.name equals school A

could be represented as:

    {
        'path': 'location.name',
        'operator': 'Equals',
        'check': 'School A'
    }

This structure contains everything needed to describe a condition.

A real-world business rule would, of course, consist of multiple conditions that all need to be satisfied. In a data structure this can simply be achieved by placing the condition in a collection of conditions:

{
    'condition' : [
        {
            'path': 'location.name',
            'operator': 'Equals',
            'check': 'School A'
        },
        {
            'path': 'contract.role',
            'operator': 'NotEquals',
            'check': 'volunteer'
        },
        {
            'path': 'contract.role',
            'operator': 'NotEquals',
            'check': 'intern'
        }
    ]
}

To make this useful, we also need a result which is returned when they all match. By adding this the structure becomes self contained:

{
    'conditions' : [
        { ... },
        { ... },
        { ... }
    ],
    'result': 'List A'
}

This is what I call a ConditionSet.

A ConditionSet contains a collection of Conditions and a Result. If all conditions match, the result can be returned. This means that within a ConditionSet AND logic is used.

The interesting part is that the complexity of the code no longer grows with the complexity of the rule. Whether a ConditionSet contains two conditions or twenty conditions, the evaluation logic remains exactly the same.

The rule becomes more complex, but the code does not.

But wait, what about that second list?

So far I’ve been able to describe one list. But the world is more complex than that, we need to be able to model more than one list. School A wants this list, School B another. And, no doubt about, there will be other requests.

And the solution is staring us in the face.

A ConditionSet is nothing more than a collection of Conditions. If that works, why wouldn’t that same idea work a level higher?

In other words: if I can have a collection of Conditions, I can also have a collection of ConditionSets. Like this:

[
    {
        "conditions": [ ... ],
        "result": "List A"
    },
    {
        "conditions": [ --- ],
        "result": "List B"
    }
]

This is what I call a RuleSet: a collection of ConditionSets.

RuleSet
|- ConditionSet
|- ConditionSet
|_ ConditionSet

Each ConditionSet is evaluated independently. If it matches , its result is returned. This gives an implied OR logic between ConditionSets.

And the nice thing is that adding a new ConditonSet does not change the evaluation logic. The code stays the same. Only the RuleSet grows by adding more ConditionSets.

And that gives us OR logic

Let’s assume both School A and School B should be considered business staff.

Instead of creating one giant rule with complicated logic, I can simply define two ConditionSets that both return the same result:

  • School A -> “Business Staff”
  • School B -> “Business Staff”

If either ConditionSet matches, the result “Business Staff” is returned.

In effect, OR logic emerges naturally from the model.

Bringing it all together

At this point I’ve got all the building blocks:

  • Conditions
  • ConditionSets
  • RuleSets

All that is needed now is a function that could evaluate a RuleSet against a piece of data and return all matching results.

That function evantually became: Invoke-StructMatcher

$results = Invoke-StructMatcher `
    -Rules $rules `
    -Data $user

It might return:

[
    "List A",
    "Bussiness Staff"
]

or even an empty array if there are no matches.

What StructMatcher is not.

Something worth mentioning is that StructMatcher is completely agnostic about what a result actually means.

To StructMatcher, a result is simply a value

It could be the name of a distribution list, like I use in the examples, but it could just as well be the name of a security group, a Team. The module evaluates a RuleSet against Data, giving meaning to the result is up to the caller.

This separation is intentional. StructMatcher is a rule evaluation engine, not a provisioning engine.

HelloID is where the problem first presented itself, but neither HelloID nor JSON is fundamental to the solution. StructMatcher operates on structured data; JSON is simply a convenient way to represent it, and native PowerShell structures work just as well. The same rule evaluation model can therefore be used anywhere the problem can be expressed as conditions and results

Learn more

StructMatcher is availlable on GitHub:

The README contains the full function documentation, supported operators, JSON (and other) schema detail and additional examples.