If you supervise people, this is the part of Workforce Intelligence you will actually use. It takes about five minutes per person, and it is deliberately different from the annual form you may be used to.
You are not scoring people out of five
The old way asks: “Rate Sipho 1–5 on problem solving.” Two supervisors answer that question completely differently, and neither can explain why.
This asks: “Which of these have you actually seen Sipho do?”
| Level | Behaviour |
|---|---|
| 1 | Recognises when something is wrong and escalates it accurately. |
| 2 | Solves routine issues using established procedures. |
| 3 | Diagnoses unfamiliar problems independently and to standard. |
| 4 | Resolves complex cross-functional problems and improves the process afterwards. |
| 5 | Creates organisational approaches that measurably improve how problems are solved. |
You pick the description that matches what you have observed. That is a much easier question to answer honestly, and a much harder one to answer arbitrarily.
Level 3 is the workplace standard — doing the job properly, unsupervised. It is not a mediocre score. Most people in most competencies should be at 3, and a role requiring 4 is asking for someone who coaches others and improves the practice.
Evidence changes how much your judgement counts
Attach evidence where you can. The engine weights an evidenced supervisor rating at 0.85 and an unevidenced one at 0.60.
That is not a criticism of your judgement. It reflects something true: a rating anyone can trace back to a specific observation is more defensible than one that rests on memory, and the person being rated has a right to know which kind they are getting.
Evidence can be an existing assessment result, a portfolio item, a logbook entry, a certificate, or just a note describing what you saw and when. The note costs you fifteen seconds and is the single highest-value thing you can add.
Write the observation note
Compare these:
Good at diagnostics.
Diagnosed an intermittent hydraulic fault on line 3 unaided on 12 August, after two other technicians had missed it.
The first is an opinion. The second is a fact that survives a disagreement six months later. Write the second one.
What happens after you submit
Your ratings join everything else known about that person — formal assessments, workplace evidence, their own self-rating, qualifications — and the engine combines them into one picture.
Three things will look odd at first, and all three are working correctly.
Most competencies will show an em dash rather than a number. A score is only published once enough verified evidence has accumulated. Your rating alone may not be enough yet. An em dash means not established, which is not the same as below standard — and the module will never report the second when it only knows the first.
Your rating and the self-rating may differ sharply. That is the point. A large gap between how someone sees themselves and how you see them is the most useful thing on the screen, and it is a conversation, not an error.
Old ratings lose weight over time. Evidence has a half-life — twelve months for safety and digital competencies, twenty-four for technical, longer for behavioural. Ageing evidence reduces the system’s confidence, not the person’s score. If you see a score described as declining because of ageing evidence, the module says so explicitly rather than implying performance has dropped.
Confidence, and why it matters to you
Every score carries a confidence figure out of 100, built from five things: how many different kinds of source agree, how much evidence is attached, how recent it is, whether the sources agree with each other, and how many ratings there are.
Anything below 30 is marked insufficient and is excluded from team averages and heatmaps entirely. Your organisation would rather show a gap in its data than a confident-looking number built on one person’s guess.
The practical implication: your rating raises confidence most when it adds something the record does not already have. An evidenced observation on a competency that has only ever been self-rated is worth far more than a fifth opinion on something already well established.
Calibration is not an accusation
After everyone has rated, your organisation may run a calibration session. It compares your ratings against the evidence-based score and against what your peers gave, and flags whether you are consistently more generous or more severe than they are.
Nearly every manager is one or the other. It is a pattern, not a failing, and the system deliberately flags it for a human to interpret rather than adjusting anything automatically. If a rating is changed, the person changing it must write down why, and your original judgement is kept alongside it rather than overwritten.
The person you rated can challenge it
They can acknowledge it, comment, disagree, ask to see the evidence, request a reassessment, or appeal. That is by design.
It is also why the observation note matters. A rating you can explain is a rating that holds. Assume everything you write will be read by the person it is about, because it can be.