Leadership · Episode 04 of 6
Velocity and peer review, running continuously.
Two objective angles: what got delivered, and how it was to work with. The team sets the weights rather than me, and neither angle goes anywhere near taste.
Creative output is notoriously awkward to measure. The traditional answer, a year-end review, suffers from recency bias. The traditional alternative, a manager’s read on “quality” and “style”, is a subjective opinion wearing a metric’s clothing. Both produce numbers that feel unfair to the person receiving them, usually because they are.
So I built something that measures two things I can actually defend, and deliberately declines to measure a third.
Angle A: monthly delivery velocity
Borrowing from agile engineering, the team classified every recurring creative and product task and assigned each a complexity weight on a Fibonacci scale.
The weights are set by democratic planning-poker polling, not handed down by me, and that was a deliberate call. Design tasks resist estimation from the outside: a manager sees a deliverable, while the designers who have already built the same thing three times know what it costs. The effort data sits with the people doing the work, so the scoring had to sit there too.
- A baseline benchmark. Each designer works within a target monthly points range: a clear expectation, not a stretch target in disguise.
- An edge-case override. Hit something genuinely bespoke and unclassified? Self-assign a value with a written justification, audited collectively at the monthly sync.
- A collaboration premium. Absorb a teammate’s workload or an unplanned high-priority task and you get a fixed +2 points on top of the base weight.
The premium came from the team, not from me
The first version of the model paid identical points to whoever completed a task. The team raised the obvious flaw, which I had missed: covering for an absent colleague scored exactly the same as your own planned work, while your actual load went up. The incentive pointed at the wrong behaviour.
A model that rewards covering for a teammate at the same rate as doing your own work is quietly asking people to stop covering.
Why the numbers stay honest
A points model is only ever as trustworthy as the allocation underneath it, and three things hold that up. Nobody picks their own queue. Work is routed by the ownership matrix from Episode 02, so cherry-picking high-weight or low-effort tasks simply isn’t available as a strategy. Nobody sets their own weights. Those come from peer polling. And the edge-case overrides that are self-assigned carry a written justification and get audited together, in the open.
Velocity measures what the allocation handed you. It does not measure what you managed to pick up.
Angle B: quarterly peer review
The purpose of this loop is habit formation, not scoring. Every quarter, the cross-functional stakeholders who work directly with my designers complete a short assessment, and I set it one explicit objective: respectful collaboration in both directions, regardless of who is senior. A junior designer should be able to expect the same professional treatment from a department head that the department head expects back. Making that a recurring, visible expectation rather than an unwritten hope turns it into a habit.
It runs on a five-point scale across three operational pillars:
- Timely delivery. Adherence to what was committed, tracked against the project timeline.
- Collaboration. Transparency, engagement across teams, and diplomacy when solving problems.
- Precision. Accuracy of execution against the constraints and the brief.
Creativity and artistic quality are left out. Aesthetic opinion is subjective, and scoring peers against it manufactures misalignment rather than measuring anything. If a designer’s taste is the problem, that is a conversation, not a rating.
What it changed
Tracking the average across four consecutive quarters turns the annual review from a guessing game into a trendline. More usefully, it creates continuous accountability: rather than correcting behaviour in the weeks before a review, the team holds a steady standard of delivery and professional communication with every department, all year.
The quieter effect is on where judgement lives. When the standard is written down and the assessment comes from the people the work actually lands on, a designer stops needing me to certify that something was good. The aim was never fewer reviews. It was fewer decisions that only one person in the room is qualified to make.
That is what makes the last episode possible: by the time the annual review arrives, there is nothing left to reconstruct.