Hire, Review and Pay by the Same Yardstick: A People System That Scales
Tom Ellsworth and Paul Williams (Valuetainment), with Matthew Stanfield (co-founder, Mattenga's Pizzeria), Chad Bennett and other Valuetainment staff; introduced by Patrick Bet-David
6:32 PM–7:23 PM ET · 3:32 PM Las Vegas · 51 minSource kED5xYsrsbuk6tGT3w4G, brphxPfiqfO9Ygo36L7G, UIjO9PIFdYtQLAvcm9YF, YuYCsF9RzleJyzFELITR
Charm is a weak signal in both interviews and reviews, so run hiring, quarterly assessment and compensation through one structured, evidence-backed process where every person is measured against the same criteria.
Tom Ellsworth (operator with 2.2 billion dollars of exits) and Paul Williams (Valuetainment's CTO) walked through how Valuetainment systematised its people operations: structured one-to-ones, a three-week quarterly assessment, a five-band rating tied to raises and a company-then-individual bonus waterfall, structured AI-screened interviews, and a nine-box view of the whole organisation. Matthew Stanfield of Mattenga's Pizzeria described replacing a full-time recruiter for an eight-location restaurant group receiving 150 applications a week, and Valuetainment staff told the story of a botched early calibration that leaked scores before they were reviewed. The two products demonstrated (Interview Now and Performance AI, sold under the Hire Metrics name) are treated here as context. The recording cuts off at 19:23:44 while post-production manager Chad Bennett is introducing himself, so the end of the staff panel and Bet-David's promised closing announcement were not captured.
Best for: Owners with one HR person and too many applicants, Multi-location operators with high hourly turnover, Founders designing their first bonus plan, Managers who run weekly one-to-ones without a fixed structure, Leaders between 15 and 100 employees who no longer know everyone, Anyone evaluating AI hiring or review tools and wanting to know what the process underneath should look like
The capabilities underneath the stories, each with a way to practise it.
Structured weekly one-to-ones
A fixed agenda you follow with every direct report every week, so the meeting tests expectations, progress and obstacles instead of drifting into small talk.
How to practise: Write a five-item agenda (goals for the year, this week's priorities, blockers, feedback both ways, one coaching point) and use it unchanged for a month. Only about a third of the room said they had one. It is working when reports arrive prepared and you can trace a coaching thread across several weeks.
18:37:32-18:38:02 / kED5xYsrsbuk6tGT3w4G
Coaching with a system, not just personality
Adapting your style to each person (an introverted engineer is not managed like a salesperson) while running everyone through the same review structure and yardstick.
How to practise: For each report, write one line on how they like to receive feedback and one line on the standard they are held to. Keep the standard identical across people; vary only the delivery. Revisit both lines each quarter. Success is a team where the quiet deliverer and the loud charmer get the grade their work earned.
A bounded review cycle: self-assessment in week one, manager assessment in week two, committee calibration and delivery by the end of week three.
How to practise: Put the three deadlines on the company calendar the day each quarter closes. Score 1-10 on five categories, require an example for every score, and deliver results like a report card. If the process runs past three weeks, cut steps rather than extend the deadline.
18:40:09-18:41:34 / kED5xYsrsbuk6tGT3w4G
Separating feedback cadence from pay cadence
Giving scored feedback every quarter but deciding raises once a year from the full-year record, so recency bias and Christmas-season good behaviour stop driving compensation.
How to practise: Publish the rule: four quarterly scores roll up to one annual score; raises are discussed only at the annual point. Hold the line when someone asks for a mid-year raise. It is working when nobody's behaviour visibly changes in the run-up to comp season.
18:42:43-18:45:41 / kED5xYsrsbuk6tGT3w4G
Designing a company-first bonus waterfall
A bonus plan where the pool only exists if the company hits minimum growth and profit targets, and the individual's share then depends on their rating.
How to practise: Set a fair target bonus percentage per job level. Define the company unlock (growth and EBITDA minimums). Map ratings to payout: meets = 100 percent of target, exceeds = a step up, outstanding = 50 percent above meets. Explain it to the whole company once a year and again at every quarterly review.
18:46:06-18:49:59 / kED5xYsrsbuk6tGT3w4G
Structured, criteria-based interviewing
Asking every candidate calibrating questions tied to the job, the culture and the non-negotiables, and scoring them side by side, rather than judging on rapport.
How to practise: For each open role, write the job spec, three to five culture attributes with a weight (for example 25 percent problem solving), and the non-negotiables you must ask aloud every time. Score every candidate against that sheet before discussing your impression. It works when the quiet candidate with the best evidence beats the fluent talker.
Turning a pile of applications into a ranked shortlist through explicit filters (skills and company scale, culture questions, non-negotiables) so a small team spends live-interview time only on the top slice.
How to practise: Define a minimum resume score and a first-round threshold (Mattenga's uses 80 percent). Each week, count applicants in, how many passed each filter, and how many reached a live interview. Aim for the 100 to 20 to 7 shape Ellsworth described. Track hires per HR hour to know it is improving.
Plotting every employee on a three-by-three grid of impact (results) versus performance or potential (the other categories), then stack-ranking across departments to find the true top 5 to 10 percent.
How to practise: After each calibration, place every person on the grid. Study the two off-diagonal corners: high potential with low impact (unblock or reassign) and high results with poor culture fit (cannot be top box). Publish the grid shape, without individual scores, so people know where they are and their trajectory.
Using calibration results to gather everyone rated exceeds or outstanding for time with the founder, so recognition reaches people the CEO would otherwise never meet.
How to practise: Each quarter, pull the list of exceeds-and-above and host one lunch or breakfast. Notice how many names you did not know. Use the time for encouragement and mentorship, not status updates.
19:17:10-19:18:08 / UIjO9PIFdYtQLAvcm9YF
Frameworks, lists & numbers
Three-week quarterly assessment timeline
After quarter close: week one, employees self-assess 1-10 on five categories with examples. Week two, managers assess independently. By end of week three, the leadership committee has calibrated and results are delivered. Feedback is quarterly; raises are annual, built from the four quarterly scores.
Five rating bands and pay consequences
Does not meet expectations and needs improvement: no raise discussion at all. Meets expectations and exceeds: average merit increase. Outstanding: above-market increase. Five bands chosen deliberately over a 10-point scale so people do not bunch in the middle.
Me, We, Us bonus waterfall
No bonus pool until the company hits minimum objectives (growth and EBITDA unlock). Target bonus percentage set fairly across job levels. Individual payout: meets = 100 percent of target; exceeds = like a salesperson level one over quota; outstanding = like level two over quota, 50 percent bigger than meets. Half the waterfall from company performance, half from personal performance.
Five manager duty questions
Have you sat down with everybody? Have you set expectations with everybody? Have you given your high performers their props? Have you given your low performers severe encouragement to get going? Have you given everybody alignment with the plan for the year?
Three hiring filter phases
1. Technical skills and comparable company size or scale. 2. Specific questions that test culture fit. 3. Non-negotiables such as on-site versus remote, location, and for restaurants sobriety and passing a drug test.
The 100 / 20 / 7 shortlist
Ellsworth's target for a single HR person: 100 resumes in, 20 ranked against the first filter, 7 brought to the hiring manager.
Mattenga's Pizzeria hiring numbers
Eight locations after twelve years in business. About three open positions per location, so 24 open roles to manage weekly. Roughly 150 applicants a week, many aged 16 to 18, with very high turnover. Live interviews only for the top 40 candidates scoring 80 percent or more. Hiring 10 to 15 people a week. After adopting structured AI screening in November, the full-time recruiter role was eliminated and hires doubled at about the same cost.
The 2026 resume problem
Five years ago, of 1,000 applicants the top five stood out, the next 200 were reasonable, the rest poor fits. In 2026 every resume has been rewritten by AI to match the posting, so 200 look perfect on paper and the candidate has applied to 150 other jobs. Roughly 99.9 percent of applicants are ghosted.
Cost of bad hires
The product video claimed bad hires represent 400 billion dollars a year in economic waste. Treat as a marketing figure, but the underlying causes it listed are sound: people interviewed well, HR liked them, or they reminded the interviewer of themselves.
EALIR score in practice
Five categories (effort, attitude, leadership, innovation, results), each out of 10, total out of 50. Demo example: employee self-scored 45, manager scored 37 without seeing the self-score, committee moved the final number up. One or more evidence questions per category.
Nine-box (three-by-three) grid
Impact (results) on one axis, performance or potential on the other. Best employees are top-right. High potential with low impact needs unblocking. High impact with poor culture fit cannot be top box by definition. Valuetainment publishes the grid shape, without scores, for the whole organisation.
Headcount thresholds
Around 15 employees you start running into people you do not know well. Around 50 you ask 'who's that?'. At 100 the question is 'who got raises and who didn't?'. Each is a cue to add structure.
Structured interview design
Train the process on the job spec, the company culture and the non-negotiables. Tailor a 5 to 30 minute conversation to each resume so no two interviews are identical. Score every candidate side by side on the same standard, with a transcript, and let everyone above a minimum resume score represent themselves.
What matters
33 insights, each with the moment it was said.
Likeability quietly inflates ratings
The upbeat, friendly report gets perceived as the stronger performer while the quieter peer who actually delivers on time and on budget is under-rated. Ellsworth's example: the deliverer should be the A-minus and the charmer the B-plus, but managers get it backwards. A structured process exists to catch exactly this reversal.
18:36:36-18:37:21 / kED5xYsrsbuk6tGT3w4G
People are the business, so the people process deserves a plan
Most of the room ran weekly one-to-ones but only about a third followed a structure. Ellsworth's challenge: if people are your commodity, why would the meeting about them be the one thing you improvise?
18:37:34-18:38:02 / kED5xYsrsbuk6tGT3w4G
A review process gives managers a duty, not just a form
The system exists so leadership can ask every manager: have you sat down with everybody, set expectations, given high performers their props, given low performers 'severe encouragement', and aligned everyone with the plan for the year? Without the process those questions never get asked.
18:38:32-18:39:08 / kED5xYsrsbuk6tGT3w4G
Treating everyone identically is as ineffective as pure personality management
Williams, an engineer, pointed out that engineers avoid conflict and daylight and will run at one pace if mismanaged. The answer is not to abandon the system but to coach each person differently inside the same system.
18:39:08-18:39:44 / kED5xYsrsbuk6tGT3w4G
Goals must be set top-down before anyone can be reviewed against them
The assessment only works if company goals exist, cascade to teams and are delivered to individuals at the start of the year. Reviewing people against goals they never received is unfair and produces arguments, not improvement.
18:39:54 / kED5xYsrsbuk6tGT3w4G
The review must fit the calendar or it becomes the distraction
Valuetainment caps the whole cycle at three weeks after quarter close: one week for self-evaluations, one for managers, then calibration. Ellsworth compared it to a teacher sending report cards home. If your review drags for a month, people resent it and stop being honest.
18:41:02-18:41:34 / kED5xYsrsbuk6tGT3w4G
High expectations are only fair when paired with outsized rewards
The deal Valuetainment makes with high performers is explicit: we will expect a lot, and you will be paid at the highest level. The reward confirms their value and keeps them, which is the point of measuring in the first place.
18:42:01-18:42:31 / kED5xYsrsbuk6tGT3w4G
Meets expectations is a great score, and most people should land there
Meets means doing the job you were hired to do excellently, on time, within budget and to your KPIs. Only a small percentage are truly outstanding. When people ask why they are not outstanding, the answer is 'you are a high meets, which is incredibly valuable, but outstanding is for over and above'.
18:43:25-18:44:23 / kED5xYsrsbuk6tGT3w4G
Everyone accepts quota tiers for salespeople; apply the same logic to everyone else
A rep 20 percent over quota gets a level-one bonus and 25 percent over gets level two, and nobody argues. Applied to non-sales roles, suddenly everybody believes they are at 125 percent of quota. Calibration exists to bring them back to reality using the same philosophy sales already accepts.
18:43:56-18:44:23 / kED5xYsrsbuk6tGT3w4G
Five bands, deliberately, not ten
Williams explained that a ten-point rating or extra middle categories get confusing and bunch everyone in the centre. Five named bands are enough to differentiate and simple enough that every employee can remember where they stand.
18:44:39-18:44:52 / kED5xYsrsbuk6tGT3w4G
Quarterly scoring defeats recency bias
When reviews happen once a year, everyone realises at Christmas that pay is about to be decided and changes how they show up. Four quarterly scores make the annual number reflect the whole year, and the last few weeks stop mattering disproportionately.
18:45:03-18:45:41 / kED5xYsrsbuk6tGT3w4G
Pay the company first so it can pay you
The 'me, we, us' bonus structure says nobody earns a bonus until the company hits minimum objectives and has the money to fund the plan. That single rule aligns every employee with company results before their own.
18:46:06-18:47:14 / kED5xYsrsbuk6tGT3w4G
A 50 percent gap between meets and outstanding is what makes the top band real
In Valuetainment's plan meets pays 100 percent of target, exceeds steps up, and outstanding pays 50 percent more than meets. Half the waterfall comes from company performance, half from personal. Because the whole company can see the maths, the outcome reads as fair even to those who did not get it.
18:48:05-18:49:59 / kED5xYsrsbuk6tGT3w4G
In 2026 every resume looks perfect, so the resume has stopped being a filter
Candidates run their resume through AI and tailor it to your posting along with 150 others. Five years ago the top five stood out; now 200 look like perfect fits on paper. The shortlist has to come from questions and evidence, not from reading resumes harder.
18:51:09-18:52:17 / brphxPfiqfO9Ygo36L7G
Ghosting applicants damages your brand
Williams estimated 99.9 percent of applicants are never told why they were passed over. Giving every applicant above a minimum score a chance to represent themselves is both fairer and better for how your company is perceived by the many you do not hire.
A great conversation tells you someone is good at conversation. Interviewer bias compounds it: engaging candidates score higher than steady, straightforward ones, and interviewers hire people who remind them of themselves. Only calibrating questions asked of everyone break the pattern.
18:52:17-18:53:20 / brphxPfiqfO9Ygo36L7G
Three filter phases, in order
First filter on technical skills and comparable company size or scale. Then test for culture with specific questions. Then confirm non-negotiables such as on-site versus remote. Doing them in this order means expensive human time is spent only on candidates who have already cleared the cheap filters.
18:53:20-18:54:00 / brphxPfiqfO9Ygo36L7G
Your one HR person cannot also be your recruiter
Most of the room had a single HR person handling payroll, benefits and onboarding. Ellsworth's point: that person will never have time to properly rank 100 resumes down to 20 and bring you 7, so the filtering step has to be systematised or outsourced to a tool.
18:54:00-18:55:05 / brphxPfiqfO9Ygo36L7G
A hiring tool should adapt the questions but standardise the scoring
The approach demonstrated trains the interview on the job, the culture and the non-negotiables, then tailors a 5 to 30 minute conversation to each resume. No two interviews are the same, but every candidate lands on the same side-by-side scorecard with a transcript. That combination is worth copying even without software.
18:56:00-18:57:39 / brphxPfiqfO9Ygo36L7G
A single interviewer's personality becomes the company's hiring filter
Mattenga's soft-spoken recruiter was intimidated by dominant personalities and never hired them, so the whole company was filtered through one temperament. A restaurant needs extroverts. Moving to structured scoring produced a wider variety of people and a noticeably better culture in the stores.
Redirect recruiting savings into job promotion and higher pay
Stanfield said the software cost roughly what the recruiter had cost, but doubled hiring capacity with the same HR headcount. His plan for the saving: advertise the jobs properly for the first time and pay people more to attract better candidates.
19:02:51-19:03:20 / brphxPfiqfO9Ygo36L7G
Weight your values and ask the non-negotiables every single time
Mattenga's assigned percentages to values (for example 25 percent problem solving, hospitality) and hard-coded non-negotiables: sober, no DUIs, able to pass a drug test. People tick the box on the form and lie, so the question must be asked out loud in every interview.
19:04:42-19:05:30 / brphxPfiqfO9Ygo36L7G
Screen for 'do you care about people' to cut paycheque-only hires
Filtering for a service mindset up front reduced the 'I need money this week and will quit next week' hire that drives restaurant turnover. Culture questions in the first filter are a retention tool, not a nicety.
19:07:05-19:07:27 / brphxPfiqfO9Ygo36L7G
No hiring process is a miracle
Stanfield was careful to say people will still come and go in any business. The goal of a filtering system is to be most effective on average, not to eliminate every miss. Set expectations accordingly with your team.
19:08:04-19:08:42 / brphxPfiqfO9Ygo36L7G
Without a system, reviews are a popularity contest
Williams relayed the brief Bet-David gave him: measure everybody by the same yardstick. Small companies typically build a review in PowerPoint or Excel and get recency and affinity bias baked in. The fix is a defined scale, evidence for every score and a calibration step, whatever tool you use.
19:09:29-19:10:45 / UIjO9PIFdYtQLAvcm9YF
Make the review conversational and mobile or people will not do it honestly
Most Valuetainment staff completed their last calibration on a phone; Williams did 42 of his by voice while driving. The lesson for any process: the less bureaucratic it feels, the more candid and complete the input.
19:11:45-19:12:20 / UIjO9PIFdYtQLAvcm9YF
No score without a specific example
The demonstrated process asks one or more evidence questions per category and prompts 'you gave me a score but no example'. Requiring evidence at input time is what makes the later committee discussion about facts instead of impressions.
19:12:20-19:14:07 / UIjO9PIFdYtQLAvcm9YF
Decide deliberately whether managers see the self-score first
Showing the self-assessment anchors the manager; hiding it gives an independent read. In the demo the employee scored 45, the manager 37 without seeing it, and the committee moved the final number up. Choose one rule and apply it to everyone.
19:12:35-19:13:41 / UIjO9PIFdYtQLAvcm9YF
Stack-rank across departments to find your real top 5 percent
Calibrated scores let you see who moved up or down, who is top quartile, and who your top ten performers are regardless of department. That cross-company view is impossible when each manager grades in isolation.
19:15:22-19:16:03 / UIjO9PIFdYtQLAvcm9YF
High results with poor culture fit cannot be a top-box employee
The nine-box separates impact from the other categories. Someone crushing goals but failing the culture cannot sit top-right by definition, and someone with clear potential but a 5 in results needs unblocking, not praise. The grid forces both conversations.
19:16:03-19:16:44 / UIjO9PIFdYtQLAvcm9YF
Headcount thresholds change what you need
Williams' rule of thumb: around 15 employees you start running into people you do not know well; around 50 you ask 'who's that?'; at 100 the question becomes 'who got raises and who didn't?'. Each threshold is a cue to formalise another layer of the people system.
19:16:44-19:17:04 / UIjO9PIFdYtQLAvcm9YF
Publish the shape of the grid, not the scores
Valuetainment posts a version of its three-by-three on campus for the whole organisation. Individual scores stay private, but everyone can see where they sit and which way they are trending. Transparency about the system builds trust without exposing anyone.
19:18:24-19:18:32 / UIjO9PIFdYtQLAvcm9YF
Releasing scores before calibration destroys trust
Valuetainment's director of studio operations, weeks into the job, entered his team's scores in an early tool and hit 'done', which sent uncalibrated scores straight to his reports. The next day he had disturbed direct reports, then months of communication and trust problems. Whatever tool you use, separate draft, review, approve and release.
19:19:38-19:22:06 / YuYCsF9RzleJyzFELITR
Claude skills from this session
Each one is a SKILL.md Claude can run with you. Open to read; copy or download; drop into ~/.claude/skills/.
Exercises & frameworks to run
Steps and the finished output for each. Speaker exercises and our written-out steps are labelled.
One-to-one agenda card30 minutes per report per weekour steps
Write five fixed items on a card: progress against this year's goals, this week's top three priorities, blockers you need me to remove, feedback in both directions, one coaching point.
Use the card in every weekly one-to-one for a month without changing it.
Keep a one-line note per meeting so coaching threads carry across weeks.
At month end, ask each report whether the meeting is more useful than before and adjust one item at most.
Finish with: A repeatable one-to-one structure and a running log per direct report.
Manager duty checklist15 minutes per manager per quarterspeaker exercise
For each manager who reports to you, ask the five questions Ellsworth listed: Have you sat down with everybody? Have you set expectations with everybody? Have you given your high performers their props? Have you given your low performers severe encouragement? Have you aligned everyone with the plan for the year?
Record a yes or no per question per manager.
Any 'no' becomes that manager's action item for the next two weeks.
Repeat at the start of each quarter.
Finish with: A simple grid showing which managers are actually managing and where to intervene.
Three-week quarterly assessment calendarThree weeks elapsed, a few hours per managerspeaker exercise
On the day the quarter closes, set three dates: self-assessments due in one week, manager assessments due one week later, calibration and delivery complete one week after that.
Employees rate themselves 1-10 on five categories with one specific example per score.
Managers rate each report on the same scale, independently.
Leadership committee reconciles the two and records a final score and band.
Managers deliver the result and the reasoning to each employee, like a report card, before the third week ends.
Finish with: A calibrated score and band for every employee, delivered within three weeks of quarter close.
Bonus waterfall on one pageHalf a day, once a yearspeaker exercise
Set the company unlock: the minimum revenue growth and EBITDA that must be hit before any bonus pool exists.
Set a target bonus percentage for each job level so targets are fair across the organisation.
Map ratings to payout: does not meet and needs improvement = 0; meets = 100 percent of target; exceeds = a defined step up (like level-one over quota); outstanding = 150 percent of meets.
Split the calculation half on company performance and half on personal performance.
Write it on one page and walk the whole company through it.
Finish with: A bonus plan every employee can compute for themselves, aligned to company results first.
Structured interview scorecardOne hour per role to build, 10 minutes per candidate to scoreour steps
For one open role, write the job spec in your own words rather than copying a posting from the internet.
List three to five culture attributes and give each a weight that adds to 100 percent (Mattenga's example: 25 percent problem solving, plus hospitality).
List the non-negotiables that must be asked aloud every time (location or on-site requirement, sobriety and drug test for restaurants, and so on).
Write two calibrating questions per attribute that ask for a specific past example.
Score every candidate on the same sheet immediately after the interview and before discussing impressions with anyone.
Finish with: A side-by-side scorecard that lets you compare candidates on evidence rather than rapport.
Recruiting funnel count10 minutes per week for a monthour steps
For four weeks, log applicants received per role per week.
Log how many pass the resume or skills filter, how many pass culture and non-negotiable screening, and how many reach a live interview.
Compare to Ellsworth's shape (100 in, 20 ranked, 7 to the boss) and Mattenga's threshold (live interviews only for the top 40 scoring 80 percent or more).
Identify the stage that eats the most human hours and systematise it first.
Finish with: A funnel you can quote in numbers and a clear target for where a tool or template saves the most time.
Your checkmarks are saved on this device and shared with the Playbook.
Every week ('How many people do a one to one with their direct reports every week?') speaker-statedEvidence: 18:37:34-18:37:49 / kED5xYsrsbuk6tGT3w4G
Every quarter ('Every quarter, the teammates complete a self assessment') speaker-statedEvidence: 18:40:14-18:41:16 / kED5xYsrsbuk6tGT3w4G
Once a year ('We don't do raises quarterly. We do feedback quarterly en route to a full annual score') speaker-statedEvidence: 18:42:47-18:42:52 / kED5xYsrsbuk6tGT3w4G
Once, before the next fiscal year begins our recommendationEvidence: 18:46:06-18:49:59 / kED5xYsrsbuk6tGT3w4G
Once a year, at the start of the year our recommendationEvidence: 18:38:28-18:39:54 / kED5xYsrsbuk6tGT3w4G
Once per role, this month for your current openings our recommendationEvidence: 18:56:00-18:57:09, 19:04:42-19:05:30 / brphxPfiqfO9Ygo36L7G
Every interview ('you have to ask them in the interview every single time') speaker-statedEvidence: 19:05:14-19:05:19 / brphxPfiqfO9Ygo36L7G
Every week for a month, then monthly our recommendationEvidence: 18:54:24-18:54:45 / brphxPfiqfO9Ygo36L7G
Once, this quarter our recommendationEvidence: 19:01:10-19:01:56 / brphxPfiqfO9Ygo36L7G
Every quarterly assessment our recommendationEvidence: 19:14:07 / UIjO9PIFdYtQLAvcm9YF
Every quarter, after calibration our recommendationEvidence: 19:15:22-19:16:44 / UIjO9PIFdYtQLAvcm9YF
Every quarter, after scores are final our recommendationEvidence: 19:17:10-19:17:49 / UIjO9PIFdYtQLAvcm9YF
Every quarter our recommendationEvidence: 19:18:24-19:18:32 / UIjO9PIFdYtQLAvcm9YF
Once, before your next review cycle our recommendationEvidence: 19:20:28-19:21:02 / YuYCsF9RzleJyzFELITR
Quotes in context
“It says for the managers, it provides a duty, a duty to objectively sit back and manage your team.”
Tom EllsworthWhat a structured review process does for managers, as opposed to employees.18:38:32 · kED5xYsrsbuk6tGT3w4G
“Have you given your low performers their, you know, severe encouragement to get going?”
Tom EllsworthOne of the five questions leadership asks every manager.18:38:53 · kED5xYsrsbuk6tGT3w4G
“But if you don't manage them correctly, don't know how to get the best out of them, they will only run at one pace.”
Paul WilliamsOn engineers, and why a system still needs to be delivered differently to different people.18:39:44 · kED5xYsrsbuk6tGT3w4G
“It's got to be an enabling process that's efficient, that fits in the calendar, but quickly gives you exactly what you need.”
Tom EllsworthWhy the quarterly review is capped at three weeks.18:41:17 · kED5xYsrsbuk6tGT3w4G
“It's like a teacher getting the report cards and sending them home.”
Tom EllsworthThe mental model for delivering quarterly scores.18:41:31 · kED5xYsrsbuk6tGT3w4G
“Those of you that are high performers, we're gonna have high expectations, but you're gonna receive outsized rewards.”
Tom EllsworthThe deal made explicitly with top performers.18:42:11 · kED5xYsrsbuk6tGT3w4G
“We do feedback quarterly en route to a full annual score, where we then would talk about raises.”
Tom EllsworthSeparating the feedback cadence from the pay cadence.18:42:52 · kED5xYsrsbuk6tGT3w4G
“You have to be firm enough in the concept of meritocracy to say no.”
Tom EllsworthOn giving no raise at all to does-not-meet and needs-improvement ratings.18:43:05 · kED5xYsrsbuk6tGT3w4G
“Well when you apply it to people, suddenly everybody thinks they're 125% of quota, and you need to bring them back down.”
Tom EllsworthWhy sales quota logic is accepted for reps but resisted everywhere else.18:44:16 · kED5xYsrsbuk6tGT3w4G
“Meets is a great score for people doing a great job within the bounds of their job.”
Tom EllsworthResetting expectations about the middle band.18:44:23 · kED5xYsrsbuk6tGT3w4G
“Well, the reason we do it quarterly, is otherwise when it comes to compensation, recency kicks in.”
Paul WilliamsWhy the assessment is quarterly rather than annual.18:45:23 · kED5xYsrsbuk6tGT3w4G
“In other words, pay the company so it can pay us.”
Tom EllsworthThe principle behind the company-first bonus unlock.18:46:46 · kED5xYsrsbuk6tGT3w4G
“But if you're not asking the right calibrating questions in the interview, then you're really just having a conversation.”
Paul WilliamsWhy rapport in an interview is not evidence of fit.18:52:38 · brphxPfiqfO9Ygo36L7G
“So stop hiring the person who interviews best, and start hiring the person that is best for the job.”
Interview Now product video (played on stage)The central hiring claim of the session, from the product video.18:55:05 · brphxPfiqfO9Ygo36L7G
“She filtered everybody in the company along that lens, and so we needed to clean that up and make it more objective.”
Matthew StanfieldOn a soft-spoken recruiter who never hired dominant personalities at Mattenga's Pizzeria.19:01:25 · brphxPfiqfO9Ygo36L7G
“Like even interview is not a miracle that everything is going to be perfect. There's no such thing.”
Matthew StanfieldSetting realistic expectations for any hiring system; 'interview' refers to the Interview Now tool.19:08:23 · brphxPfiqfO9Ygo36L7G
“I really do believe that if we don't have a system, it is literally just a popularity contest and I'm not interested in popularities. Give me something that measures everybody by the same yard stick.”
Paul Williams, relaying Patrick Bet-David's briefThe design brief for the performance review system.19:10:19 · UIjO9PIFdYtQLAvcm9YF
“Obviously, it doesn't show your scores, but it allows you to know where you are and where your trajectory is.”
Paul WilliamsOn the nine-box grid published for the whole company at Valuetainment's campus.19:18:32 · UIjO9PIFdYtQLAvcm9YF
“What I didn't realize about the done button was that it meant done done, and then it got sent out to all the employees from the studio.”
Valuetainment director of studio operationsHow uncalibrated scores leaked to a whole team in an early review tool.19:20:47 · YuYCsF9RzleJyzFELITR
Questions to ask yourself
Who on my team do I rate highly because I enjoy them, and who delivers quietly and gets a lower grade for it?
If people are my business, why is the weekly meeting about them the one thing I run without a plan?
Could every employee compute their own bonus from a one-page rule, or does it feel arbitrary to them?
Does anyone's behaviour change in the weeks before pay decisions, and what does that tell me about my review cadence?
Am I hiring the person who interviews best or the person who is best for the job?
Whose temperament is my company's real hiring filter right now?
Which non-negotiables do I assume from a ticked box instead of asking out loud?
Who are my top ten performers across all departments, and how many of them have I never spoken to?
If a score reached an employee before it was calibrated, what would break, and have I made that impossible?
Watch-outs
Rating the friendly, upbeat report above the quiet one who actually delivers on time and on budget.
Running weekly one-to-ones with no fixed structure, so they never test expectations or carry a coaching thread.
Treating everyone identically in the name of fairness; people react differently and a mismanaged engineer runs at one pace.
Letting the review cycle sprawl for a month so it becomes a distraction and people stop answering honestly.
Doing raises off a single annual review, which invites recency bias and Christmas-season good behaviour.
Handing out 'outstanding' too freely; meets is the correct score for most people doing their job well.
Copying job specs from the internet instead of writing what you actually need.
Believing the resume in 2026; every one has been tailored by AI to your posting.
Letting one interviewer's personality filter the whole company, so you never hire the extroverts a service business needs.
Trusting a ticked box on non-negotiables such as sobriety instead of asking in every interview.
Expecting any hiring system to be a miracle; people will still come and go.
Releasing scores to employees before the calibration committee has reviewed them, which caused months of trust damage at Valuetainment.
This skill produces a hiring scorecard for one role: the job spec, weighted culture criteria, spoken non-negotiables, evidence questions, filter thresholds, and a side-by-side table every candidate is scored on before anyone discusses impressions. Paul Williams' case for it starts with the resume: "in 2026, every resume you see now, I guarantee, has been rewritten perfectly for that job description ... So the list now is not three to five, the list is 200 on paper, perfect candidates for you" (18:51:34). And the interview does not rescue you: "if you're not asking the right calibrating questions in the interview, then you're really just having a conversation" (18:52:38). The product video played on stage put the goal in one line: "stop hiring the person who interviews best, and start hiring the person that is best for the job" (18:55:05).
The session demonstrated software that does this. The process underneath is what this skill encodes; it works on paper.
Inputs to gather
The role, and whether it is on-site, hybrid or remote, and where.
Weekly applicant volume and how many hires the role needs. Mattenga's: about 150 applicants a week across 24 open roles, hiring 10 to 15 people a week (19:00:25, 19:06:13).
Who screens and who interviews today. Ellsworth's picture of the room: one HR person doing payroll, benefits and onboarding, who cannot also rank 100 resumes (18:54:00).
The company's stated values, if any. Default: ask the owner to name the three behaviours they praise most and the three they fire for.
The last 20 hires and who interviewed each, for the bias audit in step 8.
Any hard requirements the law or the work imposes (certifications, background checks, the ability to lift, drive, or pass a drug test).
Steps
1. Write the job spec in your own words
Williams on how most small companies do it: "You probably find the jobs online, you copy the structure and you paste them in and you post them" (18:56:22). Instead, write four things: what this person will own, what good looks like in 90 days, the two or three skills that cannot be trained on the job, and the size or scale of place they should have worked in before. That last one is Ellsworth's first filter: "technical skills and around the same sort of size of company, same sort of scale" (18:53:20).
2. Pick three to five culture criteria and weight them
Stanfield's discovery was that the criteria could be his own: "each business values things different ... for us is like 25% problem solving or hospitality. Each business has our values, so that was a great filter" (19:04:42). Ellsworth named the families: "performance cultures, delivery cultures, accountability cultures. Every organization has those different things in different percent" (19:07:27).
Choose three to five criteria. Weight them to total 100 percent. The weights force the owner to say what matters most. Example for a restaurant: hospitality 30, problem solving 25, reliability 25, teamwork 20.
3. Write the non-negotiables and commit to asking them aloud
Non-negotiables are pass or fail and sit outside the weights. Williams' examples: on-site versus remote, location (18:56:48). Stanfield's for a restaurant: "they need to be sober with no DUIs and be able to pass a drug test. And unfortunately they lie on the resume and when they check the box, you have to ask them in the interview every single time" (19:05:14).
List them. Then write the exact sentence the interviewer will say for each, so it gets asked the same way to every candidate, every time. A ticked box on an application is not an answer.
4. Write two evidence questions per criterion
Calibrating questions ask for a specific past example, not an opinion. For each criterion, write two prompts of the form "Tell me about a time when..." with a follow-up that asks what the candidate personally did and what happened. The point is that "interview skills are not job skills" (18:52:17); a fluent talker can describe hospitality, but only an example shows it. Mattenga's culture question that cut paycheque-only hires was as simple as "do you care about people?" (19:07:05), followed by evidence.
5. Set the three filter phases and their thresholds
Ellsworth's order: "One, you get all these resumes in, you want to filter them around technical skills and around the same sort of size of company ... Then you want to get into, okay, what are some specific things that I can test for culture? ... This is an on site job. Oh, I'm looking for work from home. How do you filter those out?" (18:53:20). Cheap filters first, so human time is spent only on candidates who have cleared them.
Phase
Filter
Threshold
1
Skills and company scale, from the resume and a short screen
Minimum resume score; everyone above it gets to represent themselves (Williams, 18:57:22)
2
Culture criteria, from evidence questions
Mattenga's live-interview cut: top 40 scoring 80 percent or more (19:06:25)
3
Non-negotiables, asked aloud
Pass or fail, every candidate
Funnel target for one HR person, Ellsworth: "we got 100 resumes and these 20 have been ranked and they fit our first criteria of filter, I'm now going to spend my time on these 20 and boss, I'm going to bring you these seven" (18:54:24). Aim for that 100 to 20 to 7 shape.
6. Score every candidate on the same sheet, before discussing impressions
After each interview, the interviewer scores every criterion 1 to 5 with the example that justified it, marks each non-negotiable pass or fail, and computes the weighted total. Do this before talking to anyone. Williams described the target output as "side by side standard scoring" (18:57:06) with a transcript behind each score. A candidate with any non-negotiable fail is out regardless of total.
Weighted total = sum of (criterion score × weight) ÷ 5, giving a percentage.
7. Compare side by side, then decide
Line the shortlist up in one table. The decision goes to the highest evidence-backed total, not the best conversation. When the owner's gut disagrees with the table, write down why. Sometimes the table is missing a criterion; add it for next time rather than overriding this time. Stanfield's before and after: "we had great team members, randomly be a team member that was a little off. Now ... when we walk into stores, when we look at the team, it is noticeably different" (19:05:48).
8. Run the interviewer-bias audit
Stanfield's recruiter was "very soft spoken, quiet. And so if somebody came in with the high D personality, she would get scared and never hire them ... she filtered everybody in the company along that lens" (19:01:10). A restaurant needs extroverts; the filter was removing them. Moving to structured scoring produced "a much bigger variety of people in the company" and a culture that "lifted up" (19:01:37).
List the last 20 hires and who interviewed each.
Note each hire's broad temperament and the interviewer's.
Look for a pattern where the interviewer never hires their opposite.
If one exists, add a second scorer or the scorecard before the next hire, and check the pattern again in six months.
9. Respond to everyone above the minimum score
Williams: "99.9% of people that apply for a job only get ghosted. They never know why" (18:50:19). Beyond fairness, it is brand damage among the many you do not hire. Give every candidate above the phase-one threshold a short screen and a decision. A two-line message is enough.
10. Count the funnel weekly
For a month, log applicants in, passed phase one, passed phases two and three, live interviews, offers, hires. The stage eating the most human hours is the one to template, delegate or automate first. Stanfield's outcome after systematising: the full-time recruiter role went away and "we've been able to double the number of people we do hire" at roughly the same cost (19:00:51, 19:03:15).
Output format
Fill assets/hiring-scorecard-worksheet.md. It contains:
Job spec in four parts.
Culture criteria table with weights totalling 100.
Non-negotiables with the exact spoken question for each.
Two evidence questions per criterion.
Filter phases with thresholds.
Per-candidate scoring sheet and the side-by-side comparison table.
Interviewer-bias audit table.
Weekly funnel count.
Cadence
Activity
Cadence
Source
Ask every non-negotiable aloud
Every interview
Speaker-stated
Score on the sheet before discussing
Every interview
Our recommendation, from the side-by-side principle
Build or refresh the scorecard
Once per role, revisit when the role changes
Our recommendation
Funnel count
Weekly for a month, then monthly
Our recommendation
Interviewer-bias audit
Once now, then every six months
Our recommendation
Pitfalls the speaker warned about
Believing the resume. In 2026 it has been tailored by AI to your posting, alongside 150 others.
Mistaking rapport for fit. "You may even have a really great rapport and feel really good about the individual" and still know nothing about the job (18:52:38).
Hiring people who remind the interviewer of themselves, or never hiring their opposite. One temperament becomes the company's filter.
Trusting a ticked box on a non-negotiable. Ask it aloud, every time.
Copying a job spec from the internet.
Expecting a miracle. Stanfield: "There's no such thing. So in any business you're going to have a percentage of human beings that come and go" (19:08:23). The system improves the average; it does not remove every miss.
When not to use this
For a single senior or executive hire where the pool is a handful of known people, the funnel thresholds do not apply; keep the weighted criteria, the evidence questions and the non-negotiables, and skip the phase-one cut. Do not use the weighted total to override a legal or safety requirement; those are non-negotiables, not criteria. Check local employment law before adding any screening question that touches health, background or protected characteristics.
Companion skill: quarterly-calibration-calendar measures people after they join by the same yardstick logic.
See references/source-notes.md for verbatim quotes with timestamps.
# frontmatter
name: ai-era-hiring-scorecard
description: "Build and run a structured, side-by-side hiring scorecard for one role, as taught by Tom Ellsworth, Paul Williams and Mattenga's Pizzeria co-founder Matthew Stanfield: a job spec in your own words, three to five weighted culture criteria (Mattenga's example: 25 percent problem solving, plus hospitality), non-negotiables asked aloud in every interview because ticked boxes lie, two evidence questions per criterion, three filter phases in order (skills and company scale, culture, non-negotiables), a 100 to 20 to 7 funnel target for one HR person, and an interviewer-bias audit so one person's temperament stops filtering the whole company. Use whenever an owner or manager mentions hiring, interviews, too many applicants, AI-written resumes that all look perfect, a recruiter, a bad hire, culture fit, turnover, ghosting candidates, or says they hire people they click with — even if they never ask for a scorecard or a hiring process."
license: MIT
metadata:
source: "The Vault Conference 2026 · Hire, Review and Pay by the Same Yardstick: A People System That Scales · Tom Ellsworth and Paul Williams"
session_key: d2-performance-ai-suite
speaker: "Tom Ellsworth and Paul Williams (Valuetainment), with Matthew Stanfield (co-founder, Mattenga's Pizzeria); introduced by Patrick Bet-David"
day: "2026-09-02"
vault_url: "https://vault.chels.ai/sessions/d2-performance-ai-suite/"
attribution: "Framework as taught on stage; steps written by Cole's Notes Vault (chels.ai). Verify quotes against official recordings."
This skill produces a one-page bonus plan with three parts: the company unlock, the per-level target, and the individual payout ladder, plus a worked example and an affordability check. Tom Ellsworth's principle is that the company gets paid before anyone else: the plan "basically says we can't make a bonus together until the company hits minimum objectives this year. In other words, pay the company so it can pay us" (18:46:46). The second principle is that the plan must be visible and computable: "everybody in the organization knows it, and everybody can see the quantitative days you went through it and how fair it was. Fairness is the basis of it" (18:49:16).
The plan assumes an annual rating for each person. If the user has no rating system, run quarterly-calibration-calendar first, or agree an interim five-band rating for this year.
Inputs to gather
Revenue this year and the growth target for next year.
EBITDA (or operating profit) this year and the minimum acceptable next year.
Headcount by job level (for example individual contributor, team lead, manager, director, executive) and average salary per level.
The rating distribution from the last review, or an estimate. Ellsworth's expectation: most people at meets, "a very small percent" outstanding (18:43:25).
The total bonus pool the company can afford at target. Default: work backward from the affordability check in step 6.
Whether any existing sales commission plan is in place. The waterfall sits alongside it, not on top of it.
Steps
1. Write the company unlock
Ellsworth: "there has to be some unlock if the company achieves a minimum amount of growth and EBITDA because then we as founders and CEOs, right, now we have something in the bonus plan" (18:49:59). Set two minimums:
Minimum revenue growth for the year.
Minimum EBITDA for the year.
Below either minimum, no pool exists and no individual bonus is paid regardless of rating. Say this in one sentence on the plan. The "me" only happens "if the company has hit objectives and has made enough to fund the bonus plan" (18:46:58).
2. Set a fair target bonus percentage per job level
Ellsworth: "the targets will go across job levels, but these are the targets ... we create a fair target bonus across job levels" (18:47:26, 18:47:56). The speakers did not give percentages. Our starting defaults, to be adjusted to the business:
Level
Target bonus, percent of salary (our default)
Individual contributor
5
Team lead
8
Manager
10
Director
15
Executive
20 to 30
Target is what a meets-rated person earns when the company hits its target. Fair means the same level gets the same percentage everywhere in the company.
3. Build the individual payout ladder
Ellsworth, in order: "if you get a meets, you're at 100% of the target for your bonus. If you're at exceeds expectations, congratulations. It's like a salesperson being level one over quota. And then if you're outstanding, it's like a salesperson being level two over quota. Look how that escalates. The difference between a meets and an outstanding is 50%. The bonus is 50% bigger" (18:48:15 to 18:48:57).
Annual band
Individual multiplier
Source
Does not meet expectations
0
Speaker-stated: no raise, no bonus discussion
Needs improvement
0
Speaker-stated
Meets expectations
100 percent of target
Speaker-stated
Exceeds expectations
125 percent of target
Speaker-stated as "level one over quota"; the 125 figure is our default, sitting between meets and outstanding
Outstanding
150 percent of target
Speaker-stated: "50% bigger" than meets
The sales analogy is the explanation to use with the team. Ellsworth: a rep "way above my quota by 20%, I got level one bonus. I'm above my quota by 25%, I got my next level bonus. And everybody in the organization understands and agrees that philosophy" (18:43:56). The waterfall applies the same accepted logic to everyone else.
4. Split the waterfall half company, half personal
Ellsworth, pointing to the diagram in his book: "how the waterfall works, half is from the company performance, half is from the personal performance" (18:49:31). Our formula for that split, once the unlock is passed:
Company factor: 1.0 when the company hits its target. Our default scale: 0.5 at the unlock minimum, 1.0 at target, capped at 1.5 for a stretch result. Straight-line between.
Individual multiplier: from the ladder in step 3 (0, 0, 1.0, 1.25, 1.5).
Below the unlock, the formula is not run. Nobody is paid.
5. Work one example per level
Fill the example table so every employee can find themselves. A manager on 100,000 with a 10 percent target, company at target (factor 1.0):
Band
Calculation
Bonus
Meets
100,000 × 0.10 × (0.5 × 1.0 + 0.5 × 1.0)
10,000
Exceeds
100,000 × 0.10 × (0.5 × 1.0 + 0.5 × 1.25)
11,250
Outstanding
100,000 × 0.10 × (0.5 × 1.0 + 0.5 × 1.5)
12,500
Needs improvement
Formula not run; bottom-band exception applies
0
Note on the bottom bands: the half-company, half-personal split would mechanically pay a needs-improvement employee half their target (5,000 in this example). Ellsworth's stated rule is that the bottom two bands get nothing: "We don't even talk about raises. You have to deliver improvement" (18:42:43), and "You have to be firm enough in the concept of meritocracy to say no" (18:43:05). Apply the rule, not the arithmetic: the two bottom bands receive zero. Write that exception on the page.
Note on outstanding: with the half-and-half split, outstanding pays 25 percent more than meets at company target, not 50. If the user wants Ellsworth's "50% bigger" to hold at company target, apply the multiplier to the whole bonus instead (bonus = salary × target% × company factor × individual multiplier). Both are defensible readings of the stage material; pick one, write it down, and do not change it mid-year.
6. Run the affordability check
Multiply each level's headcount by average salary, target percent and the expected distribution across bands. Sum to get the expected payout at company target. Compare with the pool the company can fund at target EBITDA. If the payout exceeds the pool, lower target percentages or raise the unlock. Never fix affordability by rating fewer people at meets. Meets is the normal score for "doing the job excellently that you were hired to do" (18:43:32).
7. Write the one page and explain it twice a year
Ellsworth's test is that everyone can see the maths and judge it fair. Put the unlock, the level table, the ladder, the formula, the exception for bottom bands, and one worked example on a single page. Walk the whole company through it once when the plan is set, and refer back to it at every quarterly review so nobody is surprised at year end (the twice-a-year rhythm is our recommendation; the transparency is speaker-stated).
Output format
Fill assets/bonus-waterfall-one-pager.md. It contains:
Company unlock: two minimums in one sentence.
Target table by level.
Payout ladder by band with the bottom-band exception.
The formula and which reading of the split was chosen.
Worked example per level.
Affordability table: expected payout vs pool.
Communication dates.
Cadence
Activity
Cadence
Source
Set unlock, targets and ladder
Once, before the fiscal year begins
Our recommendation
Explain the plan to the whole company
At plan launch, then referenced each quarterly review
Transparency speaker-stated; rhythm ours
Confirm the unlock is or is not hit
At year-end close
Speaker-stated structure
Pay out by annual band
Annually, after the annual score
Speaker-stated
Pitfalls the speaker warned about
Paying bonuses when the company missed its minimums. The pool does not exist until the company has "made enough to fund the bonus plan."
Paying the bottom two bands anything. "Firm enough in the concept of meritocracy to say no."
Handing out outstanding freely so the 150 percent tier stops meaning anything. Outstanding is "a very small percent"; most people are, correctly, a high meets.
Making the plan opaque. If people cannot compute their own number, fairness cannot be seen and the payout reads as favouritism.
Changing the rule mid-year (our addition). The plan only works as an incentive if it is stable for the whole period it covers.
When not to use this
Do not layer the waterfall on top of an existing sales commission plan for the same people; pick one variable-pay scheme per role. Do not use it in a year with no rating process, since the individual multiplier has nothing to attach to. And treat the default percentages as starting points, not advice on what the business can afford; the affordability check decides that.
Companion skills: quarterly-calibration-calendar produces the annual band this plan pays on. quarterly-calibration (Patrick Bet-David's session) reaches the same five bands with numeric cut-offs and its own bonus multipliers; if the user already runs that rubric, keep its bands and apply this skill's company unlock and half-and-half waterfall on top.
See references/source-notes.md for verbatim quotes with timestamps.
# frontmatter
name: bonus-waterfall-designer
description: "Design a company-first bonus plan on one page using the 'me, we, us' waterfall Tom Ellsworth taught at Valuetainment: minimum company growth and EBITDA targets unlock the pool, a fair target bonus percentage per job level, then the individual's annual rating sets the payout (meets = 100 percent of target, exceeds = a step up like a salesperson one level over quota, outstanding = 150 percent of meets, nothing for the two bands below meets), with half the waterfall driven by company performance and half by personal. Includes a worked example, an affordability check and the one-page explainer every employee can compute from. Use whenever a founder, owner or finance lead mentions bonuses, incentive plans, profit sharing, variable pay, rewarding top performers, year-end payouts, or says bonuses feel arbitrary, unaffordable or cause arguments — even if they never say 'waterfall'."
license: MIT
metadata:
source: "The Vault Conference 2026 · Hire, Review and Pay by the Same Yardstick: A People System That Scales · Tom Ellsworth and Paul Williams"
session_key: d2-performance-ai-suite
speaker: "Tom Ellsworth and Paul Williams (Valuetainment); introduced by Patrick Bet-David"
day: "2026-09-02"
vault_url: "https://vault.chels.ai/sessions/d2-performance-ai-suite/"
attribution: "Framework as taught on stage; steps written by Cole's Notes Vault (chels.ai). Verify quotes against official recordings."
This skill produces a dated three-week review calendar for the coming quarter, a scoring sheet every employee and manager fills the same way, a calibrated band for each person, and a one-line rule tying four quarterly scores to one annual raise decision. Tom Ellsworth's reason for building it is a reversal he sees managers make constantly: the upbeat, outgoing report gets rated above the quieter peer who "actually delivers on time" and is "more buttoned up on budget." "This person should be an A minus and this person should be a B plus. But you get it backwards" (18:37:11). Paul Williams relayed the brief Patrick Bet-David gave him: "if we don't have a system, it is literally just a popularity contest ... Give me something that measures everybody by the same yard stick" (19:10:19).
Run this with the user before quarter close. If they are mid-quarter with no goals set, start at step 0.
Inputs to gather
Quarter close date. Default: last day of the current calendar quarter.
Headcount, and how many people each manager has. Williams' thresholds: around 15 employees you start meeting people you do not know; around 50 you ask "who's that?"; at 100 the question becomes who got raises (19:16:44).
Whether company goals exist and have been cascaded to individuals. If not, that comes first.
Who sits on the calibration committee. Valuetainment uses the leadership team: in Ellsworth's words, "Paul and his team, myself and Pat" (18:40:36). Default: CEO plus every manager's manager.
The tool: spreadsheet, form, or software. The process works in any of them; the safeguard in step 6 matters more than the tool.
Whether managers will see the self-score before scoring (step 3).
Steps
0. Confirm goals exist top-down
Ellsworth: "The company needs clear goals that are set, and so you set those, you take the time to make sure that they're top down, and now you've got everything set, but it's delivering it to the people" (18:39:54). The self-assessment asks people to rate themselves against expectations; if the expectations were never delivered, the review produces arguments instead of improvement. If goals are missing, write company goals, cascade them to teams and individuals, and run the first review a quarter later.
1. Put three dates on the calendar the day the quarter closes
Ellsworth's timeline: "Following the close of a quarter, we have self evaluations for a week. We then have managers who are given about a week, and then by the end of three weeks, it's done" (18:41:05). The constraint is deliberate: "This can't be a distraction. It's got to be an enabling process that's efficient, that fits in the calendar, but quickly gives you exactly what you need" (18:41:16).
Milestone
Due
Owner
Self-assessments complete
Quarter close + 7 days
Every employee
Manager assessments complete
Quarter close + 14 days
Every manager
Calibration done and results delivered
Quarter close + 21 days
Committee, then managers
If the cycle runs long, cut steps rather than extend the deadline (our recommendation, from the "can't be a distraction" rule).
2. Week one: self-assessment with evidence
Employees "rate themselves in one to 10 on five areas" (18:40:23). The five categories Williams showed are effort, attitude, leadership, innovation and results, giving a total out of 50 (19:12:20). Every score needs an example. Williams described the prompt that made it work: "It says, hey, you give me a score, you didn't give me a specific example" (19:14:07). A score without an example is sent back.
Make it easy to complete. Most Valuetainment staff did their last round on a phone and Williams "did 42 of my calibrations in car play mode while I was driving" (19:12:01). The less bureaucratic it feels, the more honest the input. Allow voice notes or a short conversation instead of a form if the tool permits.
3. Week two: manager assessment, independently
Managers score the same five categories on the same scale, with the same evidence rule. Decide one policy for whether the manager sees the self-score first and apply it to everyone. Williams: "We have the ability to show this to the manager before they do the manager calibration or we can keep it away. In this instance, we kept it away" (19:13:10). In that demo the employee scored 45, the manager scored 37 without seeing it, and the committee moved the final number up. Hidden gives an independent read; visible speeds the conversation. Default to hidden (our recommendation).
4. Week three: committee calibration and band
The committee reconciles the two scores per person. Ellsworth: "What did your team think, Paul? What did you think, and what do we think together? And let's process this through" (18:40:44). Record a final score out of 50 and assign one of five bands.
Band
Definition, speaker-stated
Does not meet expectations
Improvement required before any pay conversation
Needs improvement
Same
Meets expectations
"Doing the job excellently that you were hired to do on time within budget in the framework and KPIs you were given" (18:43:32)
Exceeds expectations
Over and above what was asked, like a salesperson at level one over quota
Outstanding
Well over and above, like level two over quota; "a very small percent"
Five bands is deliberate. Williams: a ten-point scale or extra middle categories "would get confusing. So we deliberately, very purposely and intentionally have five set ratings" (18:44:39). The speakers gave no numeric cut-offs from the 50-point total to a band; the committee assigns bands by definition. Check the distribution after the fact: if more than roughly one in ten land in outstanding, recalibrate (our recommendation, from Ellsworth's "very small percent").
Hold the line on meets. Ellsworth: "People say, well, why am I not an outstanding? You're a high meets, which means you're incredibly valuable, but you're doing what was asked" (18:43:53). And: "Meets is a great score for people doing a great job within the bounds of their job" (18:44:23).
5. Deliver like a report card
Managers deliver the final score, band and reasoning to each employee before day 21. Ellsworth's picture: "It's like a teacher getting the report cards and sending them home. That's exactly what we're doing here so people know where they are and what's expected of them" (18:41:31). The report card also becomes the record: when a question comes up next quarter, "you just go back to it and say, well, what did we all say last quarter? It's right here" (19:18:10).
6. Install the release safeguard
Valuetainment's director of studio operations, weeks into the job, entered his team's scores in an early tool and pressed done. "What I didn't realize about the done button was that it meant done done, and then it got sent out to all the employees from the studio" (19:20:47). Uncalibrated scores reached the whole team, and it "led to a lot of trust issues" for months (19:21:16).
Whatever the tool, map four states: draft, manager review, committee approval, released to employee. Confirm nothing can reach an employee before committee approval. Test with one dummy record before the first live cycle. Brief every new manager on the flow in their first two weeks. (Steps are ours; the lesson is the staff panel's.)
7. Ask every manager the five duty questions
The process exists partly to give managers "a duty, a duty to objectively sit back and manage your team" (18:38:32). At the start of each quarter, ask each manager Ellsworth's five questions and record yes or no:
Have you sat down with everybody?
Have you set expectations with everybody?
Have you given your high performers their props?
Have you given your low performers their severe encouragement to get going?
Have you given everybody alignment with the plan for the year?
Any no becomes that manager's action item for the next two weeks.
8. Roll four quarters into one annual pay decision
Ellsworth: "We don't do raises quarterly. We do feedback quarterly en route to a full annual score, where we then would talk about raises" (18:42:51). Williams on why: "otherwise when it comes to compensation, recency kicks in. Everybody realizes at Christmas, we're now talking about compensation changes for next year and they start to change the way they turn up" (18:45:23).
Annual pay consequences by band, speaker-stated:
Annual band
Raise
Does not meet, needs improvement
"We don't even talk about raises. You have to deliver improvement" (18:42:43). "You have to be firm enough in the concept of meritocracy to say no" (18:43:05).
Meets, exceeds
"Very average merit increase" (18:43:16)
Outstanding
"Above market increase" (18:43:16)
Publish the rule. When someone asks for a mid-year raise, point to it.
9. After calibration: nine-box, recognition, publish the shape
Three follow-throughs Williams described. Place every person on a three-by-three grid of impact (results) against performance or potential (the other four categories). Study the two off-diagonal corners: high potential with low impact needs unblocking; high results with poor culture fit "by definition can't be in this top" box (19:16:34). Stack-rank the top box across departments to find the real top 5 to 10 percent. Host a founder meal for everyone rated exceeds or above; Williams' point is that at 100 people the founder may know six of them (19:17:39). Publish the grid shape for the whole company without names or scores: "it doesn't show your scores, but it allows you to know where you are and where your trajectory is" (19:18:32).
Output format
Fill assets/quarterly-calibration-worksheet.md. In brief it contains:
The three dated milestones for this quarter with owners.
The five-category scoring sheet with an evidence column, used identically by employee and manager.
The calibration record per person: self, manager, committee final, band, delivered on.
The manager duty grid.
The annual rule, written as one sentence the whole company can read.
Rating the friendly report above the quiet deliverer. The whole system exists to catch this reversal.
Letting the cycle sprawl for a month. It becomes a distraction and people stop answering honestly.
Treating everyone identically inside the system. Williams: engineers "inherently don't like conflict, generally don't daylight ... if you don't manage them correctly, they will only run at one pace" (18:39:36). Keep the yardstick identical; vary the coaching.
Handing out outstanding freely. Everybody believes they are at "125% of quota" (18:44:16); calibration brings them back.
Deciding raises off one annual review. Recency and Christmas-season behaviour take over.
Releasing scores before calibration. Months of trust damage from one done button.
When not to use this
Do not run a review cycle before goals have been set and delivered; do step 0 first. Do not use the bands to force a fixed distribution; the check on outstanding is a sanity test, not a quota. For teams under about ten people where the owner works alongside everyone daily, the weekly one-to-one and the five duty questions may be enough; add the full cycle as headcount passes the point where the owner stops knowing everyone's work firsthand.
Companion skills: bonus-waterfall-designer uses the annual band from this process to set bonus payout. quarterly-calibration, from Patrick Bet-David's own Day Two calibration session, covers the same five categories and bands with his numeric cut-offs and a manager-spread diagnosis; use it when the user wants fixed score-to-band thresholds rather than committee judgment. This skill is the calendar and process wrapper around either rubric.
See references/source-notes.md for verbatim quotes with timestamps.
# frontmatter
name: quarterly-calibration-calendar
description: "Set up and run Valuetainment's three-week quarterly performance review as taught by Tom Ellsworth and Paul Williams: week one self-assessment, week two manager assessment, week three committee calibration and delivery, scored 1 to 10 on five categories (effort, attitude, leadership, innovation, results) with a specific example behind every score, five rating bands where 'meets expectations' is the normal and good score, quarterly feedback rolling into one annual raise decision, and a release safeguard so no score reaches an employee before calibration. Use whenever a manager or owner mentions performance reviews, annual reviews, calibration, rating scales, recency bias, one-to-ones that drift, deciding raises, a nine-box, or says reviews feel like a popularity contest — even if they never ask for a 'review process'."
license: MIT
metadata:
source: "The Vault Conference 2026 · Hire, Review and Pay by the Same Yardstick: A People System That Scales · Tom Ellsworth and Paul Williams"
session_key: d2-performance-ai-suite
speaker: "Tom Ellsworth and Paul Williams (Valuetainment), with Valuetainment staff; introduced by Patrick Bet-David"
day: "2026-09-02"
vault_url: "https://vault.chels.ai/sessions/d2-performance-ai-suite/"
attribution: "Framework as taught on stage; steps written by Cole's Notes Vault (chels.ai). Verify quotes against official recordings."