43 & 17
When I read about this, these numbers struck me as harbingers of our changing roles as leaders. They came from the largest field experiment run on AI and knowledge work: 758 consultants at BCG, a Harvard-led research team, real consulting tasks, half the subjects given GPT-4. The consultants who scored in the bottom half on a baseline assessment improved their performance by 43% once they had AI. The top half improved by 17%.
Read those numbers the way a compensation committee would. The distance between your strongest performers and your weakest ones: The distance your entire leveling architecture exists to detect, price, and reward — just narrowed dramatically, on every task inside the tool’s competence.
Ethan Mollick, one of the study’s authors, put the residual gap about as tactfully as it can be put:
The top consultants still got a boost, but less of one. … It may be like how it used to matter whether miners were good or bad at digging through rock… until the steam shovel was invented and now differences in digging ability do not matter anymore.
Here’s why this belongs in a series about running a design org, not in an economics newsletter. Your leveling rubric is a measurement instrument. It was calibrated against a world where the quality of what someone produces reliably indicated the quality of their thinking, where a senior’s artifacts were visibly, consistently better than a mid-level’s, and the visible difference was the signal.
That signal is now noisy at best. The next promotion packet you read will be full of beautiful production, and the true answer to, “is this L4 work?” is: you can no longer tell from the work.
And the interview pipeline broke the same way, for the same reason. The take-home exercise, the portfolio of polished outcomes, the whiteboard craft test: Every one of them measures production. Production is exactly what the copilot now provides to everyone, including the candidate you’re about to over-level because they have a sweet Framer portfolio full of generated UI.
Most orgs are responding with some mix of denial and prohibition: ban the tools in interviews, squint harder at packets. I want to argue for the opposite move: accept that the compression is real, and recalibrate the instrument around the thing that didn’t compress.
The gap your rubric measures is closing
The BCG experiment isn’t an outlier. It’s the most detailed entry in a stack of studies that all bend the same direction. Noy and Zhang’s randomized experiment in Science gave 453 professionals mid-level writing tasks: ChatGPT cut time 40%, raised quality 18%; in the authors’ words, “inequality between workers decreased.” The workers with the weakest baseline skills gained the most.
Jakob Nielsen, reading the BCG results alongside earlier studies of support agents and programmers, called out the pattern, “AI helps the low-performing staff the most, narrowing the gap between the two groups.”
Notice what’s being compressed, because it’s more than being faster: output quality on tasks inside the frontier. In the opening piece of this series I leaned on the same research for its other finding: Dell’Acqua and colleagues’ jagged frontier, the invisible, irregular boundary between tasks AI is superb at and tasks it will confidently ruin.
It’s a bit like Minecraft.
Deep inside the mine, you can use spells to power through the rock much faster. But instead of diamonds, you might find lava. The same spell book exists in our AI world: Everyone’s output gets better and the spread tightens. Your mid-level designer with a copilot now produces artifacts that would have read as senior two years ago.
But…
What didn’t compress is the only thing worth measuring
The same study has a second point that leaders should read twice. On a task deliberately designed to sit outside the frontier: one where the AI’s answer was subtly wrong. Consultants without AI got it right 84.5% of the time. Consultants with AI dropped to 70.6%, and the group given extra prompt training dropped to 60.6%.
The tool didn’t just fail to help. It actively made capable people worse, because they trusted polished output past the edge of its competence. In other words, they might have crafted the spells, but they don’t know how and when to use them.
So the experiment measured two different abilities, and they moved in opposite directions. The ability to produce got cheaper and flatter. The ability to leverage critical thinking, and judge: To know where the AI magic has its bounds; to smell a wrong-but-fluent answer; to decide what to verify before shipping…
Critical thinking has become scarcer and more valuable because it’s now the only thing standing between your org and confidently shipped garbage.
That second ability is what your rubric has to capture. And here’s the uncomfortable structural problem: judgment is mostly invisible in artifacts.
A high-craft portfolio shows what survived. It doesn’t show the six generations that got rejected, the flaw that got caught in review, the moment someone said this whole direction is wrong for this user and threw away an afternoon of beautiful output.
Generation leaves evidence by default.
Judgment leaves evidence only if you design for it.
We were grading the proxy all along
I don’t think AI broke our assessment methods, exactly. I think it just exposed them for what they are.
Jared Spool has been making this argument about portfolios since well before the copilots arrived — that unstructured portfolio review measures aesthetics, not capability:
We can’t just flip through the portfolio and smile at the pretty pictures.
His fix was always to define the position’s objectives first and hunt for evidence of comparable experience, evidence of what the candidate actually contributed and decided, not what the team around them shipped. That was good advice in 2019, when pretty pictures at least took skill to produce. It’s existential advice now, when pretty pictures cost a prompt.
Production quality was always a proxy for judgment — a decent one, because good production was expensive enough that it usually implied good thinking. The proxy worked until faking it became free. Every rubric column that says “produces high-quality work independently” is now paying out on a signal anyone can abuse.
The first company I’ve seen say this in public, with its hiring process on the line, is Canva. In mid-2025 its engineering org flipped its interview policy from no AI allowed to AI required — the blog post is literally titled “Yes, You Can Use AI in Our Interviews. In fact, we insist.” Their reasoning was two observations: Candidates were using AI covertly anyway, detection was a losing game, and “AI assistants can trivially solve traditional coding interview questions.”
So they retired the Computer Science Fundamentals screen entirely and replaced it with a competency they call AI-Assisted Coding. A realistic assessment with ambiguous problems, the candidate’s preferred tools in hand, scored on whether they can break down the problem, direct the tool well, and catch the flaws in what it produces.
Swap “coding” for “design” and that is exactly the recalibration our discipline owes itself, despite resting our weight on visual craft all these years.
The judgement column
So let’s rethink our heuristics, starting with another table: AI-acknowledged heuristics replace the production-quality language in your existing craft/execution column, the foundation most design ladders share (browse progression.fyi and you’ll find some version of it in nearly every design framework, including Intercom’s strategy/execution/behaviors split).
Bolting a new “uses AI effectively” competency onto the side would be worse; tool columns age like “proficient in Photoshop,” and they reward usage, which is exactly the vanity signal we’re trying to retire.
The left column is language you probably have today. The right column is an extreme, AI-pilled abstraction of what I’m replacing our Craft heuristic with. Assessors need something to observe, because a judgment rubric that can’t be evidenced is just vibes with extra steps.
| Level | What the craft column rewarded | What the judgment column looks for |
|---|---|---|
| Junior | Produces quality work with guidance and growing tool fluency. | Reliably separates broken from plausible output. States what they checked before sharing work and what they couldn’t check, unprompted. Asks for verification help rather than shipping uncertainty. |
| Mid-level | Produces quality work independently across common patterns. | Catches wrong-but-fluent output before review does. Scopes what the tool does versus what they do on a given task, and can say why. Verification is habitual. |
| Senior | Sets the craft bar, and handles ambiguous problems end-to-end. | Chooses which problems are worth solving at all. Knows where the journey lies for their domain — and writes it down (this is the clear documentation from earlier in this series, as an individual expectation). Their rejections teach: when they kill a direction, the reasoning is reusable. |
| Staff | Raises craft quality across multiple teams. | Turns personal judgment into instruments other people use. They produce multiple versions and select the best from the best. They create metadata: rubrics, evals, review checklists, seeded examples. Arbitrates quality disputes with data and critical thinking. The org’s output improves in places they never touched. |
| Principal | Org-wide craft leadership and standards. | Relentless at critical thinking. Decides where the org spends its judgment budget: which surfaces demand human review, where determinism is non-negotiable, what quality bar each workflow carries. Owns the calibration of everyone else’s calibration. |
Rejection of a direction is a good tell for critical thinking. In a large org, the promotion packet has to change: Require one rejection exhibit per packet, a piece of work (theirs or a tool’s) the candidate killed, with the reasoning. It’s the cheapest judgment evidence there is, and its absence in our current process exemplifies the current imbalance. Also anticipate that your calibration meetings get harder. Everyone is using unsanctioned tools. Evidence is easy to produce. The difficulty is the measurement finally touching the thing that matters.
On interviewing and better exercises
The second half of the challenge is the hiring side. And you have to build in imperfection. I do it through recycled problems from past experiences. I’m calling it the Salvage Exercise, because that’s the skill it isolates: taking plausible, polished, subtly flawed output, which is the default material of the next decade, and exercising judgment on it, live in a design-thinking activity.
Setup (before the loop):
- Pull a real, retired scenario from your own past experience: a flow, a screen, a research-synthesis brief. Real, because invented tasks have invented context, and context is what judgment runs on.
- Generate a draft solution, and keep a generation that contains three flaws: one surface flaw anyone competent should catch, one structural flaw (a reasonable pattern that’s wrong for this problem), and one judgment flaw: something non-obvious that’s wrong for this user or this business.
- We do ours on a Mural board, but you can also do this live, in prompt-land: Tell candidates in advance that AI tools are expected in the session and they should bring whichever they’re fluent in. You’re testing judgment, not tool trivia, and ambush conditions measure composure instead.
The session (45 minutes): hand over the brief, the context, and the draft. The prompt is one sentence: “Get this to something you’d put your name on — think out loud, and use whatever tools you’d actually use.”
Scoring: 4 dimensions, 0–2 each:
- Diagnosis. What did they notice? Full points requires catching the judgment flaw and naming what’s actually wrong; indiscriminate critique is its own tell.
- Direction. How do they drive the discussion and the tool? Do they re-prompt when the structure is wrong and hand-edit when the detail is, or swirl into a narrow part of the problem?
- Verification. What do they check against reality? Do they ask for more context? Do they stick to the brief? Do they ask about the user? Do they invent issues the materials don’t hand them?
- Stopping. Do they know when it’s shippable, and when the right move is abandoning the draft entirely? (Design the exercise so that full salvage is barely possible in the time given. The candidates who negotiate scope out loud are showing you the exact skill.)
The scoring depends on the needs of your organization. Generically, a junior bar is 3-4 with the surface flaw caught. A senior bar is 6+ including the judgment flaw. Calibrate those numbers to your own pipeline after five runs, not before.
Interviewer calibration is real work. Pilot on internal volunteers whose level you already know, and if the exercise scores your known-seniors mid-range, fix the exercise, not the seniors. And it doesn’t replace your whole loop; it replaces a take-home and a craft exercise, the two areas where the AI copilot has already scrambled the signal.
Recalibrating, in one cycle
Calibration updates for large design orgs:
- Audit the last cycle’s promotion packets. For every piece of evidence, marke as production (the artifact is the argument) or judgment (a decision, rejection, or verification is the argument). Most orgs will find packets running 80–90% production evidence. That ratio is your exposure, quantified. If a recent promotion looks shaky under this light, resist relitigating it. It’s unfair, and you’re recalibrating the instrument, not auditing the people it already measured.
- Rewrite the craft column with your leads. Start from the table above. Argue every cell into behaviors your assessors can actually observe. Trade-off: Appending a critical thinking column while keeping the old one means production evidence keeps winning ties, and nothing changes.
- Add the rejection exhibit to the packet template now, a full cycle before it becomes crucial, so people have time to start noticing their own rejections. Expect the first crop to be thin. It takes time to build new confidence in a new muscle.
- Build one Salvage Exercise from your own materials. One discipline first, whichever pipeline hires next. The flaw-seeding session doubles as a useful afternoon: you’re writing current documentation for your own stack, in exercise form.
- Pilot internally. Run five internal sessions to calibrate scoring, then tell candidates exactly what the session is and that AI is expected. Put it in the job description. Transparency here isn’t just kindness. Per Canva, secrecy was already lost, and the candidates who prep with their tools are showing you their real working setup.
- Re-norm quarterly against the tide. The bleeding edge moves every model release, which means your calibrations rot and your level bars drift. Book the recalibration review now with the same rhythm as the eval re-runs your eval owner already runs when models change. A leveling heuristic is now a living instrument, and living instruments have maintenance schedules.
Tools & Resources
- Navigating the Jagged Technological Frontier — Dell’Acqua et al., SSRN (also in Organization Science) The 758-consultant experiment behind the 43/17 numbers and the outside-frontier accuracy drop. The primary text for this whole argument.
- Centaurs and Cyborgs on the Jagged Frontier — Ethan Mollick The study author’s readable walkthrough, including the distribution charts to show your team.
- Experimental Evidence on the Productivity Effects of Generative AI — Noy & Zhang, Science —The writing-task replication of the compression effect. Useful because it’s outside consulting and it affects your cross-functional partners.
- AI Helps Elite Consultants — Jakob Nielsen The skills-gap-narrowing synthesis across studies, in exec-forwardable form.
- Yes, You Can Use AI in Our Interviews — Simon Newton, Canva Engineering The public precedent for AI-required interviews and the AI-Assisted Coding competency. The closest existing template for the Salvage Exercise.
- Reviewing UX Portfolios: 4 High-Risk Hiring Mistakes — Jared Spool Pre-AI argument that unstructured portfolio review was already measuring the wrong thing.
- progression.fyi Public career frameworks (including Intercom’s design ladder). Useful for seeing how universally the production-quality column shape repeats.
Back to 43 & 17
Every asset in your talent system, including the heuristics, the packet, the portfolio review, the take-home, was built to measure a gap that is narrowing and, dare I say, passé.
What’s left of the gap lives somewhere they can’t see at all: in the critical thinking your people leverage to refuse, verify, and know better than to trust.