Skip to content

Jun 19, 2026 · 17 min read

Org Design for an AI-Native Design Org

Meet the skill librarians, eval owners, and agent-ops. They’re the connective tissue your org is missing to prevent it from descending into AI-induced chaos.

The most important markings at an orchestra concert are in pencil. By the time the house lights go down, someone the audience will never see has already sourced the right edition of the score, marked the bowings into every string part so that forty arms move in the same standardized direction, fixed the page turns that would force a player to drop a bar, corrected the errata, and set a performance-ready folder on every stand. The job is called the performance librarian. It’s real and salaried, with its own international professional association more than 560 members strong, and its mandate is plain: acquire, prepare, catalog, and maintain the music.

No one buys a ticket to watch the librarian.

And yet, the orchestra cannot play without one.

That role has been stuck in my head all year, because our industry just chose the orchestra as its official metaphor for design’s future — and aspired to always be the conductor.

Nobody budgets for a librarian

The consensus metaphor for the AI-era designer is already set: orchestration. Hervé Mischler, a principal product designer at Salesforce, put it like this:

Our role is evolving from interface architects to experience orchestrators and system builders — we’re building the systems we design.
Hervé Mischler · Salesforce

Kat Holmes, who was Salesforce’s EVP and Chief Experience & Design Officer until earlier this year, made the same move one level up — arguing that everyone becomes an agent orchestrator, managing multiple agents daily the way we manage tabs. And her advice to design leaders was blunt about the org-chart consequences: “Hire conversation designers and context engineers.”

Cool. I broadly buy the orchestration framing. But notice which chair every one of these pieces reaches for. The conductor is the glamorous lead of the metaphor — vision, taste, interpretation, the arms everyone watches. Nobody asks who else an orchestra actually employs. The Chicago Symphony’s librarians describe their job as making sure the music reaches the players in the best possible state to be performed — and if you read that sentence as an org designer instead of a concertgoer, it snaps into focus: sourcing, standardizing, quality-checking, and distributing the material everyone else performs from.

The new roles emerging in AI-native design orgs — the people who curate the workflow library, own quality evaluation, and keep a growing pile of tools and agents from becoming a junk drawer…

They are not headcount inflation. They’re connective tissue, the small set of functions that turn a hundred individually-augmented designers back into one org. But you already know the question that greets them at a budget review: which three roles are you giving up? On a reorg slide, connective tissue reads as overhead. Your HR system doesn’t even have job codes for it yet.

In the last piece in this series I argued that AI adoption is an ops problem and that your workflow library needs an owner. This piece is about who — and about arming you for that budget review with the artifact that settles it: a draft RACI for skills, evals, and tool governance you can mark up with your leads this week.

The 2016 org chart

The book most of us used to structure our orgs is Merholz and Skinner’s Org Design for Design Orgs — the centralized-partnership model, the team-of-teams, the twelve qualities of an effective design organization. It has aged remarkably well, and quality number twelve, manage operations effectively, is exactly what we’re talking about.

But the book is from 2016. It could not have anticipated a world where the material of design work is itself generated, evaluated, and operated.

Merholz himself sees the shift. On his podcast this spring, reporting back from a listening tour of design executives, he described what’s actually changing: “We’re collapsing the distance between the work of design and the work of development.” One leader he spoke with had spun up a two-person AI enablement squad that grew into a platform design team. The episode’s hottest take is that many orgs cut design operations right before they were needed most.

Kristin Skinner, his co-author, made the underlying mechanism explicit in a recent essay on why AI deployments stall:

AI runs on whatever it finds.
Kristin Skinner · Velocity, on Substack

Her argument is that AI doesn’t fix an operating model — it inherits it, and accelerates whatever dysfunction it lands on. Faster analysis piles up at the same unchanged approval gates. Insight gets generated with no structural owner waiting to receive it. That maps exactly onto what I see in design orgs: the tooling arrived years before the roles that were supposed to catch its output.

And the DesignOps canon confirms the gap by omission. NN/g’s DesignOps 101, whose definition literally opens with the word orchestration, contains no AI line item at all.

It couldn’t. It’s from 2019.

NN/g’s own recent guidance has caught up: Laura Klein’s piece on AI and ops leads is titled, essentially, as an order: “Stop expecting everyone to figure this out on their own.” Her prescription is curated prompt repositories, standardized workflows, security policies that make the compliant path the easiest path, and piloting before scaling.

Notice what all of that is. It’s not tooling. It’s roles — persistent ownership of shared material.

So let’s give these roles a name.

Three new roles, and why each one kind of already exists

I want to be precise about what I’m claiming. I’m not claiming your org needs three new directors. At 100 people, these start as hats — fractional assignments, maybe 1.5 to 3 FTE of total capacity across all three. What I’m claiming is that each is a distinct function with a distinct failure mode when unowned, and that they can’t be smeared across everyone as “part of the job.” Somebody marks the bowings, or nobody does.

The skill librarian

The librarian owns the org’s library of reusable AI capability: prompts, workflows, custom agents, and increasingly skills in the formal sense — Anthropic’s Agent Skills are literally folders of instructions and resources an agent loads on demand, designed so that “organizations can distribute skills across teams,” and Anthropic’s engineers describe writing one as putting together an onboarding guide for a new hire.

Read that twice: the artifact your practitioners are now producing is organizational knowledge, packaged. Someone has to run intake on contributions, verify them against current models, document what each is good for and where it fails, and, this is the part everyone skips, retire the ones that rot.

A skill is a directory containing a SKILL.md file that contains organized folders of instructions, scripts, and resources that give agents additional capabilities.
A “skill” is a concrete, versionable artifact that someone on your team can own. It ain’t a vibe. Anthropic

If you doubt the retirement problem is real, look at Moderna, the poster child for enthusiastic adoption: more than 750 custom GPTs deployed across the company, with 40% of weekly active users building their own. That’s an impressive crowd. It is also, without a librarian, seven hundred and fifty uncatalogued scores of unknown edition. (“Prompt librarian” already exists as a term in library science, fittingly. “Skills librarian” sounds cooler — use whatever your HR system will tolerate.)

Who takes the hat: whoever ran your design-system contribution model. Same instincts, new material.

The eval owner

If the librarian owns what we reuse, the eval owner owns how we know it’s good. Hamel Husain’s much-cited diagnosis of failed AI products is that they share one root cause : “a failure to create robust evaluation systems“. And the discipline that grew around that insight has become, in Lenny Rachitsky’s framing, the hottest new skill for product builders: “Evals are the new PRDs.” Eugene Yan’s version of the point is the one I’d put on the wall: evals are a practice, not an artifact — an ongoing application of the scientific method to your AI-assisted output, not a test suite you build once.

For a design org, the eval owner defines the quality bar for each standardized workflow — what a good research synthesis, previz sequence, or UI-copy batch looks like, and how you’d check. And re-runs that check every time a model version changes underneath you. This is design-quality work playing a new instrument. It’s also already happening at the companies furthest along: when Salesforce restructured around agents, it doubled its evaluations team by redeploying support staff to assess agent answer quality. And NN/g’s newest taxonomy of the design jobs AI created includes a whole orientation: designing the AI itself, whose day-to-day is shaping model behavior and evaluation criteria.

Who takes the hat: Pull from research. Your researchers already own rigor, sampling, and “how would we know?” — this is the single highest-leverage redeployment of a research background I can think of right now.

AgentOps

The third role is the one that sounds most like invented jargon and is most rapidly becoming real. AgentOps, the operations discipline for AI agents, already has academic surveys defining it: the monitoring, anomaly detection, and maintenance of agent systems. In a design org, agent-ops owns the tool registry (what’s approved, for which data), the pilots, the spend, and the question nobody owns today: when an agent-assisted workflow fails, who runs the postmortem?

The market data says this role is arriving whether we plan for it or not. Microsoft’s 2025 Work Trend Index, the “Frontier Firm” report, found 28% of managers already considering hiring AI workforce managers, 32% planning to hire AI agent specialists within eighteen months, and 78% of leaders considering AI-specific roles overall. Salesforce’s own newsroom put out a press release: “The company is also hiring net-new roles like deployment strategists, AI conversation designers, and AI architects that didn’t exist just months prior.”

An illustration depicting how many agents per worker is feasible, and how to manage it.
Major-employer research is already modeling org structures around agents, not just tools Microsoft

Who takes the hat: a design technologist or a DesignOps program manager with a systems streak — someone who reads an API bill without flinching.

But isn’t this exactly how bureaucracies are born?

The skeptic in you has a fair point, and I’d rather steal it than dodge it. Yes — every platform shift spawns a priesthood, and some of it calcifies. “Head of AI” as a title tripled in five years per LinkedIn’s data, and McKinsey finds organizations hiring AI compliance and ethics specialists in growing numbers. Not all of those roles will earn their keep.

Here’s the test I’d apply, and it comes straight from the metaphor. The librarian exists because the alternative is worse and more expensive: a hundred musicians each sourcing their own edition, marking their own bowings, discovering the bad page turn mid-concert. The connective roles pay for themselves when the duplicated, invisible coordination work they absorb, currently smeared across every one of your designers in five-minute increments, exceeds their cost.

At a team of eight, it doesn’t, of course. Don’t build this at a team of eight. At a hundred-plus people across five disciplines, the smeared version is already costing you more than three salaries. You’re just paying for it in a currency your CFO’s spreadsheet doesn’t yet acknowledge.

The other half of the answer is that these roles need decision rights, not just titles. A librarian who can’t retire a rotten skill is a wiki gardener. An eval owner who can’t block a workflow from being standardized is a QA theater troupe. It’s more than three new job descriptions:

It’s a RACI.

Who owns what: a hypothetical RACI

Quick translation for anyone who’s escaped this particular acronym: a RACI names, for each decision, who is Responsible for doing the work, who is Accountable for the outcome — exactly one name, the same rule Atlassian’s DACI framework enforces on its Approver, plus who’s Consulted before and Informed after. It’s unglamorous. So are bowings.

This is a draft, deliberately. It encodes my defaults; the workshop where your leads argue with it is where it becomes yours.

Decision VP Design Skill librarian Eval owner Agent-ops Discipline leads Practitioners Security / Legal / IT
Contribute a candidate skill or workflow to the library I A I C R
Publish a skill to the org library I A / R C C R (contributor)
Set the quality bar + evals for a standardized workflow I C A / R R C
Re-run evals when a model or tool version changes I C A R I
Retire or quarantine a failing skill I A / R C I C I
Approve a new tool for pilot (non-sensitive data) A I I R C I C
Approve a tool for production / customer data A I C R I I C (hard gate)
Monitor agent usage, cost, and failure modes I I C A / R I I
Run the postmortem when an AI-assisted deliverable ships wrong I I C A / R R (owning pod) R C (if customer-facing)
Onboard designers to the library and its norms I A C I R

Four things to notice about the shape of it, because the shape is the argument:

  • Practitioners are R on contribution. The crowd creates, the librarian curates. If your designers aren’t Responsible anywhere on this chart, you’ve built a ministry, not a library.
  • The eval owner is Consulted on publishing and Accountable on quality. Publishing without a quality gate is how libraries rot. Quality gates without a librarian are how nothing ships. The tension between those two roles is designed in. Don’t merge them to save a hat.
  • Security and legal are Consulted on pilots and a hard gate only at production data. That asymmetry is the whole game — per Klein, make the compliant option the easy option, and keep the pilot path fast enough that people don’t route around it.
  • You appear mostly as A on tool approval and I everywhere else. If you’re Consulted on every row, you’ve built a bottleneck with your own name on it. The connective roles exist precisely so you can stop being the connective tissue yourself.

Designing your librarian

Here’s how to get started, with the decision points called out:

  1. Assign the three hats before opening any reqs. Fractional, named, in writing — the design-system person takes librarian, a senior researcher takes evals, a design technologist takes agent-ops, each at roughly 20% time to start. Decision point: hats vs. heads. Start with hats, but write the conversion triggers now, while nobody’s defensive: the librarian hat becomes a head when the library crosses ~25 live skills or intake exceeds a weekly hour budget. Evals when more than three workflows are standardized per discipline. Agent-ops when agent spend becomes a real budget line. A trigger agreed in advance is a promotion: the same conversation two quarters late is an “ownership” mess.

  2. Run the RACI workshop — 90 minutes, discipline leads plus the three hat-wearers. Start from the table above and argue only about the A’s. One A per row, no exceptions. If two names feel right, the row is really two decisions — split it. Trade-off: every A you keep for yourself is real friction you’re adding to someone else’s week. Keep tool-approval-for-production. Consider delegating pilots entirely.

  3. Wire the vetoes with security and legal before they wire themselves. Bring them the two-tier structure: consulted on pilots, hard gate on production data as a proposal, not a request. In my experience, they’ll trade speed on tier one for teeth on tier two, and what they most want is what the RACI gives them: a named person (agent-ops) to call, instead of a hundred designers to chase. Atlassian’s responsible-tech review process, a lightweight cross-functional template with a traffic-light protocol, is a good starting point.

  4. Give each role one artifact and one cadence. Librarian: the library changelog, weekly. Eval owner: a scorecard per standardized workflow, refreshed on every model change. Agent-ops: the tool registry and a monthly usage-cost-failures report. One page each. Trade-off: the temptation is dashboards; resist it until someone’s actually reading the pages. A role with no recurring artifact will evaporate by Q3 — the artifact is how a hat survives contact with the roadmap.

  5. Centralize the roles, federate the contribution. These three sit together, in DesignOps if you have it, not embedded one-per-discipline. This is the centralized-partnership logic from Merholz and Skinner applied one level down: central ownership of standards and infrastructure, distributed creation by the pods who feel the problems. Decision point: discipline leads will ask for their own embedded librarian. At 100 people, say no; you’ll get five incompatible libraries — 2014’s five shades of blue, again, in prompt form.

  6. Start the HR paperwork now, cynically. Map each hat onto your existing job architecture and draft the eventual job posting even though you’re not posting it. When a conversion trigger fires, you want a two-week req, not a two-quarter job-architecture project. Your HRIS almost certainly has no code for any of these titles yet. Job codes trail reality by about eighteen months, and the Microsoft numbers above say they’re coming. Be ready before your recruiter is.

Tools & Resources

Everything referenced, and what each is for:

Those essential markings

Back to the pencil. Every player in the hall performs better because one person did the unseen work for all of them — and when the concert goes well, nobody thinks about the librarian at all. That is exactly the fate you’re signing these three roles up for, which is why they’ll be the hardest line on your reorg slide to defend. When the question comes, “what do these people actually make?,” the honest answer is: nothing the customer will ever look at, and everything the org performs from.

The job codes will catch up. Assign the hats now.