Product managers now make decisions on evidence that AI gathered, summarized, or analyzed. They hand specs and briefs to AI agents that build exactly what's written. And they ship features that behave differently when the AI model behind them is updated. The AI Fluency Assessment for Product Managers measures how well they handle AI inside that work.
We built it in two parts. The first decides what to measure: which skills earn a place, and what evidence backs them. The second decides how to measure them. Each scenario tests one skill on its own, follows the same set of rules, and goes through several rounds of independent review before it's used. Every scenario is calibrated against how real candidates score, and we keep revising it for measurement accuracy.
Part one: deciding what to measure
The evidence base
The skills rest on two kinds of evidence: published research on AI in product work, and a job analysis of the role itself.
The research looks at the work product managers actually do: judging evidence that AI produces, handing work to an AI agent, and running AI inside a live feature. It also shows which problems only exist because AI is involved. Every claim we make about AI in that work rests on two kinds of study. One shows the human tendency behind the problem, from any period of research. The other shows the same problem with AI, and was published in 2023 or later.
The job analysis
For the job analysis, we studied live product manager job postings from US employers, all posted within the last year. They cover the three main kinds of product role in equal measure: technical product managers who work on platforms and systems, growth product managers who work on acquisition and retention, and product managers who own the core product and its features.
We read each posting for three things: the work the role produces, the AI tools it names, and any AI responsibility it spells out. That showed us which outputs every product manager produces, like roadmaps, success metrics, and stakeholder updates, and which ones belong to a particular kind of role. Experiment plans, for example, sit at the heart of growth roles but rarely show up anywhere else.
For each output, we then set out two things.
Strong performance. What good work on that output looks like when AI is part of it. For a roadmap built on AI-summarized customer research, that means knowing which parts of the summary the decision rests on, and checking those first.
Failure modes. A failure mode is a specific way AI makes that work go wrong. Each one captures three things: what goes wrong, why nobody notices, and what a capable product manager without AI fluency does instead. Say an AI summary of customer interviews leaves out a view that only one or two customers hold. The summary still reads as complete, so nothing looks off, and a capable product manager without AI fluency takes it into the roadmap review exactly as it is.
The exclusion test
Every failure mode then faces one test: take the AI out of the work. If the problem still happens, it belongs to product management in general, and we exclude it. We only keep the failure modes that exist because AI is involved. That's what keeps the assessment measuring AI fluency inside the role, rather than the role itself.
Five gates a skill has to pass
Next, we group the failure modes by the capability that would have prevented them. Each group becomes a candidate skill, and every candidate skill has to pass the same five gates, in order. If it fails any one of them, it's cut.
Gate | The test |
|---|---|
Specific to the role | Only a product manager runs into this failure. If it would read the same for any professional, it belongs in the general AI Fluency Assessment instead. |
Not already covered | The general AI Fluency Assessment doesn't already measure it, or doesn't measure it at this depth. A role skill has to add something new. |
Measurable in a work scenario | A single scenario from the role's work can measure it, with a graded score from the best response to the weakest. |
Needs AI fluency to answer | A strong product manager with little AI experience gets it wrong. If ordinary professional judgment gets it right, the skill is measuring the job, not AI fluency. |
Enough evidence to measure | The skill can be measured through several different scenarios, each drawn from separate research, so no score ever rests on a single example. |
The three dimensions of AI fluency in product management
This is where part one lands. Of all the candidate skills the research turned up, three passed all five gates. Each one became a dimension: a distinct skill the assessment measures and scores on its own. Together, with their definitions and the evidence behind each, the three dimensions make up the rubric for the AI Fluency Assessment for Product Managers.
Each dimension covers a different way AI now shows up in product work: the evidence a product manager decides on, the work they hand to AI agents, and the AI features they ship. Every scenario in the assessment is written against this rubric, and every score traces back to it.
Critical eye on AI output. Judges whether the facts AI gathered, summarized, or analyzed are solid enough to decide on. It also covers owning work written with AI: standing behind it, with confidence that matches the evidence.
Delegating to AI agents. Writes the specs, acceptance criteria, documented research, and briefs that an AI agent acts on, with no person there to fill the gaps.
Designing reliable AI features. Specifies an AI feature so it holds up in use. That means setting a measurable standard for the quality of its output, and deciding what users see when the output is wrong. It also means checking that the feature works for every group of users, and planning for what changes when the model behind it is updated.
Part two: how each scenario is built
Specifications set before writing
Before we write a scenario, we fix its specifications. They apply to the scenario and to every one of its options.
Difficulty. Every scenario sits at one of three levels, defined by what it asks the candidate to do. A beginner scenario turns on a single AI behavior. An intermediate one asks the candidate to weigh two considerations against each other. An advanced one rewards a response that goes against common practice.
Neutral language. Plain English that someone outside product management can follow on a first read, with every product term explained where it appears.
Fairness. Nothing in the scenario gives anyone an edge for reasons that have nothing to do with AI fluency, such as their first language, cultural background, access to a paid AI tool, or experience at one kind of company.
Tool-agnostic. No scenario depends on how a particular AI product works. If a scenario needs a product to behave a certain way, it says so, and the best response rests on how AI works in general.
Weighted options
Each scenario comes from a product manager's own work, with AI in the picture, and offers four options. Each option carries a weight, and the weights follow the same logic in every scenario.
The best option splits the work correctly between the person and the AI, in proportion to what's at stake. It carries the full weight.
A workable option gets to an acceptable result, but at a real cost, like an extra round of work or a check aimed at the wrong risk. It carries part of the weight.
A plausible option sounds reasonable, but rests on a misunderstanding of how AI works. It carries a small part of the weight.
The AI-naive option is what a capable product manager without AI fluency would do. It sounds sensible, and it carries no weight at all.
The best option and the AI-naive option always differ on one fact about how AI behaves. That might be something the AI can't know, information it never sees, or a gap between how reliable its output looks and how reliable it really is. Every option comes with a written reason for its weight, naming the AI behavior behind it.
What earns the full weight also changes from one scenario to the next. Sometimes it's handing more of the work to the AI than instinct suggests. Sometimes it's changing what the AI is given, rather than checking what it returns. And sometimes it's not using AI for that step at all. That variety means no single habit, like always picking the most cautious option, leads to a high score.
Gates for every scenario and option
Every scenario and every option then goes through the same gates. Anything that fails gets rewritten or cut.
Gate | The test |
|---|---|
Needs the AI | Take the AI out, and the scenario stops working. If it still works, it's measuring the job, not AI fluency. |
Specific to the role | Swap in another role, and the scenario stops working. If it still works, it isn't really about product management. |
No clue in the options | With the scenario hidden, nobody can pick the best option just from how the options are written. |
Every wrong option is a real choice | Each weaker option is one a capable product manager would actually pick. None of them can be ruled out without knowing how AI works. |
One checkable fact | The gap between the best option and the AI-naive option comes down to a checkable fact about AI, never to style or tone. |
What each scenario measures
Each scenario drops the candidate into a real moment from product work, one where AI is already part of the job. An AI-generated summary of customer interviews might be about to shape the roadmap. A feature brief might be about to go to an AI agent that will build exactly what it says, nothing more and nothing less. A new AI-powered feature might be about to launch to customers. In every case, the candidate has to decide what happens next.
Every scenario is backed by more than one published study. The scenarios about briefing AI agents, for example, build on research presented at ICLR 2026 showing that AI coding agents have trouble telling a complete brief from an incomplete one. So when a product manager leaves a gap in a brief, the agent may simply fill it in on its own.
The assessment doesn't test theory, definitions, or technical knowledge of how AI models are built. Knowing about AI and working well with it are two different skills, and only the second one counts here. What we score is the decision a product manager makes when AI is part of the work, and whether that decision holds up given how AI really behaves.
How scoring works
Each option's weight counts toward the score for its dimension. Each of the three dimensions has five scenarios, and its total is reported as a percentage. The overall score is the average of the three, so each dimension counts for a third. There's no pass mark.
Because each dimension is scored on its own, the report shows where someone is strong and where they need support. That makes the assessment useful beyond screening, too. With current employees, a talent team can see which dimension needs work for each person, and focus training there instead of across the board.
The assessment takes about nine minutes. Every scenario has a fixed sixty-second timer, so a sitting never runs longer than fifteen minutes.
Independent review
Once a scenario clears the gates, it goes through two rounds of independent review, and each round catches things the other can't.
First, Bryq's in-house industrial and organizational (I/O) psychologists review it for readability. They keep revising until someone outside product management can follow it on a first read.
Then an automated review applies every rule to every scenario, several times over, so no scenario depends on a single reading. It also tests each scenario against the toughest objection a senior product manager would raise.
How the assessment is calibrated
Every scenario was calibrated before launch. We set its difficulty during design, then checked it in independent review against the dimension it measures and against the other scenarios in the set. Market experts in product management also took part in the calibration.
Calibration doesn't stop at launch. As candidates take the assessment, we compare each scenario's results with the difficulty it was set at, and with how well it separates product managers who work well with AI from those who don't. Any scenario that doesn't measure as intended gets revised or replaced.
Where the assessment fits
The AI Fluency Assessment for Product Managers sits on top of the general Bryq AI Fluency Assessment. The general one covers five dimensions of working with AI in any role, while the one for product managers covers three dimensions of product work. They're two assessments, with two separate scores.
The page AI Fluency Assessment for Product Managers describes what candidates see and what employers get back. The page How to hire a product manager shows where it fits in a full hiring process.
Frequently asked questions
How do you know the scenarios measure what they claim to? Every dimension traces back to documented ways AI goes wrong in product work, and each one is backed by published research. Every scenario passes a set of gates and two rounds of independent review, and was calibrated before launch with input from market experts in product management.
Does it measure knowledge of AI or use of AI? Use. Each scenario puts the candidate in real product work, with AI part of the task, and challenges them to decide what happens next. No scenario tests what a model is, how it works, or what a term means.
How does a skill earn a place in the assessment? It has to pass a set of gates. It must be specific to the role, and measurable in a single work scenario with a graded score. It must also be something a strong product manager without AI experience gets wrong, and it needs enough research behind it to support several different scenarios.
How is a scenario checked before it's used? First, we fix its specifications: difficulty, neutral language, fairness, and tool-agnostic wording. Then the scenario and every option pass a set of gates, such as only working with the AI in it, and giving nothing away in how the options are written. Next come two review rounds: readability, by Bryq's in-house I/O psychologists, and an automated review that applies every rule to every scenario. Finally, it's calibrated with market experts in product management.
Which product roles does it cover? Product managers and product owners, across technical, growth, and core product roles.
Where can I read more about AI fluency in hiring? In How to assess AI skills in hiring, and on the AI Fluency Hub.







