Why we rebuilt our AI Fluency Assessment as a simulation
We had a working AI fluency test. Multiple choice, five dimensions, built on solid research. It scored well and customers used it. And it still bothered me, because a multiple-choice test has a ceiling it cannot get past: it measures what a candidate can recognize, not what they can do.
Here is the moment that decided it. We watched two candidates score almost identically on the knowledge version. Same recall, same “correct” answers about evaluating AI output. Then we put them in front of a real task with a real AI draft that had a quietly wrong number buried in the third paragraph. One candidate caught it in about forty seconds. The other shipped it. On paper they looked the same. In the work they were not remotely the same hire.
So we rebuilt the assessment around simulations.
What changed
Candidates now work through a realistic task instead of picking answers from a list. They brief the AI, look at what it gives back, decide what to trust, fix what is wrong, and turn it into something usable. We score the behavior as it happens, across the same five dimensions: task strategy, prompting, critical evaluation, ethical use, and workflow integration. Still about 15 minutes. Still tool-agnostic. Still 0 to 100, no pass or fail.
We kept a multiple-choice format for high-volume screening, because when you are sorting thousands of applicants you sometimes need speed over depth. Both formats feed the same profile, so the scores stay comparable.
What we learned building it
Three things surprised us.
First, critical evaluation is where most people fall down, and it is invisible in a quiz. People are good at describing the idea of checking AI output. Far fewer actually do it when there is a plausible-looking draft in front of them and a clock running. The simulation exposes that gap in a way no question bank can.
Second, prompting skill and evaluation skill do not travel together. We assumed strong prompters would also be strong evaluators. They are not. Some of the most fluent prompters were the least skeptical of what came back, which is a specific and hireable risk to know about.
Third, tool-agnostic matters more than we thought. Candidates showed up comfortable with different tools. Measuring the transferable behavior, rather than proficiency with one product, meant the scores held up regardless of what they used day to day.
Why this is the right way to hire for AI
AI fluency is a hard skill now. It belongs next to Excel and data analysis on the list of things that decide whether a new hire performs. And like those skills, it is best measured by watching someone do it, not by asking them to talk about it.
The industry is not there yet. Most AI assessments on the market still test definitions. That is the easy thing to build and the wrong thing to measure. We would rather show a candidate a messy, realistic task and see how they handle it, because that is the job.
If you want to see the simulation and a sample five-dimension profile, book a demo. The AI Fluency Assessment is included in every Bryq plan and turns on from your dashboard inside your existing ATS.










