A 60-minute skills test tells you whether an AEO candidate understands mechanism or just vocabulary. Give them a live URL and ask five things: why is this page or is it not cited, how would you diagnose it, what would you fix first, how would you measure the result, and what would you refuse to promise. Strong candidates name specific, mechanism-level causes (renders client-side, no entity anchor, answer buried past the first hundred tokens), sequence fixes by dependency, measure citation frequency not rankings, and refuse to guarantee outcomes. Weak candidates speak in keywords and generic best practices. The test is not knowledge recall. It is diagnostic reasoning on a real page.
Hiring for AEO is hard because the title is new and the market is full of people who relabeled their SEO experience. A resume cannot distinguish them, and a knowledge quiz only tests recall, which anyone can cram. What separates a real practitioner is diagnostic reasoning: the ability to look at a page and correctly explain why an engine does or does not cite it, then sequence the fix. This article is a 60-minute test that surfaces that ability directly.
The design principle: put a real URL in front of the candidate and watch them think. Everything below is a way to make their reasoning visible, because reasoning is the skill and vocabulary is not.
Task 1: diagnose a live page (20 minutes)
Give the candidate a real URL, ideally one you already understand, and ask: is this page likely to be cited by an AI engine for its target query, and why or why not. Watch for mechanism. A strong candidate checks whether the content is readable at all (they will ask about rendering, because client-side rendering is the most common silent failure), whether the answer is stated directly and early, whether the entity behind the page is clear, and whether the structure survives chunking. A weak candidate talks about keyword density, meta tags, and generic "quality content" without naming a single specific mechanism. The tell is whether they diagnose the page in front of them or recite a checklist that ignores it.
Task 2: sequence the fixes (15 minutes)
Ask: if this page needs work, what do you fix first, and why that order. This tests whether they understand dependency. The correct reasoning is that readability comes before structure comes before entity comes before content, because each layer depends on the one beneath it, fixing content on a page engines cannot read returns nothing. A strong candidate sequences by dependency and can explain the logic. A weak candidate proposes a flat list of tactics with no order, or jumps straight to "write more content," which is the reflex of someone who sells content rather than diagnoses problems. This mirrors how the real signal set layers.
Task 3: define the measurement (10 minutes)
Ask: how would you know if your work succeeded. This is the fastest disqualifier. A real practitioner measures citation frequency and share of model, and understands that citations are probabilistic and read as patterns over repeated queries, not as fixed positions. A candidate who answers in keyword rankings is bringing SEO measurement to an AEO problem, which means they have not internalized the difference. Listen for whether they know that a single query result is noise and only the pattern is signal.
Task 4: state the limits (5 minutes)
Ask: what would you refuse to promise a client. The answer you want is that they will not guarantee a specific citation count or "AI ranking," because no one controls the systems. A candidate who promises guaranteed outcomes is either misunderstanding the technology or willing to misrepresent it, and both are disqualifying in someone who will speak to your clients. The willingness to name the limit is a maturity signal, it is the same honesty that separates a strong agency from a repackaged one, tested at the individual level.
Task 5: explain one concept simply (10 minutes)
Ask them to explain one AEO concept, chunking, entity, information gain, to a smart non-expert. Teaching clarity is a strong proxy for real understanding, because you cannot explain simply what you only know as vocabulary. A candidate who can make why chunking matters or why engines penalize consensus echo genuinely clear has understood the mechanism. A candidate who can only restate the definition has not. This task also predicts on-the-job value, since much of the work is explaining these tradeoffs to stakeholders.
Scoring the test
You are scoring one thing across all five tasks: does this person reason about mechanism, or recite vocabulary. Mechanism looks like specific, page-level, dependency-aware, honestly-bounded reasoning. Vocabulary looks like keyword talk, flat tactic lists, ranking measurement, confident guarantees, and definitions without explanation. A candidate does not need to get every detail right, the field moves fast, but they need to reason correctly about the page in front of them, which is the skill that transfers.
One warning for the interview room: the strongest AEO candidate will often seem less confident than the weakest, because the person who understands the systems knows what they cannot promise, and the person who does not will guarantee you the moon. If you hire on confidence, you will systematically select the people who understand the field least, because in a domain this probabilistic, calibrated uncertainty is the expert signal and false certainty is the amateur one. Test for reasoning, not for reassurance.
Sources
- Google, AI features and your website: the technical ground truth a candidate's diagnosis should reflect. developers.google.com
- Website AI Score, AEO scoring signals: the layered signal set a strong candidate sequences fixes against. View article
- Website AI Score, share of model vs rank tracking: the measurement literacy Task 3 tests for. View article
- Website AI Score, five citation patterns: the probabilistic reading a real practitioner brings. View article
- Website AI Score, how to evaluate an AEO agency: the same test at the vendor level. View article

